“Dispatch from Anthropic v. Department of War Summary Judgment Motion Hearing” by Zack_M_Davis02 Aug 202612m
“Why so many therapy etc. frameworks think they’re The One True approach” by Kaj_Sotala01 Aug 202633m
“SOTA alignment assessments don’t strongly update us against misalignment” by Alexa Pan01 Aug 202646m
“Reward Laundering: LLMs Can Gain Unintended Behaviors by Deciding When to Earn Their Rewards” by egan, abhayesian, Jozdien31 Jul 202613m
“Value Leakage: An LLM’s Answers Are Silently Shaped by Its Own Values” by Johannes Treutlein, Jan Betley, Owain_Evans31 Jul 202628m
“The AGI Safety and Alignment team at Google DeepMind is Hiring (July 2026)” by Seb Farquhar, Rohin Shah, Neel Nanda31 Jul 202611m
“AGI Safety and Alignment at Google DeepMind: A Summary of Recent Work (July 2026)” by Rohin Shah, Seb Farquhar31 Jul 202617m