“Recommendations for People Getting into Technical AI Governance Research” by Aaron_Scher, yams, peterbarnett, Naci Cankaya09 Sep 20261m
[Linkpost] “Estimating GPT-6 Astra’s no-CoT Time Horizon” by Francis Rhys Ward, Dewi Gould09 Sep 202619m
“Exploration Hacking in AI Debate: Initial Empirics and Generalisation Splitting” by Jason R Brown, Nathalie Kirch, Joschka Braun, hyannakoudakis, David Lindner09 Sep 202627m
“A Conceptual Framework for Reasoning about Exploration Hacking” by Jason R Brown, Nathalie Kirch, Joschka Braun, hyannakoudakis, David Lindner09 Sep 202635m
“Pausing AI ASAP is preferable to agreeing to pause at some future time” by Connor Williams09 Sep 20262m
“OpenAI have solved the Navier-Stokes Problem with a substantially more powerful model than Astra.” by fluxxrider08 Sep 20260m
[Linkpost] “Frontier models still hack on simple variations of alignment evals from early 2025” by Dean Valentine08 Sep 20263m
“Psychological Support for AI Safety Researchers Is Neglected and Easy to Provide” by Ihor Kendiukhov08 Sep 202615m