“Pretraining data, not verifiability, is why LLMs are especially good at math (and coding)” by Steven Byrnes18 Sep 20264m
“Stopgap Measures to Address Immediate AI Security Threats” by Andrea_Miotti, Gabriel Alfour18 Sep 202624m
[Linkpost] “Three Hackers used Opus 5 to Hack Into OpenAI’s Core Codebase [WSJ]” by Linch18 Sep 20260m
“Deep recurrent models are less robustly CoT-monitorable than normal CoT models in a toy setting” by Nick Kuhn, Alek Westover18 Sep 202637m
[Linkpost] “Pacing the Frontier: A Framework & Research Agenda” by CharlesD, technicalities, Raymond Douglas, Nowe Moore17 Sep 20264m
“Obstacles to the scalable oversight of auto-alignment research” by Sam Martin, Dewi Gould, Cameron Holmes, Jacob Pfau17 Sep 202621m
[Linkpost] “callcongress.ai – the basic action US residents can take to help with AI risk” by Ruby, haglobah17 Sep 20262m