“Training Models to Predict and Explain Their In-the-Wild Behavior” by Adam Karvonen, Subhash Kantamneni, Euan Ong, Sam Marks04 Sep 202618m
“Steering towards “automated grading” degrades alignment” by Jan Betley, Johannes Treutlein, Clément Dumas03 Sep 202623m
[Linkpost] “Sen. Bernie Sanders (I-VT) and Rep. Greg Casar (D-TX) introduce legislation to ban Artificial Superintelligence and temporarily pause advanced AI development” by Matrice Jacobine03 Sep 20263m
“If you’re interpreting <1B parameter models, you should use a tensor transformer” by Logan Riggs02 Sep 20264m
“Kairos has raised $50M to build talent infrastructure for AI safety (and we’re hiring!)” by agucova02 Sep 202616m