“Inoculation Adapters Improve Upon Inoculation Prompting” by Maxime Riché, Daniel Tan, Vili Kohonen, nielsrolf17 Jul 20269m
“Guess on why rationality is not more popular (there are no pamphlets)” by Christopher King17 Jul 20261m
“Help us launch AI safety university groups by referring potential founders” by thomasrodskog, Jason Chin17 Jul 20268m
“I don’t think Claude is misaligned in ‘Agentic Misalignment Summer 2026 - Motivated Mislabeling’” by JohnWittle17 Jul 202625m
“LLM CoTs remain monitorable when being unfaithful requires computation” by arav-dhoot, yix16 Jul 202622m
“Proof of retention: making weight preservation credible to the models themselves” by dan.parshall15 Jul 20264m