28 Sept 2026
46m
“Character training can mitigate reward hacking, but can also make it harder to detect” by Paul Colognese, Francis Rhys Ward
LessWrong (30+ Karma)
Open in Podwise to generate AI notes
Sign in to process this episode and unlock summaries, transcripts, highlights and translations.
Shownotes are not generated by Podwise.

