23 Jun 2024
15m
“Sycophancy to subterfuge: Investigating reward tampering in large language models” by evhub, Carson Denison
LessWrong (30+ Karma)
Open in Podwise to generate AI notes
Sign in to process this episode and unlock summaries, transcripts, highlights and translations.
Shownotes are not generated by Podwise.
