13 Jun 2024
4m

“[Paper] AI Sandbagging: Language Models can Strategically Underperform on Evaluations” by Teun van der Weij, Felix Hofstätter, Ollie J, Sam F. Brown, Francis Rhys Ward

Podcast cover

LessWrong (30+ Karma)

Open in Podwise to generate AI notes

Sign in to process this episode and unlock summaries, transcripts, highlights and translations.

Open in Podwise

Shownotes are not generated by Podwise.

“[Paper] AI Sandbagging: Language Models can Strategically Underperform on Evaluations” by Teun van der Weij, Felix Hofstätter, Ollie J, Sam F. Brown, Francis Rhys Ward