YouTube08 Jul 2026
13m

Your LLM Deception Monitor Is Broken. The Fix Is in the Training Data - Sachin Kumar, LexisNexis

Podcast cover

AI Engineer

Sleeper agents in fine-tuned large language models present a critical security risk, as backdoors remain dormant during standard behavioral evaluations and only activate upon specific, benign triggers. Traditional monitoring, including joint feature analysis, fails to isolate these malicious signals because they are diluted by the model's overall semantic representation. A more effective approach involves calculating the activation delta between base and fine-tuned models, then applying a sparse autoencoder—termed DiffSAE—to this difference. This method isolates the backdoor as a distinct, interpretable directional shift in activations, providing a high signal-to-noise ratio. By integrating this technique into development pipelines, developers can detect backdoors with near-zero false positives, offering a robust, computationally efficient defense against hidden malicious behaviors that evade conventional safety testing.

Outlines

Sign in to continue reading, translating and more.

Open full episode in Podwise