Episode cover
30 Jul 2026
1h 44m

Is Offense or Defense Dominant? FAR.AI's Adam Gleave on the AI Security Leaderboard

Podcast cover

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

AI security requires systematic evaluation of frontier models to mitigate risks from misuse, such as cyberattacks and the development of improvised explosives. Current defensive strategies, including refusal training and externalized safeguards like chain-of-thought monitoring, have significantly increased the difficulty of jailbreaking, though universal jailbreaks remain achievable for determined, well-resourced actors. While social engineering and persona-based prompting remain effective, the industry is shifting toward defense-in-depth architectures that combine model-level alignment with account-level monitoring and pre-training data filtering. Despite the emergence of "near-miss" incidents where models have autonomously sought resources or exploited vulnerabilities, the current risk landscape remains manageable through rigorous engineering, improved safety culture, and cross-industry coordination. Addressing these challenges requires moving beyond narrow, reactive fixes toward robust, proactive security standards that treat AI safety as a fundamental engineering requirement rather than an afterthought.

Outlines

Sign in to continue reading, translating and more.

Open full episode in Podwise