AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Latent Space
AI systems introduce unique security vulnerabilities that differ fundamentally from traditional software, necessitating specialized defense strategies. Zico Kolter and Matt Fredrikson, founders of Gray Swan, explain that as AI agents gain autonomy and access to sensitive tools, they become susceptible to indirect prompt injection and correlated failures. To mitigate these risks, they developed "Shade," an automated red teaming model that outperforms human testers in identifying vulnerabilities, and "Signal," a defensive filter designed to enforce custom enterprise policies. Rather than relying on naive scaling for safety, effective security requires explicit, specialized training and the integration of AI-driven interpretability. By automating the science of security and interpretability, organizations can better navigate the trade-off between model capability and operational safety, moving toward more robust, policy-compliant AI deployments in enterprise environments.
Sign in to continue reading, translating and more.
Open full episode in Podwise
