Podcast cover
Salim Virji · Technology

Google SRE Prodcast

SRE Prodcast brings Google's experience with Site Reliability Engineering together with special guests and exciting topics to discuss the present and future of reliable production engineering!

Episodes

Episode cover

Adam Kramer discusses Incident Response

08 Jul 2026
15m
AI processed

Managing large-scale cloud infrastructure requires evolving beyond simple virtual machine provisioning to orchestrating complex, integrated systems including networks, firewalls, and load balancers. Google Compute Engine exemplifies this shift, operating as a "data center in a box" that necessitates specialized SRE tea...

Episode cover

Courtney Nash on Complex Systems

16 Jun 2026
7m
AI processed

The necessity of human expertise increases alongside the adoption of artificial intelligence in complex systems. Courtney Nash, co-founder of The Void and an expert in cognitive science, argues that while AI is often marketed as a tool to simplify operations, the resolution of critical failures—such as metastable failu...

Episode cover

John Allspaw at SREcon Part 1

09 Jun 2026
8m
AI processed

Expertise in software engineering functions similarly to gravity—it is a ubiquitous yet often overlooked force that drives both success and failure. While organizations invest heavily in interviewing to gauge candidate expertise, they frequently fail to analyze how that knowledge is maintained or transferred beyond for...

Episode cover

John Allspaw at SREcon Part 2

09 Jun 2026
9m
AI processed

SREcon Americas serves as a critical practitioner-focused gathering where the messy, non-linear reality of production software is prioritized over idealized theoretical models. John Allspaw, a resilience engineering expert and former Velocity Conference chair, highlights that the "no bullshit" culture of the SRE commun...

Episode cover

Ricard Bejarano: The SRE At Home

02 Jun 2026
7m
AI processed

Home labbing serves as a powerful catalyst for professional growth and technical mastery within the site reliability engineering field. Ricard Bejarano, a lead SRE at Cisco’s Thousand Eyes and author of The Home Lab Handbook, explains how transitioning from Docker to Kubernetes in a personal environment directly transl...

Episode cover

Matt Zelesko and the Future of SRE

26 May 2026
23m
AI processed

The rapid integration of AI and agentic workflows is fundamentally reshaping the role of Site Reliability Engineers (SREs), shifting the profession from human-centric operations to human-supervised automation. This transition demands a move toward broader, generalist capabilities, as agents increasingly handle code cre...

Episode cover

Handling Burnout with Sam Anderson

21 May 2026
10m
AI processed

Burnout in the SRE industry manifests through symptoms like diminished motivation, emotional exhaustion, and a persistent inability to accept help. Addressing this requires applying systems engineering principles to personal well-being: treat the self as a system by monitoring indicators of health—similar to SLIs—and i...

Episode cover

The One with Crisis Engineering and Mikey Dickerson

15 May 2026
43m
AI processed

Crisis management requires recognizing that organizations only abandon entrenched, ineffective habits when faced with an existential threat. During these brief windows of total disintegration, leaders can steer teams toward necessary structural adaptations. Site Reliability Engineers are uniquely positioned to act as f...

Episode cover

This is Fine! With Colette Alexander and Clint Byrum

12 May 2026
9m
AI processed

Mean Time To Repair (MTTR) remains a problematic metric in site reliability engineering because it lacks statistical significance and fails to capture actual business impact. Incidents of varying durations often result in vastly different outcomes, rendering time-based averages misleading for operational decision-makin...

Episode cover

The One With Damion Yates and Building AI systems

26 Feb 2026
31m
AI processed

The podcast explores the evolution of reliability engineering within Google DeepMind, focusing on the unique challenges and adaptations required in a research-oriented environment. Damion Yates, based in London at Google DeepMind, recounts his initial role in introducing reliability concepts to a team primarily focused...

Episode cover

The One With Carla Geisser and Crisis Engineering

11 Feb 2026
25m
AI processed

The podcast explores the concept of "crisis engineering," distinguishing it from standard incident response with its focus on novel, high-stakes situations lacking established playbooks. Carla Geisser from LayerEleph, a crisis engineering firm, defines a crisis by five elements: fundamental surprise, broken critical fu...

Episode cover

The One with Parker Barnes, Felipe Tiengo Ferreira, and AI

05 Feb 2026
36m
AI processed

The podcast explores the challenges and novel tooling involved in ensuring the safety and reliability of AI models in production. It highlights the shift from well-defined safety boundaries in traditional systems to the open-ended interactions of modern AI, making safety a squishy continuum encompassing quality and use...

Episode cover

The One With Shannon Brady and Operating Systems

28 Jan 2026
24m
AI processed

The podcast explores the gLinux platform at Google, focusing on managing and securing a large fleet of devices. Shannon Neufeld-Brady from the gLinux team discusses the team's shift from traditional SRE roles while maintaining SRE principles. Key strategies include rigorous testing, canarying, and monitoring, along wit...

Episode cover

The One With Denia Del Cid and AI

21 Jan 2026
29m
AI processed

The podcast explores the application of AI in site reliability engineering (SRE), particularly focusing on how AI can be used to reduce toil and improve incident management. Denia del Cid, an SRE at Google, discusses her work on using AI to analyze support cases and ticket queues, enabling earlier detection of outages ...

Episode cover

The One With Heather Adkins and Security (and AI)

14 Jan 2026
24m
AI processed

The podcast explores the intersection of site reliability engineering (SRE) and cybersecurity, particularly during incidents. Heather Adkins from Google's Office of Cybersecurity Resilience, highlights the key difference: security incidents are often instigated by malicious actors, requiring anticipation of their behav...

Episode cover

The One With SLOs

07 Jan 2026
38m
AI processed

The podcast explores the practical application of Service Level Objectives (SLOs) in diverse organizational settings, addressing common challenges and misconceptions. It emphasizes that successful SLO adoption requires ownership and participation from the teams responsible for both writing and running the code, rather ...

Episode cover

The One With Steph Hippo and Observability

16 Dec 2025
33m
AI processed

In this episode of the Google SRE podcast, Matt Siegler and Florian Rathgeber interview Steph Hippo, Platform Engineering Director at Honeycomb, about the influence of AI on observability and SRE practices. Steph discusses how AI and observability create a self-reinforcing loop, with AI providing insights that enhance ...

Episode cover

The One with Ben Good and Our Kubernetes Friends

30 Jul 2025
32m
AI processed

In this episode of Google's podcast about site reliability engineering, Steve McGhee and Kaslin Fields host Ben Good to discuss platform engineering and its relationship to Kubernetes. They explore how Kubernetes serves as a foundation for building platforms, emphasizing that it's a key tool but not the only requiremen...

Episode cover

The One With AI Agents, Ramón Llamas, and Swapnil Haria

23 Jul 2025
42m
AI processed

This episode of Google's SRE podcast, "Prodcast," features host Steve McGhee, Matt Siegler, and guests Ramon and Swapnil Haria, discussing the emerging trend of using AI agents in site reliability engineering and production software. The conversation covers the spectrum of AI agents, from simple LLM prompts to complex ...

Episode cover

The One with Technical Program Managers and Karanveer Anand

16 Jul 2025
27m
AI processed

In this episode of Google's podcast on SRE and production software, host Steve McGhee and Jordan Greenberg interview Karanveer Anand, a technical program manager (TPM) with a background in SRE. They discuss the role of TPMs in SRE, highlighting the importance of translating technical concepts into business terms and ma...

Page 1

Follow this podcast in Podwise

Sign in to get AI summaries, transcripts and mind maps for any episode, including new ones.

Open in Podwise