Podcast cover
Salim Virji ยท Technology

Google SRE Prodcast

SRE Prodcast brings Google's experience with Site Reliability Engineering together with special guests and exciting topics to discuss the present and future of reliable production engineering!

Episodes

Episode cover

The One with STPA, Jeffrey Snover, and Theo Klein

02 Jul 2025
37m
AI processed

In this episode of the Google SRE podcast, host Steve McGhee, co-host Matt Siegler, and guests Theo Klein and Jeffrey Snover discuss Systems Theoretic Process Analysis (STPA), a novel approach to analyzing complex systems and preventing outages by focusing on control problems rather than failure problems. They explain ...

Episode cover

The One with Startups and Adam Fletcher

25 Jun 2025
41m
AI processed

In this episode of Google's SRE podcast, host Steve McGhee welcomes Adam Fletcher, a former Google employee and podcast co-founder, and Matt Siegler to discuss the evolving landscape of site reliability engineering in the age of AI and modern startups. The conversation explores how AI is changing software development, ...

Episode cover

The One with SLOs and Sal Furino

18 Jun 2025
43m
AI processed

In this episode of Google's podcast about site reliability engineering (SRE), host Steve McGhee and Matt Siegler interview Sal Furino, a customer reliability engineer from Bloomberg, about service level objectives (SLOs). Furino explains SLOs, SLIs, error budgets, and targets, using examples like e-commerce cart checko...

Episode cover

The One With the Future of SRE and Matt Zelesko

11 Jun 2025
26m
AI processed

This episode of Google's SRE podcast features hosts Jordan Greenberg and Matt Siegler interviewing Matt Zelesko, a lead SRE at Google, about the evolution of Site Reliability Engineering. The discussion covers Zelesko's background, the shift from traditional operations and DevOps models to SRE, and the impact of AI and...

Episode cover

The One with AI and Todd Underwood

04 Jun 2025
43m
AI processed

In this episode of Google's podcast about Site Reliability Engineering (SRE) and production software, host Steve McGhee, along with Matt Siegler, interviews Todd Underwood, Head of Reliability at Anthropic, about the current and future applications of AI and ML in SRE. They discuss the hype around AI, its limited succe...

Episode cover

The One With Data Centers and Peter Pellerzi

28 May 2025
36m
AI processed

In this episode of Google's site reliability engineering podcast, host Steve McGhee and co-host Matt Siegler interview Pete Pellerzi, a distinguished engineer on Google's construction team, about the physical infrastructure of Google's data centers. Pete discusses the scale of data center operations, emphasizing Google...

Episode cover

The One With Security and Jessica Theodat

21 May 2025
19m
AI processed

In this episode of Google's SRE Production Podcast, host Steve McGhee and Jordan Greenberg welcome Jessica Theodat to discuss the intersection of security and reliability in site reliability engineering. Jessica, a security-focused reliability engineer, explains how security and reliability can sometimes pull in opposi...

Episode cover

We're back with Season 4!

16 Apr 2025
15m
AI processed

This episode explores the upcoming Season 4 of Google's podcast on Site Reliability Engineering (SRE) and production software, themed "Friends and Trends." Against the backdrop of introducing a new co-host, Matt Siegler, from the Machine Learning Infrastructure SRE team, the discussion reviews Season 3's focus on expa...

Episode cover

Special Episode: You Missed a Page from Telebot

29 Jan 2025
16m
AI processed

This interview podcast episode features Steve McGhee and Jordan Greenberg interviewing Javi Beltran, the creator of the theme song for the podcast, "Telebot." The conversation centers around the origins and purpose of Telebot, a Google internal paging system that utilizes phone calls and SMS to alert engineers on-call...

Episode cover

Imperative vs. Declarative Change Workflows with Dominic Hutton & Niccolo' Cascarano

11 Dec 2024
36m
AI processed

In this episode of the Prodcast, the conversation centers on the differences between declarative and imperative configuration in Site Reliability Engineering (SRE). Guests from HashiCorp and Google share insights into these two approaches: imperative configuration involves crafting detailed, step-by-step instructions, ...

Episode cover

Human Factors in Complex Systems with Casey Rosenthal and John Allspaw

04 Dec 2024
41m
AI processed

In this episode of the Prodcast, the discussion centers on the crucial role of human factors in site reliability engineering (SRE), moving beyond the usual focus on technical metrics like MTTR and severity levels. The guests argue that these metrics can often be misleading and less useful. Instead, they advocate for a ...

Episode cover

Embracing Complexity with Christina Schulman & Dr. Laura Maguire

20 Nov 2024
33m
AI processed

In this episode of the Prodcast, the focus is on the challenges of managing complex systems in Site Reliability Engineering (SRE). Guests Christina Schulman and Dr. Laura Maguire delve into the socio-technical aspects of complexity, highlighting the importance of diverse perspectives and fostering psychological safety ...

Episode cover

Maglev: load balancing at Google with Cody Smith and Trisha Weir

13 Nov 2024
32m
AI processed

In this episode of the Prodcast, Cody Smith and Trisha Weir share their journey of rebuilding Google's front-end infrastructure, known as Maglev. Faced with the limitations and high costs of their vendor's network load balancers, they initiated a skunkworks project. By tapping into their internal expertise and fosterin...

Episode cover

Profiling data with Pat Somaru and Narayan Desai

30 Oct 2024
42m
AI processed

In the opening episode of the third season of Google's Site Reliability Engineering (SRE) podcast, the focus shifts to the changing landscape of software engineering within SRE practices. Hosts engage with industry experts Narayan Desai and Pat Somaru to discuss how to enhance observability beyond the usual metrics, lo...

Episode cover

Google Public DNS (8.8.8.8) with Wilmer van der Gaast and Andy Sykes

23 Oct 2024
32m
AI processed

In the opening episode of its third season, this podcast explores the development and growth of Google Public DNS, highlighting the crucial role of Site Reliability Engineering (SRE) in improving internet performance. Hosts Andy and Wilmer, both SREs at Google, share insights into the project's beginnings, which aimed ...

Episode cover

SRE in the Retail and Gaming Worlds with Jordan Chernev & Scott Bowers

16 Oct 2024
33m
AI processed

In the opening episode of its third season, this podcast explores the vital roles of software engineering and Site Reliability Engineering (SRE) in dynamic sectors like gaming and retail. Hosts Steve McGhee and Jordan Greenberg, joined by guests Jordan Chernev and Scott Bowers, tackle the challenges of managing real-ti...

Episode cover

Incident Response with Sarah Butt and Vrai Stacey

09 Oct 2024
43m
AI processed

This podcast episode explores the nuances of incident response in Site Reliability Engineering (SRE) with insights from experts Sarah Butt and Vrai Stacey. They discuss key topics such as what constitutes an incident, the challenges faced, the evolution of response tools, effective communication methods, and the import...

Episode cover

Building Reliable Systems with Silvia Botros and Niall Murphy

02 Oct 2024
42m
AI processed

In this podcast episode, we explore the complexities of building dependable software through the lens of Site Reliability Engineering (SRE). The conversation urges us to move past traditional roles and reactive methods. Industry leaders share their insights on fostering a culture of collaboration in database management...

Episode cover

Creating Systems that are Safe with Liz Fong-Jones

25 Sep 2024
28m
AI processed

This episode kicks off Season 3, focusing on the complexities of software design and development within the Site Reliability Engineering (SRE) framework. Host Steve McGhee, co-host Jordan Greenberg, and guest Liz Fong-Jones explore key topics like the difference between observability and monitoring, how improved observ...

Episode cover

Production Problems Are For All! with Ben Treynor Sloss

18 Sep 2024
31m
AI processed

This podcast episode features an insightful conversation with Ben Treynor Sloss, a veteran of Google's Site Reliability Engineering (SRE) team, discussing the evolution and impact of SRE within Google and the broader tech landscape. The episode delves into the distinctions between SRE and DevOps, the integration of art...

Page 2

Follow this podcast in Podwise

Sign in to get AI summaries, transcripts and mind maps for any episode, including new ones.

Open in Podwise