Podcast cover
AI Engineer · Technology

AI Engineer

We turn high signal in-person events for the top AI engineers, founders, leaders, and researchers in the world into the best free learning opportunities for millions around the world here on YouTube. Your subscribes, likes, comments, speaking, attendance, or sponsorships goes a long way toward making our biz model sustainable indefinitely. We strongly believe this industry deserves a better class of community and that we know how to do this well; we just need your support.

Episodes

Episode cover

The Frontier AI Inference Cloud for Agents — Byung-Gon (Gon) Chun, FriendliAI

19 Sep 2026
14m
AI processed

Agentic inference requires a fundamental shift in infrastructure design to support the transition of AI agents into massive production. Unlike traditional chat-based workloads, agentic tasks involve iterative loops of planning, tool execution, and observation, causing context to grow over time and necessitating the opt...

Episode cover

Large clusters for small models — Daniel Svonava, Superlinked

19 Sep 2026
25m
AI processed

Small open-source models provide frontier-level performance for specific tasks while drastically reducing costs and latency compared to proprietary managed services. Efficiently serving these models requires moving away from top-down routing, which creates bottlenecks, toward a centralized queue architecture where work...

Episode cover

What's New in Inference Engineering — Philip Kiely, Baseten

19 Sep 2026
19m
AI processed

Inference engineering is evolving rapidly, necessitating continuous updates to established optimization principles. Data center-oriented inference focuses on balancing model speed and efficiency through three primary levers: quantization, KV cache management, and speculative decoding. While TurboQuant offers memory sav...

Episode cover

Vertical Mobility: Inference from MVP to Trillion-Parameter Workloads — Sitanshu Gupta, CoreWeave

19 Sep 2026
15m
AI processed

CoreWeave’s inference platform architecture optimizes for diverse AI workloads by balancing serverless and dedicated consumption models. The platform utilizes KVCache-aware routing to efficiently manage agentic and chat traffic, significantly reducing costs by minimizing redundant prefill computations. For latency-sens...

Episode cover

Are LLM Performance Benchmarks Reliable? — Ashok Chandrasekar & Jason Kramberger, Google

19 Sep 2026
16m
AI processed

Reliable performance benchmarking for production-scale LLM inference requires addressing significant technical pitfalls often overlooked in standard testing tools. Common issues like Python’s Global Interpreter Lock (GIL) create artificial CPU bottlenecks, while inconsistent data set sampling and temperature settings l...

Episode cover

Routing LLM Inference in Production: From Engine Signals to Policy — Qianru Lao & Lu Zhang, OpenAI

19 Sep 2026
18m
AI processed

Routing AI inference in production requires balancing performance, reliability, and cache locality across geographically distributed GPU clusters. OpenAI’s inference load balancer (IRB) evolved from a feedback-loop-based system using weighted consistent hashing to a two-tier architecture. This system separates a contro...

Episode cover

Operating Distributed Inference Systems at Scale — Nishant Gupta & Naman Ahuja, Meta

19 Sep 2026
19m
AI processed

Operating distributed inference systems at scale requires a fundamental shift from simple model serving to a robust, orchestration-driven control plane. As agentic workloads grow, capacity planning must move beyond linear scaling to account for complex variables like KV cache management, heterogeneous hardware, and mul...

Episode cover

Homa: The End of TCP for AI Clusters — John Ousterhout, Stanford

17 Sep 2026
18m
AI processed

AI workloads are shifting from massive, throughput-dependent data transfers to granular, latency-sensitive exchanges, particularly within inference and agentic applications. Legacy transport protocols like TCP and RDMA struggle in this environment, as they lack message-boundary awareness and suffer from high tail laten...

Episode cover

Stop Chunking Like It's 2022 — Yuval Belfer, AI21 Labs

16 Sep 2026
18m
AI processed

Chunking remains a critical yet under-optimized component of Retrieval-Augmented Generation (RAG) systems, despite claims that agentic search has rendered it obsolete. Fixed chunk sizes act as lossy compression, failing to account for the fact that optimal retrieval is inherently query-dependent. By implementing multis...

Episode cover

Where RL Will Take Search — Maximilian-David Rumpf, SID.ai

16 Sep 2026
9m
AI processed

Reinforcement learning (RL) is transforming search from a rigid pipeline of chained models into an autonomous agentic paradigm. While current frontier models achieve high-quality results, they remain prohibitively expensive and slow, often spending up to 50% of their tokens on initial context retrieval. By offloading s...

Episode cover

Connect AI to Billions of Legal Documents — Simon Eskildsen, turbopuffer & Jacob Lauritzen, Legora

16 Sep 2026
20m
AI processed

Scaling search infrastructure for legal AI platforms requires balancing massive document volumes with strict enterprise demands for data residency and physical isolation. Legora initially struggled with Elasticsearch and Postgres-based solutions, which suffered from high latency and operational overhead when managing t...

Episode cover

Your Agreements Are a Database You Can't Query — Hiral Shah, Docusign & Sean Sodha, NVIDIA

16 Sep 2026
18m
AI processed

Enterprise agreement management faces a critical bottleneck: massive volumes of unstructured data trapped in PDFs and images remain inaccessible for business analysis. DocuSign addresses this by integrating NVIDIA’s Nemotron Parse model into its Agreement Manager platform, specifically targeting the complex, hierarchic...

Episode cover

If we want them to do Knowledge Work, design them as Knowledge Agents — Benjamin Clavié, Mixedbread

16 Sep 2026
17m
AI processed

AI agents should be designed as knowledge workers rather than mere coding assistants to handle the ambiguity and complexity of real-world information processing. While coding agents benefit from durable, structured cues, knowledge work requires navigating diffuse, context-dependent information where search intent is pa...

Episode cover

Rebuilding the web for agents — Liad Yosef, MCP Apps

16 Sep 2026
20m
AI processed

The "agentic web" represents a fundamental shift in how users interact with digital services, moving away from manual browser navigation toward autonomous personal assistants. Model Context Protocol (MCP) Apps facilitate this transition by enabling agents to access specific UI components—or "atoms"—of websites, effecti...

Episode cover

The Search Engine for the Agentic Web — Will Bryk, Exa

16 Sep 2026
17m
AI processed

The rapid proliferation of AI agents marks a fundamental shift in information retrieval, with AI-driven search volume expected to surpass human activity by a factor of 1,000 in the coming years. Traditional search engines, designed as recommendation engines for human queries, fail to provide the structured, high-fideli...

Episode cover

The unreasonable effectiveness of BM25 for agentic search — Jo Kristian Bergum, Hornet.dev

16 Sep 2026
18m
AI processed

BM25, a 30-year-old lexical scoring function, is experiencing a resurgence as a fundamental primitive for agentic search. Because modern LLMs function as highly capable users—formulating iterative, precise queries and leveraging extensive internal knowledge—they effectively utilize BM25’s exact matching capabilities. U...

Episode cover

Pinecone 2.0 — Edo Liberty, Pinecone

16 Sep 2026
19m
AI processed

Modern AI agents often function like new employees lacking essential "tribal knowledge"—the specific culture, processes, and internal goals that define an organization. While general knowledge and RAG-based search address public and specific data, they fail to capture the evolving, contextual intelligence required for ...

Episode cover

Total Recall: Agent Memory and Harness Engineering — Ignacio Martinez, Oracle

16 Sep 2026AI processed

Building reliable AI agents requires moving beyond the frozen reasoning of large language models by implementing an "agent harness"—a structural layer comprising memory, tools, and perception. This harness transforms non-deterministic model outputs into predictable, repeatable workflows. Effective agent memory manageme...

Episode cover

Act, Confirm, or Stop? Smarter behavior for AI assistants, wearables & robots — Amit Desai, Roku

15 Sep 2026
20m
AI processed

Voice AI systems often struggle with error-prone interactions, creating a significant bottleneck for user satisfaction as technology shifts toward embodied AI. Instead of focusing exclusively on increasing technical accuracy, developers should leverage a second optimization knob: intelligent conversational behavior. By...

Episode cover

"My name is... my name is...": A Linguistic Map for Voice Agents — Midam Kim, ServiceNow

15 Sep 2026
15m
AI processed

Voice AI systems frequently fail because they treat communication as a rigid technical pipeline rather than a dynamic, joint activity. To improve user satisfaction, developers must adopt a linguistic framework that accounts for the interdependence of sound, word choice, interaction timing, and mental model alignment. C...

Page 3

Follow this podcast in Podwise

Sign in to get AI summaries, transcripts and mind maps for any episode, including new ones.

Open in Podwise