
System Design for Next-Gen Frontier Models — Dylan Patel, SemiAnalysis
This podcast focuses on the challenges of large language model (LLM) inference, particularly for models with trillions of parameters. The speaker discusses the computational demands of pre-fill and decode processes, highlighting the need for techniques like continuous batching and disaggregated pre-fill to improve eff...



















