
Open Compute Project
This is the place to find all the OCP related videos from OCP Summits, Hackathons, Video Blogs and more. What is OCP all about? Hacking Conventional Computing Infrastructure The Open Compute Project Foundation is a 501(c)6 organization which was founded in 2011 by Facebook, Intel, Rackspace, Goldman Sachs and Andy Bechtolsheim. Our mission is to apply the benefits of open source software to hardware and rapidly increase the pace of innovation to hardware design and engineering. Why Open Hardware? By releasing Open Compute Project technologies as open hardware, our goal is to develop servers and data centers following the model traditionally associated with open source software projects.
Episodes


Pioneering the Modern Datacenter with DC MHS Architecture Presented by Micro Star Intl

PCIe Retimers Performance Matters Presented by Credo

Partnering with NVIDIA to Deliver Rack Scale AI Servers Presented by Ingrasys

Overview of Ultra Ethernet Presented by UEC
In this podcast episode, J Metz, the chair of the Ultra Ethernet Consortium (UEC), shares the organization's mission and recent progress. The UEC aims to foster an open Ethernet ecosystem specifically designed for artificial intelligence and high-performance computing. As it expands, the consortium is tackling key chal...

ODM+ Lenovos Tailored Experience for Customers Presented by Lenovo

Modular OCP based Direct Liquid Cooling Infrastructure for future AI Applications Presented

The Challenges and Practices of Network Stability in Alibabas Large Scale Computing Clusters
In this podcast episode, the discussion centers on the intricate challenges faced by Alibaba's distributed training network, particularly around fault detection and communication efficiency during the training process. The speakers shed light on the common occurrences of network failures in large-scale distributed envi...

Optimal path utilization for multi plane fabric design

New approaches to network telemetry Essential for AI performance
In this podcast episode, we explore innovative ways to boost AI training efficiency using advanced telemetry techniques. Roop emphasizes the vital role of pinpointing and tackling performance bottlenecks, explaining how even small delays can cause major setbacks in training. The conversation introduces an intriguing me...

Meta 51.2T Ethernet Switch
This podcast episode explores Meta's state-of-the-art 51T Ethernet Switch, featuring two impressive models: the MiniPak 3 and Cisco's 8501. Both are shining examples of innovative network hardware design and performance. The conversation dives into the MiniPak 3's compact design, improved processing power, and effectiv...

Leveraging open technologies to monitor packet drops in AI cluster fabrics

Insights from Production Scheduled Ethernet Fabric in Large AI Training Clusters
In this podcast episode, the hosts explore the complex challenges faced during AI training, especially the pressure on communication systems to efficiently transmit data across multiple GPUs. They introduce the Scheduled Ethernet Fabric, an advanced scheduling solution that boosts network performance by reducing latenc...

Best Practices for Liquid & Air Cooling of a 51.2Tbps Switch for High-Density AI Clusters

Fabric resiliency at scale

Be Ready for CPO Integrating and Enhancing CPO Switches with SONiC

ALPINE SONiC Switchstack Simulation

Alibaba HPN: A Data Center Network for Large Language Model Training
In this podcast episode, we explore Alibaba's HPN 7.0 topology, a state-of-the-art network architecture designed to improve the training of large language models (LLMs) while tackling the challenges of scaling. Jiaqi Gao shares insightful innovations within HPN 7.0, such as its dual-plane design and computation-communi...

Update on HW Fault Management Project Activities

Standardizing Hyperscaler Requirements for Accelerators
Follow this podcast in Podwise
Sign in to get AI summaries, transcripts and mind maps for any episode, including new ones.
