Podcast cover
Open Compute Project ยท Technology

Open Compute Project

This is the place to find all the OCP related videos from OCP Summits, Hackathons, Video Blogs and more. What is OCP all about? Hacking Conventional Computing Infrastructure The Open Compute Project Foundation is a 501(c)6 organization which was founded in 2011 by Facebook, Intel, Rackspace, Goldman Sachs and Andy Bechtolsheim. Our mission is to apply the benefits of open source software to hardware and rapidly increase the pace of innovation to hardware design and engineering. Why Open Hardware? By releasing Open Compute Project technologies as open hardware, our goal is to develop servers and data centers following the model traditionally associated with open source software projects.

Episodes

Episode cover

Power Solutions for AI Data Centers presented by Delta

22 Oct 2024
14m
Episode cover

Pioneering the Modern Datacenter with DC MHS Architecture Presented by Micro Star Intl

22 Oct 2024
12m
Episode cover

PCIe Retimers Performance Matters Presented by Credo

22 Oct 2024
14m
Episode cover

Partnering with NVIDIA to Deliver Rack Scale AI Servers Presented by Ingrasys

22 Oct 2024
15m
Episode cover

Overview of Ultra Ethernet Presented by UEC

22 Oct 2024
15m
AI processed

In this podcast episode, J Metz, the chair of the Ultra Ethernet Consortium (UEC), shares the organization's mission and recent progress. The UEC aims to foster an open Ethernet ecosystem specifically designed for artificial intelligence and high-performance computing. As it expands, the consortium is tackling key chal...

Episode cover

ODM+ Lenovos Tailored Experience for Customers Presented by Lenovo

22 Oct 2024
15m
Episode cover

Modular OCP based Direct Liquid Cooling Infrastructure for future AI Applications Presented

22 Oct 2024
11m
Episode cover

The Challenges and Practices of Network Stability in Alibabas Large Scale Computing Clusters

22 Oct 2024
18m
AI processed

In this podcast episode, the discussion centers on the intricate challenges faced by Alibaba's distributed training network, particularly around fault detection and communication efficiency during the training process. The speakers shed light on the common occurrences of network failures in large-scale distributed envi...

Episode cover

Optimal path utilization for multi plane fabric design

22 Oct 2024
20m
Episode cover

New approaches to network telemetry Essential for AI performance

22 Oct 2024
11m
AI processed

In this podcast episode, we explore innovative ways to boost AI training efficiency using advanced telemetry techniques. Roop emphasizes the vital role of pinpointing and tackling performance bottlenecks, explaining how even small delays can cause major setbacks in training. The conversation introduces an intriguing me...

Episode cover

Meta 51.2T Ethernet Switch

22 Oct 2024
14m
AI processed

This podcast episode explores Meta's state-of-the-art 51T Ethernet Switch, featuring two impressive models: the MiniPak 3 and Cisco's 8501. Both are shining examples of innovative network hardware design and performance. The conversation dives into the MiniPak 3's compact design, improved processing power, and effectiv...

Episode cover

Leveraging open technologies to monitor packet drops in AI cluster fabrics

22 Oct 2024
20m
Episode cover

Insights from Production Scheduled Ethernet Fabric in Large AI Training Clusters

22 Oct 2024
23m
AI processed

In this podcast episode, the hosts explore the complex challenges faced during AI training, especially the pressure on communication systems to efficiently transmit data across multiple GPUs. They introduce the Scheduled Ethernet Fabric, an advanced scheduling solution that boosts network performance by reducing latenc...

Episode cover

Best Practices for Liquid & Air Cooling of a 51.2Tbps Switch for High-Density AI Clusters

22 Oct 2024
18m
Episode cover

Fabric resiliency at scale

22 Oct 2024
9m
Episode cover

Be Ready for CPO Integrating and Enhancing CPO Switches with SONiC

22 Oct 2024
11m
Episode cover

ALPINE SONiC Switchstack Simulation

22 Oct 2024
19m
Episode cover

Alibaba HPN: A Data Center Network for Large Language Model Training

22 Oct 2024
19m
AI processed

In this podcast episode, we explore Alibaba's HPN 7.0 topology, a state-of-the-art network architecture designed to improve the training of large language models (LLMs) while tackling the challenges of scaling. Jiaqi Gao shares insightful innovations within HPN 7.0, such as its dual-plane design and computation-communi...

Episode cover

Update on HW Fault Management Project Activities

22 Oct 2024
30m
Episode cover

Standardizing Hyperscaler Requirements for Accelerators

22 Oct 2024
27m
Page 16

Follow this podcast in Podwise

Sign in to get AI summaries, transcripts and mind maps for any episode, including new ones.

Open in Podwise