YouTube12 Nov 2023
1h 10m

Trends in Deep Learning Hardware: Bill Dally (NVIDIA)

Podcast cover

Paul G. Allen School

Bill Dally's lecture explores hardware advancements driving deep learning, particularly focusing on NVIDIA's contributions and future directions. He highlights that deep learning's current revolution was enabled by hardware, emphasizing algorithms, data, and computing power as key ingredients. Dally introduces Huang's Law, noting the doubling of inference performance every year for the last decade, largely due to smaller number usage, complex instructions, and process technology improvements. He also touches on the importance of software in making hardware useful, referencing NVIDIA's CUDNN. Looking ahead, Dally discusses optimizing number representation through logarithmic numbers, optimal clipping, and scaling granularity, as well as exploring sparsity and improving memory bandwidth with 3D memory. He also shares insights from accelerator projects like Magnet and Magnetic BERT, demonstrating significant efficiency gains. The lecture concludes with a Q&A session, addressing network optimization techniques, energy savings in complex instructions, and the role of software in clipping techniques.

Outlines

Part 1: Introduction, Background

Part 2: LLM Training, Inference, Methods

Part 3: Architecture, Performance, Huang's Law

Part 4: Scaling, Infrastructure, Software

Part 5: Energy, Efficiency, Hardware Optimization

Part 6: Number Systems, Math, Representations

Part 7: Clipping, Scaling, Sparsity

Part 8: Memory, Overhead, Future Outlook

Sign in to continue reading, translating and more.

Open full episode in Podwise