Meta’s advertising business relies on a massive, vertically integrated infrastructure that balances retrieval and ranking models to process over three billion daily active users. Retrieval systems like Andromeda and ranking models like Lattice and the generative ads recommendation model (GEM) function within strict sub-second latency budgets. To improve prediction accuracy, adaptive ranking models dynamically allocate compute resources based on the length of a user’s interaction history, allowing for deeper personalization. This system requires constant hardware-software co-design, where custom silicon and specialized hardware SKUs are optimized for specific workloads. By leveraging large language models to generate performance-optimized software kernels, the infrastructure team efficiently manages a heterogeneous fleet of hardware, ensuring that complex machine learning models remain cost-effective while delivering high-precision ad targeting. Matt Steiner, VP of Monetization Infrastructure at Meta, details this ongoing effort to align hardware evolution with sophisticated AI training and inference requirements.
Sign in to continue reading, translating and more.
Open full episode in Podwise
