
GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model
Meta has improved the training efficiency of its Generative Ads Recommendation Model (GEM), the foundation model for ads on Instagram and Facebook. They doubled end-to-end training efficiency to 20–25% Model FLOPs Utilization (MFU) while scaling training FLOPs 4x in 12 months.
Why it matters
Improving the efficiency of the AI powering Instagram and Facebook ads allows Meta to scale its recommendation systems more effectively. This enables the delivery of more accurate ad recommendations for users and advertisers.
The details
- Custom kernels like Jagged Flash Attention eliminate compute waste from variable sequence lengths.
- MXFP8 ultra-low-precision training increases throughput without regressing the quality of ad predictions.
- 5D parallelism matches communication volume to available bandwidth across Meta's network hierarchy.
Get the weekly recap
The stories like this one, picked and explained — once a week, straight to your inbox.