From EPYC to Helios, AMD and Meta are Scaling the Future of AI Together
Three things you need to know:
Meta and AMD are co-engineering AI infrastructure across the full stack – spanning AMD Helios™ rackscale solutions, AMD EPYC™ CPUs, AMD Instinct™ GPUs, AMD Pensando™ networking and ROCm™ software, all optimized for Meta’s AI workloads.
The companies are designing for gigawatt-scale deployments, jointly engineering complete AI systems that integrate compute, memory, networking, power, cooling, and software to accelerate large-scale AI training and inference.
Meta is expanding deployments across multiple AMD generations, serving as a lead customer for 6th Gen AMD EPYC™ “Venice” while advancing from AMD Instinct™ MI300X to MI350X to a custom MI450-based GPU.
AI is transforming not only what data centers do, but how they are designed. At Meta, demand for AI infrastructure is growing rapidly across training, inference, recommendation and ranking, content, and new AI experiences serving billions of people. Meeting it requires more than adding GPUs or servers; it requires scaling AI capacity across the data center at a pace the industry has never seen.
In a conversation with AMD Chair and CEO Dr. Lisa Su at Advancing AI 2026, Meta Head of Infrastructure Santosh Janardhan outlined how Meta is approaching this next phase of growth, and why deep collaboration with partners like AMD is increasingly critical.
Designing AI Infrastructure for Gigawatt Scale
A traditional data center architecture has typically optimized individual servers or components. AI has changed this dynamic. The performance of an AI platform now depends on how effectively compute, networking, memory, power, cooling and software operate together. And with Meta’s scale, these elements can’t be designed independently.
The entire infrastructure must be engineered as one integrated AI system, running efficiently and reliably at gigawatt scale.
That shift is opening the door to far deeper collaboration between infrastructure operators and technology partners. Rather than selecting finished components at the end of a development cycle, Meta and AMD are aligning requirements and engineering decisions from silicon through software. Flexibility matters as much as scale: Training, inference, recommendation systems and emerging agentic applications each place different demands on infrastructure.
Scaling AI Together with AMD EPYC
An open, heterogeneous architecture allows Meta to match the right hardware to each workload while maintaining high utilization across the fleet. This is the foundation of Meta’s growing work with AMD: co-engineering the silicon, systems and software needed to support a diverse and rapidly evolving portfolio of AI workloads.
The collaboration spans multiple generations of AMD EPYC™ processors – Milan, Bergamo, Turin and now Venice. Meta has deployed millions of EPYC CPUs across its infrastructure and is working closely with AMD as a lead customer for “Venice,” which is expected to deliver more cores, greater performance, improved efficiency and customization in areas especially important to Meta’s workloads. Meta is now validating “Venice”-based platforms in its labs, where early results have been promising, and the teams are preparing for deployment at scale.
The role of the CPU is also expanding as AI systems become more agentic. GPUs remain central to training and inference, but CPUs are increasingly responsible for orchestrating agents, executing tools, running code, managing data and coordinating the broader systems around AI models. As these workloads grow, the balance between CPU and GPU performance becomes even more consequential. Meta plans to build on its “Venice” deployments while continuing to collaborate with AMD on “Verano” and future generations of the EPYC roadmap.
Co-Engineering AI Infrastructure with Instinct GPUs
The partnership has also grown significantly across GPUs and rack-scale systems. Earlier this year, Meta introduced the Open Rack Wide (ORW) specification through the Open Compute Project, defining an open rack architecture purpose-built for AI-scale infrastructure.
AMD Helios rackscale solutions are built on the ORW standard, extending AMD’s commitment to openness from silicon and software to systems, racks and large-scale clusters. Helios integrates AMD Instinct™ GPUs, EPYC processors, AMD Pensando™ networking and ROCm™ software into a fully engineered rack-scale platform designed for the performance, power efficiency and serviceability required at gigawatt scale.
That shared architectural vision has allowed the two companies to collaborate more deeply with each successive generation of Instinct accelerators. The work began with MI300X, which gave both companies experience running AMD Instinct at hyperscale while optimizing production workloads on AMD ROCm software. The MI350 series extended the partnership into AI training, including the recommendation and ranking models central to Meta’s platforms.
Now Meta plans to deploy a custom Instinct MI450-based GPU, illustrating how the collaboration has progressed from product deployment to deep, end-to-end platform co-design.
Meta has early access to Helios and is already testing the platform with engineers from both companies working side by side across hardware, software, systems integration building toward workload.
This work will support Meta’s plans to deploy up to 6 gigawatts of infrastructure based on AMD GPUs. At that scale, early collaboration is essential as decisions about power, cooling, networking, memory and software can't wait until final qualification; they must be considered together from the start.
Building the Next Era of AI Infrastructure Together
The next generation of AI infrastructure will be built by treating the data center as a complete system and aligning workloads, silicon, networking, systems and software much earlier in the design process.
That collaboration is already informing future AMD platforms, including AMD Instinct™ MI500 accelerators and next-generation rack-scale systems. By combining Meta’s experience operating infrastructure for billions of people with AMD’s CPU, GPU and systems roadmaps, the companies aim to bring AI capacity online faster, more efficiently and with the flexibility future workloads will require.
By combining Meta's experience operating infrastructure for billions of people with AMD's CPU, GPU, and systems roadmaps, the two companies aim to bring more AI capacity online faster, more efficiently, and with the flexibility that the future workloads will demand.
Press inquiries: corporate.pressinquiry@amd.com