Newsletter Subscribe
Enter your email address below and subscribe to our newsletter
[forminator_form id="25163"]

cerebras.ai+1cerebras.ai+1Business Insider+1AMD Advanced Micro Devices, Inc. and Cerebras Systems announced a technical partnership on Tuesday to build a disaggregated AI inference platform that combines AMD's Helios rack-scale infrastructure with Cerebras' Wafer-Scale Engine technology. The companies said the combined system is expected to deliver up to five times more tokens per second per watt compared with a Cerebras WSE-only configuration, based on internal modeling conducted in July 2026.cerebras.ai+1
The joint offering assigns different portions of an inference workload to the hardware best suited for each stage. AMD Helios will handle high-throughput prefill processing — the computationally intensive step of reading a user's prompt — while Cerebras' WSE technology handles token generation, where its memory bandwidth advantages are most pronounced.Investing.com+1
The platform is scheduled to become available through Cerebras Cloud in the second half of 2026.cerebras.ai+1
The announcement was made during AMD CEO Lisa Su's keynote at the company's Advancing AI 2026 event in San Francisco on July 23. Business Insider reported that the partnership reflects AMD's broader bet that disaggregated inference — splitting AI workloads across specialized hardware rather than running everything on a single chip — represents the next phase of AI infrastructure.Business Insider+2
The approach has gained momentum across the industry. AWS and Cerebras announced a similar collaboration in March 2026, pairing Amazon's Trainium chips for prefill with Cerebras hardware for decode on Amazon Bedrock. Nvidia acquired Groq to pursue a comparable architecture. UBS wrote in June that limitations of current systems "are driving a shift toward disaggregated inference," according to Business Insider.youtube+2
The Cerebras deal came amid a flurry of announcements at Advancing AI 2026. AMD unveiled its full Helios rack system, which bundles 72 Instinct MI455X accelerators with 31TB of HBM4 memory and delivers approximately three AI exaflops per rack. Su claimed Helios delivers up to 30% more inference tokens per dollar than Nvidia's Vera Rubin NVL72 rack.DCDNoticias+2
AMD also announced a multibillion-dollar infrastructure partnership with Anthropic on the day prior, adding to existing relationships with OpenAI, Meta , Microsoft , and Oracle .Business Insider
For Cerebras, the AMD deal extends a string of major partnerships following its IPO earlier this year, which raised $6.4 billion in what the company called the largest semiconductor IPO of all time. Cerebras already powers OpenAI's GPT-5.4 model under a multibillion-dollar agreement signed in January.The New York Times+2