Newsletter Subscribe
Enter your email address below and subscribe to our newsletter
[forminator_form id="25163"]

servethehomeservethehomeservethehomeNVIDIA dominated the technical program at Hot Chips 2026, presenting a pair of specialized processors that extend its Vera Rubin platform beyond GPUs: the Groq 3 Language Processing Unit and the BlueField-4 Data Processing Unit. The conference, held August 23–25 at Stanford University's Memorial Auditorium, also featured NVIDIA presentations on its 88-core Vera CPU and Rubin GPU.vdura+1
The Groq 3 LPU presentation detailed how NVIDIA is integrating technology from Groq, whose team largely joined NVIDIA following a $20 billion acquihire announced in late December 2025. A single LPX rack, containing 256 LPUs, can decode 11,000 tokens per second on a 31-billion-parameter model (Gemma 4), offering a 4x higher token output rate than the next-closest public competitor in third-party benchmarks conducted by Artificial Analysis.servethehome+2
The chip relies on large amounts of on-die SRAM — 128 GB across a rack — delivering 40 PB/s of aggregate SRAM bandwidth and 315 PFLOPS of FP8 compute. Its fully deterministic architecture eliminates branching and hardware scheduling, pushing all instruction scheduling to software. Power consumption is equally predictable, enabling look-ahead techniques that cut voltage droop by 60 percent.servethehome
When paired with Rubin GPUs, the LPX rack extends the performance frontier by up to 5x at the highest per-user token rates, handling the decode portion of inference while GPUs focus on prefill.servethehome
On the networking side, NVIDIA presented BlueField-4 as the first "AI-native DPU," combining 64 Arm Neoverse V2 cores at 1.7 GHz with ConnectX-9 networking delivering 800G Ethernet. In a Vera Rubin system, the DPU manages 7 Tb/s of aggregate bandwidth per compute tray while enforcing zero-trust security at line rate.servethehome
Storage performance numbers stood out: NVMe over fabrics running on BlueField-4 reaches 1.6 Tb/s with eight cores and 20 million IOPS with 16 cores, providing what NVIDIA called a 2x faster path to data for GPUs.servethehome
Separately, Groq — which continues to operate its GroqCloud inference service independently — hired Ola Kasali, Meta's former network deployment manager, as head of network and data center engineering. The company raised $350 million at a $3.5 billion valuation in August, with plans to scale from 54 megawatts of compute to over 200 MW by 2027.sdxcentral