Newsletter Subscribe
Enter your email address below and subscribe to our newsletter
[forminator_form id="25163"]

startupfortune+1x+1finance.yahoo+1New benchmark data from SemiAnalysis's InferenceX platform shows AMD's Advanced Micro Devices, Inc. Instinct MI355X GPU running Moonshot AI's 2.8-trillion-parameter Kimi K3 model at a lower cost per token than Nvidia's B300, giving AMD fresh ammunition in the data center AI chip race just days before its Q2 earnings report.
The InferenceX platform, an open-source vendor-neutral benchmark that continuously measures AI inference performance across GPUs and software stacks, published day-zero benchmarks for Kimi K3 across the latest Nvidia and AMD GPUs. According to reporting on the results, the MI355X is running Kimi K3 at less than half the hourly cloud rental cost of the B300, a gap driven largely by the MI355X's substantially lower per-GPU-hour pricing in the market. Cloud providers currently list MI355X instances starting around $2.50 per GPU-hour compared to roughly $6.00 for the B300.news.ycombinator+4
Kimi K3, released by Beijing-based Moonshot AI on July 16, is a Mixture-of-Experts model with 2.8 trillion total parameters and a one-million-token context window. Its sheer size — requiring 288 GB of HBM per GPU — means it cannot fit on Nvidia's older B200 nodes at FP4 precision, limiting deployment to B300, GB300 NVL72, or MI355X systems. That hardware constraint effectively creates a head-to-head comparison between AMD's and Nvidia's latest 288 GB accelerators.amd+3
The Kimi K3 result extends a pattern visible across SemiAnalysis's earlier InferenceX comparisons. On the smaller Kimi K2.6 model, the B300 delivered 29 percent more raw throughput per GPU, but the MI355X was 23 percent cheaper per token. A similar dynamic held on MiniMax M3, where the MI355X was 24 percent cheaper per token despite lower absolute throughput. An independent evaluation by Signal65 in July found the MI355X achieved 1.4 to 2.1 times more tokens per dollar than the B200 across multiple models.signal65+2
The timing is notable. AMD reports Q2 2026 earnings on August 4, with consensus expectations of roughly $11.2 billion in revenue, representing 46 percent year-over-year growth. The data center segment, which hit $5.8 billion in Q1 on continued Instinct GPU shipments, is the focal point for investors. Wedbush recently raised its 2026 revenue estimate to $49.2 billion, citing faster-than-expected data center growth.seekingalpha+3
Whether inference cost benchmarks translate into market share gains remains the central question. Nvidia's B300 still leads on raw throughput, and its CUDA software ecosystem retains deep enterprise entrenchment. But as open-weight models like Kimi K3 push hardware requirements to their limits, AMD's pricing advantage offers hyperscalers and cloud providers a compelling alternative calculus.