Newsletter Subscribe
Enter your email address below and subscribe to our newsletter
[forminator_form id="25163"]

news.futunn+1officechai+1memeburn+1Alibaba's Qwen3.8-Max has claimed the top position on Artificial Analysis's Agentic Index with a score of 55.4, narrowly edging out Anthropic's Claude Opus 5 at 55.3 points and OpenAI's GPT 5.6, according to the benchmarking organization's latest rankings published this week. The result marks the first time a Chinese AI model has held the top spot on the agentic capability leaderboard, which measures a model's ability to solve complex, multi-step problems autonomously.news.futunn+1
Alibaba's Hong Kong-listed shares jumped about 7% following the model's strong showing on multiple leaderboards.memeburn
The 2.4-trillion-parameter model, unveiled on August 3, uses a mixture-of-experts architecture that activates roughly 95 billion parameters per forward pass. It also scored 56 on the broader Artificial Analysis Intelligence Index, placing it level with Claude Opus 4.8 and ahead of models from Google Alphabet Inc. , Meta , and xAI.officechai+2
Much of Qwen3.8-Max's agentic gains come from a shift in how it approaches tasks. The model averages 64 turns on the GDPval-AA evaluation, compared with 14 for its predecessor Qwen3.7 Max, with input token usage rising roughly 15 times. That extra computation drives higher scores but also raises per-task costs to $1.14, more than double Qwen3.7 Max's $0.53.officechai
Alibaba has said it will release the model's weights next week, which would make Qwen3.8-Max the largest open-weight model available at roughly six times the size of Alibaba's previous biggest open release. Per-token API pricing is set at $2.00 per million input tokens and $6.00 per million output tokens, cheaper than its predecessor.venturebeat+1
The release intensifies a pattern in which Chinese labs are closing the gap with US frontier developers. Moonshot AI's Kimi K3, at 2.8 trillion parameters, still leads Chinese models on the Intelligence Index with a score of 57, but Qwen3.8-Max now sits just one point behind. VentureBeat noted the model also outperforms GPT-5.6 Sol Max and Fable 5 on OSWorld-Verified, a desktop computer-use benchmark.officechai+1
The gains are not uniform. Qwen3.8-Max's hallucination rate climbed from 23% to 40% on the AA-Omniscience evaluation, meaning it now attempts questions it previously declined and gets many wrong. Anthropic and OpenAI retain a lead at the very top of the overall Intelligence Index, and per-task costs for agentic work remain higher than several competitors in the same performance bracket.officechai