Newsletter Subscribe
Enter your email address below and subscribe to our newsletter
[forminator_form id="25163"]

techcrunch+1stable-learncryptobriefingThe mystery that consumed AI developers for nearly a week has been solved. Z.ai, the Chinese AI company formerly known as Zhipu AI, confirmed on Wednesday that the anonymous "Ox Alpha" model that appeared on OpenRouter and OpenCode is GLM-5.3-Flash, the first natively multimodal model in its GLM-5 series.africa.businessinsider+1
The model, which had been served free and without attribution since August 20, drew widespread attention from developers and tech leaders alike. Stripe CEO Patrick Collison called it "very impressive" on X. Before the reveal, Ox Alpha had accumulated over 503,000 unique users and processed 44 trillion tokens on OpenCode alone.techcrunch+1
GLM-5.3-Flash uses a mixture-of-experts architecture with 320 billion total parameters but activates only 18 billion per token, cutting compute costs dramatically compared to dense models of similar capability. The model supports a one-million-token context window and accepts text, image, and video inputs natively.cryptobriefing+1
On the DeepSWE v1.1 coding benchmark, GLM-5.3-Flash scored 63.4, up from 46.2 for its predecessor GLM-5.2. Z.ai claims it approaches Claude Opus 4.8 on its internal coding benchmark, scoring 29.0 versus Opus 4.8's 29.5. API pricing sits at $0.15 per million input tokens and $0.50 per million output tokens — roughly one-tenth the cost of comparable proprietary alternatives.officechai+1
The weights are available under the MIT license on Hugging Face and ModelScope, allowing unrestricted commercial use.testingcatalog+1
Perhaps the most consequential aspect of the release is the infrastructure behind it. Z.ai says GLM-5.3-Flash runs entirely on domestically manufactured AI accelerators, with a custom SGLang-based serving stack that delivered a threefold improvement in end-to-end performance. The company framed this as evidence that Chinese chips can support frontier-model inference at competitive costs, a claim that carries weight amid ongoing U.S. export controls restricting access to Nvidia's most advanced processors.testingcatalog+1
Z.ai had previously trained its GLM-5 base model on Huawei Ascend chips, but GLM-5.3-Flash extends that independence into the inference layer at scale.siliconrepublic
The anonymous launch followed a pattern Z.ai has used before — it previously tested an earlier model under the name "Pony Alpha" on OpenRouter. This time, the gambit generated far more attention, with developers dissecting tokenizer fingerprints and API error codes to identify the maker before the official announcement.kingy+2
As TechCrunch noted, the release "adds to the burgeoning threat of cheap, capable models from China that could take real market share away from expensive frontier model companies like OpenAI and Anthropic."techcrunch