Newsletter Subscribe
Enter your email address below and subscribe to our newsletter
[forminator_form id="25163"]

trendingtopics+1trendingtopicstrendingtopics+1OpenAI's GPT-6 Astra, released on September 3, 2026, has produced a rare split among independent evaluators: one ranks it as the top AI model in the world, while another places it level with its predecessor and behind rival Anthropic.
Epoch AI's composite Epoch Capabilities Index gives Astra a score of 169, placing it first among 267 models — ahead of Anthropic's Claude Fable 5.1 at 163. Artificial Analysis, a San Francisco-based benchmarking firm, reached a different conclusion. Its Intelligence Index rates Astra at 61 points, matching OpenAI's Microsoft Corporation own GPT-5.6 Sol exactly, with Anthropic's Claude Fable 5.1 five points ahead at 66. On coding tasks, the picture is more favorable: Astra scored 67 on the Artificial Analysis Coding Agent Index, roughly on par with Claude Opus 5 and Meta's Muse Spark 1.3.trendingtopics+4
The disagreement stems partly from what each index measures. Epoch AI aggregates scores across reasoning, mathematics, science, and agent benchmarks into a single number, while Artificial Analysis weights knowledge, long-document reasoning, and factual accuracy differently. Epoch AI noted on X that Astra's record was "within uncertainty range" of prior leaders, suggesting the gap may be narrower than the headline number implies.latent+1
The ARC Prize organization provided its own evaluation, testing Astra on ARC-AGI-3, an interactive benchmark that requires agents to explore unfamiliar environments, infer goals, and plan actions without explicit instructions. Under the standard harness, Astra scored 62.7% at a cost of roughly $26,000 per run. Under OpenAI's Provider Adapter harness, which preserves the model's reasoning state between requests, the score jumped to 99.9% for about $19,000.arcprize
A standout finding was Astra's action efficiency. Tested against a baseline drawn from approximately 500 human participants, Astra used fewer actions than the median human on 96% of levels and averaged 51.7% fewer actions per level overall. "This is a material milestone," ARC Prize wrote, adding that the result "matched and surpassed human parity" by their measure of action efficiency.arcprize
Epoch AI also reported that on its new FrontierMath Erdős benchmark — a set of 68 open problems posed by the mathematician Paul Erdős — Astra was the only model to solve any, producing Lean-verified proofs for two problems on a budget of $300 per attempt. Three additional solutions emerged from costlier non-standardized runs exceeding $220,000, which Epoch excluded from the official score.the-decoder+1
For developers, the performance comes at a price. OpenAI set Astra's API rates at $10 per million input tokens and $50 per million output tokens, a 2.5-fold increase over GPT-5.6 Sol. Artificial Analysis found that Astra's improved token efficiency offsets the higher pricing in coding tasks but leaves the model about 75% more expensive per task for general intelligence work.trendingtopics