Gradient Ascent Logo
Update

GPT-6 Astra outperforms GPT-5.6 Sol in various benchmarks

Originally posted on LinkedIn. View original post →

GPT-6 Astra scores 61 on the Artificial Analysis Intelligence Index. That is the same score as GPT-5.6 Sol. Many people are looking at the composite benchmark number and are confused about why Astra is getting talked about so much.

Other benchmarks tell a different story. On computer use, Astra reaches 72.6% on OSWorld 2.0 versus 65.7% for Sol, and finishes tasks in roughly half the time (about 40 minutes versus 75 minutes). On Agents’ Last Exam, it scores 59.3%, ahead of both Sol and Claude Opus 5, while using significantly fewer tokens. AutomationBench more than doubles, from 18.1% to 41.4%. BenchCAD rises from 83.3% to 95.9%. Terminal-Bench Science jumps from 22.4% to 64.6%. FrontierMath Tier 4 saturates at 97.6%, and ExploitBench hits 100%. On ARC-AGI-3, Astra scores 62.7% under the Standard harness (versus 7.8% for Sol) and 99.9% with the Provider Adapter harness.

The Intelligence Index is heavily weighted toward a mix of established evaluations. The larger gains appear in interactive computer use, multi-step professional workflows, scientific terminal tasks, hard mathematics, and cybersecurity; often with meaningful improvements in both score and efficiency (fewer tokens or less time per task). Those are the areas where Astra shows the clearest step forward relative to its predecessor.

Full results from OpenAI:
https://openai.com/index/gpt-6-astra/

GPT-6 Astra selected benchmark gains compared to GPT-5.6 Sol

Performance comparison between GPT-6 Astra and GPT-5.6 Sol across OSWorld 2.0, Agents' Last Exam, AutomationBench, Terminal-Bench Science, FrontierMath Tier 4, ExploitBench, and ARC-AGI-3.