Gradient Ascent Logo
Update

GPT-6 Astra Surpasses Human Baseline in Reasoning Task

Originally posted on LinkedIn. View original post →

GPT-6 Astra scores 62.7% on ARC-AGI-3 Semi-Private under the Standard harness. It also surpasses the human baseline in action efficiency on the vast majority of levels. All this with the Standard harness; the same minimal, provider-neutral setup that has been highly controversial in model evaluation. When using the model’s own scaffolding, the score jumps to a near-perfect 99.9%.

I remember when OpenAI released o3, its first major “reasoning” model (Hard to think that it has been less than 2 years). It crushed ARC-AGI-1. That was the first time I truly believed LLMs were here to stay and would take us to AGI. GPT-6 Astra feels like the continuation of that belief.

Read more: https://arcprize.org/blog/astra

ARC-AGI-3 Leaderboard score versus cost curve

ARC-AGI-3 Leaderboard: Score (%) versus Cost ($) verified by ARC Prize, showcasing GPT-6 Astra under Standard and Provider Adapter harnesses.