Gradient Ascent Logo
Update

Google Gemini 3.8 Leads DeepSWE Benchmark at Low Cost

Originally posted on LinkedIn. View original post →

This is very interesting. New release from Google DeepMind, Gemini 3.8 Flash now leads the DeepSWE benchmark (widely considered the most reliable benchmark for measuring a model's coding capability) at a very low cost (~5x lower than Opus 5, ~10x lower than Fable 5, and ~3x lower than GPT 5.6 Sol). Has Google really cooked with this new release, or have they benchmaxxed?

Even if the model behaviour aligns with the benchmark score, Google still has a long way to go in terms of token efficiency or the number of steps to complete a task.

DeepSWE benchmark results showing Gemini 3.8 Flash pass rate and cost

DeepSWE benchmark results showing Gemini 3.8 Flash [high] achieving a 74% pass@1 rate at $2.36 average cost, significantly undercutting frontier alternatives.