Update
Google Gemini 3.8 Leads DeepSWE Benchmark at Low Cost
Originally posted on LinkedIn. View original post →
This is very interesting. New release from Google DeepMind, Gemini 3.8 Flash now leads the DeepSWE benchmark (widely considered the most reliable benchmark for measuring a model's coding capability) at a very low cost (~5x lower than Opus 5, ~10x lower than Fable 5, and ~3x lower than GPT 5.6 Sol). Has Google really cooked with this new release, or have they benchmaxxed?
Even if the model behaviour aligns with the benchmark score, Google still has a long way to go in terms of token efficiency or the number of steps to complete a task.

DeepSWE benchmark results showing Gemini 3.8 Flash [high] achieving a 74% pass@1 rate at $2.36 average cost, significantly undercutting frontier alternatives.