Gradient Ascent Logo
Update

Claude Fable Beats SimpleBench with Gemini Close Behind

Originally posted on LinkedIn. View original post →

Claude Fable 5.1 becomes the first model to beat human baseline score on SimpleBench with Gemini 3.8 Flash very close behind.

"SimpleBench includes over 200 questions covering spatio-temporal reasoning, social intelligence, and what we call linguistic adversarial robustness (or trick questions)." (See details at https://simple-bench.com/)

These questions have been historically very tricky for LLMs to solve and hence this benchmark has stood for longer time than most of the other benchmarks. It is humbling to see SimpleBench finally got beaten.

SimpleBench Leaderboard showing Claude Fable 5.1 and Gemini 3.8 Flash above human baseline

SimpleBench leaderboard rankings showing Claude Fable 5.1 (86.6%) surpassing the human baseline (83.7%), with Gemini 3.8 Flash close behind (82.4%). Source: SimpleBench.