GPT-6 Astra at max effort scores 95/100 on EyeBench-V3, a visual-perception benchmark on which no other model has cleared 60. The result comes from adi (@adonis_singh), the independent evaluator who built EyeBench and has run every version of it. Second place on his leaderboard is OpenAI's previous flagship, GPT-5.6 Sol at max effort, on 58 — and Astra got there while spending roughly half as much money and emitting about a quarter as many output tokens.

The full EyeBench-V3 leaderboard. Credit: adi (@adonis_singh).
| # | Model | Score | Est. cost | Est. $/correct | Output tokens |
|---|---|---|---|---|---|
| 1 | gpt-6-astra (max) | 95/100 | $25.71 | $0.27 | 444k |
| 2 | gpt-5.6-sol (max) | 58/100 | $54.79 | $0.94 | 1.77M |
| 3 | gpt-5.6-terra (max) | 42/100 | $44.80 | $1.07 | 2.93M |
| 4 | gpt-5.5 (xhigh) | 41/100 | $44.65 | $1.09 | 1.43M |
| 5 | gpt-5.5-pro (medium) | 41/100 | $244 | $5.94 | 1.30M |
| 6 | claude-fable-5.1 (max) | 40/100 | $166 | $4.15 | 3.25M |
| 7 | qwen3.8-max (xhigh) | 39/100 | $4.88 | $0.13 | 725k |
| 8 | muse-spark-1.3 (xhigh) | 37/100 | $4.22 | $0.11 | 907k |
Further down: gemini-3.8-flash 31, gpt-5.6-luna 30, kimi-k3 29, gemini-3.1-pro 28, claude-fable-5 and gemini-3.7-flash 27, claude-opus-5 26, claude-sonnet-5 and glm-5.3-flash 21, grok-4.6 19, deepseek-v4-flash-vision 13, and Thinking Machines' inkling last at 9.
Two things about adi's framing are worth checking against his own numbers. "Half the cost" holds: $25.71 versus $54.79. "Nearly double" is generous — 95 against 58 is 1.6x, not 2x, though the 37-point gap is the largest any single model has opened on this benchmark. The token claim, "~3.8x less," reads as 444k against 1.77M in the output-token column, closer to 4x.
OpenAI priced Astra at $10/$50 per million input/output tokens, 2.5x GPT-5.6 Sol's $4/$20, per Artificial Analysis. So a run that costs half as much despite paying 2.5x per token is a pure token-efficiency result: Astra is answering these questions in far fewer tokens than Sol needs. Artificial Analysis saw the same pattern elsewhere, measuring roughly a 3x token reduction against Sol at max effort in the Codex harness. On EyeBench the compounding effect is stark in the cost-per-correct-answer column — $0.27 for Astra against $0.94 for Sol and $4.15 for claude-fable-5.1.
EyeBench exists because adi argues mainstream multimodal evals let language ability stand in for sight. V3 launched in March 2026 with harder and more varied questions, and at launch the best model in the world scored 35% (gpt-5.4-pro). Later, adi noted that claude-opus-4.7 scored 16% — Anthropic's best result on it at the time, and "still pretty blind in comparison." A 35 → 58 → 95 progression on a fixed set of visual questions in six months is a fast saturation curve for a benchmark built specifically to be unsaturated.
The caveats are the usual ones for a one-person benchmark: the question set and grading code are not public, these are single runs without error bars, and the models are not all tested at the same effort level (max, xhigh, high and medium all appear in the same ranking), so the cross-vendor comparisons are looser than the within-OpenAI ones. Nothing here has been independently reproduced.
adi's EyeBench-V3 resultEyeBench-V3 launch postadi on claude-opus-4.7Artificial Analysis: Benchmarking GPT-6 Astra
GPT-6 Astra takes 10 of 16 RuneBench records, at $15 a task
GPT-6 Astra writes the first error-free chorale on the Bach Benchmark

OpenAI ships GPT-6 Astra and declares the AGI era

MazeBench: the best coding agent collects 13 gems, four frontier models get zero

OpenAI cuts GPT-5.6 Sol API pricing by 20%

Claude Opus 5: near-Fable intelligence at half the price