Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.
373
Tracked models
34
Providers
310
Benchmarked
30.2
Avg. index
373 models
| Rank | Model | Provider | Score | Benchmarks | Inference | Agentic | Programming | Value | Price |
|---|---|---|---|---|---|---|---|---|---|
| 21 | Muse Spark 1.1 muse-spark-1.1 multimodalvisionmulti-input reasoning | Meta | 65.8 overall | 66.4 | 81.9 | 75.6 | 55.1 | 38.8 | $1.25 in / $4.25 out |
| 22 | Seed 2.1 Pro seed-2.1-pro multimodalvisionmulti-input reasoning | ByteDance | 65.4 overall | 67.2 | 0.0 | 68.4 | 59.9 | 0.0 | N/A |
| 23 | GLM-5.3 glm-5.3 textinference | Zhipu AI | 64.7 overall | 68.8 | 81.9 | 59.9 | 0.0 | 37.4 | $1.4 in / $4.4 out |
| 24 | Qwen3.8 Flash qwen3.8-flash multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 63.5 overall | 61.1 | 58.6 | 64.5 | 60.0 | 83.5 | $0.15 in / $0.47 out |
| 25 | Claude Fable 5 claude-fable-5 multimodalvisionmulti-input reasoning | Anthropic | 62.5 overall | 69.5 | 58.6 | 0.0 | 84.2 | 1.9 | |
| 26 | Qwen3.8-Flash-Next qwen3.8-flash-next multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 61.9 overall | 61.1 | 0.0 | 64.5 | 60.0 | 0.0 | N/A |
| 27 | MiMo-V2-Pro mimo-v2-pro codeprogrammingtool use | Xiaomi | 61.7 overall | 0.0 | 0.0 | 0.0 | 61.7 | 0.0 | N/A |
| 28 | GPT-5.1 Codex High gpt-5.1-codex-high multimodalvisionmulti-input reasoning | OpenAI | 61.4 overall | 61.4 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 29 | GPT-5 High gpt-5-high-2025-08-07 multimodalvisionmulti-input reasoning | OpenAI | 60.8 overall | 60.8 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 30 | GPT-5.5 gpt-5.5 multimodalvisionmulti-input reasoning | OpenAI | 60.6 overall | 74.2 | 95.2 | 54.9 | 49.2 | 5.8 | $5 in / $30 out |
| 31 | Claude Opus 4.8 claude-opus-4-8 multimodalvisionmulti-input reasoning | Anthropic | 60.0 overall | 72.5 | 26.7 | 68.7 | 81.8 | 9.5 | |
| 32 | Claude Opus 5 claude-opus-5 multimodalvisionmulti-input reasoning | Anthropic | 59.9 overall | 70.5 | 58.6 | 69.4 | 0.0 | 9.2 | |
| 33 | DeepSeek-V3.2 (Non-thinking) deepseek-chat textinference | DeepSeek | 59.7 overall | 0.0 | 48.7 | 0.0 | 0.0 | 77.3 | $0.28 in / $0.42 out |
| 34 | Qwen3.7 Max qwen3.7-max multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 59.3 overall | 64.0 | 58.6 | 48.0 | 74.4 | 40.8 | $1.25 in / $3.75 out |
| 35 | Seed 2.1 Turbo seed-2.1-turbo multimodalvisionmulti-input reasoning | ByteDance | 59.1 overall | 63.6 | 0.0 | 57.6 | 55.1 | 0.0 | N/A |
| 36 | Kimi K2-Thinking-0905 kimi-k2-thinking-0905 codeprogrammingtool use | Moonshot AI | 59.0 overall | 64.9 | 0.0 | 51.2 | 59.9 | 0.0 | |
| 37 | GPT-5.5 Pro gpt-5.5-pro multimodalvisionmulti-input reasoning | OpenAI | 58.9 overall | 59.9 | 0.0 | 67.1 | 48.8 | 0.0 | N/A |
| 38 | GLM-5.2 glm-5.2 codeprogrammingtool use | Zhipu AI | 58.7 overall | 66.8 | 81.9 | 39.1 | 57.9 | 47.6 | $0.95 in / $3 out |
| 39 | Gemini 3.8 Flash gemini-3.8-flash multimodalvisionmulti-input reasoning | Google | 58.5 overall | 50.6 | 81.9 | 0.0 | 0.0 | 43.2 | |
| 40 | Gemini 3 Pro gemini-3-pro-preview multimodalvisionmulti-input reasoning | Google | 58.3 overall | 68.8 | 0.0 | 51.0 | 52.9 | 0.0 |
Muse Spark 1.1
Meta
65.8
$1.25 in / $4.25 out
Seed 2.1 Pro
ByteDance
65.4
N/A
GLM-5.3
Zhipu AI
64.7
$1.4 in / $4.4 out
Want benchmark charts, model comparison, and pricing analytics?
Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.
Open full leaderboardRankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.
| $10 in / $50 out |
| N/A |
| N/A |
| $5 in / $25 out |
| $5 in / $25 out |
| N/A |
| $0.75 in / $3.75 out |
| N/A |
Qwen3.8 Flash
Alibaba Cloud / Qwen Team
63.5
$0.15 in / $0.47 out
Claude Fable 5
Anthropic
62.5
$10 in / $50 out
Qwen3.8-Flash-Next
Alibaba Cloud / Qwen Team
61.9
N/A
MiMo-V2-Pro
Xiaomi
61.7
N/A
GPT-5.1 Codex High
OpenAI
61.4
N/A
GPT-5 High
OpenAI
60.8
N/A
GPT-5.5
OpenAI
60.6
$5 in / $30 out
Claude Opus 4.8
Anthropic
60.0
$5 in / $25 out
Claude Opus 5
Anthropic
59.9
$5 in / $25 out
DeepSeek-V3.2 (Non-thinking)
DeepSeek
59.7
$0.28 in / $0.42 out
Qwen3.7 Max
Alibaba Cloud / Qwen Team
59.3
$1.25 in / $3.75 out
Seed 2.1 Turbo
ByteDance
59.1
N/A
Kimi K2-Thinking-0905
Moonshot AI
59.0
N/A
GPT-5.5 Pro
OpenAI
58.9
N/A
GLM-5.2
Zhipu AI
58.7
$0.95 in / $3 out
Gemini 3.8 Flash
58.5
$0.75 in / $3.75 out
Gemini 3 Pro
58.3
N/A