Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.
334
Tracked models
29
Providers
286
Benchmarked
29.3
Avg. index
334 models
| Rank | Model | Provider | Score | Benchmarks | Inference | Agentic | Programming | Value | Price |
|---|---|---|---|---|---|---|---|---|---|
| 161 | Qwen3 235B A22B qwen3-235b-a22b multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 29.3 overall | 29.3 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 162 | GPT-5.2 Codex gpt-5.2-codex multimodalvisionmulti-input reasoning | OpenAI | 29.1 overall | 0.0 | 31.1 | 0.0 | 30.9 | 22.0 | $1.75 in / $14 out |
| 163 | Nemotron 3 Nano (30B A3B) nemotron-3-nano-30b-a3b codeprogrammingtool use | NVIDIA | 29.0 overall | 43.6 | 32.9 | 3.0 | 3.8 | 100.0 | $0.06 in / $0.24 out |
| 164 | GPT-4.5 gpt-4.5 multimodalvisionmulti-input reasoning | OpenAI | 28.7 overall | 41.1 | 0.0 | 35.8 | 5.2 | 0.0 | N/A |
| 165 | GPT-4.1 mini gpt-4.1-mini-2025-04-14 multimodalvisionmulti-input reasoning | OpenAI | 28.5 overall | 19.2 | 84.6 | 8.9 | 2.2 | 69.5 | |
| 166 | Mistral Large 3 (675B Instruct 2512) mistral-large-latest multimodalvisionmulti-input reasoning | Mistral AI | 28.1 overall | 21.9 | 22.0 | 0.0 | 0.0 | 55.6 | |
| 167 | Claude 3.5 Sonnet claude-3-5-sonnet-20241022 multimodalvisionmulti-input reasoning | Anthropic | 28.0 overall | 32.0 | 0.0 | 38.7 | 11.1 | 0.0 | |
| 168 | Hermes 3 70B hermes-3-70b textinference | Nous Research | 27.6 overall | 27.6 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 169 | Llama 4 Scout llama-4-scout multimodalvisionmulti-input reasoning | Meta | 27.6 overall | 27.6 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 170 | GPT-4o gpt-4o-2024-05-13 multimodalvisionmulti-input reasoning | OpenAI | 27.6 overall | 20.5 | 38.0 | 0.0 | 0.0 | 30.7 | $2.5 in / $10 out |
| 171 | Qwen3 VL 8B Thinking qwen3-vl-8b-thinking multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 27.4 overall | 32.8 | 0.0 | 21.1 | 0.0 | 0.0 | N/A |
| 172 | MAI-Code-1-Flash mai-code-1-flash codeprogrammingtool use | Microsoft | 27.4 overall | 32.2 | 0.0 | 0.0 | 21.4 | 0.0 | N/A |
| 173 | Pixtral Large pixtral-large multimodalvisionmulti-input reasoning | Mistral AI | 27.3 overall | 27.3 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 174 | Command A+ command-a-plus-05-2026 multimodalvisionmulti-input reasoning | Cohere | 26.4 overall | 35.2 | 0.0 | 0.0 | 15.3 | 0.0 | |
| 175 | Qwen3 VL 30B A3B Thinking qwen3-vl-30b-a3b-thinking multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 26.4 overall | 32.7 | 0.0 | 19.2 | 0.0 | 0.0 | |
| 176 | DeepSeek R1 Distill Llama 70B deepseek-r1-distill-llama-70b textinference | DeepSeek | 26.4 overall | 26.4 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 177 | QwQ-32B qwq-32b textinference | Alibaba Cloud / Qwen Team | 26.4 overall | 26.4 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 178 | QwQ-32B-Preview qwq-32b-preview textinference | Alibaba Cloud / Qwen Team | 26.4 overall | 26.4 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 179 | Nemotron 3 Super (120B A12B) nemotron-3-super-120b-a12b codeprogrammingtool use | NVIDIA | 26.2 overall | 45.5 | 0.0 | 7.6 | 22.0 | 0.0 | N/A |
| 180 | Gemini 1.5 Pro gemini-1.5-pro multimodalvisionmulti-input reasoning | Google | 26.2 overall | 26.2 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
Qwen3 235B A22B
Alibaba Cloud / Qwen Team
29.3
N/A
GPT-5.2 Codex
OpenAI
29.1
$1.75 in / $14 out
Nemotron 3 Nano (30B A3B)
NVIDIA
29.0
$0.06 in / $0.24 out
Want benchmark charts, model comparison, and pricing analytics?
Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.
Open full leaderboardRankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.
| $0.4 in / $1.6 out |
| $0.5 in / $1.5 out |
| N/A |
| N/A |
| N/A |
| N/A |
GPT-4.5
OpenAI
28.7
N/A
GPT-4.1 mini
OpenAI
28.5
$0.4 in / $1.6 out
Mistral Large 3 (675B Instruct 2512)
Mistral AI
28.1
$0.5 in / $1.5 out
Claude 3.5 Sonnet
Anthropic
28.0
N/A
Hermes 3 70B
Nous Research
27.6
N/A
Llama 4 Scout
Meta
27.6
N/A
GPT-4o
OpenAI
27.6
$2.5 in / $10 out
Qwen3 VL 8B Thinking
Alibaba Cloud / Qwen Team
27.4
N/A
MAI-Code-1-Flash
Microsoft
27.4
N/A
Pixtral Large
Mistral AI
27.3
N/A
Command A+
Cohere
26.4
N/A
Qwen3 VL 30B A3B Thinking
Alibaba Cloud / Qwen Team
26.4
N/A
DeepSeek R1 Distill Llama 70B
DeepSeek
26.4
N/A
QwQ-32B
Alibaba Cloud / Qwen Team
26.4
N/A
QwQ-32B-Preview
Alibaba Cloud / Qwen Team
26.4
N/A
Nemotron 3 Super (120B A12B)
NVIDIA
26.2
N/A
Gemini 1.5 Pro
26.2
N/A