Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.
334
Tracked models
29
Providers
286
Benchmarked
28.3
Avg. index
334 models
| Rank | Model | Provider | Score | Benchmarks | Inference | Agentic | Programming | Value | Price |
|---|---|---|---|---|---|---|---|---|---|
| 141 | Claude Haiku 4.5 claude-haiku-4-5-20251001 multimodalvisionmulti-input reasoning | Anthropic | 30.8 Benchmarks | 30.8 | 53.3 | 50.8 | 53.9 | 45.6 | $1 in / $5 out |
| 142 | DeepSeek-V3 0324 deepseek-v3-0324 textinference | DeepSeek | 30.4 Benchmarks | 30.4 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 143 | Qwen3.5-4B qwen3.5-4b multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 29.9 Benchmarks | 29.9 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 144 | MiniMax M2 minimax-m2 codeprogrammingtool use | MiniMax | 29.8 Benchmarks | 29.8 | 62.8 | 40.2 | 40.0 | 73.2 | $0.3 in / $1.2 out |
| 145 | Ministral 3 (8B Reasoning 2512) ministral-8b-latest multimodalvisionmulti-input reasoning | Mistral AI | 29.4 Benchmarks | 29.4 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 146 | Phi 4 Reasoning Plus phi-4-reasoning-plus textinference | Microsoft | 29.4 Benchmarks | 29.4 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 147 | Qwen3 235B A22B qwen3-235b-a22b multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 29.3 Benchmarks | 29.3 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 148 | GPT-4o gpt-4o-2024-08-06 multimodalvisionmulti-input reasoning | OpenAI | 29.1 Benchmarks | 29.1 | 39.6 | 14.9 | 3.7 | 31.2 | |
| 149 | Qwen3 Max qwen3-max codeprogrammingtool use | Alibaba Cloud / Qwen Team | 28.0 Benchmarks | 28.0 | 0.0 | 0.0 | 32.1 | 0.0 | N/A |
| 150 | Hermes 3 70B hermes-3-70b textinference | Nous Research | 27.6 Benchmarks | 27.6 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 151 | Llama 4 Scout llama-4-scout multimodalvisionmulti-input reasoning | Meta | 27.6 Benchmarks | 27.6 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 152 | Qwen3-Next-80B-A3B-Instruct qwen3-next-80b-a3b-instruct textinference | Alibaba Cloud / Qwen Team | 27.4 Benchmarks | 27.4 | 0.0 | 17.9 | 0.0 | 0.0 | N/A |
| 153 | Pixtral Large pixtral-large multimodalvisionmulti-input reasoning | Mistral AI | 27.3 Benchmarks | 27.3 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 154 | GPT-4.1 gpt-4.1-2025-04-14 multimodalvisionmulti-input reasoning | OpenAI | 27.2 Benchmarks | 27.2 | 73.2 | 32.8 | 14.7 | 40.7 | |
| 155 | Qwen3 VL 32B Instruct qwen3-vl-32b-instruct multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 26.5 Benchmarks | 26.5 | 0.0 | 25.1 | 0.0 | 0.0 | |
| 156 | DeepSeek R1 Distill Llama 70B deepseek-r1-distill-llama-70b textinference | DeepSeek | 26.4 Benchmarks | 26.4 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 157 | QwQ-32B qwq-32b textinference | Alibaba Cloud / Qwen Team | 26.4 Benchmarks | 26.4 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 158 | QwQ-32B-Preview qwq-32b-preview textinference | Alibaba Cloud / Qwen Team | 26.4 Benchmarks | 26.4 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 159 | Gemini 1.5 Pro gemini-1.5-pro multimodalvisionmulti-input reasoning | Google | 26.2 Benchmarks | 26.2 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 160 | DeepSeek-V3 deepseek-v3 codeprogrammingtool use | DeepSeek | 26.1 Benchmarks | 26.1 | 0.0 | 0.0 | 8.8 | 0.0 | N/A |
Claude Haiku 4.5
Anthropic
30.8
$1 in / $5 out
DeepSeek-V3 0324
DeepSeek
30.4
N/A
Qwen3.5-4B
Alibaba Cloud / Qwen Team
29.9
N/A
Want benchmark charts, model comparison, and pricing analytics?
Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.
Open full leaderboardRankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.
| N/A |
| $2.5 in / $10 out |
| N/A |
| $2 in / $8 out |
| N/A |
| N/A |
MiniMax M2
MiniMax
29.8
$0.3 in / $1.2 out
Ministral 3 (8B Reasoning 2512)
Mistral AI
29.4
N/A
Phi 4 Reasoning Plus
Microsoft
29.4
N/A
Qwen3 235B A22B
Alibaba Cloud / Qwen Team
29.3
N/A
GPT-4o
OpenAI
29.1
$2.5 in / $10 out
Qwen3 Max
Alibaba Cloud / Qwen Team
28.0
N/A
Hermes 3 70B
Nous Research
27.6
N/A
Llama 4 Scout
Meta
27.6
N/A
Qwen3-Next-80B-A3B-Instruct
Alibaba Cloud / Qwen Team
27.4
N/A
Pixtral Large
Mistral AI
27.3
N/A
GPT-4.1
OpenAI
27.2
$2 in / $8 out
Qwen3 VL 32B Instruct
Alibaba Cloud / Qwen Team
26.5
N/A
DeepSeek R1 Distill Llama 70B
DeepSeek
26.4
N/A
QwQ-32B
Alibaba Cloud / Qwen Team
26.4
N/A
QwQ-32B-Preview
Alibaba Cloud / Qwen Team
26.4
N/A
Gemini 1.5 Pro
26.2
N/A
DeepSeek-V3
DeepSeek
26.1
N/A