Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.
334
Tracked models
29
Providers
286
Benchmarked
28.3
Avg. index
334 models
| Rank | Model | Provider | Score | Benchmarks | Inference | Agentic | Programming | Value | Price |
|---|---|---|---|---|---|---|---|---|---|
| 181 | LongCat-Flash-Lite longcat-flash-lite codeprogrammingtool use | Meituan | 22.8 Benchmarks | 22.8 | 72.8 | 30.1 | 23.9 | 95.6 | $0.1 in / $0.4 out |
| 182 | MiniMax M1 80K minimax-m1-80k codeprogrammingtool use | MiniMax | 22.8 Benchmarks | 22.8 | 0.0 | 20.9 | 16.2 | 0.0 | N/A |
| 183 | DeepSeek R1 Distill Qwen 14B deepseek-r1-distill-qwen-14b textinference | DeepSeek | 22.7 Benchmarks | 22.7 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 184 | Magistral Small 2506 magistral-small-2506 textinference | Mistral AI | 22.7 Benchmarks | 22.7 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 185 | Qwen2.5 VL 72B Instruct qwen2.5-vl-72b multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 22.7 Benchmarks | 22.7 | 0.0 | 5.2 | 0.0 | 0.0 | N/A |
| 186 | Mistral Large 3 (675B Base) mistral-large-3-675b-base-2512 multimodalvisionmulti-input reasoning | Mistral AI | 21.9 Benchmarks | 21.9 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 187 | Mistral Large 3 (675B Instruct 2512 Eagle) mistral-large-3-675B-instruct-2512-eagle multimodalvisionmulti-input reasoning | Mistral AI | 21.9 Benchmarks | 21.9 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 188 | Mistral Large 3 (675B Instruct 2512 NVFP4) mistral-large-3-675b-instruct-2512-nvfp4 multimodalvisionmulti-input reasoning | Mistral AI | 21.9 Benchmarks | 21.9 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 189 | Mistral Large 3 (675B Instruct 2512) mistral-large-latest multimodalvisionmulti-input reasoning | Mistral AI | 21.9 Benchmarks | 21.9 | 22.0 | 0.0 | 0.0 | 55.6 | |
| 190 | Gemini 1.5 Flash gemini-1.5-flash multimodalvisionmulti-input reasoning | Google | 21.8 Benchmarks | 21.8 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 191 | Llama-3.3 Nemotron Super 49B v1 llama-3.3-nemotron-super-49b-v1 textinference | NVIDIA | 21.3 Benchmarks | 21.3 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 192 | MiniMax M1 40K minimax-m1-40k codeprogrammingtool use | MiniMax | 21.3 Benchmarks | 21.3 | 0.0 | 26.8 | 15.5 | 0.0 | N/A |
| 193 | Phi 4 Reasoning phi-4-reasoning textinference | Microsoft | 21.3 Benchmarks | 21.3 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 194 | Magistral Medium magistral-medium multimodalvisionmulti-input reasoning | Mistral AI | 20.6 Benchmarks | 20.6 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 195 | GPT-4o gpt-4o-2024-05-13 multimodalvisionmulti-input reasoning | OpenAI | 20.5 Benchmarks | 20.5 | 38.0 | 0.0 | 0.0 | 30.7 | |
| 196 | Min istral 3 (3B Reasoning 2512) ministral-3b-latest multimodalvisionmulti-input reasoning | Mistral AI | 20.4 Benchmarks | 20.4 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 197 | Gemini 2.5 Flash-Lite gemini-2.5-flash-lite multimodalvisionmulti-input reasoning | Google | 20.3 Benchmarks | 20.3 | 0.0 | 0.0 | 2.9 | 0.0 | |
| 198 | Qwen3 VL 4B Thinking qwen3-vl-4b-thinking multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 20.2 Benchmarks | 20.2 | 32.9 | 17.0 | 0.0 | 79.3 | |
| 199 | Qwen3 32B qwen3-32b textinference | Alibaba Cloud / Qwen Team | 20.1 Benchmarks | 20.1 | 2.2 | 0.0 | 0.0 | 78.0 | $0.1 in / $0.3 out |
| 200 | Phi 4 Mini Reasoning phi-4-mini-reasoning textinference | Microsoft | 19.9 Benchmarks | 19.9 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
LongCat-Flash-Lite
Meituan
22.8
$0.1 in / $0.4 out
MiniMax M1 80K
MiniMax
22.8
N/A
DeepSeek R1 Distill Qwen 14B
DeepSeek
22.7
N/A
Want benchmark charts, model comparison, and pricing analytics?
Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.
Open full leaderboardRankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.
| N/A |
| N/A |
| N/A |
| $0.5 in / $1.5 out |
| N/A |
| N/A |
| $2.5 in / $10 out |
| N/A |
| N/A |
| $0.1 in / $1 out |
Magistral Small 2506
Mistral AI
22.7
N/A
Qwen2.5 VL 72B Instruct
Alibaba Cloud / Qwen Team
22.7
N/A
Mistral Large 3 (675B Base)
Mistral AI
21.9
N/A
Mistral Large 3 (675B Instruct 2512 Eagle)
Mistral AI
21.9
N/A
Mistral Large 3 (675B Instruct 2512 NVFP4)
Mistral AI
21.9
N/A
Mistral Large 3 (675B Instruct 2512)
Mistral AI
21.9
$0.5 in / $1.5 out
Gemini 1.5 Flash
21.8
N/A
Llama-3.3 Nemotron Super 49B v1
NVIDIA
21.3
N/A
MiniMax M1 40K
MiniMax
21.3
N/A
Phi 4 Reasoning
Microsoft
21.3
N/A
Magistral Medium
Mistral AI
20.6
N/A
GPT-4o
OpenAI
20.5
$2.5 in / $10 out
Min istral 3 (3B Reasoning 2512)
Mistral AI
20.4
N/A
Gemini 2.5 Flash-Lite
20.3
N/A
Qwen3 VL 4B Thinking
Alibaba Cloud / Qwen Team
20.2
$0.1 in / $1 out
Qwen3 32B
Alibaba Cloud / Qwen Team
20.1
$0.1 in / $0.3 out
Phi 4 Mini Reasoning
Microsoft
19.9
N/A