Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.
334
Tracked models
29
Providers
286
Benchmarked
28.3
Avg. index
334 models
| Rank | Model | Provider | Score | Benchmarks | Inference | Agentic | Programming | Value | Price |
|---|---|---|---|---|---|---|---|---|---|
| 221 | Qwen2.5 14B Instruct qwen-2.5-14b-instruct textinference | Alibaba Cloud / Qwen Team | 13.4 Benchmarks | 13.4 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 222 | Qwen3.5-2B qwen3.5-2b multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 13.2 Benchmarks | 13.2 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 223 | Mistral Small 3 24B Instruct mistral-small-24b-instruct-2501 textinference | Mistral AI | 13.0 Benchmarks | 13.0 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 224 | Mistral Small 3.1 24B Base mistral-small-3.1-24b-base-2503 multimodalvisionmulti-input reasoning | Mistral AI | 12.9 Benchmarks | 12.9 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 225 | Nova Lite nova-lite multimodalvisionmulti-input reasoning | Amazon | 12.9 Benchmarks | 12.9 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 226 | GPT-4.1 nano gpt-4.1-nano-2025-04-14 multimodalvisionmulti-input reasoning | OpenAI | 11.6 Benchmarks | 11.6 | 87.8 | 0.0 | 0.0 | 94.9 | |
| 227 | Qwen2 72B Instruct qwen2-72b-instruct textinference | Alibaba Cloud / Qwen Team | 11.0 Benchmarks | 11.0 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 228 | Llama 3.1 70B Instruct llama-3.1-70b-instruct textinference | Meta | 10.3 Benchmarks | 10.3 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 229 | Claude 3.5 Haiku claude-3-5-haiku-20241022 codeprogrammingtool use | Anthropic | 9.9 Benchmarks | 9.9 | 0.0 | 3.0 | 6.6 | 0.0 | |
| 230 | Gemini 1.5 Flash 8B gemini-1.5-flash-8b multimodalvisionmulti-input reasoning | Google | 9.9 Benchmarks | 9.9 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 231 | Grok-1.5V grok-1.5v multimodalvisionmulti-input reasoning | xAI | 9.7 Benchmarks | 9.7 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 232 | Claude 3 Sonnet claude-3-sonnet-20240229 multimodalvisionmulti-input reasoning | Anthropic | 9.2 Benchmarks | 9.2 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 233 | Qwen2.5 VL 7B Instruct qwen2.5-vl-7b multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 9.0 Benchmarks | 9.0 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 234 | Mistral Large 3 mistral-large-3-2509 multimodalvisionmulti-input reasoning | Mistral AI | 8.8 Benchmarks | 8.8 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 235 | Gemma 3 12B gemma-3-12b-it multimodalvisionmulti-input reasoning | Google | 8.6 Benchmarks | 8.6 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 236 | Gemma 4 E2B gemma-4-e2b-it multimodalvisionmulti-input reasoning | Google | 8.5 Benchmarks | 8.5 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 237 | Nova Micro nova-micro textinference | Amazon | 8.4 Benchmarks | 8.4 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 238 | Grok-1.5 grok-1.5 multimodalvisionmulti-input reasoning | xAI | 8.2 Benchmarks | 8.2 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 239 | Phi-4-multimodal-instruct phi-4-multimodal-instruct multimodalvisionmulti-input reasoning | Microsoft | 8.0 Benchmarks | 8.0 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 240 | Pixtral-12B pixtral-12b-2409 multimodalvisionmulti-input reasoning | Mistral AI | 8.0 Benchmarks | 8.0 | 0.0 | 0.0 | 0.0 | 0.0 |
Qwen2.5 14B Instruct
Alibaba Cloud / Qwen Team
13.4
N/A
Qwen3.5-2B
Alibaba Cloud / Qwen Team
13.2
N/A
Mistral Small 3 24B Instruct
Mistral AI
13.0
N/A
Want benchmark charts, model comparison, and pricing analytics?
Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.
Open full leaderboardRankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.
| N/A |
| $0.1 in / $0.4 out |
| N/A |
| N/A |
| N/A |
| N/A |
| N/A |
Mistral Small 3.1 24B Base
Mistral AI
12.9
N/A
Nova Lite
Amazon
12.9
N/A
GPT-4.1 nano
OpenAI
11.6
$0.1 in / $0.4 out
Qwen2 72B Instruct
Alibaba Cloud / Qwen Team
11.0
N/A
Llama 3.1 70B Instruct
Meta
10.3
N/A
Claude 3.5 Haiku
Anthropic
9.9
N/A
Gemini 1.5 Flash 8B
9.9
N/A
Grok-1.5V
xAI
9.7
N/A
Claude 3 Sonnet
Anthropic
9.2
N/A
Qwen2.5 VL 7B Instruct
Alibaba Cloud / Qwen Team
9.0
N/A
Mistral Large 3
Mistral AI
8.8
N/A
Gemma 3 12B
8.6
N/A
Gemma 4 E2B
8.5
N/A
Nova Micro
Amazon
8.4
N/A
Grok-1.5
xAI
8.2
N/A
Phi-4-multimodal-instruct
Microsoft
8.0
N/A
Pixtral-12B
Mistral AI
8.0
N/A