Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.
334
Tracked models
29
Providers
286
Benchmarked
29.3
Avg. index
334 models
| Rank | Model | Provider | Score | Benchmarks | Inference | Agentic | Programming | Value | Price |
|---|---|---|---|---|---|---|---|---|---|
| 221 | Nova Pro nova-pro multimodalvisionmulti-input reasoning | Amazon | 19.0 overall | 19.0 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 222 | North Mini Code 1.0 north-mini-code-1.0 codeprogrammingtool use | Cohere | 18.6 overall | 0.0 | 0.0 | 0.0 | 18.6 | 0.0 | N/A |
| 223 | DeepSeek-V3 deepseek-v3 codeprogrammingtool use | DeepSeek | 18.5 overall | 26.1 | 0.0 | 0.0 | 8.8 | 0.0 | N/A |
| 224 | Mistral Small 3.2 24B Instruct mistral-small-3.2-24b-instruct-2506 multimodalvisionmulti-input reasoning | Mistral AI | 18.4 overall | 18.4 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 225 | Llama 3.1 405B Instruct llama-3.1-405b-instruct textinference | Meta | 18.3 overall | 18.3 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 226 | Llama 3.3 70B Instruct llama-3.3-70b-instruct textinference | Meta | 18.0 overall | 18.0 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 227 | Gemma 4 E4B gemma-4-e4b-it multimodalvisionmulti-input reasoning | Google | 17.8 overall | 17.8 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 228 | Claude 3 Opus claude-3-opus-20240229 multimodalvisionmulti-input reasoning | Anthropic | 17.7 overall | 17.7 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 229 | Kimi K2 Instruct kimi-k2-instruct codeprogrammingtool use | Moonshot AI | 17.3 overall | 23.2 | 0.0 | 13.5 | 14.0 | 0.0 | N/A |
| 230 | Qwen2.5 32B Instruct qwen-2.5-32b-instruct textinference | Alibaba Cloud / Qwen Team | 17.1 overall | 17.1 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 231 | o3-mini o3-mini codeprogrammingtool use | OpenAI | 16.9 overall | 25.3 | 0.0 | 11.9 | 11.6 | 0.0 | N/A |
| 232 | DeepSeek R1 Distill Qwen 7B deepseek-r1-distill-qwen-7b textinference | DeepSeek | 16.8 overall | 16.8 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 233 | DeepSeek R1 Distill Llama 8B deepseek-r1-distill-llama-8b textinference | DeepSeek | 16.3 overall | 16.3 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 234 | Qwen2.5 72B Instruct qwen-2.5-72b-instruct textinference | Alibaba Cloud / Qwen Team | 16.3 overall | 16.3 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 235 | Kimi K2-Instruct-0905 kimi-k2-instruct-0905 codeprogrammingtool use | Moonshot AI | 15.8 overall | 23.2 | 0.0 | 6.0 | 17.1 | 0.0 | |
| 236 | GPT OSS 20B gpt-oss-20b textinference | OpenAI | 15.5 overall | 23.7 | 0.0 | 6.0 | 0.0 | 0.0 | N/A |
| 237 | Qwen3 VL 8B Instruct qwen3-vl-8b-instruct multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 15.4 overall | 8.0 | 0.0 | 24.0 | 0.0 | 0.0 | N/A |
| 238 | Mistral Small 3.1 24B Instruct mistral-small-3.1-24b-instruct-2503 multimodalvisionmulti-input reasoning | Mistral AI | 15.1 overall | 15.1 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 239 | Llama 3.1 Nemotron Nano 8B V1 llama-3.1-nemotron-nano-8b-v1 textinference | NVIDIA | 15.0 overall | 15.0 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 240 | Llama 3.2 90B Instruct llama-3.2-90b-instruct multimodalvisionmulti-input reasoning | Meta | 14.9 overall | 14.9 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
Nova Pro
Amazon
19.0
N/A
North Mini Code 1.0
Cohere
18.6
N/A
DeepSeek-V3
DeepSeek
18.5
N/A
Want benchmark charts, model comparison, and pricing analytics?
Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.
Open full leaderboardRankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.
| N/A |
| N/A |
| N/A |
| N/A |
Mistral Small 3.2 24B Instruct
Mistral AI
18.4
N/A
Llama 3.1 405B Instruct
Meta
18.3
N/A
Llama 3.3 70B Instruct
Meta
18.0
N/A
Gemma 4 E4B
17.8
N/A
Claude 3 Opus
Anthropic
17.7
N/A
Kimi K2 Instruct
Moonshot AI
17.3
N/A
Qwen2.5 32B Instruct
Alibaba Cloud / Qwen Team
17.1
N/A
o3-mini
OpenAI
16.9
N/A
DeepSeek R1 Distill Qwen 7B
DeepSeek
16.8
N/A
DeepSeek R1 Distill Llama 8B
DeepSeek
16.3
N/A
Qwen2.5 72B Instruct
Alibaba Cloud / Qwen Team
16.3
N/A
Kimi K2-Instruct-0905
Moonshot AI
15.8
N/A
GPT OSS 20B
OpenAI
15.5
N/A
Qwen3 VL 8B Instruct
Alibaba Cloud / Qwen Team
15.4
N/A
Mistral Small 3.1 24B Instruct
Mistral AI
15.1
N/A
Llama 3.1 Nemotron Nano 8B V1
NVIDIA
15.0
N/A
Llama 3.2 90B Instruct
Meta
14.9
N/A