Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.
334
Tracked models
29
Providers
286
Benchmarked
28.3
Avg. index
334 models
| Rank | Model | Provider | Score | Benchmarks | Inference | Agentic | Programming | Value | Price |
|---|---|---|---|---|---|---|---|---|---|
| 261 | Llama 3.1 8B Instruct llama-3.1-8b-instruct textinference | Meta | 3.0 Benchmarks | 3.0 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 262 | Phi-3.5-mini-instruct phi-3.5-mini-instruct multimodalvisionmulti-input reasoning | Microsoft | 2.4 Benchmarks | 2.4 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 263 | GPT-3.5 Turbo gpt-3.5-turbo-0125 multimodalvisionmulti-input reasoning | OpenAI | 2.3 Benchmarks | 2.3 | 30.1 | 0.0 | 0.0 | 60.7 | |
| 264 | Phi-3.5-vision-instruct phi-3.5-vision-instruct multimodalvisionmulti-input reasoning | Microsoft | 2.3 Benchmarks | 2.3 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 265 | Qwen2 7B Instruct qwen2-7b-instruct textinference | Alibaba Cloud / Qwen Team | 2.2 Benchmarks | 2.2 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 266 | Phi 4 Mini phi-4-mini textinference | Microsoft | 1.9 Benchmarks | 1.9 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 267 | Gemma 3n E4B Instructed gemma-3n-e4b-it multimodalvisionmulti-input reasoning | Google | 1.2 Benchmarks | 1.2 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 268 | Gemma 3n E4B Instructed LiteRT Preview gemma-3n-e4b-it-litert-preview multimodalvisionmulti-input reasoning | Google | 1.2 Benchmarks | 1.2 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 269 | DeepSeek VL2 Tiny deepseek-vl2-tiny multimodalvisionmulti-input reasoning | DeepSeek | 1.1 Benchmarks | 1.1 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 270 | Gemma 3n E2B Instructed gemma-3n-e2b-it multimodalvisionmulti-input reasoning | Google | 1.0 Benchmarks | 1.0 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 271 | Gemma 3n E2B Instructed LiteRT (Preview) gemma-3n-e2b-it-litert-preview multimodalvisionmulti-input reasoning | Google | 1.0 Benchmarks | 1.0 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 272 | Gemma 3 1B gemma-3-1b-it textinference | Google | 0.9 Benchmarks | 0.9 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 273 | Codestral-22B codestral-22b textinference | Mistral AI | 0.0 Benchmarks | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 274 | Command R+ command-r-plus-04-2024 textinference | Cohere | 0.0 Benchmarks | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 275 | DeepSeek-V3.2 (Non-thinking) deepseek-chat textinference | DeepSeek | 0.0 Benchmarks | 0.0 | 52.0 | 0.0 | 0.0 | 82.7 | $0.28 in / $0.42 out |
| 276 | DeepSeek-R1 deepseek-r1 textinference | DeepSeek | 0.0 Benchmarks | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 277 | DeepSeek-V2.5 deepseek-v2.5 codeprogrammingtool use | DeepSeek | 0.0 Benchmarks | 0.0 | 0.0 | 0.0 | 0.7 | 0.0 | N/A |
| 278 | Devstral Medium devstral-medium-2507 codeprogrammingtool use | Mistral AI | 0.0 Benchmarks | 0.0 | 0.0 | 0.0 | 20.6 | 0.0 | N/A |
| 279 | Devstral Small 1.1 devstral-small-2507 codeprogrammingtool use | Mistral AI | 0.0 Benchmarks | 0.0 | 0.0 | 0.0 | 12.5 | 0.0 | N/A |
| 280 | Gemini 3.5 Flash Cyber gemini-3.5-flash-cyber textinference | Google | 0.0 Benchmarks | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
Llama 3.1 8B Instruct
Meta
3.0
N/A
Phi-3.5-mini-instruct
Microsoft
2.4
N/A
GPT-3.5 Turbo
OpenAI
2.3
$0.5 in / $1.5 out
Want benchmark charts, model comparison, and pricing analytics?
Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.
Open full leaderboardRankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.
| $0.5 in / $1.5 out |
| N/A |
| N/A |
| N/A |
| N/A |
| N/A |
Phi-3.5-vision-instruct
Microsoft
2.3
N/A
Qwen2 7B Instruct
Alibaba Cloud / Qwen Team
2.2
N/A
Phi 4 Mini
Microsoft
1.9
N/A
Gemma 3n E4B Instructed
1.2
N/A
Gemma 3n E4B Instructed LiteRT Preview
1.2
N/A
DeepSeek VL2 Tiny
DeepSeek
1.1
N/A
Gemma 3n E2B Instructed
1.0
N/A
Gemma 3n E2B Instructed LiteRT (Preview)
1.0
N/A
Gemma 3 1B
0.9
N/A
Codestral-22B
Mistral AI
0.0
N/A
Command R+
Cohere
0.0
N/A
DeepSeek-V3.2 (Non-thinking)
DeepSeek
0.0
$0.28 in / $0.42 out
DeepSeek-R1
DeepSeek
0.0
N/A
DeepSeek-V2.5
DeepSeek
0.0
N/A
Devstral Medium
Mistral AI
0.0
N/A
Devstral Small 1.1
Mistral AI
0.0
N/A
Gemini 3.5 Flash Cyber
0.0
N/A