Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.
334
Tracked models
29
Providers
286
Benchmarked
28.3
Avg. index
334 models
| Rank | Model | Provider | Score | Benchmarks | Inference | Agentic | Programming | Value | Price |
|---|---|---|---|---|---|---|---|---|---|
| 241 | Qwen3 VL 8B Instruct qwen3-vl-8b-instruct multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 8.0 Benchmarks | 8.0 | 0.0 | 24.0 | 0.0 | 0.0 | N/A |
| 242 | Gemma 3 27B gemma-3-27b-it multimodalvisionmulti-input reasoning | Google | 7.6 Benchmarks | 7.6 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 243 | Jamba 1.5 Large jamba-1.5-large textinference | AI21 Labs | 7.5 Benchmarks | 7.5 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 244 | Phi-3.5-MoE-instruct phi-3.5-moe-instruct multimodalvisionmulti-input reasoning | Microsoft | 7.5 Benchmarks | 7.5 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 245 | Qwen2.5-Omni-7B qwen2.5-omni-7b multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 7.2 Benchmarks | 7.2 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 246 | DeepSeek VL2 deepseek-vl2 multimodalvisionmulti-input reasoning | DeepSeek | 6.8 Benchmarks | 6.8 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 247 | Qwen2.5 7B Instruct qwen-2.5-7b-instruct textinference | Alibaba Cloud / Qwen Team | 6.8 Benchmarks | 6.8 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 248 | Qwen2-VL-72B-Instruct qwen2-vl-72b multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 6.8 Benchmarks | 6.8 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 249 | Gemini Diffusion gemini-diffusion codeprogrammingtool use | Google | 6.5 Benchmarks | 6.5 | 0.0 | 0.0 | 1.5 | 0.0 | N/A |
| 250 | GPT-4 gpt-4-0613 multimodalvisionmulti-input reasoning | OpenAI | 6.2 Benchmarks | 6.2 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 251 | Mistral Small 3 24B Base mistral-small-24b-base-2501 multimodalvisionmulti-input reasoning | Mistral AI | 5.9 Benchmarks | 5.9 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 252 | DeepSeek R1 Distill Qwen 1.5B deepseek-r1-distill-qwen-1.5b textinference | DeepSeek | 5.6 Benchmarks | 5.6 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 253 | Claude 3 Haiku claude-3-haiku-20240307 multimodalvisionmulti-input reasoning | Anthropic | 5.3 Benchmarks | 5.3 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 254 | Llama 3.2 3B Instruct llama-3.2-3b-instruct textinference | Meta | 4.8 Benchmarks | 4.8 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 255 | DeepSeek VL2 Small deepseek-vl2-small multimodalvisionmulti-input reasoning | DeepSeek | 4.6 Benchmarks | 4.6 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 256 | Gemma 3 4B gemma-3-4b-it multimodalvisionmulti-input reasoning | Google | 4.3 Benchmarks | 4.3 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 257 | Jamba 1.5 Mini jamba-1.5-mini textinference | AI21 Labs | 4.3 Benchmarks | 4.3 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 258 | GPT-5.1 Codex Mini gpt-5.1-codex-mini multimodalvisionmulti-input reasoning | OpenAI | 3.8 Benchmarks | 3.8 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 259 | Llama 3.2 11B Instruct llama-3.2-11b-instruct multimodalvisionmulti-input reasoning | Meta | 3.8 Benchmarks | 3.8 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 260 | Gemini 1.0 Pro gemini-1.0-pro multimodalvisionmulti-input reasoning | Google | 3.0 Benchmarks | 3.0 | 0.0 | 0.0 | 0.0 | 0.0 |
Qwen3 VL 8B Instruct
Alibaba Cloud / Qwen Team
8.0
N/A
Gemma 3 27B
7.6
N/A
Jamba 1.5 Large
AI21 Labs
7.5
N/A
Want benchmark charts, model comparison, and pricing analytics?
Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.
Open full leaderboardRankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.
| N/A |
| N/A |
| N/A |
| N/A |
| N/A |
Phi-3.5-MoE-instruct
Microsoft
7.5
N/A
Qwen2.5-Omni-7B
Alibaba Cloud / Qwen Team
7.2
N/A
DeepSeek VL2
DeepSeek
6.8
N/A
Qwen2.5 7B Instruct
Alibaba Cloud / Qwen Team
6.8
N/A
Qwen2-VL-72B-Instruct
Alibaba Cloud / Qwen Team
6.8
N/A
Gemini Diffusion
6.5
N/A
GPT-4
OpenAI
6.2
N/A
Mistral Small 3 24B Base
Mistral AI
5.9
N/A
DeepSeek R1 Distill Qwen 1.5B
DeepSeek
5.6
N/A
Claude 3 Haiku
Anthropic
5.3
N/A
Llama 3.2 3B Instruct
Meta
4.8
N/A
DeepSeek VL2 Small
DeepSeek
4.6
N/A
Gemma 3 4B
4.3
N/A
Jamba 1.5 Mini
AI21 Labs
4.3
N/A
GPT-5.1 Codex Mini
OpenAI
3.8
N/A
Llama 3.2 11B Instruct
Meta
3.8
N/A
Gemini 1.0 Pro
3.0
N/A