Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.
334
Tracked models
29
Providers
286
Benchmarked
12.2
Avg. index
334 models
| Rank | Model | Provider | Score | Benchmarks | Inference | Agentic | Programming | Value | Price |
|---|---|---|---|---|---|---|---|---|---|
| 101 | Nemotron 3 Ultra (550B A55B) nemotron-3-ultra-550b-a55b codeprogrammingtool use | NVIDIA | 11.5 Agentic | 54.7 | 0.0 | 11.5 | 40.6 | 0.0 | N/A |
| 102 | Qwen3.6-35B-A3B qwen3.6-35b-a3b multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 9.8 Agentic | 51.1 | 0.0 | 9.8 | 25.2 | 0.0 | N/A |
| 103 | GLM-4.7-Flash glm-4.7-flash codeprogrammingtool use | Zhipu AI | 9.0 Agentic | 36.4 | 0.0 | 9.0 | 17.7 | 0.0 | N/A |
| 104 | GPT-4.1 mini gpt-4.1-mini-2025-04-14 multimodalvisionmulti-input reasoning | OpenAI | 8.9 Agentic | 19.2 | 84.6 | 8.9 | 2.2 | 69.5 | |
| 105 | Nemotron 3 Super (120B A12B) nemotron-3-super-120b-a12b codeprogrammingtool use | NVIDIA | 7.6 Agentic | 45.5 | 0.0 | 7.6 | 22.0 | 0.0 | N/A |
| 106 | GPT-5.4 nano gpt-5.4-nano multimodalvisionmulti-input reasoning | OpenAI | 6.9 Agentic | 41.8 | 44.5 | 6.9 | 8.2 | 76.8 | $0.2 in / $1.25 out |
| 107 | Sarvam-30B sarvam-30b codeprogrammingtool use | Sarvam AI | 6.4 Agentic | 44.9 | 0.0 | 6.4 | 4.4 | 0.0 | N/A |
| 108 | GPT OSS 20B gpt-oss-20b textinference | OpenAI | 6.0 Agentic | 23.7 | 0.0 | 6.0 | 0.0 | 0.0 | N/A |
| 109 | Kimi K2-Instruct-0905 kimi-k2-instruct-0905 codeprogrammingtool use | Moonshot AI | 6.0 Agentic | 23.2 | 0.0 | 6.0 | 17.1 | 0.0 | |
| 110 | Qwen2.5 VL 72B Instruct qwen2.5-vl-72b multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 5.2 Agentic | 22.7 | 0.0 | 5.2 | 0.0 | 0.0 | N/A |
| 111 | DeepSeek-V3.2-Speciale deepseek-v3.2-speciale codeprogrammingtool use | DeepSeek | 5.0 Agentic | 50.7 | 0.0 | 5.0 | 42.0 | 0.0 | |
| 112 | Claude 3.5 Haiku claude-3-5-haiku-20241022 codeprogrammingtool use | Anthropic | 3.0 Agentic | 9.9 | 0.0 | 3.0 | 6.6 | 0.0 | |
| 113 | Nemotron 3 Nano (30B A3B) nemotron-3-nano-30b-a3b codeprogrammingtool use | NVIDIA | 3.0 Agentic | 43.6 | 32.9 | 3.0 | 3.8 | 100.0 | $0.06 in / $0.24 out |
| 114 | Qwen2.5 VL 32B Instruct qwen2.5-vl-32b multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 1.5 Agentic | 19.4 | 0.0 | 1.5 | 0.0 | 0.0 | N/A |
| 115 | ChatGPT-4o Latest chatgpt-4o-latest multimodalvisionmulti-input reasoning | OpenAI | 0.0 Agentic | 53.0 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 116 | Claude 3.5 Sonnet claude-3-5-sonnet-20240620 multimodalvisionmulti-input reasoning | Anthropic | 0.0 Agentic | 23.3 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 117 | Claude 3 Haiku claude-3-haiku-20240307 multimodalvisionmulti-input reasoning | Anthropic | 0.0 Agentic | 5.3 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 118 | Claude 3 Opus claude-3-opus-20240229 multimodalvisionmulti-input reasoning | Anthropic | 0.0 Agentic | 17.7 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 119 | Claude 3 Sonnet claude-3-sonnet-20240229 multimodalvisionmulti-input reasoning | Anthropic | 0.0 Agentic | 9.2 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 120 | Claude Fable 5 claude-fable-5 multimodalvisionmulti-input reasoning | Anthropic | 0.0 Agentic | 70.8 | 62.8 | 0.0 | 84.2 | 0.0 |
Nemotron 3 Ultra (550B A55B)
NVIDIA
11.5
N/A
Qwen3.6-35B-A3B
Alibaba Cloud / Qwen Team
9.8
N/A
GLM-4.7-Flash
Zhipu AI
9.0
N/A
Want benchmark charts, model comparison, and pricing analytics?
Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.
Open full leaderboardRankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.
| $0.4 in / $1.6 out |
| N/A |
| N/A |
| N/A |
| N/A |
| N/A |
| N/A |
| N/A |
| N/A |
| $10 in / $50 out |
GPT-4.1 mini
OpenAI
8.9
$0.4 in / $1.6 out
Nemotron 3 Super (120B A12B)
NVIDIA
7.6
N/A
GPT-5.4 nano
OpenAI
6.9
$0.2 in / $1.25 out
Sarvam-30B
Sarvam AI
6.4
N/A
GPT OSS 20B
OpenAI
6.0
N/A
Kimi K2-Instruct-0905
Moonshot AI
6.0
N/A
Qwen2.5 VL 72B Instruct
Alibaba Cloud / Qwen Team
5.2
N/A
DeepSeek-V3.2-Speciale
DeepSeek
5.0
N/A
Claude 3.5 Haiku
Anthropic
3.0
N/A
Nemotron 3 Nano (30B A3B)
NVIDIA
3.0
$0.06 in / $0.24 out
Qwen2.5 VL 32B Instruct
Alibaba Cloud / Qwen Team
1.5
N/A
ChatGPT-4o Latest
OpenAI
0.0
N/A
Claude 3.5 Sonnet
Anthropic
0.0
N/A
Claude 3 Haiku
Anthropic
0.0
N/A
Claude 3 Opus
Anthropic
0.0
N/A
Claude 3 Sonnet
Anthropic
0.0
N/A
Claude Fable 5
Anthropic
0.0
$10 in / $50 out