Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.
334
Tracked models
29
Providers
286
Benchmarked
29.3
Avg. index
334 models
| Rank | Model | Provider | Score | Benchmarks | Inference | Agentic | Programming | Value | Price |
|---|---|---|---|---|---|---|---|---|---|
| 181 | o1-preview o1-preview codeprogrammingtool use | OpenAI | 26.1 overall | 40.2 | 0.0 | 0.0 | 8.1 | 0.0 | N/A |
| 182 | Kimi K2 Base kimi-k2-base textinference | Moonshot AI | 26.0 overall | 26.0 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 183 | Qwen3 VL 32B Instruct qwen3-vl-32b-instruct multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 25.9 overall | 26.5 | 0.0 | 25.1 | 0.0 | 0.0 | |
| 184 | DeepSeek-V3.1 deepseek-v3.1 codeprogrammingtool use | DeepSeek | 25.8 overall | 36.7 | 0.0 | 13.6 | 25.4 | 0.0 | N/A |
| 185 | MiniCPM-SALA minicpm-sala textinference | OpenBMB | 25.8 overall | 25.8 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 186 | GPT-4 Turbo gpt-4-turbo-2024-04-09 textinference | OpenAI | 25.8 overall | 15.5 | 50.2 | 0.0 | 0.0 | 15.4 | $10 in / $30 out |
| 187 | Grok-2 grok-2 multimodalvisionmulti-input reasoning | xAI | 25.7 overall | 25.7 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 188 | Qwen3 32B qwen3-32b textinference | Alibaba Cloud / Qwen Team | 25.5 overall | 20.1 | 2.2 | 0.0 | 0.0 | 78.0 | $0.1 in / $0.3 out |
| 189 | DeepSeek R1 Distill Qwen 32B deepseek-r1-distill-qwen-32b textinference | DeepSeek | 24.4 overall | 24.4 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 190 | Gemini 2.0 Flash-Lite gemini-2.0-flash-lite multimodalvisionmulti-input reasoning | Google | 24.4 overall | 24.4 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 191 | Sarvam-105B sarvam-105b codeprogrammingtool use | Sarvam AI | 24.0 overall | 41.1 | 0.0 | 16.7 | 10.3 | 0.0 | N/A |
| 192 | Qwen3 VL 30B A3B Instruct qwen3-vl-30b-a3b-instruct multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 23.7 overall | 25.6 | 0.0 | 21.4 | 0.0 | 0.0 | |
| 193 | o1-mini o1-mini textinference | OpenAI | 23.6 overall | 23.6 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 194 | Claude 3.5 Sonnet claude-3-5-sonnet-20240620 multimodalvisionmulti-input reasoning | Anthropic | 23.3 overall | 23.3 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 195 | Nemotron Nano 9B v2 nvidia-nemotron-nano-9b-v2 textinference | NVIDIA | 23.1 overall | 23.1 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 196 | Qwen3-Next-80B-A3B-Instruct qwen3-next-80b-a3b-instruct textinference | Alibaba Cloud / Qwen Team | 23.0 overall | 27.4 | 0.0 | 17.9 | 0.0 | 0.0 | N/A |
| 197 | DiffusionGemma 26B-A4B diffusiongemma-26b-a4b-it multimodalvisionmulti-input reasoning | Google | 22.9 overall | 22.9 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 198 | ERNIE 4.5 ernie-4.5 textinference | Baidu | 22.8 overall | 22.8 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 199 | Grok-2 mini grok-2-mini multimodalvisionmulti-input reasoning | xAI | 22.8 overall | 22.8 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 200 | DeepSeek R1 Distill Qwen 14B deepseek-r1-distill-qwen-14b textinference | DeepSeek | 22.7 overall | 22.7 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
o1-preview
OpenAI
26.1
N/A
Kimi K2 Base
Moonshot AI
26.0
N/A
Qwen3 VL 32B Instruct
Alibaba Cloud / Qwen Team
25.9
N/A
Want benchmark charts, model comparison, and pricing analytics?
Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.
Open full leaderboardRankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.
| N/A |
| N/A |
| N/A |
| N/A |
| N/A |
DeepSeek-V3.1
DeepSeek
25.8
N/A
MiniCPM-SALA
OpenBMB
25.8
N/A
GPT-4 Turbo
OpenAI
25.8
$10 in / $30 out
Grok-2
xAI
25.7
N/A
Qwen3 32B
Alibaba Cloud / Qwen Team
25.5
$0.1 in / $0.3 out
DeepSeek R1 Distill Qwen 32B
DeepSeek
24.4
N/A
Gemini 2.0 Flash-Lite
24.4
N/A
Sarvam-105B
Sarvam AI
24.0
N/A
Qwen3 VL 30B A3B Instruct
Alibaba Cloud / Qwen Team
23.7
N/A
o1-mini
OpenAI
23.6
N/A
Claude 3.5 Sonnet
Anthropic
23.3
N/A
Nemotron Nano 9B v2
NVIDIA
23.1
N/A
Qwen3-Next-80B-A3B-Instruct
Alibaba Cloud / Qwen Team
23.0
N/A
DiffusionGemma 26B-A4B
22.9
N/A
ERNIE 4.5
Baidu
22.8
N/A
Grok-2 mini
xAI
22.8
N/A
DeepSeek R1 Distill Qwen 14B
DeepSeek
22.7
N/A