Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.
334
Tracked models
29
Providers
286
Benchmarked
28.3
Avg. index
334 models
| Rank | Model | Provider | Score | Benchmarks | Inference | Agentic | Programming | Value | Price |
|---|---|---|---|---|---|---|---|---|---|
| 81 | LongCat-Flash-Thinking longcat-flash-thinking codeprogrammingtool use | Meituan | 48.2 Benchmarks | 48.2 | 0.0 | 0.0 | 18.4 | 0.0 | N/A |
| 82 | DeepSeek-R1-0528 deepseek-r1-0528 codeprogrammingtool use | DeepSeek | 47.9 Benchmarks | 47.9 | 0.0 | 0.0 | 5.7 | 0.0 | N/A |
| 83 | MiMo-V2.5 mimo-v2.5 multimodalvisionmulti-input reasoning | Xiaomi | 47.7 Benchmarks | 47.7 | 84.8 | 0.0 | 27.2 | 92.7 | $0.168 in / $0.336 out |
| 84 | Step3-VL-10B step3-vl-10b multimodalvisionmulti-input reasoning | StepFun | 46.9 Benchmarks | 46.9 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 85 | o4-mini o4-mini multimodalvisionmulti-input reasoning | OpenAI | 46.2 Benchmarks | 46.2 | 0.0 | 36.1 | 28.7 | 0.0 | N/A |
| 86 | Claude Opus 4.1 claude-opus-4-1-20250805 multimodalvisionmulti-input reasoning | Anthropic | 46.0 Benchmarks | 46.0 | 0.0 | 67.4 | 60.8 | 0.0 | |
| 87 | Nemotron 3 Super (120B A12B) nemotron-3-super-120b-a12b codeprogrammingtool use | NVIDIA | 45.5 Benchmarks | 45.5 | 0.0 | 7.6 | 22.0 | 0.0 | N/A |
| 88 | Nova 2 Pro nova-2-pro multimodalvisionmulti-input reasoning | Amazon | 45.3 Benchmarks | 45.3 | 0.0 | 57.2 | 49.6 | 0.0 | N/A |
| 89 | Gemini 2.0 Flash Thinking gemini-2.0-flash-thinking multimodalvisionmulti-input reasoning | Google | 44.9 Benchmarks | 44.9 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 90 | Sarvam-30B sarvam-30b codeprogrammingtool use | Sarvam AI | 44.9 Benchmarks | 44.9 | 0.0 | 6.4 | 4.4 | 0.0 | N/A |
| 91 | GLM-4.6 glm-4.6 multimodalvisionmulti-input reasoning | Zhipu AI | 44.5 Benchmarks | 44.5 | 0.0 | 35.4 | 43.2 | 0.0 | N/A |
| 92 | Qwen3-235B-A22B-Thinking-2507 qwen3-235b-a22b-thinking-2507 textinference | Alibaba Cloud / Qwen Team | 44.4 Benchmarks | 44.4 | 0.0 | 26.8 | 0.0 | 0.0 | N/A |
| 93 | GPT OSS 120B High gpt-oss-120b-high multimodalvisionmulti-input reasoning | OpenAI | 44.1 Benchmarks | 44.1 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 94 | o1-pro o1-pro multimodalvisionmulti-input reasoning | OpenAI | 44.1 Benchmarks | 44.1 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 95 | K-EXAONE-236B-A23B k-exaone-236b-a23b multimodalvisionmulti-input reasoning | LG AI Research | 43.9 Benchmarks | 43.9 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 96 | Gemma 4 26B-A4B gemma-4-26b-a4b-it multimodalvisionmulti-input reasoning | Google | 43.8 Benchmarks | 43.8 | 32.9 | 0.0 | 0.0 | 90.2 | |
| 97 | Nemotron 3 Nano (30B A3B) nemotron-3-nano-30b-a3b codeprogrammingtool use | NVIDIA | 43.6 Benchmarks | 43.6 | 32.9 | 3.0 | 3.8 | 100.0 | $0.06 in / $0.24 out |
| 98 | o3 o3-2025-04-16 multimodalvisionmulti-input reasoning | OpenAI | 42.9 Benchmarks | 42.9 | 0.0 | 17.9 | 27.7 | 0.0 | N/A |
| 99 | Mercury 2 mercury-2 codeprogrammingtool use | Inception | 42.7 Benchmarks | 42.7 | 68.9 | 0.0 | 15.3 | 84.4 | $0.25 in / $0.75 out |
| 100 | Gemini 2.5 Pro gemini-2.5-pro multimodalvisionmulti-input reasoning | Google | 42.5 Benchmarks | 42.5 | 51.0 | 0.0 | 21.4 | 29.8 |
LongCat-Flash-Thinking
Meituan
48.2
N/A
DeepSeek-R1-0528
DeepSeek
47.9
N/A
MiMo-V2.5
Xiaomi
47.7
$0.168 in / $0.336 out
Want benchmark charts, model comparison, and pricing analytics?
Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.
Open full leaderboardRankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.
| N/A |
| N/A |
| N/A |
| $0.13 in / $0.4 out |
| $1.25 in / $10 out |
Step3-VL-10B
StepFun
46.9
N/A
o4-mini
OpenAI
46.2
N/A
Claude Opus 4.1
Anthropic
46.0
N/A
Nemotron 3 Super (120B A12B)
NVIDIA
45.5
N/A
Nova 2 Pro
Amazon
45.3
N/A
Gemini 2.0 Flash Thinking
44.9
N/A
Sarvam-30B
Sarvam AI
44.9
N/A
GLM-4.6
Zhipu AI
44.5
N/A
Qwen3-235B-A22B-Thinking-2507
Alibaba Cloud / Qwen Team
44.4
N/A
GPT OSS 120B High
OpenAI
44.1
N/A
o1-pro
OpenAI
44.1
N/A
K-EXAONE-236B-A23B
LG AI Research
43.9
N/A
Gemma 4 26B-A4B
43.8
$0.13 in / $0.4 out
Nemotron 3 Nano (30B A3B)
NVIDIA
43.6
$0.06 in / $0.24 out
o3
OpenAI
42.9
N/A
Mercury 2
Inception
42.7
$0.25 in / $0.75 out
Gemini 2.5 Pro
42.5
$1.25 in / $10 out