Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.
373
Tracked models
34
Providers
310
Benchmarked
14.1
Avg. index
373 models
| Rank | Model | Provider | Score | Benchmarks | Inference | Agentic | Programming | Value | Price |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Nemotron 3.5 Lightning (30B A3B) nemotron-3.5-lightning-30b-a3b codeprogrammingtool use | NVIDIA | 100.0 Value / Price | 26.6 | 30.5 | 6.9 | 9.0 | 100.0 | $0.05 in / $0.2 out |
| 2 | DeepSeek-V4-Flash-0731 deepseek-v4-flash-0731 textinference | DeepSeek | 99.0 Value / Price | 0.0 | 81.9 | 56.3 | 0.0 | 99.0 | $0.09 in / $0.18 out |
| 3 | Nemotron 3 Nano (30B A3B) nemotron-3-nano-30b-a3b codeprogrammingtool use | NVIDIA | 98.1 Value / Price | 42.4 | 30.5 | 3.0 | 5.9 | 98.1 | |
| 4 | DeepSeek-V4-Flash-0423 deepseek-v4-flash-0423 codeprogrammingtool use | DeepSeek | 95.1 Value / Price | 51.4 | 81.9 | 23.2 | 40.6 | 95.1 | |
| 5 | Laguna S 2.1 laguna-s-2.1 codeprogrammingtool use | Poolside | 95.1 Value / Price | 0.0 | 81.9 | 32.7 | 52.2 | 95.1 | $0.1 in / $0.2 out |
| 6 | Laguna XS 2.1 laguna-xs-2.1 codeprogrammingtool use | Poolside | 95.1 Value / Price | 0.0 | 30.5 | 0.0 | 22.3 | 95.1 | $0.1 in / $0.2 out |
| 7 | Muse Spark 1.2 muse-spark-1.2 multimodalvisionmulti-input reasoning | Meta | 95.1 Value / Price | 0.0 | 81.9 | 0.0 | 0.0 | 95.1 | $0.1 in / $0.2 out |
| 8 | Muse Spark 1.3 muse-spark-1.3 multimodalvisionmulti-input reasoning | Meta | 95.1 Value / Price | 70.4 | 81.9 | 0.0 | 0.0 | 95.1 | $0.1 in / $0.2 out |
| 9 | DeepSeek-V4-Flash-Max deepseek-v4-flash-max codeprogrammingtool use | DeepSeek | 91.3 Value / Price | 55.3 | 81.9 | 32.2 | 44.1 | 91.3 | |
| 10 | LongCat-Flash-Lite longcat-flash-lite codeprogrammingtool use | Meituan | 91.0 Value / Price | 22.3 | 71.9 | 30.1 | 24.1 | 91.0 | |
| 11 | GPT-4.1 nano gpt-4.1-nano-2025-04-14 multimodalvisionmulti-input reasoning | OpenAI | 90.3 Value / Price | 11.1 | 85.9 | 0.0 | 0.0 | 90.3 | |
| 12 | Step-3.5-Flash step-3.5-flash codeprogrammingtool use | StepFun | 89.3 Value / Price | 63.2 | 60.2 | 36.5 | 48.8 | 89.3 | $0.1 in / $0.4 out |
| 13 | MiMo-V2.5 mimo-v2.5 multimodalvisionmulti-input reasoning | Xiaomi | 87.4 Value / Price | 46.5 | 81.9 | 0.0 | 29.7 | 87.4 | $0.168 in / $0.336 out |
| 14 | Gemma 4 31B gemma-4-31b-it multimodalvisionmulti-input reasoning | Google | 86.4 Value / Price | 52.9 | 30.5 | 0.0 | 0.0 | 86.4 | |
| 15 | Gemma 4 26B-A4B gemma-4-26b-a4b-it multimodalvisionmulti-input reasoning | Google | 85.4 Value / Price | 41.6 | 30.5 | 0.0 | 0.0 | 85.4 | |
| 16 | Qwen3.8 Flash qwen3.8-flash multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 83.5 Value / Price | 61.1 | 58.6 | 64.5 | 60.0 | 83.5 | $0.15 in / $0.47 out |
| 17 | GLM-5.3-Flash glm-5.3-flash multimodalvisionmulti-input reasoning | Zhipu AI | 82.5 Value / Price | 64.0 | 81.9 | 72.6 | 0.0 | 82.5 | $0.15 in / $0.5 out |
| 18 | Mercury 2 mercury-2 codeprogrammingtool use | Inception | 79.8 Value / Price | 42.1 | 68.7 | 0.0 | 21.0 | 79.8 | $0.25 in / $0.75 out |
| 19 | Qwen3 VL 4B Instruct qwen3-vl-4b-instruct multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 79.6 Value / Price | 18.3 | 30.5 | 16.1 | 0.0 | 79.6 | |
| 20 | DeepSeek-V3.2 (Non-thinking) deepseek-chat textinference | DeepSeek | 77.3 Value / Price | 0.0 | 48.7 | 0.0 | 0.0 | 77.3 | $0.28 in / $0.42 out |
Nemotron 3.5 Lightning (30B A3B)
NVIDIA
100.0
$0.05 in / $0.2 out
DeepSeek-V4-Flash-0731
DeepSeek
99.0
$0.09 in / $0.18 out
Nemotron 3 Nano (30B A3B)
NVIDIA
98.1
$0.06 in / $0.24 out
Page 1 of 19 · 373 models
Want benchmark charts, model comparison, and pricing analytics?
Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.
Open full leaderboardRankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.
| $0.06 in / $0.24 out |
| $0.1 in / $0.2 out |
| $0.14 in / $0.28 out |
| $0.1 in / $0.4 out |
| $0.1 in / $0.4 out |
| $0.13 in / $0.38 out |
| $0.13 in / $0.4 out |
| $0.1 in / $0.6 out |
DeepSeek-V4-Flash-0423
DeepSeek
95.1
$0.1 in / $0.2 out
Laguna S 2.1
Poolside
95.1
$0.1 in / $0.2 out
Laguna XS 2.1
Poolside
95.1
$0.1 in / $0.2 out
Muse Spark 1.2
Meta
95.1
$0.1 in / $0.2 out
Muse Spark 1.3
Meta
95.1
$0.1 in / $0.2 out
DeepSeek-V4-Flash-Max
DeepSeek
91.3
$0.14 in / $0.28 out
LongCat-Flash-Lite
Meituan
91.0
$0.1 in / $0.4 out
GPT-4.1 nano
OpenAI
90.3
$0.1 in / $0.4 out
Step-3.5-Flash
StepFun
89.3
$0.1 in / $0.4 out
MiMo-V2.5
Xiaomi
87.4
$0.168 in / $0.336 out
Gemma 4 31B
86.4
$0.13 in / $0.38 out
Gemma 4 26B-A4B
85.4
$0.13 in / $0.4 out
Qwen3.8 Flash
Alibaba Cloud / Qwen Team
83.5
$0.15 in / $0.47 out
GLM-5.3-Flash
Zhipu AI
82.5
$0.15 in / $0.5 out
Mercury 2
Inception
79.8
$0.25 in / $0.75 out
Qwen3 VL 4B Instruct
Alibaba Cloud / Qwen Team
79.6
$0.1 in / $0.6 out
DeepSeek-V3.2 (Non-thinking)
DeepSeek
77.3
$0.28 in / $0.42 out