Skytells
HomeModelsCLIChangelog
  • Home
  • Models
  • CLI
  • Changelog
Skytells

Addressing the world's greatest challenges with AI. Enterprise research, foundation models, and infrastructure trusted by organizations worldwide since 2012.

Get Started

  • Console
  • Learn
  • Documentation
  • API Reference
  • Pricing
  • ModelsNew

Platform

  • Cloud AgentsNew
  • AI Solutions
  • Infrastructure
  • Edge Network
  • Trust Center
  • CLI

Resources

  • Blog
  • Changelog
  • AI Leaderboard
  • Research
  • Status

Company

  • About
  • Careers
  • Legal
  • Privacy Policy

© 2012–2026 Skytells, Inc. All rights reserved.

Live rankings

AI Model Leaderboard

Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.

Explore full leaderboardBrowse model catalog

334

Tracked models

29

Providers

286

Benchmarked

28.3

Avg. index

OverallBenchmarksInferenceAgenticProgrammingValue / Price

334 models

RankModelProviderScoreBenchmarksInferenceAgenticProgrammingValuePrice
141

Claude Haiku 4.5

claude-haiku-4-5-20251001

multimodalvisionmulti-input reasoning
Anthropic

30.8

Benchmarks

30.853.350.853.945.6$1 in / $5 out
142

DeepSeek-V3 0324

deepseek-v3-0324

textinference
DeepSeek

30.4

Benchmarks

30.40.00.00.00.0N/A
143

Qwen3.5-4B

qwen3.5-4b

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

29.9

Benchmarks

29.90.00.00.00.0N/A
144

MiniMax M2

minimax-m2

codeprogrammingtool use
MiniMax

29.8

Benchmarks

29.862.840.240.073.2$0.3 in / $1.2 out
145

Ministral 3 (8B Reasoning 2512)

ministral-8b-latest

multimodalvisionmulti-input reasoning
Mistral AI

29.4

Benchmarks

29.40.00.00.00.0
146

Phi 4 Reasoning Plus

phi-4-reasoning-plus

textinference
MMicrosoft

29.4

Benchmarks

29.40.00.00.00.0N/A
147

Qwen3 235B A22B

qwen3-235b-a22b

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

29.3

Benchmarks

29.30.00.00.00.0N/A
148

GPT-4o

gpt-4o-2024-08-06

multimodalvisionmulti-input reasoning
OpenAI

29.1

Benchmarks

29.139.614.93.731.2
149

Qwen3 Max

qwen3-max

codeprogrammingtool use
AAlibaba Cloud / Qwen Team

28.0

Benchmarks

28.00.00.032.10.0N/A
150

Hermes 3 70B

hermes-3-70b

textinference
NNous Research

27.6

Benchmarks

27.60.00.00.00.0N/A
151

Llama 4 Scout

llama-4-scout

multimodalvisionmulti-input reasoning
MMeta

27.6

Benchmarks

27.60.00.00.00.0N/A
152

Qwen3-Next-80B-A3B-Instruct

qwen3-next-80b-a3b-instruct

textinference
AAlibaba Cloud / Qwen Team

27.4

Benchmarks

27.40.017.90.00.0N/A
153

Pixtral Large

pixtral-large

multimodalvisionmulti-input reasoning
Mistral AI

27.3

Benchmarks

27.30.00.00.00.0
154

GPT-4.1

gpt-4.1-2025-04-14

multimodalvisionmulti-input reasoning
OpenAI

27.2

Benchmarks

27.273.232.814.740.7
155

Qwen3 VL 32B Instruct

qwen3-vl-32b-instruct

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

26.5

Benchmarks

26.50.025.10.00.0
156

DeepSeek R1 Distill Llama 70B

deepseek-r1-distill-llama-70b

textinference
DeepSeek

26.4

Benchmarks

26.40.00.00.00.0N/A
157

QwQ-32B

qwq-32b

textinference
AAlibaba Cloud / Qwen Team

26.4

Benchmarks

26.40.00.00.00.0N/A
158

QwQ-32B-Preview

qwq-32b-preview

textinference
AAlibaba Cloud / Qwen Team

26.4

Benchmarks

26.40.00.00.00.0N/A
159

Gemini 1.5 Pro

gemini-1.5-pro

multimodalvisionmulti-input reasoning
Google

26.2

Benchmarks

26.20.00.00.00.0
160

DeepSeek-V3

deepseek-v3

codeprogrammingtool use
DeepSeek

26.1

Benchmarks

26.10.00.08.80.0N/A
141

Claude Haiku 4.5

Anthropic

30.8

$1 in / $5 out

142

DeepSeek-V3 0324

DeepSeek

30.4

N/A

143
A

Qwen3.5-4B

Alibaba Cloud / Qwen Team

29.9

N/A

144

Page 8 of 17 · 334 models

PreviousNext

Want benchmark charts, model comparison, and pricing analytics?

Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.

Open full leaderboard

Rankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.

N/A
$2.5 in / $10 out
N/A
$2 in / $8 out
N/A
N/A

MiniMax M2

MiniMax

29.8

$0.3 in / $1.2 out

145

Ministral 3 (8B Reasoning 2512)

Mistral AI

29.4

N/A

146
M

Phi 4 Reasoning Plus

Microsoft

29.4

N/A

147
A

Qwen3 235B A22B

Alibaba Cloud / Qwen Team

29.3

N/A

148

GPT-4o

OpenAI

29.1

$2.5 in / $10 out

149
A

Qwen3 Max

Alibaba Cloud / Qwen Team

28.0

N/A

150
N

Hermes 3 70B

Nous Research

27.6

N/A

151
M

Llama 4 Scout

Meta

27.6

N/A

152
A

Qwen3-Next-80B-A3B-Instruct

Alibaba Cloud / Qwen Team

27.4

N/A

153

Pixtral Large

Mistral AI

27.3

N/A

154

GPT-4.1

OpenAI

27.2

$2 in / $8 out

155
A

Qwen3 VL 32B Instruct

Alibaba Cloud / Qwen Team

26.5

N/A

156

DeepSeek R1 Distill Llama 70B

DeepSeek

26.4

N/A

157
A

QwQ-32B

Alibaba Cloud / Qwen Team

26.4

N/A

158
A

QwQ-32B-Preview

Alibaba Cloud / Qwen Team

26.4

N/A

159

Gemini 1.5 Pro

Google

26.2

N/A

160

DeepSeek-V3

DeepSeek

26.1

N/A