Skytells
HomeModelsCLIChangelog
  • Home
  • Models
  • CLI
  • Changelog
Skytells

Addressing the world's greatest challenges with AI. Enterprise research, foundation models, and infrastructure trusted by organizations worldwide since 2012.

Get Started

  • Console
  • Learn
  • Documentation
  • API Reference
  • Pricing
  • ModelsNew

Platform

  • Cloud AgentsNew
  • AI Solutions
  • Infrastructure
  • Edge Network
  • Trust Center
  • CLI

Resources

  • Blog
  • Changelog
  • AI Leaderboard
  • Research
  • Status

Company

  • About
  • Careers
  • Legal
  • Privacy Policy

© 2012–2026 Skytells, Inc. All rights reserved.

Live rankings

AI Model Leaderboard

Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.

Explore full leaderboardBrowse model catalog

334

Tracked models

29

Providers

286

Benchmarked

29.3

Avg. index

OverallBenchmarksInferenceAgenticProgrammingValue / Price

334 models

RankModelProviderScoreBenchmarksInferenceAgenticProgrammingValuePrice
161

Qwen3 235B A22B

qwen3-235b-a22b

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

29.3

overall

29.30.00.00.00.0N/A
162

GPT-5.2 Codex

gpt-5.2-codex

multimodalvisionmulti-input reasoning
OpenAI

29.1

overall

0.031.10.030.922.0$1.75 in / $14 out
163

Nemotron 3 Nano (30B A3B)

nemotron-3-nano-30b-a3b

codeprogrammingtool use
NNVIDIA

29.0

overall

43.632.93.03.8100.0$0.06 in / $0.24 out
164

GPT-4.5

gpt-4.5

multimodalvisionmulti-input reasoning
OpenAI

28.7

overall

41.10.035.85.20.0N/A
165

GPT-4.1 mini

gpt-4.1-mini-2025-04-14

multimodalvisionmulti-input reasoning
OpenAI

28.5

overall

19.284.68.92.269.5
166

Mistral Large 3 (675B Instruct 2512)

mistral-large-latest

multimodalvisionmulti-input reasoning
Mistral AI

28.1

overall

21.922.00.00.055.6
167

Claude 3.5 Sonnet

claude-3-5-sonnet-20241022

multimodalvisionmulti-input reasoning
Anthropic

28.0

overall

32.00.038.711.10.0
168

Hermes 3 70B

hermes-3-70b

textinference
NNous Research

27.6

overall

27.60.00.00.00.0N/A
169

Llama 4 Scout

llama-4-scout

multimodalvisionmulti-input reasoning
MMeta

27.6

overall

27.60.00.00.00.0N/A
170

GPT-4o

gpt-4o-2024-05-13

multimodalvisionmulti-input reasoning
OpenAI

27.6

overall

20.538.00.00.030.7$2.5 in / $10 out
171

Qwen3 VL 8B Thinking

qwen3-vl-8b-thinking

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

27.4

overall

32.80.021.10.00.0N/A
172

MAI-Code-1-Flash

mai-code-1-flash

codeprogrammingtool use
MMicrosoft

27.4

overall

32.20.00.021.40.0N/A
173

Pixtral Large

pixtral-large

multimodalvisionmulti-input reasoning
Mistral AI

27.3

overall

27.30.00.00.00.0
174

Command A+

command-a-plus-05-2026

multimodalvisionmulti-input reasoning
Cohere

26.4

overall

35.20.00.015.30.0
175

Qwen3 VL 30B A3B Thinking

qwen3-vl-30b-a3b-thinking

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

26.4

overall

32.70.019.20.00.0
176

DeepSeek R1 Distill Llama 70B

deepseek-r1-distill-llama-70b

textinference
DeepSeek

26.4

overall

26.40.00.00.00.0N/A
177

QwQ-32B

qwq-32b

textinference
AAlibaba Cloud / Qwen Team

26.4

overall

26.40.00.00.00.0N/A
178

QwQ-32B-Preview

qwq-32b-preview

textinference
AAlibaba Cloud / Qwen Team

26.4

overall

26.40.00.00.00.0N/A
179

Nemotron 3 Super (120B A12B)

nemotron-3-super-120b-a12b

codeprogrammingtool use
NNVIDIA

26.2

overall

45.50.07.622.00.0N/A
180

Gemini 1.5 Pro

gemini-1.5-pro

multimodalvisionmulti-input reasoning
Google

26.2

overall

26.20.00.00.00.0N/A
161
A

Qwen3 235B A22B

Alibaba Cloud / Qwen Team

29.3

N/A

162

GPT-5.2 Codex

OpenAI

29.1

$1.75 in / $14 out

163
N

Nemotron 3 Nano (30B A3B)

NVIDIA

29.0

$0.06 in / $0.24 out

164

Page 9 of 17 · 334 models

PreviousNext

Want benchmark charts, model comparison, and pricing analytics?

Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.

Open full leaderboard

Rankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.

$0.4 in / $1.6 out
$0.5 in / $1.5 out
N/A
N/A
N/A
N/A

GPT-4.5

OpenAI

28.7

N/A

165

GPT-4.1 mini

OpenAI

28.5

$0.4 in / $1.6 out

166

Mistral Large 3 (675B Instruct 2512)

Mistral AI

28.1

$0.5 in / $1.5 out

167

Claude 3.5 Sonnet

Anthropic

28.0

N/A

168
N

Hermes 3 70B

Nous Research

27.6

N/A

169
M

Llama 4 Scout

Meta

27.6

N/A

170

GPT-4o

OpenAI

27.6

$2.5 in / $10 out

171
A

Qwen3 VL 8B Thinking

Alibaba Cloud / Qwen Team

27.4

N/A

172
M

MAI-Code-1-Flash

Microsoft

27.4

N/A

173

Pixtral Large

Mistral AI

27.3

N/A

174

Command A+

Cohere

26.4

N/A

175
A

Qwen3 VL 30B A3B Thinking

Alibaba Cloud / Qwen Team

26.4

N/A

176

DeepSeek R1 Distill Llama 70B

DeepSeek

26.4

N/A

177
A

QwQ-32B

Alibaba Cloud / Qwen Team

26.4

N/A

178
A

QwQ-32B-Preview

Alibaba Cloud / Qwen Team

26.4

N/A

179
N

Nemotron 3 Super (120B A12B)

NVIDIA

26.2

N/A

180

Gemini 1.5 Pro

Google

26.2

N/A