Skytells
HomeModelsCLIChangelog
  • Home
  • Models
  • CLI
  • Changelog
Skytells

Addressing the world's greatest challenges with AI. Enterprise research, foundation models, and infrastructure trusted by organizations worldwide since 2012.

Get Started

  • Console
  • Learn
  • Documentation
  • API Reference
  • Pricing
  • ModelsNew

Platform

  • Cloud AgentsNew
  • AI Solutions
  • Infrastructure
  • Edge Network
  • Trust Center
  • CLI

Resources

  • Blog
  • Changelog
  • AI Leaderboard
  • Research
  • Status

Company

  • About
  • Careers
  • Legal
  • Privacy Policy

© 2012–2026 Skytells, Inc. All rights reserved.

Live rankings

AI Model Leaderboard

Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.

Explore full leaderboardBrowse model catalog

334

Tracked models

29

Providers

286

Benchmarked

29.3

Avg. index

OverallBenchmarksInferenceAgenticProgrammingValue / Price

334 models

RankModelProviderScoreBenchmarksInferenceAgenticProgrammingValuePrice
121

Nemotron 3 Ultra (550B A55B)

nemotron-3-ultra-550b-a55b

codeprogrammingtool use
NNVIDIA

36.5

overall

54.70.011.540.60.0N/A
122

Gemini 2.5 Pro

gemini-2.5-pro

multimodalvisionmulti-input reasoning
Google

36.5

overall

42.551.00.021.429.8
123

Qwen3 VL 32B Thinking

qwen3-vl-32b-thinking

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

36.5

overall

41.20.031.10.00.0
124

Qwen3-235B-A22B-Thinking-2507

qwen3-235b-a22b-thinking-2507

textinference
AAlibaba Cloud / Qwen Team

36.3

overall

44.40.026.80.00.0N/A
125

Nova 2 Lite

nova-2-lite

multimodalvisionmulti-input reasoning
AAmazon

36.3

overall

41.162.813.026.059.1$0.3 in / $2.5 out
126

Qwen3.5-9B

qwen3.5-9b

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

36.3

overall

36.30.00.00.00.0N/A
127

LongCat-Flash-Chat

longcat-flash-chat

codeprogrammingtool use
Meituan

36.3

overall

26.00.048.136.60.0N/A
128

Nova 2 Omni

nova-2-omni

multimodalvisionmulti-input reasoning
AAmazon

36.1

overall

36.10.00.00.00.0N/A
129

Grok 4 Fast

grok-4-fast

multimodalvisionmulti-input reasoning
xAI

35.9

overall

55.70.012.80.00.0N/A
130

Qwen3 VL 235B A22B Thinking

qwen3-vl-235b-a22b-thinking

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

35.7

overall

35.20.036.20.00.0
131

DeepSeek-V3.2 (Thinking)

deepseek-reasoner

codeprogrammingtool use
DeepSeek

35.4

overall

49.80.012.442.00.0N/A
132

Ministral 3 (14B Reasoning 2512)

ministral-14b-latest

multimodalvisionmulti-input reasoning
Mistral AI

35.4

overall

35.40.00.00.00.0
133

LongCat-Flash-Thinking

longcat-flash-thinking

codeprogrammingtool use
Meituan

35.1

overall

48.20.00.018.40.0
134

Kimi-k1.5

kimi-k1.5

multimodalvisionmulti-input reasoning
Moonshot AI

34.7

overall

34.70.00.00.00.0N/A
135

Qwen3 30B A3B

qwen3-30b-a3b

textinference
AAlibaba Cloud / Qwen Team

34.7

overall

23.726.60.00.078.5$0.1 in / $0.44 out
136

GPT-4.1

gpt-4.1-2025-04-14

multimodalvisionmulti-input reasoning
OpenAI

34.5

overall

27.273.232.814.740.7
137

GLM-4.5

glm-4.5

codeprogrammingtool use
ZZhipu AI

34.4

overall

31.50.036.036.10.0N/A
138

GPT-4.1 nano

gpt-4.1-nano-2025-04-14

multimodalvisionmulti-input reasoning
OpenAI

34.3

overall

11.687.80.00.094.9
139

GPT-5.4 Mini

gpt-5.4-mini

texttext-to-textlanguage
OpenAI

33.7

overall

51.144.515.020.042.7
140

Mistral Medium 3.5

mistral-medium-3-5

multimodalvisionmulti-input reasoning
Mistral AI

33.5

overall

34.623.215.459.034.8
121
N

Nemotron 3 Ultra (550B A55B)

NVIDIA

36.5

N/A

122

Gemini 2.5 Pro

Google

36.5

$1.25 in / $10 out

123
A

Qwen3 VL 32B Thinking

Alibaba Cloud / Qwen Team

36.5

N/A

124

Page 7 of 17 · 334 models

PreviousNext

Want benchmark charts, model comparison, and pricing analytics?

Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.

Open full leaderboard

Rankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.

$1.25 in / $10 out
N/A
N/A
N/A
N/A
$2 in / $8 out
$0.1 in / $0.4 out
$0.75 in / $4.5 out
$1.5 in / $7.5 out
A

Qwen3-235B-A22B-Thinking-2507

Alibaba Cloud / Qwen Team

36.3

N/A

125
A

Nova 2 Lite

Amazon

36.3

$0.3 in / $2.5 out

126
A

Qwen3.5-9B

Alibaba Cloud / Qwen Team

36.3

N/A

127

LongCat-Flash-Chat

Meituan

36.3

N/A

128
A

Nova 2 Omni

Amazon

36.1

N/A

129

Grok 4 Fast

xAI

35.9

N/A

130
A

Qwen3 VL 235B A22B Thinking

Alibaba Cloud / Qwen Team

35.7

N/A

131

DeepSeek-V3.2 (Thinking)

DeepSeek

35.4

N/A

132

Ministral 3 (14B Reasoning 2512)

Mistral AI

35.4

N/A

133

LongCat-Flash-Thinking

Meituan

35.1

N/A

134

Kimi-k1.5

Moonshot AI

34.7

N/A

135
A

Qwen3 30B A3B

Alibaba Cloud / Qwen Team

34.7

$0.1 in / $0.44 out

136

GPT-4.1

OpenAI

34.5

$2 in / $8 out

137
Z

GLM-4.5

Zhipu AI

34.4

N/A

138

GPT-4.1 nano

OpenAI

34.3

$0.1 in / $0.4 out

139

GPT-5.4 Mini

OpenAI

33.7

$0.75 in / $4.5 out

140

Mistral Medium 3.5

Mistral AI

33.5

$1.5 in / $7.5 out