Skytells
HomeModelsCLIChangelog
  • Home
  • Models
  • CLI
  • Changelog
Skytells

Addressing the world's greatest challenges with AI. Enterprise research, foundation models, and infrastructure trusted by organizations worldwide since 2012.

Get Started

  • Console
  • Learn
  • Documentation
  • API Reference
  • Pricing
  • ModelsNew

Platform

  • Cloud AgentsNew
  • AI Solutions
  • Infrastructure
  • Edge Network
  • Trust Center
  • CLI

Resources

  • Blog
  • Changelog
  • AI Leaderboard
  • Research
  • Status

Company

  • About
  • Careers
  • Legal
  • Privacy Policy

© 2012–2026 Skytells, Inc. All rights reserved.

Live rankings

AI Model Leaderboard

Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.

Explore full leaderboardBrowse model catalog

334

Tracked models

29

Providers

286

Benchmarked

29.3

Avg. index

OverallBenchmarksInferenceAgenticProgrammingValue / Price

334 models

RankModelProviderScoreBenchmarksInferenceAgenticProgrammingValuePrice
101

Qwen3-Next-80B-A3B-Thinking

qwen3-next-80b-a3b-thinking

textinference
AAlibaba Cloud / Qwen Team

42.0

overall

42.30.041.70.00.0N/A
102

Qwen3 VL 235B A22B Instruct

qwen3-vl-235b-a22b-instruct

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

42.0

overall

34.30.051.00.00.0
103

Qwen3-Coder 480B A35B Instruct

qwen3-coder-480b-a35b-instruct

codeprogrammingtool use
AAlibaba Cloud / Qwen Team

41.8

overall

0.00.050.732.10.0
104

Qwen3.5-35B-A3B

qwen3.5-35b-a3b

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

41.6

overall

53.40.038.330.20.0N/A
105

GLM-4.6

glm-4.6

multimodalvisionmulti-input reasoning
ZZhipu AI

41.2

overall

44.50.035.443.20.0N/A
106

Kimi K2 0905

kimi-k2-0905

textinference
Moonshot AI

41.0

overall

41.00.00.00.00.0N/A
107

Qwen3-235B-A22B-Instruct-2507

qwen3-235b-a22b-instruct-2507

textinference
AAlibaba Cloud / Qwen Team

40.7

overall

40.70.00.00.00.0N/A
108

LongCat-Flash-Lite

longcat-flash-lite

codeprogrammingtool use
Meituan

40.1

overall

22.872.830.123.995.6$0.1 in / $0.4 out
109

Gemini 2.5 Pro Preview 06-05

gemini-2.5-pro-preview-06-05

multimodalvisionmulti-input reasoning
Google

39.4

overall

50.20.00.025.80.0
110

DeepSeek-V3.2-Exp

deepseek-v3.2-exp

codeprogrammingtool use
DeepSeek

39.1

overall

50.20.027.238.00.0N/A
111

LongCat-Flash-Thinking-2601

longcat-flash-thinking-2601

codeprogrammingtool use
Meituan

38.3

overall

52.90.025.633.50.0
112

Mistral Small 4

mistral-small-latest

multimodalvisionmulti-input reasoning
Mistral AI

38.2

overall

31.323.20.00.081.7
113

Grok Code Fast 1

grok-code-fast-1

codeprogrammingtool use
xAI

38.2

overall

0.027.20.036.160.5$0.2 in / $1.5 out
114

o4-mini

o4-mini

multimodalvisionmulti-input reasoning
OpenAI

37.7

overall

46.20.036.128.70.0N/A
115

QvQ-72B-Preview

qvq-72b-preview

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

37.6

overall

37.60.00.00.00.0N/A
116

Gemini 3.5 Flash-Lite

gemini-3.5-flash-lite

multimodalvisionmulti-input reasoning
Google

37.6

overall

18.984.80.017.259.1
117

MiMo-V2-Flash

mimo-v2-flash

codeprogrammingtool use
Xiaomi

37.4

overall

50.60.023.535.80.0N/A
118

DeepSeek-V3.2

deepseek-v3.2

codeprogrammingtool use
DeepSeek

37.3

overall

55.00.012.442.00.0N/A
119

GLM-5

glm-5

codeprogrammingtool use
ZZhipu AI

37.3

overall

0.06.935.160.440.0$1 in / $3.2 out
120

DeepSeek R1 Zero

deepseek-r1-zero

textinference
DeepSeek

36.8

overall

36.80.00.00.00.0N/A
101
A

Qwen3-Next-80B-A3B-Thinking

Alibaba Cloud / Qwen Team

42.0

N/A

102
A

Qwen3 VL 235B A22B Instruct

Alibaba Cloud / Qwen Team

42.0

N/A

103
A

Qwen3-Coder 480B A35B Instruct

Alibaba Cloud / Qwen Team

41.8

N/A

104

Page 6 of 17 · 334 models

PreviousNext

Want benchmark charts, model comparison, and pricing analytics?

Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.

Open full leaderboard

Rankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.

N/A
N/A
N/A
N/A
$0.15 in / $0.6 out
$0.3 in / $2.5 out
A

Qwen3.5-35B-A3B

Alibaba Cloud / Qwen Team

41.6

N/A

105
Z

GLM-4.6

Zhipu AI

41.2

N/A

106

Kimi K2 0905

Moonshot AI

41.0

N/A

107
A

Qwen3-235B-A22B-Instruct-2507

Alibaba Cloud / Qwen Team

40.7

N/A

108

LongCat-Flash-Lite

Meituan

40.1

$0.1 in / $0.4 out

109

Gemini 2.5 Pro Preview 06-05

Google

39.4

N/A

110

DeepSeek-V3.2-Exp

DeepSeek

39.1

N/A

111

LongCat-Flash-Thinking-2601

Meituan

38.3

N/A

112

Mistral Small 4

Mistral AI

38.2

$0.15 in / $0.6 out

113

Grok Code Fast 1

xAI

38.2

$0.2 in / $1.5 out

114

o4-mini

OpenAI

37.7

N/A

115
A

QvQ-72B-Preview

Alibaba Cloud / Qwen Team

37.6

N/A

116

Gemini 3.5 Flash-Lite

Google

37.6

$0.3 in / $2.5 out

117

MiMo-V2-Flash

Xiaomi

37.4

N/A

118

DeepSeek-V3.2

DeepSeek

37.3

N/A

119
Z

GLM-5

Zhipu AI

37.3

$1 in / $3.2 out

120

DeepSeek R1 Zero

DeepSeek

36.8

N/A