Skytells
HomeModelsCLIChangelog
  • Home
  • Models
  • CLI
  • Changelog
Skytells

Addressing the world's greatest challenges with AI. Enterprise research, foundation models, and infrastructure trusted by organizations worldwide since 2012.

Get Started

  • Console
  • Learn
  • Documentation
  • API Reference
  • Pricing
  • ModelsNew

Platform

  • Cloud AgentsNew
  • AI Solutions
  • Infrastructure
  • Edge Network
  • Trust Center
  • CLI

Resources

  • Blog
  • Changelog
  • AI Leaderboard
  • Research
  • Status

Company

  • About
  • Careers
  • Legal
  • Privacy Policy

© 2012–2026 Skytells, Inc. All rights reserved.

Live rankings

AI Model Leaderboard

Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.

Explore full leaderboardBrowse model catalog

334

Tracked models

29

Providers

286

Benchmarked

28.3

Avg. index

OverallBenchmarksInferenceAgenticProgrammingValue / Price

334 models

RankModelProviderScoreBenchmarksInferenceAgenticProgrammingValuePrice
161

Kimi K2 Base

kimi-k2-base

textinference
Moonshot AI

26.0

Benchmarks

26.00.00.00.00.0N/A
162

LongCat-Flash-Chat

longcat-flash-chat

codeprogrammingtool use
Meituan

26.0

Benchmarks

26.00.048.136.60.0N/A
163

MiniCPM-SALA

minicpm-sala

textinference
OOpenBMB

25.8

Benchmarks

25.80.00.00.00.0N/A
164

GLM-4.5-Air

glm-4.5-air

codeprogrammingtool use
ZZhipu AI

25.7

Benchmarks

25.70.024.216.00.0N/A
165

Grok-2

grok-2

multimodalvisionmulti-input reasoning
xAI

25.7

Benchmarks

25.70.00.00.00.0N/A
166

Qwen3 VL 30B A3B Instruct

qwen3-vl-30b-a3b-instruct

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

25.6

Benchmarks

25.60.021.40.00.0
167

o3-mini

o3-mini

codeprogrammingtool use
OpenAI

25.3

Benchmarks

25.30.011.911.60.0N/A
168

GPT-5 nano

gpt-5-nano-2025-08-07

multimodalvisionmulti-input reasoning
OpenAI

25.2

Benchmarks

25.20.00.013.30.0
169

DeepSeek R1 Distill Qwen 32B

deepseek-r1-distill-qwen-32b

textinference
DeepSeek

24.4

Benchmarks

24.40.00.00.00.0N/A
170

Gemini 2.0 Flash-Lite

gemini-2.0-flash-lite

multimodalvisionmulti-input reasoning
Google

24.4

Benchmarks

24.40.00.00.00.0
171

GPT OSS 20B

gpt-oss-20b

textinference
OpenAI

23.7

Benchmarks

23.70.06.00.00.0N/A
172

Qwen3 30B A3B

qwen3-30b-a3b

textinference
AAlibaba Cloud / Qwen Team

23.7

Benchmarks

23.726.60.00.078.5$0.1 in / $0.44 out
173

o1-mini

o1-mini

textinference
OpenAI

23.6

Benchmarks

23.60.00.00.00.0N/A
174

Claude 3.5 Sonnet

claude-3-5-sonnet-20240620

multimodalvisionmulti-input reasoning
Anthropic

23.3

Benchmarks

23.30.00.00.00.0
175

Kimi K2 Instruct

kimi-k2-instruct

codeprogrammingtool use
Moonshot AI

23.2

Benchmarks

23.20.013.514.00.0N/A
176

Kimi K2-Instruct-0905

kimi-k2-instruct-0905

codeprogrammingtool use
Moonshot AI

23.2

Benchmarks

23.20.06.017.10.0
177

Nemotron Nano 9B v2

nvidia-nemotron-nano-9b-v2

textinference
NNVIDIA

23.1

Benchmarks

23.10.00.00.00.0N/A
178

DiffusionGemma 26B-A4B

diffusiongemma-26b-a4b-it

multimodalvisionmulti-input reasoning
Google

22.9

Benchmarks

22.90.00.00.00.0
179

ERNIE 4.5

ernie-4.5

textinference
BBaidu

22.8

Benchmarks

22.80.00.00.00.0N/A
180

Grok-2 mini

grok-2-mini

multimodalvisionmulti-input reasoning
xAI

22.8

Benchmarks

22.80.00.00.00.0N/A
161

Kimi K2 Base

Moonshot AI

26.0

N/A

162

LongCat-Flash-Chat

Meituan

26.0

N/A

163
O

MiniCPM-SALA

OpenBMB

25.8

N/A

164

Page 9 of 17 · 334 models

PreviousNext

Want benchmark charts, model comparison, and pricing analytics?

Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.

Open full leaderboard

Rankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.

N/A
N/A
N/A
N/A
N/A
N/A
Z

GLM-4.5-Air

Zhipu AI

25.7

N/A

165

Grok-2

xAI

25.7

N/A

166
A

Qwen3 VL 30B A3B Instruct

Alibaba Cloud / Qwen Team

25.6

N/A

167

o3-mini

OpenAI

25.3

N/A

168

GPT-5 nano

OpenAI

25.2

N/A

169

DeepSeek R1 Distill Qwen 32B

DeepSeek

24.4

N/A

170

Gemini 2.0 Flash-Lite

Google

24.4

N/A

171

GPT OSS 20B

OpenAI

23.7

N/A

172
A

Qwen3 30B A3B

Alibaba Cloud / Qwen Team

23.7

$0.1 in / $0.44 out

173

o1-mini

OpenAI

23.6

N/A

174

Claude 3.5 Sonnet

Anthropic

23.3

N/A

175

Kimi K2 Instruct

Moonshot AI

23.2

N/A

176

Kimi K2-Instruct-0905

Moonshot AI

23.2

N/A

177
N

Nemotron Nano 9B v2

NVIDIA

23.1

N/A

178

DiffusionGemma 26B-A4B

Google

22.9

N/A

179
B

ERNIE 4.5

Baidu

22.8

N/A

180

Grok-2 mini

xAI

22.8

N/A