Skytells
HomeModelsCLIChangelog
  • Home
  • Models
  • CLI
  • Changelog
Skytells

Addressing the world's greatest challenges with AI. Enterprise research, foundation models, and infrastructure trusted by organizations worldwide since 2012.

Get Started

  • Console
  • Learn
  • Documentation
  • API Reference
  • Pricing
  • ModelsNew

Platform

  • Cloud AgentsNew
  • AI Solutions
  • Infrastructure
  • Edge Network
  • Trust Center
  • CLI

Resources

  • Blog
  • Changelog
  • AI Leaderboard
  • Research
  • Status

Company

  • About
  • Careers
  • Legal
  • Privacy Policy

© 2012–2026 Skytells, Inc. All rights reserved.

Live rankings

AI Model Leaderboard

Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.

Explore full leaderboardBrowse model catalog

334

Tracked models

29

Providers

286

Benchmarked

29.3

Avg. index

OverallBenchmarksInferenceAgenticProgrammingValue / Price

334 models

RankModelProviderScoreBenchmarksInferenceAgenticProgrammingValuePrice
1

Claude Mythos Preview

claude-mythos-preview

multimodalvisionmulti-input reasoning
Anthropic

76.5

overall

80.00.066.682.70.0N/A
2

Kimi K3

kimi-k3

multimodalvisionmulti-input reasoning
Moonshot AI

72.2

overall

76.184.884.20.012.2
3

GPT-5.6 Sol

gpt-5.6-sol

multimodalvisionmulti-input reasoning
OpenAI

71.2

overall

79.894.571.473.73.7
4

Grok-4 Heavy

grok-4-heavy

multimodalvisionmulti-input reasoning
xAI

69.5

overall

69.50.00.00.00.0N/A
5

Seed 2.1 Pro

seed-2.1-pro

multimodalvisionmulti-input reasoning
BByteDance

68.6

overall

69.20.075.660.20.0N/A
6

Grok-4.1 Fast Non-Reasoning

grok-4-1-fast-non-reasoning

multimodalvisionmulti-input reasoning
xAI

68.5

overall

0.063.00.00.077.3
7

Grok-4.1 Fast Reasoning

grok-4-1-fast-reasoning

multimodalvisionmulti-input reasoning
xAI

68.5

overall

0.063.00.00.077.3
8

Grok-4 Fast Reasoning

grok-4-fast-reasoning

multimodalvisionmulti-input reasoning
xAI

68.5

overall

0.063.00.00.077.3
9

Muse Spark 1.1

muse-spark-1.1

multimodalvisionmulti-input reasoning
MMeta

68.4

overall

69.784.876.658.141.5$1.25 in / $4.25 out
10

GPT-5.6 Terra

gpt-5.6-terra

multimodalvisionmulti-input reasoning
OpenAI

67.8

overall

74.694.561.569.617.1
11

GPT-5.1 High

gpt-5.1-high-2025-11-12

multimodalvisionmulti-input reasoning
OpenAI

67.1

overall

67.10.00.00.00.0
12

GPT-5.6 Luna

gpt-5.6-luna

multimodalvisionmulti-input reasoning
OpenAI

64.4

overall

63.994.555.065.537.8$1 in / $6 out
13

DeepSeek-V3.2 (Non-thinking)

deepseek-chat

textinference
DeepSeek

63.8

overall

0.052.00.00.082.7$0.28 in / $0.42 out
14

Claude Fable 5

claude-fable-5

multimodalvisionmulti-input reasoning
Anthropic

63.7

overall

70.862.80.084.20.0
15

GPT-5.5

gpt-5.5

multimodalvisionmulti-input reasoning
OpenAI

62.9

overall

76.794.561.051.23.7$5 in / $30 out
16

Claude Opus 4.8

claude-opus-4-8

multimodalvisionmulti-input reasoning
Anthropic

61.9

overall

74.428.374.581.38.0
17

MiMo-V2-Pro

mimo-v2-pro

codeprogrammingtool use
Xiaomi

61.9

overall

0.00.00.061.90.0N/A
18

GLM-5.2

glm-5.2

codeprogrammingtool use
ZZhipu AI

61.7

overall

68.784.844.159.951.2$0.95 in / $3 out
19

GPT-5 High

gpt-5-high-2025-08-07

multimodalvisionmulti-input reasoning
OpenAI

61.6

overall

61.60.00.00.00.0
20

Seed 2.1 Turbo

seed-2.1-turbo

multimodalvisionmulti-input reasoning
BByteDance

61.5

overall

66.30.063.153.80.0N/A
1

Claude Mythos Preview

Anthropic

76.5

N/A

2

Kimi K3

Moonshot AI

72.2

$3 in / $15 out

3

GPT-5.6 Sol

OpenAI

71.2

$5 in / $30 out

Page 1 of 17 · 334 models

Next

Want benchmark charts, model comparison, and pricing analytics?

Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.

Open full leaderboard

Rankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.

$3 in / $15 out
$5 in / $30 out
$0.2 in / $0.5 out
$0.2 in / $0.5 out
$0.2 in / $0.5 out
$2.5 in / $15 out
N/A
$10 in / $50 out
$5 in / $25 out
N/A
4

Grok-4 Heavy

xAI

69.5

N/A

5
B

Seed 2.1 Pro

ByteDance

68.6

N/A

6

Grok-4.1 Fast Non-Reasoning

xAI

68.5

$0.2 in / $0.5 out

7

Grok-4.1 Fast Reasoning

xAI

68.5

$0.2 in / $0.5 out

8

Grok-4 Fast Reasoning

xAI

68.5

$0.2 in / $0.5 out

9
M

Muse Spark 1.1

Meta

68.4

$1.25 in / $4.25 out

10

GPT-5.6 Terra

OpenAI

67.8

$2.5 in / $15 out

11

GPT-5.1 High

OpenAI

67.1

N/A

12

GPT-5.6 Luna

OpenAI

64.4

$1 in / $6 out

13

DeepSeek-V3.2 (Non-thinking)

DeepSeek

63.8

$0.28 in / $0.42 out

14

Claude Fable 5

Anthropic

63.7

$10 in / $50 out

15

GPT-5.5

OpenAI

62.9

$5 in / $30 out

16

Claude Opus 4.8

Anthropic

61.9

$5 in / $25 out

17

MiMo-V2-Pro

Xiaomi

61.9

N/A

18
Z

GLM-5.2

Zhipu AI

61.7

$0.95 in / $3 out

19

GPT-5 High

OpenAI

61.6

N/A

20
B

Seed 2.1 Turbo

ByteDance

61.5

N/A