Skytells
HomeModelsCLIChangelog
  • Home
  • Models
  • CLI
  • Changelog
Skytells

Addressing the world's greatest challenges with AI. Enterprise research, foundation models, and infrastructure trusted by organizations worldwide since 2012.

Get Started

  • Console
  • Learn
  • Documentation
  • API Reference
  • Pricing
  • ModelsNew

Platform

  • Cloud AgentsNew
  • AI Solutions
  • Infrastructure
  • Edge Network
  • Trust Center
  • CLI

Resources

  • Blog
  • Changelog
  • AI Leaderboard
  • Research
  • Status

Company

  • About
  • Careers
  • Legal
  • Privacy Policy

© 2012–2026 Skytells, Inc. All rights reserved.

Live rankings

AI Model Leaderboard

Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.

Explore full leaderboardBrowse model catalog

334

Tracked models

29

Providers

286

Benchmarked

28.3

Avg. index

OverallBenchmarksInferenceAgenticProgrammingValue / Price

334 models

RankModelProviderScoreBenchmarksInferenceAgenticProgrammingValuePrice
1

Claude Mythos Preview

claude-mythos-preview

multimodalvisionmulti-input reasoning
Anthropic

80.0

Benchmarks

80.00.066.682.70.0N/A
2

GPT-5.6 Sol

gpt-5.6-sol

multimodalvisionmulti-input reasoning
OpenAI

79.8

Benchmarks

79.894.571.473.73.7
3

GPT-5.5

gpt-5.5

multimodalvisionmulti-input reasoning
OpenAI

76.7

Benchmarks

76.794.561.051.23.7
4

Kimi K3

kimi-k3

multimodalvisionmulti-input reasoning
Moonshot AI

76.1

Benchmarks

76.184.884.20.012.2$3 in / $15 out
5

Claude Opus 4.7

claude-opus-4-7

multimodalvisionmulti-input reasoning
Anthropic

75.1

Benchmarks

75.128.354.277.98.0
6

Claude Opus 4.6

claude-opus-4-6

multimodalvisionmulti-input reasoning
Anthropic

74.6

Benchmarks

74.628.349.871.58.0
7

GPT-5.6 Terra

gpt-5.6-terra

multimodalvisionmulti-input reasoning
OpenAI

74.6

Benchmarks

74.694.561.569.617.1
8

Claude Opus 4.8

claude-opus-4-8

multimodalvisionmulti-input reasoning
Anthropic

74.4

Benchmarks

74.428.374.581.38.0
9

Gemini 3.1 Pro

gemini-3.1-pro-preview

multimodalvisionmulti-input reasoning
Google

72.3

Benchmarks

72.358.352.660.922.7
10

GPT-5.2

gpt-5.2-2025-12-11

multimodalvisionmulti-input reasoning
OpenAI

70.9

Benchmarks

70.962.837.066.331.5
11

Claude Fable 5

claude-fable-5

multimodalvisionmulti-input reasoning
Anthropic

70.8

Benchmarks

70.862.80.084.20.0
12

GPT-5.4

gpt-5.4

texttext-to-textlanguage
OpenAI

70.5

Benchmarks

70.537.248.950.318.5
13

Gemini 3 Pro

gemini-3-pro-preview

multimodalvisionmulti-input reasoning
Google

70.3

Benchmarks

70.30.058.052.30.0
14

Muse Spark 1.1

muse-spark-1.1

multimodalvisionmulti-input reasoning
MMeta

69.7

Benchmarks

69.784.876.658.141.5$1.25 in / $4.25 out
15

Grok-4 Heavy

grok-4-heavy

multimodalvisionmulti-input reasoning
xAI

69.5

Benchmarks

69.50.00.00.00.0N/A
16

Grok 4.5

grok-4.5

multimodalvisionmulti-input reasoning
xAI

69.3

Benchmarks

69.338.20.070.835.6$2 in / $6 out
17

Seed 2.1 Pro

seed-2.1-pro

multimodalvisionmulti-input reasoning
BByteDance

69.2

Benchmarks

69.20.075.660.20.0N/A
18

GLM-5.2

glm-5.2

codeprogrammingtool use
ZZhipu AI

68.7

Benchmarks

68.784.844.159.951.2$0.95 in / $3 out
19

Gemini 3 Flash

gemini-3-flash-preview

multimodalvisionmulti-input reasoning
Google

68.5

Benchmarks

68.562.833.261.954.9
20

Claude Sonnet 5

claude-sonnet-5

multimodalvisionmulti-input reasoning
Anthropic

67.5

Benchmarks

67.528.360.775.412.0
1

Claude Mythos Preview

Anthropic

80.0

N/A

2

GPT-5.6 Sol

OpenAI

79.8

$5 in / $30 out

3

GPT-5.5

OpenAI

76.7

$5 in / $30 out

4

Page 1 of 17 · 334 models

Next

Want benchmark charts, model comparison, and pricing analytics?

Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.

Open full leaderboard

Rankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.

$5 in / $30 out
$5 in / $30 out
$5 in / $25 out
$5 in / $25 out
$2.5 in / $15 out
$5 in / $25 out
$2.5 in / $15 out
$1.75 in / $14 out
$10 in / $50 out
$2.5 in / $15 out
N/A
$0.5 in / $3 out
$3 in / $15 out

Kimi K3

Moonshot AI

76.1

$3 in / $15 out

5

Claude Opus 4.7

Anthropic

75.1

$5 in / $25 out

6

Claude Opus 4.6

Anthropic

74.6

$5 in / $25 out

7

GPT-5.6 Terra

OpenAI

74.6

$2.5 in / $15 out

8

Claude Opus 4.8

Anthropic

74.4

$5 in / $25 out

9

Gemini 3.1 Pro

Google

72.3

$2.5 in / $15 out

10

GPT-5.2

OpenAI

70.9

$1.75 in / $14 out

11

Claude Fable 5

Anthropic

70.8

$10 in / $50 out

12

GPT-5.4

OpenAI

70.5

$2.5 in / $15 out

13

Gemini 3 Pro

Google

70.3

N/A

14
M

Muse Spark 1.1

Meta

69.7

$1.25 in / $4.25 out

15

Grok-4 Heavy

xAI

69.5

N/A

16

Grok 4.5

xAI

69.3

$2 in / $6 out

17
B

Seed 2.1 Pro

ByteDance

69.2

N/A

18
Z

GLM-5.2

Zhipu AI

68.7

$0.95 in / $3 out

19

Gemini 3 Flash

Google

68.5

$0.5 in / $3 out

20

Claude Sonnet 5

Anthropic

67.5

$3 in / $15 out