Skytells
HomeModelsCLIChangelog
  • Home
  • Models
  • CLI
  • Changelog
Skytells

Addressing the world's greatest challenges with AI. Enterprise research, foundation models, and infrastructure trusted by organizations worldwide since 2012.

Get Started

  • Console
  • Learn
  • Documentation
  • API Reference
  • Pricing
  • ModelsNew

Platform

  • Cloud AgentsNew
  • AI Solutions
  • Infrastructure
  • Edge Network
  • Trust Center
  • CLI

Resources

  • Blog
  • Changelog
  • AI Leaderboard
  • Research
  • Status

Company

  • About
  • Careers
  • Legal
  • Privacy Policy

© 2012–2026 Skytells, Inc. All rights reserved.

Live rankings

AI Model Leaderboard

Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.

Explore full leaderboardBrowse model catalog

334

Tracked models

29

Providers

286

Benchmarked

28.3

Avg. index

OverallBenchmarksInferenceAgenticProgrammingValue / Price

334 models

RankModelProviderScoreBenchmarksInferenceAgenticProgrammingValuePrice
21

GPT-5.1 High

gpt-5.1-high-2025-11-12

multimodalvisionmulti-input reasoning
OpenAI

67.1

Benchmarks

67.10.00.00.00.0N/A
22

Muse Spark

muse-spark

multimodalvisionmulti-input reasoning
MMeta

67.1

Benchmarks

67.10.064.436.20.0N/A
23

Qwen3.6 Plus

qwen3.6-plus

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

66.9

Benchmarks

66.962.833.652.854.9$0.5 in / $3 out
24

Seed 2.0 Pro

seed-2.0-pro

multimodalvisionmulti-input reasoning
BByteDance

66.8

Benchmarks

66.823.244.856.054.9$0.5 in / $3 out
25

Seed 2.1 Turbo

seed-2.1-turbo

multimodalvisionmulti-input reasoning
BByteDance

66.3

Benchmarks

66.30.063.153.80.0N/A
26

Kimi K2-Thinking-0905

kimi-k2-thinking-0905

codeprogrammingtool use
Moonshot AI

66.2

Benchmarks

66.20.050.858.00.0
27

Hy3

hy3

codeprogrammingtool use
TTencent

65.9

Benchmarks

65.90.040.459.20.0N/A
28

DeepSeek-V4-Pro-Max

deepseek-v4-pro-max

codeprogrammingtool use
DeepSeek

64.9

Benchmarks

64.984.851.052.045.1
29

Qwen3.7 Max

qwen3.7-max

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

64.7

Benchmarks

64.762.849.674.243.9$1.25 in / $3.75 out
30

GPT-5.2 Pro

gpt-5.2-pro-2025-12-11

multimodalvisionmulti-input reasoning
OpenAI

64.4

Benchmarks

64.40.046.10.00.0
31

GPT-5.1 Medium

gpt-5.1-medium-2025-11-12

multimodalvisionmulti-input reasoning
OpenAI

64.1

Benchmarks

64.144.90.00.032.0
32

GPT-5.6 Luna

gpt-5.6-luna

multimodalvisionmulti-input reasoning
OpenAI

63.9

Benchmarks

63.994.555.065.537.8
33

Kimi K2.5

kimi-k2.5

multimodalvisionmulti-input reasoning
Moonshot AI

63.7

Benchmarks

63.70.041.041.80.0N/A
34

Kimi K2.6

kimi-k2.6

texttext-to-textlanguage
Moonshot AI

63.5

Benchmarks

63.532.947.269.446.3
35

Qwen3.7-Plus

qwen3.7-plus

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

62.9

Benchmarks

62.962.848.161.169.5$0.32 in / $1.28 out
36

Step-3.5-Flash

step-3.5-flash

codeprogrammingtool use
SStepFun

62.8

Benchmarks

62.859.836.548.693.9$0.1 in / $0.4 out
37

GLM-5.1

glm-5.1

codeprogrammingtool use
ZZhipu AI

62.6

Benchmarks

62.615.935.247.240.2$1.4 in / $4.4 out
38

GPT-5.1

gpt-5.1-2025-11-13

multimodalvisionmulti-input reasoning
OpenAI

62.4

Benchmarks

62.462.80.051.437.3
39

GPT-5.1 Instant

gpt-5.1-instant-2025-11-12

multimodalvisionmulti-input reasoning
OpenAI

62.4

Benchmarks

62.462.80.051.437.3
40

GPT-5.1 Thinking

gpt-5.1-thinking-2025-11-12

multimodalvisionmulti-input reasoning
OpenAI

62.4

Benchmarks

62.40.00.051.40.0
21

GPT-5.1 High

OpenAI

67.1

N/A

22
M

Muse Spark

Meta

67.1

N/A

23
A

Qwen3.6 Plus

Alibaba Cloud / Qwen Team

66.9

$0.5 in / $3 out

24

Page 2 of 17 · 334 models

PreviousNext

Want benchmark charts, model comparison, and pricing analytics?

Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.

Open full leaderboard

Rankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.

N/A
$1.6 in / $3.2 out
N/A
$1.25 in / $10 out
$1 in / $6 out
$0.75 in / $3.5 out
$1.25 in / $10 out
$1.25 in / $10 out
N/A
B

Seed 2.0 Pro

ByteDance

66.8

$0.5 in / $3 out

25
B

Seed 2.1 Turbo

ByteDance

66.3

N/A

26

Kimi K2-Thinking-0905

Moonshot AI

66.2

N/A

27
T

Hy3

Tencent

65.9

N/A

28

DeepSeek-V4-Pro-Max

DeepSeek

64.9

$1.6 in / $3.2 out

29
A

Qwen3.7 Max

Alibaba Cloud / Qwen Team

64.7

$1.25 in / $3.75 out

30

GPT-5.2 Pro

OpenAI

64.4

N/A

31

GPT-5.1 Medium

OpenAI

64.1

$1.25 in / $10 out

32

GPT-5.6 Luna

OpenAI

63.9

$1 in / $6 out

33

Kimi K2.5

Moonshot AI

63.7

N/A

34

Kimi K2.6

Moonshot AI

63.5

$0.75 in / $3.5 out

35
A

Qwen3.7-Plus

Alibaba Cloud / Qwen Team

62.9

$0.32 in / $1.28 out

36
S

Step-3.5-Flash

StepFun

62.8

$0.1 in / $0.4 out

37
Z

GLM-5.1

Zhipu AI

62.6

$1.4 in / $4.4 out

38

GPT-5.1

OpenAI

62.4

$1.25 in / $10 out

39

GPT-5.1 Instant

OpenAI

62.4

$1.25 in / $10 out

40

GPT-5.1 Thinking

OpenAI

62.4

N/A