Skytells
HomeModelsCLIChangelog
  • Home
  • Models
  • CLI
  • Changelog
Skytells

Addressing the world's greatest challenges with AI. Enterprise research, foundation models, and infrastructure trusted by organizations worldwide since 2012.

Get Started

  • Console
  • Learn
  • Documentation
  • API Reference
  • Pricing
  • ModelsNew

Platform

  • Cloud AgentsNew
  • AI Solutions
  • Infrastructure
  • Edge Network
  • Trust Center
  • CLI

Resources

  • Blog
  • Changelog
  • AI Leaderboard
  • Research
  • Status

Company

  • About
  • Careers
  • Legal
  • Privacy Policy

© 2012–2026 Skytells, Inc. All rights reserved.

Live rankings

AI Model Leaderboard

Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.

Explore full leaderboardBrowse model catalog

334

Tracked models

29

Providers

286

Benchmarked

15.1

Avg. index

OverallBenchmarksInferenceAgenticProgrammingValue / Price

334 models

RankModelProviderScoreBenchmarksInferenceAgenticProgrammingValuePrice
41

GPT-5.1 Thinking

gpt-5.1-thinking-2025-11-12

multimodalvisionmulti-input reasoning
OpenAI

51.4

Programming

62.40.00.051.40.0N/A
42

GPT-5.5

gpt-5.5

multimodalvisionmulti-input reasoning
OpenAI

51.2

Programming

76.794.561.051.23.7$5 in / $30 out
43

MiMo-V2-Omni

mimo-v2-omni

multimodalvisionmulti-input reasoning
Xiaomi

50.9

Programming

0.00.00.050.90.0N/A
44

MiniMax M2.5

minimax-m2.5

codeprogrammingtool use
MiniMax

50.4

Programming

0.068.943.650.472.9$0.3 in / $1.2 out
45

GPT-5.4

gpt-5.4

texttext-to-textlanguage
OpenAI

50.3

Programming

70.537.248.950.318.5
46

GPT-5 Codex

gpt-5-codex-2025-09-15

codeprogrammingtool use
OpenAI

49.8

Programming

0.00.00.049.80.0N/A
47

Nova 2 Pro

nova-2-pro

multimodalvisionmulti-input reasoning
AAmazon

49.6

Programming

45.30.057.249.60.0N/A
48

GPT-5.5 Pro

gpt-5.5-pro

multimodalvisionmulti-input reasoning
OpenAI

48.8

Programming

61.10.069.248.80.0N/A
49

Step-3.5-Flash

step-3.5-flash

codeprogrammingtool use
SStepFun

48.6

Programming

62.859.836.548.693.9$0.1 in / $0.4 out
50

MiniMax M2.1

minimax-m2.1

codeprogrammingtool use
MiniMax

47.6

Programming

39.168.945.747.672.9$0.3 in / $1.2 out
51

GLM-5.1

glm-5.1

codeprogrammingtool use
ZZhipu AI

47.2

Programming

62.615.935.247.240.2$1.4 in / $4.4 out
52

GPT-5.1 Codex

gpt-5.1-codex

multimodalvisionmulti-input reasoning
OpenAI

47.2

Programming

0.00.00.047.20.0
53

GPT-5

gpt-5-2025-08-07

multimodalvisionmulti-input reasoning
OpenAI

47.1

Programming

59.90.024.347.10.0
54

Claude Opus 4

claude-opus-4-20250514

multimodalvisionmulti-input reasoning
Anthropic

47.0

Programming

35.80.057.447.00.0
55

Seed 2.0 Lite

seed-2.0-lite

multimodalvisionmulti-input reasoning
BByteDance

46.1

Programming

56.50.00.046.10.0N/A
56

GLM-4.6

glm-4.6

multimodalvisionmulti-input reasoning
ZZhipu AI

43.2

Programming

44.50.035.443.20.0N/A
57

Claude Sonnet 4

claude-sonnet-4-20250514

multimodalvisionmulti-input reasoning
Anthropic

42.8

Programming

39.30.049.442.80.0
58

GLM-4.7

glm-4.7

multimodalvisionmulti-input reasoning
ZZhipu AI

42.5

Programming

60.30.026.642.50.0N/A
59

DeepSeek-V3.2 (Thinking)

deepseek-reasoner

codeprogrammingtool use
DeepSeek

42.0

Programming

49.80.012.442.00.0
60

DeepSeek-V3.2

deepseek-v3.2

codeprogrammingtool use
DeepSeek

42.0

Programming

55.00.012.442.00.0N/A
41

GPT-5.1 Thinking

OpenAI

51.4

N/A

42

GPT-5.5

OpenAI

51.2

$5 in / $30 out

43

MiMo-V2-Omni

Xiaomi

50.9

N/A

44

Page 3 of 17 · 334 models

PreviousNext

Want benchmark charts, model comparison, and pricing analytics?

Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.

Open full leaderboard

Rankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.

$2.5 in / $15 out
N/A
N/A
N/A
N/A
N/A

MiniMax M2.5

MiniMax

50.4

$0.3 in / $1.2 out

45

GPT-5.4

OpenAI

50.3

$2.5 in / $15 out

46

GPT-5 Codex

OpenAI

49.8

N/A

47
A

Nova 2 Pro

Amazon

49.6

N/A

48

GPT-5.5 Pro

OpenAI

48.8

N/A

49
S

Step-3.5-Flash

StepFun

48.6

$0.1 in / $0.4 out

50

MiniMax M2.1

MiniMax

47.6

$0.3 in / $1.2 out

51
Z

GLM-5.1

Zhipu AI

47.2

$1.4 in / $4.4 out

52

GPT-5.1 Codex

OpenAI

47.2

N/A

53

GPT-5

OpenAI

47.1

N/A

54

Claude Opus 4

Anthropic

47.0

N/A

55
B

Seed 2.0 Lite

ByteDance

46.1

N/A

56
Z

GLM-4.6

Zhipu AI

43.2

N/A

57

Claude Sonnet 4

Anthropic

42.8

N/A

58
Z

GLM-4.7

Zhipu AI

42.5

N/A

59

DeepSeek-V3.2 (Thinking)

DeepSeek

42.0

N/A

60

DeepSeek-V3.2

DeepSeek

42.0

N/A