Skytells
  • Home
  • Changelog
Skytells

Addressing the world's greatest challenges with AI. Enterprise research, foundation models, and infrastructure trusted by organizations worldwide since 2012.

Get Started

  • Console
  • Learn
  • Pricing
  • ModelsNew
  • EveNew
  • CLI

Platform

  • DropsNew
  • Cognition
  • Cloud AgentsNew
  • AI Solutions
  • Infrastructure
  • Edge Network

Resources

  • Documentation
  • Blog
  • Changelog
  • AI Leaderboard
  • Research
  • Status

Company

  • About
  • Careers
  • Trust Center
  • Legal
  • Privacy Policy

© 2012–2026 Skytells, Inc. All rights reserved.

Live rankings

AI Model Leaderboard

Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.

Explore full leaderboardBrowse model catalog

373

Tracked models

34

Providers

310

Benchmarked

30.2

Avg. index

OverallBenchmarksInferenceAgenticProgrammingValue / Price

373 models

RankModelProviderScoreBenchmarksInferenceAgenticProgrammingValuePrice
1

Muse Spark 1.2

muse-spark-1.2

multimodalvisionmulti-input reasoning
MMeta

87.0

overall

0.081.90.00.095.1$0.1 in / $0.2 out
2

DeepSeek-V4-Flash-Vision-Exp

deepseek-v4-flash-vision-exp

multimodalvisionmulti-input reasoning
DeepSeek

79.5

overall

0.081.90.00.075.7
3

Muse Spark 1.3

muse-spark-1.3

multimodalvisionmulti-input reasoning
MMeta

78.4

overall

70.481.90.00.095.1$0.1 in / $0.2 out
4

Claude Mythos Preview

claude-mythos-preview

multimodalvisionmulti-input reasoning
Anthropic

75.8

overall

79.50.064.883.00.0
5

DeepSeek-V4-Flash-0731

deepseek-v4-flash-0731

textinference
DeepSeek

73.0

overall

0.081.956.30.099.0$0.09 in / $0.18 out
6

GLM-5.3-Flash

glm-5.3-flash

multimodalvisionmulti-input reasoning
ZZhipu AI

72.7

overall

64.081.972.60.082.5$0.15 in / $0.5 out
7

DeepSeek-V4-Pro-0813

deepseek-v4-pro-0813

textinference
DeepSeek

71.2

overall

67.481.968.10.072.3$0.435 in / $0.87 out
8

Kimi K3

kimi-k3

multimodalvisionmulti-input reasoning
Moonshot AI

69.6

overall

75.381.978.20.013.1$3 in / $15 out
9

GPT-6 Astra

gpt-6-astra

multimodalvisionmulti-input reasoning
OpenAI

69.6

overall

74.895.274.60.01.9$10 in / $50 out
10

Hy4 preview

hy4-preview

codeprogrammingtool use
TTencent

69.2

overall

67.20.070.370.60.0N/A
11

GPT-5.6 Sol

gpt-5.6-sol

multimodalvisionmulti-input reasoning
OpenAI

68.7

overall

77.695.263.172.75.8$5 in / $30 out
12

Grok-4 Heavy

grok-4-heavy

multimodalvisionmulti-input reasoning
xAI

68.1

overall

68.10.00.00.00.0N/A
13

Qwen3.8 Max

qwen3.8-max

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

67.2

overall

70.381.962.772.035.0$1.65 in / $4.951 out
14

Grok-4.1 Fast Non-Reasoning

grok-4-1-fast-non-reasoning

multimodalvisionmulti-input reasoning
xAI

66.9

overall

0.063.20.00.072.7
15

Grok-4.1 Fast Reasoning

grok-4-1-fast-reasoning

multimodalvisionmulti-input reasoning
xAI

66.9

overall

0.063.20.00.072.7
16

Grok-4 Fast Reasoning

grok-4-fast-reasoning

multimodalvisionmulti-input reasoning
xAI

66.9

overall

0.063.20.00.072.7
17

GPT-5.1 High

gpt-5.1-high-2025-11-12

multimodalvisionmulti-input reasoning
OpenAI

66.2

overall

66.20.00.00.00.0
18

Gemini 3.7 Flash

gemini-3.7-flash

multimodalvisionmulti-input reasoning
Google

66.1

overall

65.381.90.00.043.2
19

GPT-5.6 Terra

gpt-5.6-terra

multimodalvisionmulti-input reasoning
OpenAI

66.0

overall

72.595.254.669.321.4
20

GPT-5.6 Luna

gpt-5.6-luna

multimodalvisionmulti-input reasoning
OpenAI

66.0

overall

62.795.248.565.870.4$0.2 in / $1.2 out
1
M

Muse Spark 1.2

Meta

87.0

$0.1 in / $0.2 out

2

DeepSeek-V4-Flash-Vision-Exp

DeepSeek

79.5

$0.22 in / $0.66 out

3
M

Muse Spark 1.3

Meta

78.4

$0.1 in / $0.2 out

4

Page 1 of 19 · 373 models

Next

Want benchmark charts, model comparison, and pricing analytics?

Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.

Open full leaderboard

Rankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.

$0.22 in / $0.66 out
N/A
$0.2 in / $0.5 out
$0.2 in / $0.5 out
$0.2 in / $0.5 out
N/A
$0.75 in / $3.75 out
$2 in / $12 out

Claude Mythos Preview

Anthropic

75.8

N/A

5

DeepSeek-V4-Flash-0731

DeepSeek

73.0

$0.09 in / $0.18 out

6
Z

GLM-5.3-Flash

Zhipu AI

72.7

$0.15 in / $0.5 out

7

DeepSeek-V4-Pro-0813

DeepSeek

71.2

$0.435 in / $0.87 out

8

Kimi K3

Moonshot AI

69.6

$3 in / $15 out

9

GPT-6 Astra

OpenAI

69.6

$10 in / $50 out

10
T

Hy4 preview

Tencent

69.2

N/A

11

GPT-5.6 Sol

OpenAI

68.7

$5 in / $30 out

12

Grok-4 Heavy

xAI

68.1

N/A

13
A

Qwen3.8 Max

Alibaba Cloud / Qwen Team

67.2

$1.65 in / $4.951 out

14

Grok-4.1 Fast Non-Reasoning

xAI

66.9

$0.2 in / $0.5 out

15

Grok-4.1 Fast Reasoning

xAI

66.9

$0.2 in / $0.5 out

16

Grok-4 Fast Reasoning

xAI

66.9

$0.2 in / $0.5 out

17

GPT-5.1 High

OpenAI

66.2

N/A

18

Gemini 3.7 Flash

Google

66.1

$0.75 in / $3.75 out

19

GPT-5.6 Terra

OpenAI

66.0

$2 in / $12 out

20

GPT-5.6 Luna

OpenAI

66.0

$0.2 in / $1.2 out