Skytells
HomeModelsCLIChangelog
  • Home
  • Models
  • CLI
  • Changelog
Skytells

Addressing the world's greatest challenges with AI. Enterprise research, foundation models, and infrastructure trusted by organizations worldwide since 2012.

Get Started

  • Console
  • Learn
  • Documentation
  • API Reference
  • Pricing
  • ModelsNew

Platform

  • Cloud AgentsNew
  • AI Solutions
  • Infrastructure
  • Edge Network
  • Trust Center
  • CLI

Resources

  • Blog
  • Changelog
  • AI Leaderboard
  • Research
  • Status

Company

  • About
  • Careers
  • Legal
  • Privacy Policy

© 2012–2026 Skytells, Inc. All rights reserved.

Live rankings

AI Model Leaderboard

Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.

Explore full leaderboardBrowse model catalog

334

Tracked models

29

Providers

286

Benchmarked

28.3

Avg. index

OverallBenchmarksInferenceAgenticProgrammingValue / Price

334 models

RankModelProviderScoreBenchmarksInferenceAgenticProgrammingValuePrice
41

Claude Sonnet 4.6

claude-sonnet-4-6

multimodalvisionmulti-input reasoning
Anthropic

62.3

Benchmarks

62.312.641.065.612.0$3 in / $15 out
42

GPT-5 High

gpt-5-high-2025-08-07

multimodalvisionmulti-input reasoning
OpenAI

61.6

Benchmarks

61.60.00.00.00.0
43

GPT-5.5 Pro

gpt-5.5-pro

multimodalvisionmulti-input reasoning
OpenAI

61.1

Benchmarks

61.10.069.248.80.0N/A
44

GPT-5.1 Codex High

gpt-5.1-codex-high

multimodalvisionmulti-input reasoning
OpenAI

61.0

Benchmarks

61.00.00.00.00.0
45

Qwen3.5-122B-A10B

qwen3.5-122b-a10b

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

60.5

Benchmarks

60.50.044.638.30.0N/A
46

GLM-4.7

glm-4.7

multimodalvisionmulti-input reasoning
ZZhipu AI

60.3

Benchmarks

60.30.026.642.50.0N/A
47

Gemini 3.5 Flash

gemini-3.5-flash

multimodalvisionmulti-input reasoning
Google

60.2

Benchmarks

60.284.867.321.831.7
48

MAI-Thinking-1

mai-thinking-1

codeprogrammingtool use
MMicrosoft

60.1

Benchmarks

60.10.00.032.20.0N/A
49

GPT-5

gpt-5-2025-08-07

multimodalvisionmulti-input reasoning
OpenAI

59.9

Benchmarks

59.90.024.347.10.0N/A
50

Grok-3

grok-3

multimodalvisionmulti-input reasoning
xAI

58.4

Benchmarks

58.450.20.00.024.1$3 in / $15 out
51

Gemini 3.6 Flash

gemini-3.6-flash

multimodalvisionmulti-input reasoning
Google

57.8

Benchmarks

57.884.80.052.634.8
52

Qwen3.5-27B

qwen3.5-27b

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

57.8

Benchmarks

57.832.941.339.161.0$0.3 in / $2.4 out
53

ERNIE 5.0

ernie-5.0

multimodalvisionmulti-input reasoning
BBaidu

56.8

Benchmarks

56.80.00.00.00.0N/A
54

Seed 2.0 Lite

seed-2.0-lite

multimodalvisionmulti-input reasoning
BByteDance

56.5

Benchmarks

56.50.00.046.10.0N/A
55

DeepSeek-V4-Flash-Max

deepseek-v4-flash-max

codeprogrammingtool use
DeepSeek

56.2

Benchmarks

56.284.835.341.498.8
56

Grok 4 Fast

grok-4-fast

multimodalvisionmulti-input reasoning
xAI

55.7

Benchmarks

55.70.012.80.00.0N/A
57

Gemma 4 31B

gemma-4-31b-it

multimodalvisionmulti-input reasoning
Google

55.4

Benchmarks

55.432.90.00.091.5
58

DeepSeek-V3.2

deepseek-v3.2

codeprogrammingtool use
DeepSeek

55.0

Benchmarks

55.00.012.442.00.0N/A
59

Nemotron 3 Ultra (550B A55B)

nemotron-3-ultra-550b-a55b

codeprogrammingtool use
NNVIDIA

54.7

Benchmarks

54.70.011.540.60.0N/A
60

Qwen3.6-27B

qwen3.6-27b

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

54.7

Benchmarks

54.732.90.039.449.4$0.6 in / $3.6 out
41

Claude Sonnet 4.6

Anthropic

62.3

$3 in / $15 out

42

GPT-5 High

OpenAI

61.6

N/A

43

GPT-5.5 Pro

OpenAI

61.1

N/A

44

Page 3 of 17 · 334 models

PreviousNext

Want benchmark charts, model comparison, and pricing analytics?

Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.

Open full leaderboard

Rankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.

N/A
N/A
$1.5 in / $9 out
$1.5 in / $7.5 out
$0.1 in / $0.2 out
$0.13 in / $0.38 out

GPT-5.1 Codex High

OpenAI

61.0

N/A

45
A

Qwen3.5-122B-A10B

Alibaba Cloud / Qwen Team

60.5

N/A

46
Z

GLM-4.7

Zhipu AI

60.3

N/A

47

Gemini 3.5 Flash

Google

60.2

$1.5 in / $9 out

48
M

MAI-Thinking-1

Microsoft

60.1

N/A

49

GPT-5

OpenAI

59.9

N/A

50

Grok-3

xAI

58.4

$3 in / $15 out

51

Gemini 3.6 Flash

Google

57.8

$1.5 in / $7.5 out

52
A

Qwen3.5-27B

Alibaba Cloud / Qwen Team

57.8

$0.3 in / $2.4 out

53
B

ERNIE 5.0

Baidu

56.8

N/A

54
B

Seed 2.0 Lite

ByteDance

56.5

N/A

55

DeepSeek-V4-Flash-Max

DeepSeek

56.2

$0.1 in / $0.2 out

56

Grok 4 Fast

xAI

55.7

N/A

57

Gemma 4 31B

Google

55.4

$0.13 in / $0.38 out

58

DeepSeek-V3.2

DeepSeek

55.0

N/A

59
N

Nemotron 3 Ultra (550B A55B)

NVIDIA

54.7

N/A

60
A

Qwen3.6-27B

Alibaba Cloud / Qwen Team

54.7

$0.6 in / $3.6 out