Skytells
HomeModelsCLIChangelog
  • Home
  • Models
  • CLI
  • Changelog
Skytells

Addressing the world's greatest challenges with AI. Enterprise research, foundation models, and infrastructure trusted by organizations worldwide since 2012.

Get Started

  • Console
  • Learn
  • Documentation
  • API Reference
  • Pricing
  • ModelsNew

Platform

  • Cloud AgentsNew
  • AI Solutions
  • Infrastructure
  • Edge Network
  • Trust Center
  • CLI

Resources

  • Blog
  • Changelog
  • AI Leaderboard
  • Research
  • Status

Company

  • About
  • Careers
  • Legal
  • Privacy Policy

© 2012–2026 Skytells, Inc. All rights reserved.

Live rankings

AI Model Leaderboard

Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.

Explore full leaderboardBrowse model catalog

334

Tracked models

29

Providers

286

Benchmarked

28.3

Avg. index

OverallBenchmarksInferenceAgenticProgrammingValue / Price

334 models

RankModelProviderScoreBenchmarksInferenceAgenticProgrammingValuePrice
61

GPT-5 Medium

gpt-5-medium-2025-08-07

multimodalvisionmulti-input reasoning
OpenAI

54.6

Benchmarks

54.60.00.00.00.0N/A
62

Claude Opus 4.5

claude-opus-4-5-20251101

multimodalvisionmulti-input reasoning
Anthropic

54.5

Benchmarks

54.50.035.272.20.0
63

Qwen3.5-397B-A17B

qwen3.5-397b-a17b

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

54.5

Benchmarks

54.50.024.755.30.0N/A
64

Qwen3.5-35B-A3B

qwen3.5-35b-a3b

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

53.4

Benchmarks

53.40.038.330.20.0N/A
65

ChatGPT-4o Latest

chatgpt-4o-latest

multimodalvisionmulti-input reasoning
OpenAI

53.0

Benchmarks

53.00.00.00.00.0
66

LongCat-Flash-Thinking-2601

longcat-flash-thinking-2601

codeprogrammingtool use
Meituan

52.9

Benchmarks

52.90.025.633.50.0
67

Gemini 3.1 Flash-Lite

gemini-3.1-flash-lite-preview

multimodalvisionmulti-input reasoning
Google

52.6

Benchmarks

52.662.80.00.067.1
68

GPT OSS 20B High

gpt-oss-20b-high

textinference
OpenAI

52.3

Benchmarks

52.30.00.00.00.0N/A
69

Claude Sonnet 4.5

claude-sonnet-4-5-20250929

multimodalvisionmulti-input reasoning
Anthropic

51.4

Benchmarks

51.412.669.974.612.0
70

GPT-5.4 Mini

gpt-5.4-mini

texttext-to-textlanguage
OpenAI

51.1

Benchmarks

51.144.515.020.042.7
71

Qwen3.6-35B-A3B

qwen3.6-35b-a3b

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

51.1

Benchmarks

51.10.09.825.20.0N/A
72

Grok-3 Mini

grok-3-mini

multimodalvisionmulti-input reasoning
xAI

51.0

Benchmarks

51.00.00.00.00.0N/A
73

DeepSeek-V3.2-Speciale

deepseek-v3.2-speciale

codeprogrammingtool use
DeepSeek

50.7

Benchmarks

50.70.05.042.00.0
74

MiMo-V2-Flash

mimo-v2-flash

codeprogrammingtool use
Xiaomi

50.6

Benchmarks

50.60.023.535.80.0N/A
75

DeepSeek-V3.2-Exp

deepseek-v3.2-exp

codeprogrammingtool use
DeepSeek

50.2

Benchmarks

50.20.027.238.00.0N/A
76

Gemini 2.5 Pro Preview 06-05

gemini-2.5-pro-preview-06-05

multimodalvisionmulti-input reasoning
Google

50.2

Benchmarks

50.20.00.025.80.0
77

DeepSeek-V3.2 (Thinking)

deepseek-reasoner

codeprogrammingtool use
DeepSeek

49.8

Benchmarks

49.80.012.442.00.0
78

MiniMax M3

minimax-m3

multimodalvisionmulti-input reasoning
MiniMax

49.6

Benchmarks

49.662.837.568.373.2$0.3 in / $1.2 out
79

GPT-5.5 Instant

gpt-5.5-instant

multimodalvisionmulti-input reasoning
OpenAI

49.5

Benchmarks

49.562.80.00.017.3
80

Grok-4

grok-4

multimodalvisionmulti-input reasoning
xAI

49.1

Benchmarks

49.10.00.00.00.0N/A
61

GPT-5 Medium

OpenAI

54.6

N/A

62

Claude Opus 4.5

Anthropic

54.5

N/A

63
A

Qwen3.5-397B-A17B

Alibaba Cloud / Qwen Team

54.5

N/A

64

Page 4 of 17 · 334 models

PreviousNext

Want benchmark charts, model comparison, and pricing analytics?

Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.

Open full leaderboard

Rankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.

N/A
N/A
N/A
$0.25 in / $1.5 out
$3 in / $15 out
$0.75 in / $4.5 out
N/A
N/A
N/A
$5 in / $30 out
A

Qwen3.5-35B-A3B

Alibaba Cloud / Qwen Team

53.4

N/A

65

ChatGPT-4o Latest

OpenAI

53.0

N/A

66

LongCat-Flash-Thinking-2601

Meituan

52.9

N/A

67

Gemini 3.1 Flash-Lite

Google

52.6

$0.25 in / $1.5 out

68

GPT OSS 20B High

OpenAI

52.3

N/A

69

Claude Sonnet 4.5

Anthropic

51.4

$3 in / $15 out

70

GPT-5.4 Mini

OpenAI

51.1

$0.75 in / $4.5 out

71
A

Qwen3.6-35B-A3B

Alibaba Cloud / Qwen Team

51.1

N/A

72

Grok-3 Mini

xAI

51.0

N/A

73

DeepSeek-V3.2-Speciale

DeepSeek

50.7

N/A

74

MiMo-V2-Flash

Xiaomi

50.6

N/A

75

DeepSeek-V3.2-Exp

DeepSeek

50.2

N/A

76

Gemini 2.5 Pro Preview 06-05

Google

50.2

N/A

77

DeepSeek-V3.2 (Thinking)

DeepSeek

49.8

N/A

78

MiniMax M3

MiniMax

49.6

$0.3 in / $1.2 out

79

GPT-5.5 Instant

OpenAI

49.5

$5 in / $30 out

80

Grok-4

xAI

49.1

N/A