Skytells
HomeModelsCLIChangelog
  • Home
  • Models
  • CLI
  • Changelog
Skytells

Addressing the world's greatest challenges with AI. Enterprise research, foundation models, and infrastructure trusted by organizations worldwide since 2012.

Get Started

  • Console
  • Learn
  • Documentation
  • API Reference
  • Pricing
  • ModelsNew

Platform

  • Cloud AgentsNew
  • AI Solutions
  • Infrastructure
  • Edge Network
  • Trust Center
  • CLI

Resources

  • Blog
  • Changelog
  • AI Leaderboard
  • Research
  • Status

Company

  • About
  • Careers
  • Legal
  • Privacy Policy

© 2012–2026 Skytells, Inc. All rights reserved.

Live rankings

AI Model Leaderboard

Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.

Explore full leaderboardBrowse model catalog

334

Tracked models

29

Providers

286

Benchmarked

28.3

Avg. index

OverallBenchmarksInferenceAgenticProgrammingValue / Price

334 models

RankModelProviderScoreBenchmarksInferenceAgenticProgrammingValuePrice
261

Llama 3.1 8B Instruct

llama-3.1-8b-instruct

textinference
MMeta

3.0

Benchmarks

3.00.00.00.00.0N/A
262

Phi-3.5-mini-instruct

phi-3.5-mini-instruct

multimodalvisionmulti-input reasoning
MMicrosoft

2.4

Benchmarks

2.40.00.00.00.0N/A
263

GPT-3.5 Turbo

gpt-3.5-turbo-0125

multimodalvisionmulti-input reasoning
OpenAI

2.3

Benchmarks

2.330.10.00.060.7
264

Phi-3.5-vision-instruct

phi-3.5-vision-instruct

multimodalvisionmulti-input reasoning
MMicrosoft

2.3

Benchmarks

2.30.00.00.00.0N/A
265

Qwen2 7B Instruct

qwen2-7b-instruct

textinference
AAlibaba Cloud / Qwen Team

2.2

Benchmarks

2.20.00.00.00.0N/A
266

Phi 4 Mini

phi-4-mini

textinference
MMicrosoft

1.9

Benchmarks

1.90.00.00.00.0N/A
267

Gemma 3n E4B Instructed

gemma-3n-e4b-it

multimodalvisionmulti-input reasoning
Google

1.2

Benchmarks

1.20.00.00.00.0
268

Gemma 3n E4B Instructed LiteRT Preview

gemma-3n-e4b-it-litert-preview

multimodalvisionmulti-input reasoning
Google

1.2

Benchmarks

1.20.00.00.00.0
269

DeepSeek VL2 Tiny

deepseek-vl2-tiny

multimodalvisionmulti-input reasoning
DeepSeek

1.1

Benchmarks

1.10.00.00.00.0
270

Gemma 3n E2B Instructed

gemma-3n-e2b-it

multimodalvisionmulti-input reasoning
Google

1.0

Benchmarks

1.00.00.00.00.0
271

Gemma 3n E2B Instructed LiteRT (Preview)

gemma-3n-e2b-it-litert-preview

multimodalvisionmulti-input reasoning
Google

1.0

Benchmarks

1.00.00.00.00.0
272

Gemma 3 1B

gemma-3-1b-it

textinference
Google

0.9

Benchmarks

0.90.00.00.00.0N/A
273

Codestral-22B

codestral-22b

textinference
Mistral AI

0.0

Benchmarks

0.00.00.00.00.0N/A
274

Command R+

command-r-plus-04-2024

textinference
Cohere

0.0

Benchmarks

0.00.00.00.00.0N/A
275

DeepSeek-V3.2 (Non-thinking)

deepseek-chat

textinference
DeepSeek

0.0

Benchmarks

0.052.00.00.082.7$0.28 in / $0.42 out
276

DeepSeek-R1

deepseek-r1

textinference
DeepSeek

0.0

Benchmarks

0.00.00.00.00.0N/A
277

DeepSeek-V2.5

deepseek-v2.5

codeprogrammingtool use
DeepSeek

0.0

Benchmarks

0.00.00.00.70.0N/A
278

Devstral Medium

devstral-medium-2507

codeprogrammingtool use
Mistral AI

0.0

Benchmarks

0.00.00.020.60.0N/A
279

Devstral Small 1.1

devstral-small-2507

codeprogrammingtool use
Mistral AI

0.0

Benchmarks

0.00.00.012.50.0N/A
280

Gemini 3.5 Flash Cyber

gemini-3.5-flash-cyber

textinference
Google

0.0

Benchmarks

0.00.00.00.00.0N/A
261
M

Llama 3.1 8B Instruct

Meta

3.0

N/A

262
M

Phi-3.5-mini-instruct

Microsoft

2.4

N/A

263

GPT-3.5 Turbo

OpenAI

2.3

$0.5 in / $1.5 out

264

Page 14 of 17 · 334 models

PreviousNext

Want benchmark charts, model comparison, and pricing analytics?

Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.

Open full leaderboard

Rankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.

$0.5 in / $1.5 out
N/A
N/A
N/A
N/A
N/A
M

Phi-3.5-vision-instruct

Microsoft

2.3

N/A

265
A

Qwen2 7B Instruct

Alibaba Cloud / Qwen Team

2.2

N/A

266
M

Phi 4 Mini

Microsoft

1.9

N/A

267

Gemma 3n E4B Instructed

Google

1.2

N/A

268

Gemma 3n E4B Instructed LiteRT Preview

Google

1.2

N/A

269

DeepSeek VL2 Tiny

DeepSeek

1.1

N/A

270

Gemma 3n E2B Instructed

Google

1.0

N/A

271

Gemma 3n E2B Instructed LiteRT (Preview)

Google

1.0

N/A

272

Gemma 3 1B

Google

0.9

N/A

273

Codestral-22B

Mistral AI

0.0

N/A

274

Command R+

Cohere

0.0

N/A

275

DeepSeek-V3.2 (Non-thinking)

DeepSeek

0.0

$0.28 in / $0.42 out

276

DeepSeek-R1

DeepSeek

0.0

N/A

277

DeepSeek-V2.5

DeepSeek

0.0

N/A

278

Devstral Medium

Mistral AI

0.0

N/A

279

Devstral Small 1.1

Mistral AI

0.0

N/A

280

Gemini 3.5 Flash Cyber

Google

0.0

N/A