CompareLLM.ai
Live
Leaderboard
Models
Compare
Best of & Stacks
Research & News
…
CompareLLM.ai
Precision Benchmarks

Programmatic, dated AI model benchmarks, head-to-head comparisons, and Stack Engine presets.

Daily ingest · 06:00 UTC

Analytics & Benchmarks

  • AI Model Leaderboard
  • Head-to-Head Compare Hub
  • Models Directory
  • Stack Engine Presets
  • Frontier Models
  • Open Weights Catalog

Guides & Intent Lists

  • Best LLM Lists (2026)
  • Best Coding LLM
  • Best Cheap LLM
  • Fastest Low-Latency LLM
  • Claude vs GPT Benchmark
  • What is Elo?
  • Methodology Guides
  • News & Dispatches

Transparency & API

  • Evaluation Methodology
  • Benchmark Changelog
  • Public JSON API
  • llms.txt Specification
  • Privacy Policy
  • Sign In / Account

© 2026 CompareLLM. Public benchmark data aggregated from Arena Elo, LiveBench, SWE-bench & OpenRouter.

Every score has a dated snapshot.

Theme:
Currency:
Updated Aug 16, 2026·100 Models·Live Ingest

Compare AI Models on Verified Benchmarks, Speed & Real Cost

Independent crowd preference Elo, SWE-bench coding tests, and live $/1M API pricing with zero synthetic bias.

Best codingBest cheap
Intelligence vs API Price Tradeoff

Quality vs. Price Pareto Frontier

Top-Left = Best Value

Plotted by intelligence (Elo) vs cost (Out $). The teal line connects models with unbeatable price-to-performance.

Leaderboard TableTable↓
30/71 models shown
Better
↑
Crowd vote (Elo)↓
Lower
↑ Up: Crowd vote (Elo)← Left: Cheaper

Vertical ↑

Crowd vote (Elo)

Lost more votes ↓People liked it more ↑

Two hidden answers. A person picks the one they like. Higher Elo means more wins — not a school test or a coding exam.

Horizontal →

Output (answer) price / 1 million tokens (USD)

← CheaperMore expensive →

Cost to generate the answer. Provider list prices are USD

Best-Value Frontier Line (7 Models)Models defining the teal efficiency boundary
Brand:
💬 LLMs & ReasoningLLMs🎨 Text-to-Image ModelsText-to-ImageNew
Sortable LLM Leaderboard

Live AI Model Leaderboard

Top 10 of 33 models.Top 30 of 33 models. Sortable rankings from dated verified benchmarks.

Compare every row vs:
All ModelsFrontierOpen Weights🇨🇳 China
1Claude Opus 5Frontier

Anthropic

Elo1,624
SWE79.2%
In /1M$5
Out /1M$25
Benchmark baselineSpecs
2Claude Fable 5Frontier

Anthropic

Elo1,616
SWE80%
In /1M$10
Out /1M$50
vs Claude Opus 5Specs
3GPT-5.6 SolFrontier

OpenAI

Elo1,608
SWE77.6%
In /1M$5
Out /1M$30
vs Claude Opus 5Specs
4Claude Opus 4.8Frontier

Anthropic

Elo1,598
SWE78.4%
In /1M$5
Out /1M$25
vs Claude Opus 5Specs
5Grok 4.6Frontier

xAI

Elo1,592
SWE69.1%
In /1M$2
Out /1M$6
vs Claude Opus 5Specs
6Claude Opus 4.6Frontier

Anthropic

Elo1,574
SWE75.9%
In /1M$15
Out /1M$75
vs Claude Opus 5Specs
7Gemini 3.6 ProFrontier

Google

Elo1,570
SWE74.6%
In /1M$1.25
Out /1M$5
vs Claude Opus 5Specs
8Claude Opus 4.5Frontier

Anthropic

Elo1,568
SWE76.8%
In /1M$15
Out /1M$75
vs Claude Opus 5Specs
9OpenAI o3-miniFrontier

OpenAI

Elo1,560
SWE78.5%
In /1M$1.1
Out /1M$4.4
vs Claude Opus 5Specs
10GPT-5Frontier

OpenAI

Elo1,558
SWE68.4%
In /1M$5
Out /1M$20
vs Claude Opus 5Specs
#Model Elo LiveBench SWE-bench tok/s In $/1M Out $/1M Compare vs
1
Claude Opus 5Frontier

Anthropic

1,62474.8%79.2%72 tok/s$5.00 / 1M$25.00 / 1M
2
Claude Fable 5Frontier

Anthropic

1,61675.2%80%58 tok/s$10.00 / 1M$50.00 / 1M
3
GPT-5.6 SolFrontier

OpenAI

1,60873.9%77.6%84 tok/s$5.00 / 1M$30.00 / 1M
4
Claude Opus 4.8Frontier

Anthropic

1,59873.6%78.4%66 tok/s$5.00 / 1M$25.00 / 1M
5
Grok 4.6Frontier

xAI

1,59271.4%69.1%118 tok/s$2.00 / 1M$6.00 / 1M
6
Claude Opus 4.6Frontier

Anthropic

1,57472%75.9%64 tok/s$15.00 / 1M$75.00 / 1M
7
Gemini 3.6 ProFrontier

Google

1,57071.8%74.6%102 tok/s$1.25 / 1M$5.00 / 1M
8
Claude Opus 4.5Frontier

Anthropic

1,56871.2%76.8%62 tok/s$15.00 / 1M$75.00 / 1M
9
OpenAI o3-miniFrontier

OpenAI

1,56072.4%78.5%92 tok/s$1.10 / 1M$4.40 / 1M
10
GPT-5Frontier

OpenAI

1,55869.8%68.4%78 tok/s$5.00 / 1M$20.00 / 1M
11
GLM-5.3🇨🇳 CNFrontier

Zhipu

1,55870.8%76.4%90 tok/s$0.50 / 1M$1.80 / 1M
12
Claude Sonnet 5Frontier

Anthropic

1,55670.4%73.8%118 tok/s$2.00 / 1M$10.00 / 1M
13
Gemini 3 ProFrontier

Google

1,55270.6%72.4%92 tok/s$1.25 / 1M$5.00 / 1M
14
GPT-5.6 TerraFrontier

OpenAI

1,54869.1%70.4%128 tok/s$2.00 / 1M$12.00 / 1M
15
GPT-4.5 OrionFrontier

OpenAI

1,54068.5%71%60 tok/s$75.00 / 1M$150.00 / 1M
16
DeepSeek V4 Pro🇨🇳 CNOpenFrontier

DeepSeek

1,53670.1%71.6%70 tok/s$0.40 / 1M$1.60 / 1M
17
Gemini 3.7 FlashFrontier

Google

1,53069.6%72.4%215 tok/s$0.75 / 1M$3.75 / 1M
18
Kimi K3🇨🇳 CNFrontier

Moonshot

1,52869.5%73.8%68 tok/s$0.60 / 1M$2.50 / 1M
19
Claude Sonnet 4.5Frontier

Anthropic

1,52467.4%70.1%88 tok/s$3.00 / 1M$15.00 / 1M
20
Qwen 3 Max🇨🇳 CNFrontier

Alibaba

1,51869.4%67.8%80 tok/s$1.20 / 1M$4.80 / 1M
21
Kimi K2.5 Max🇨🇳 CNFrontier

Moonshot

1,51567.2%71.3%72 tok/s$0.50 / 1M$2.00 / 1M
22
Gemini 3.6 FlashFrontier

Google

1,50667.2%70.8%198 tok/s$0.75 / 1M$3.75 / 1M
23
Gemini 3 FlashFrontier

Google

1,49566.8%75.8%190 tok/s$0.15 / 1M$0.60 / 1M
24
Qwen QwQ 32B🇨🇳 CNOpenFrontier

Alibaba

1,49568%67.5%68 tok/s$0.30 / 1M$1.20 / 1M
25
GLM-5.2🇨🇳 CNOpenFrontier

Zhipu

1,49265.2%69.3%84 tok/s$0.50 / 1M$1.80 / 1M
26
Gemini 2.0 Flash ThinkingFrontier

Google

1,49067%66.8%115 tok/s$0.10 / 1M$0.40 / 1M
27
DeepSeek V4 Flash🇨🇳 CNOpenFrontier

DeepSeek

1,49067.8%71.2%142 tok/s$0.14 / 1M$0.28 / 1M
28
MiniMax M2.5🇨🇳 CNOpenFrontier

MiniMax

1,47864.4%75.8%90 tok/s$0.30 / 1M$1.20 / 1M
29
Yi-Lightning🇨🇳 CNFrontier

01.AI

1,47565%63.4%130 tok/s$0.14 / 1M$0.14 / 1M
30
GLM-5V Turbo🇨🇳 CNOpenFrontier

Zhipu

1,47563.8%65.2%112 tok/s$0.15 / 1M$0.60 / 1M

3 more not shown

33 models in this filter · sorted by elo
All benchmark leaderboardsIntent rankings
US vs China AI FrontierGap Closed & Surpassing on Cost

How Close is China's AI to US Flagships?

Quality is virtually tied on independent benchmarks, while open-weight Chinese frontiers offer substantial price-to-performance advantages.

🇺🇸 US LeaderClaude Opus 5
vs
🇨🇳 China RivalGLM-5.3
General IntelligenceChatbot Arena human preference score
Claude Opus 51624 Elo
vs
GLM-5.31558 Elo
Near Parity (96%)
Software EngineeringSWE-bench real GitHub bug resolution
Claude Opus 579.2%
vs
GLM-5.376.4%
Near Parity (96%)
API Pricing & CostOutput price per 1 Million tokens
Claude Opus 5$25.00
vs
GLM-5.3$1.80
13.9× Cheaper
Model AvailabilityWeights download & self-hosting
Claude Opus 5Closed API
vs
GLM-5.3Open Weights
Open Source
Key Takeaway:GLM-5.3 matches 96% of Claude Opus 5's intelligence at 93% lower cost.
Full Showdown Fact Sheet

Leaders Matchup

Open Claude Opus 5 vs GPT-5.6 Sol fact sheetOpen showdown fact sheet|View all comparisons
Multi-Dimensional Capability Radar

Frontier Trio Capability Radar

Comparing top Western standard models with China's leading frontier rival across 6 skill dimensions. Tap any spoke or dot to inspect.

Tap any node to inspect
Percentile 0–100
Verified Intelligence & Dispatches

Recent AI News & Benchmark Briefings

Independent analysis on model promotions, price reductions, and verified score movements.

View all dispatches (50)
Haiku 4.5 vs Llama 4 Scout: cheap Claude versus cheap open Llama
compare
CompareLLM Intelligence DeskAug 16, 2026

Haiku 4.5 vs Llama 4 Scout: cheap Claude versus cheap open Llama

$1/$5 hosted Claude vs a cheap long-context open sibling. The alternative-to-Sonnet pair that is honest.

Shipped & Updated

New and Refreshed Models

View all 100 catalog models
Zhipu2026-08-14

GLM-5.3

Z.ai Aug 14 2026 post-train of the GLM-5.2 744B base. Coding-plan live; open weights promised after a two-week safety review (z.ai/blog/glm-5.3).

Alibaba2026-08-14

Qwen3.8 27B

Alibaba open-weight 27B drop dated Aug 14 2026. Dense enough to self-host; not a frontier MoE.

How to read this leaderboard

Plain-English methodology and leaderboard answers

Preference Elo is a crowd vote from LMArena / Arena. People see two hidden answers and pick the one they like more. The model that wins more often gets a higher Elo. That means people preferred it — not that it passed a school test. It is not SWE-bench, not accuracy, and not a number we invent.
Fastest
Opus 5 vs GPT-5.6
Open weights
Top Verified Frontiers
Top Preference Elo
1,624Claude Opus 5
Fastest TTFT
82 msGemini 3.7 Flash
Best Value (Elo ≥ 1450)
$0.14/1MYi-Lightning
Indexable Models
7163 Active

Head-to-Head Comparison Studio

All Comparisons Matrix
Trending Comparisons & RivalriesVerified Telemetry
Frontier Titan

Claude Opus 5 vs GPT-5.6 Sol

Top flagship duel for #1 general intelligence

East vs West

Claude Sonnet 5 vs DeepSeek V4

Western standard vs Chinese frontier price/perf leader

Open Weights

Qwen3.8 vs Llama 4 Scout

Top self-hostable open-weight models compared

High-Speed

GLM-5.3 vs Gemini 3.7 Flash

Fast reasoning & latency vs low $/1M API cost

Verified briefRead dispatch
Qwen3.8 27B (Aug 14): dense enough to self-host, not a frontier MoE
launch
CompareLLM Intelligence DeskAug 14, 2026

Qwen3.8 27B (Aug 14): dense enough to self-host, not a frontier MoE

Alibaba open-weight 27B drop. Local/open-weight lists should see it. It will not win frontier-agents.

Verified briefRead dispatch
GLM-5.3 (Aug 14): Z.ai post-train of the 5.2 744B base
launch
CompareLLM Intelligence DeskAug 14, 2026

GLM-5.3 (Aug 14): Z.ai post-train of the 5.2 744B base

Coding-plan live; open weights promised after a two-week safety review. Newest Zhipu row.

Verified briefRead dispatch
Google2026-08-13

Gemini 3.7 Flash

Google Aug 13 2026 workhorse. Official intro price $0.75/$3.75 per 1M through Dec 31 2026; 1,048,576-token context (Google blog).

xAI2026-08-12

Grok 4.6

xAI Aug 12 2026 post-training refresh of Grok 4.5. Same $2/$6 API price, 500k context, stronger agentic traces.

ByteDance2026-08-10

Seed 2.1 Turbo

ByteDance Seed 2.1 Turbo, listed on public model timelines as an Aug 10 2026 API drop.

DeepSeek2026-07-31

DeepSeek V4 Flash

DeepSeek Jul 31 2026 price-performance SKU. Public listings put it near $0.14/$0.28 per 1M tokens.