CompareLLM
CompareLLM.ai
Live
LeaderboardModelsCompareStacksBest ofGuidesNewsMethod
…
CompareLLM
CompareLLM.ai
Precision Benchmarks

Programmatic, dated AI model benchmarks, head-to-head comparisons, and Stack Engine presets.

Daily ingest · 06:00 UTC

Analytics & Benchmarks

  • AI Model Leaderboard
  • Head-to-Head Compare Hub
  • Models Directory
  • Stack Engine Presets
  • Frontier Models
  • Open Weights Catalog

Guides & Intent Lists

  • Best LLM Lists (2026)
  • Best Coding LLM
  • Best Cheap LLM
  • Fastest Low-Latency LLM
  • Claude vs GPT Benchmark
  • What is Elo?
  • Methodology Guides
  • News & Dispatches

Transparency & API

  • Evaluation Methodology
  • Benchmark Changelog
  • Public JSON API
  • llms.txt Specification
  • Privacy Policy
  • Sign In / Account

© 2026 CompareLLM. Public benchmark data aggregated from Arena Elo, LiveBench, SWE-bench & OpenRouter.

Every score has a dated snapshot.

Updated Aug 16, 2026·70 Models Evaluated·Live Ingest

Compare AI Models on Verified Benchmarks, Speed & Real Cost

Independent crowd preference Elo, SWE-bench coding tests, and live $/1M API pricing with zero synthetic bias.

Best codingBest cheap
Top Preference Elo1,624Claude Opus 5
Fastest TTFT82 msGemini 3.7 Flash

Find a high-scoring model that costs less

Top-Left = Best Value

Up = People liked it more. Left = Cheaper. Teal line = best deals. Gold ring = newest.

Two hidden answers. A person picks the one they like. Higher Elo means more wins — not a school test or a coding exam.

ECB · USD
30/70
Better
↑
Crowd vote (Elo)↓
Lower
↑ Up: Crowd vote (Elo)← Left: Cheaper

Vertical ↑

Crowd vote (Elo)

Lost more votes ↓People liked it more ↑

Two hidden answers. A person picks the one they like. Higher Elo means more wins — not a school test or a coding exam.

Horizontal →

Output (answer) price / 1 million tokens (USD)

← CheaperMore expensive →

Cost to generate the answer. Provider list prices are USD

AnthropicOn the best-deal line

Claude Opus 5

Elo

1,624

Quality score

Prompt / in

$10.00

per 1M in

Answer / out

$50.00

per 1M out

In plain numbers vs Claude Fable 5

Scores 100% of Claude Fable 5's Elo at 0% lower price ($50.00 vs $50.00 / 1M)

Compare Opus 5 vs Fable 5View full specs
Brand:

Leaders Matchup

Open Claude Opus 5 vs GPT-5.6 Sol fact sheetOpen showdown fact sheet|View all comparisons
Capability Matchup Radar

Multi-Dimensional Capability Radar

Each spoke is a skill. Farther from the center is better (0–100th percentile). Tap any dot to inspect details.

Tap any node to inspect
Percentile 0–100
Dimensional AdvantageEven 3–3
Sortable Leaderboard

Live AI Model Leaderboard

Top 10 of 16 models.Top 16 of 16 models. Sortable rankings from dated snapshots.

Compare every row against

Open the Compare column to pick a different model for that row.

AllAllFrontierFrontierOpenOpen weights🇨🇳 CN🇨🇳 CN Models
1Llama 3.1 70B
MetaOpen
Elo1,240
SWE40.2%
In /1M$0.4/1M
Out /1M$0.4/1M
View Llama 3.1 70B specs
2Llama 3.3 70B
MetaOpen
Elo1,285
SWE45.1%
In /1M$0.1/1M
Out /1M$0.32/1M
View Llama 3.3 70B specs
3DeepSeek V3
DeepSeek🇨🇳 CNOpen
Elo1,310
SWE48.6%
In /1M$0.26/1M
Out /1M$1.03/1M
View DeepSeek V3 specs
4DeepSeek R1
DeepSeek🇨🇳 CNOpen
Elo1,358
SWE65.2%
In /1M$0.7/1M
Out /1M$2.5/1M
View DeepSeek R1 specs
5DeepSeek Coder V2
DeepSeek🇨🇳 CNOpen
Elo1,365
SWE60.5%
In /1M$0.14/1M
Out /1M$0.28/1M
View DeepSeek Coder V2 specs
6Llama 3.1 405B
MetaOpen
Elo1,370
SWE58.4%
In /1M$1.25/1M
Out /1M$3.5/1M
View Llama 3.1 405B specs
7Qwen3.8 27B
Alibaba🇨🇳 CNOpen
Elo1,398
SWE58.8%
In /1M$0.45/1M
Out /1M$3.2/1M
View Qwen3.8 27B specs
8Llama 4 Scout
MetaOpen
Elo1,412
SWE52.4%
In /1M$0.1/1M
Out /1M$0.3/1M
View Llama 4 Scout specs
9Qwen 2.5 Coder 32B
Alibaba🇨🇳 CNOpen
Elo1,425
SWE65.2%
In /1M$0.18/1M
Out /1M$0.72/1M
View Qwen 2.5 Coder 32B specs
10Qwen3 235B
Alibaba🇨🇳 CNOpen
Elo1,440
SWE59.4%
In /1M$0.45/1M
Out /1M$1.82/1M
View Qwen3 235B specs
#Model Elo LiveBench SWE-bench tok/s In $/1M Out $/1M Compare vs
1Llama 3.1 70B
MetaOpen
1,24046.8%40.2%125 tok/s$0.4/1M$0.4/1M
2Llama 3.3 70B
MetaOpen
1,28549.8%45.1%140 tok/s$0.1/1M$0.32/1M
3DeepSeek V3
DeepSeek🇨🇳 CNOpen
1,31055.4%48.6%66 tok/s$0.26/1M$1.03/1M
4DeepSeek R1
DeepSeek🇨🇳 CNOpen
1,35859%65.2%62 tok/s$0.7/1M$2.5/1M
5DeepSeek Coder V2
DeepSeek🇨🇳 CNOpen
1,36558.6%60.5%75 tok/s$0.14/1M$0.28/1M
6Llama 3.1 405B
MetaOpen
1,37061.2%58.4%42 tok/s$1.25/1M$3.5/1M
7Qwen3.8 27B
Alibaba🇨🇳 CNOpen
1,39860.6%58.8%152 tok/s$0.45/1M$3.2/1M
8Llama 4 Scout
MetaOpen
1,41259.6%52.4%155 tok/s$0.1/1M$0.3/1M
9Qwen 2.5 Coder 32B
Alibaba🇨🇳 CNOpen
1,42564.8%65.2%110 tok/s$0.18/1M$0.72/1M
10Qwen3 235B
Alibaba🇨🇳 CNOpen
1,44063.5%59.4%74 tok/s$0.45/1M$1.82/1M
11Llama 4 Maverick
MetaOpen
1,46862.8%57.1%130 tok/s$0.2/1M$0.8/1M
12MiniMax M2.5
MiniMax🇨🇳 CNOpen
1,47864.4%75.8%90 tok/s$0.22/1M$0.9/1M
13DeepSeek V4 Flash
DeepSeek🇨🇳 CNOpen
1,49067.8%71.2%142 tok/s$0.14/1M$0.28/1M
14GLM-5.2
Zhipu🇨🇳 CNOpen
1,49265.2%69.3%84 tok/s$0.31/1M$0.97/1M
15Qwen QwQ 32B
Alibaba🇨🇳 CNOpen
1,49568%67.5%68 tok/s$0.3/1M$1.2/1M
16DeepSeek V4 Pro
DeepSeek🇨🇳 CNOpen
1,53670.1%71.6%70 tok/s$0.44/1M$0.87/1M
16 models in this filter · sorted by elo
All benchmark leaderboardsIntent rankings
Automated Audit TrailLive Ingest

Recent Benchmark Movements & Changelog

Live updates captured when scores shift across Chatbot Arena, LiveBench, SWE-bench, or when pricing and models change.

RSSFull changelog (508)
New Model
Aug 16, 2026
Verified Intelligence & Dispatches

Recent AI News & Benchmark Briefings

Independent analysis on model promotions, price reductions, and verified score movements.

View all dispatches (51)
GLM-5.2 vs GLM 5V Turbo Price and SWE-bench 2026
news
AIEval EditorialAug 16, 2026

GLM-5.2 vs GLM 5V Turbo Price and SWE-bench 2026

GLM-5.2 undercuts GLM 5V Turbo by 74.2% on input at $0.31/1M, with 1,492 Elo and 69.3% SWE-bench in dated 2026 snapshots.

Verified briefRead dispatch

Shipped this month

New and Refreshed Models

View all 431 catalog models

Zhipu

GLM-5.3

Z.ai Aug 14 2026 post-train of the GLM-5.2 744B base. Coding-plan live; open weights promised after a two-week safety review (z.ai/blog/glm-5.3).

2026-08-14

Alibaba

Qwen3.8 27B

Alibaba open-weight 27B drop dated Aug 14 2026. Dense enough to self-host; not a frontier MoE.

2026-08-14

Google

Gemini 3.7 Flash

Google Aug 13 2026 workhorse. Official intro price $0.75/$3.75 per 1M through Dec 31 2026; 1,048,576-token context (Google blog).

2026-08-13

How to read this leaderboard

Plain-English methodology and leaderboard answers

Preference Elo is a crowd vote from LMArena / Arena. People see two hidden answers and pick the one they like more. The model that wins more often gets a higher Elo. That means people preferred it — not that it passed a school test. It is not SWE-bench, not accuracy, and not a number we invent.
Fastest
Claude vs GPT
Open weights
How Scores Work

Elo is a crowd vote, not a test

Blind crowd votes on real prompts — higher Elo means people preferred its answers over competitors.

A person is shown two answers without knowing which AI wrote them. They pick the one they like more.

After thousands of votes, the AI that wins more often gets a higher Elo rating.

  • Higher Elo — people liked its answers more.
  • Not synthetic — zero test prompt contamination.
Read complete Elo methodology guide

Direct Head-to-Head Comparison

Pick any two models to compare preference Elo, coding solve rate, speed, and real $/1M token pricing.

Elo (votes)1,568higher is better
SWE-bench76.8%coding solve
Elo (votes)1,558higher is better
SWE-bench68.4%coding solve
Popular:
Opus 4.6 vs GPT-5Grok 4.6 vs OpusGemini Flash vs GPT-5 miniDeepSeek V4 vs Sonnet
Best Value (Elo ≥ 1450)$0.14/1MYi-Lightning
Indexable Models7062 Active

Catalog refresh: Claude Opus 5 / Fable 5 / Sonnet 5, GPT-5.6 Sol·Terra·Luna, Gemini 3.6/3.7 Flash, Grok 4.6 Aug 12 prices, GLM-5.3 and Qwen3.8 27B (Aug 14).

Inspect audit logDate audit →
New Model
Aug 16, 2026

Organic/GEO layer: What is Elo guide, answer-first boxes, FAQ schema, AI crawlers allowed, admin ingest/alias/promote desk. DeepSeek V4-Pro-0813 alias added. No official flagship after Aug 14.

Inspect audit logDate audit →
New Model
Aug 16, 2026

Catalog add: Claude Haiku 4.5 ($1/$5 official) and Llama 4 Scout (open-weight long-context) for cheap-alternative compares.

Inspect audit logDate audit →
New Model
Aug 16, 2026

Same-class predecessors: Claude 3 Opus / Opus 4 / 3.5 Sonnet / Sonnet 4 / 3.5 Haiku, GPT-4 Turbo, Grok 3, DeepSeek V3, Llama 3.1 70B, Gemini 1.5 Pro.

Inspect audit logDate audit →
Score Shiftopenrouter
Aug 16, 2026
GoogleGoogle: Gemini 3.1 Flash Lite

New Context window for gemini-3-1-flash-lite: 1048576 (openrouter).

View factsheetDate audit →
Price Updateopenrouter
Aug 16, 2026
OpenAIOpenAI: GPT-5.6 Terra (batch)

New Output price for gpt-5-6-terra-batch: 6 (openrouter).

View factsheetDate audit →
Haiku 4.5 vs Llama 4 Scout: cheap Claude versus cheap open Llama
compare
CompareLLM Intelligence DeskAug 16, 2026

Haiku 4.5 vs Llama 4 Scout: cheap Claude versus cheap open Llama

$1/$5 hosted Claude vs a cheap long-context open sibling. The alternative-to-Sonnet pair that is honest.

Verified briefRead dispatch
Qwen3.8 27B (Aug 14): dense enough to self-host, not a frontier MoE
launch
CompareLLM Intelligence DeskAug 14, 2026

Qwen3.8 27B (Aug 14): dense enough to self-host, not a frontier MoE

Alibaba open-weight 27B drop. Local/open-weight lists should see it. It will not win frontier-agents.

Verified briefRead dispatch

xAI

Grok 4.6

xAI Aug 12 2026 post-training refresh of Grok 4.5. Same $2/$6 API price, 500k context, stronger agentic traces.

2026-08-12

ByteDance

Seed 2.1 Turbo

ByteDance Seed 2.1 Turbo, listed on public model timelines as an Aug 10 2026 API drop.

2026-08-10

DeepSeek

DeepSeek V4 Flash

DeepSeek Jul 31 2026 price-performance SKU. Public listings put it near $0.14/$0.28 per 1M tokens.

2026-07-31