CompareLLM
CompareLLM.ai
Live
Leaderboard
Quality & Reasoning
Overall Arena EloPrimary

LMSYS crowd human preference ranking

Coding Elo & SWE-benchCode

Real GitHub issue software solve rate

LiveBench Reasoning

Contamination-free automated tests

GPQA Diamond

PhD-level science & domain knowledge

Text-to-Image LeaderboardArena

FLUX.2, GPT Image 2, Nano Banana, Midjourney

Speed & Token Cost
Throughput (tok/s)Speed

Output generation token rate

Time to First Token (TTFT)

Response latency for voice & chat loops

Output Price ($/1M tokens)

Cost per million generated tokens

Live Pareto Frontier Scatter

Quality vs Cost efficiency boundary

Models
Model Classes
Frontier ModelsProprietary

Opus 4.5, GPT-5, Gemini 2.0 Pro

Open-Weight CatalogApache/MIT

Llama 3.3, DeepSeek, Qwen 2.5

🇨🇳 Chinese LLMsCN

DeepSeek V3, Qwen, GLM-5, MiniMax

Browse All 40+ Models
Top Providers
Anthropic

Claude Opus 4.5, Sonnet 4.5, Haiku

OpenAI

GPT-5, GPT-4.5, GPT-4o, o3

DeepSeek

DeepSeek V3, R1 Reasoning

Google

Gemini 2.0 Pro, Flash, Thinking

Compare
Popular Head-to-Head ShowdownsView all 48+ pairs →
Claude Opus 5 vs GPT-5.6 Sol

Flagship frontier titan duel

Sonnet 5 vs 🇨🇳 DeepSeek V4

East vs West coding favorite

🇨🇳 Qwen3.8 vs 🇺🇸 Llama 4 Scout

Latest open-weights clash

GPT-5.6 Sol vs GPT-5

Generational upgrade (1 step back)

Opus 5 vs Opus 4.5

Anthropic generational delta

🎨 FLUX.2 Max vs GPT Image 2

Text-to-Image frontier leader

Open Interactive Comparison Matrix
Best of & Stacks
Best LLM Lists (2026)
Best Coding LLM

SWE-bench verified repository tests

Best Cheap LLM

Sub-$1/1M token value powerhouses

Fastest Low-Latency LLM

Sub-200ms TTFT for voice & live chat

Claude vs GPT Benchmark

Anthropic vs OpenAI head-to-head

Stack Engine Presets
Cheap Coding Agents

Budget repo bots with high SWE-bench

Lowest-Latency Chat

Fast interactive support loops

Open-Weight Reasoning

Self-hostable reasoning power

Explore All 11 Presets
Research & News
Intelligence & Telemetry
News & Benchmark BriefingsDispatches

Verified model promotions & price shifts

Hourly Benchmark ChangelogLive

Dated ingest audit trail with exact diffs

Evaluation Methodology

Standardized scoring formulas & harnesses

What is Arena Elo?

Understanding blind pairwise human ratings

…
CompareLLM
CompareLLM.ai
Precision Benchmarks

Programmatic, dated AI model benchmarks, head-to-head comparisons, and Stack Engine presets.

Daily ingest · 06:00 UTC

Analytics & Benchmarks

  • AI Model Leaderboard
  • Head-to-Head Compare Hub
  • Models Directory
  • Stack Engine Presets
  • Frontier Models
  • Open Weights Catalog

Guides & Intent Lists

  • Best LLM Lists (2026)
  • Best Coding LLM
  • Best Cheap LLM
  • Fastest Low-Latency LLM
  • Claude vs GPT Benchmark
  • What is Elo?
  • Methodology Guides
  • News & Dispatches

Transparency & API

  • Evaluation Methodology
  • Benchmark Changelog
  • Public JSON API
  • llms.txt Specification
  • Privacy Policy
  • Sign In / Account

© 2026 CompareLLM. Public benchmark data aggregated from Arena Elo, LiveBench, SWE-bench & OpenRouter.

Every score has a dated snapshot.

Display Currency:
  1. Home
  2. News Desk
Editorial Intelligence & Release Desk

AI Model News

51 dispatches published

Deep benchmark briefings, price reductions, and architecture analysis. Every dispatch links verified snapshot data and head-to-head showdowns.

AllLaunchesPricesVersusAnalysis
GPT-5.6 Sol: OpenAI’s $5/$30 flagship tier on this leaderboardlaunch
Jul 9, 2026·CompareLLM Intelligence Desk

GPT-5.6 Sol: OpenAI’s $5/$30 flagship tier on this leaderboard

Jul 9 2026 family; Jul 30 price update left Sol unchanged. This is the GPT side of Claude vs GPT.

gpt-5-6-sol
Read briefing
PrevPage 3 of 6Next
Previous
1234…6
Next
Keep Opus 4.8 on the board: a baseline, not a flagship
launch
Jun 20, 2026·CompareLLM Intelligence Desk

Keep Opus 4.8 on the board: a baseline, not a flagship

Opus 4.8 still bills $5/$25. We keep it so Opus 5 vs 4.8 is a real generation delta, not marketing.

claude-opus-4-8claude-opus-5
Read briefing
GLM-5.2 remains the prior Zhipu flagship for upgrade pairslaunch
Jun 18, 2026·CompareLLM Intelligence Desk

GLM-5.2 remains the prior Zhipu flagship for upgrade pairs

Strong Chinese/English coding. 5.3 is the new row; 5.2 stays so the delta is visible.

glm-5-2glm-5-3
Read briefing
Claude Fable 5: the $10/$50 long-horizon Anthropic SKUlaunch
Jun 15, 2026·CompareLLM Intelligence Desk

Claude Fable 5: the $10/$50 long-horizon Anthropic SKU

Fable 5 is not Sonnet with extra adjectives. It is the expensive long-horizon row. Read SWE-bench and output price together.

claude-fable-5
Read briefing
DeepSeek vs Claude: the cheap-versus-frontier question, remappedcompare
Jun 13, 2026·CompareLLM Intelligence Desk

DeepSeek vs Claude: the cheap-versus-frontier question, remapped

Top DeepSeek vs top Anthropic Elo. Usually V4 Pro vs Opus 5. License and invoice decide as much as Elo.

deepseek-v4-proclaude-opus-5
Read briefing
DeepSeek V4 Pro: high reasoning density per dollarlaunch
Jun 12, 2026·CompareLLM Intelligence Desk

DeepSeek V4 Pro: high reasoning density per dollar

2026-06 open-weight-adjacent flagship. The usual cheap-vs-Claude question starts here.

deepseek-v4-pro
Read briefing
Kimi K3: Moonshot’s long-context open-weight competitive rowlaunch
Jun 8, 2026·CompareLLM Intelligence Desk

Kimi K3: Moonshot’s long-context open-weight competitive row

Jun 2026. The Kimi vs GPT hub remaps to the top Moonshot Elo vs top OpenAI Elo.

kimi-k3
Read briefing
Claude Opus 4.6: the May follow-on that still shows up in old bookmarkslaunch
May 10, 2026·CompareLLM Intelligence Desk

Claude Opus 4.6: the May follow-on that still shows up in old bookmarks

Slightly behind 4.5 on one official SWE-bench sweep; stronger long-horizon traces. We keep the row for old vs URLs.

claude-opus-4-6
Read briefing
Qwen 3 Max: Alibaba’s flagship for math and multilingual codelaunch
Apr 18, 2026·CompareLLM Intelligence Desk

Qwen 3 Max: Alibaba’s flagship for math and multilingual code

Closed flagship. Strong math. Compare to Sol and Opus on GPQA and SWE-bench, not on brand heat.

qwen-3-max
Read briefing
Llama vs Claude: open-weight Meta versus the Anthropic crowncompare
Apr 9, 2026·CompareLLM Intelligence Desk

Llama vs Claude: open-weight Meta versus the Anthropic crown

Hub remaps to top Meta vs top Anthropic Elo. Usually Maverick vs Opus 5.

llama-4-maverickclaude-opus-5
Read briefing