CompareLLM
CompareLLM.ai
Live
LeaderboardModelsCompareStacksBest ofGuidesNewsMethod
…
CompareLLM
CompareLLM.ai
Precision Benchmarks

Programmatic, dated AI model benchmarks, head-to-head comparisons, and Stack Engine presets.

Daily ingest · 06:00 UTC

Analytics & Benchmarks

  • AI Model Leaderboard
  • Head-to-Head Compare Hub
  • Models Directory
  • Stack Engine Presets
  • Frontier Models
  • Open Weights Catalog

Guides & Intent Lists

  • Best LLM Lists (2026)
  • Best Coding LLM
  • Best Cheap LLM
  • Fastest Low-Latency LLM
  • Claude vs GPT Benchmark
  • What is Elo?
  • Methodology Guides
  • News & Dispatches

Transparency & API

  • Evaluation Methodology
  • Benchmark Changelog
  • Public JSON API
  • llms.txt Specification
  • Privacy Policy
  • Sign In / Account

© 2026 CompareLLM. Public benchmark data aggregated from Arena Elo, LiveBench, SWE-bench & OpenRouter.

Every score has a dated snapshot.

  1. Home
  2. Benchmarks
  3. Preference Elo (LMArena / Arena)
Benchmark Specification
High (Blind Prompts)

Preference Elo (LMArena / Arena)

Quick answer

Preference Elo is a crowd vote from LMArena / Arena. People see two hidden answers and pick the one they like more. The model that wins more often gets a higher Elo. That means people preferred it — not that it passed a school test. It is not SWE-bench, not accuracy, and not a number we invent.

Crowdsourced blind pairwise human preference battles. Users prompt two anonymous models and vote on the superior response; Bradley-Terry Elo is computed over hundreds of thousands of comparisons. Ingested from official verified public snapshots.

Full explainer: What is Elo on an AI leaderboard?

Primary EvaluatorLMSYS / Chatbot Arena
Update CadenceDaily snapshots
Leaderboard LinkView full table

Top Performing Models

Top 15 verified
#1
Claude Opus 5Anthropic
seed-bootstrap · Aug 16, 2026
1,624
#2
Claude Fable 5Anthropic
seed-bootstrap · Aug 16, 2026
1,616vs #1
#3
GPT-5.6 SolOpenAI
seed-bootstrap · Aug 16, 2026
1,608vs #1
#4
Claude Opus 4.8Anthropic
seed-bootstrap · Aug 16, 2026
1,598vs #1
#5
Grok 4.6xAI
seed-bootstrap · Aug 16, 2026
1,592vs #1
#6
Claude Opus 4.6Anthropic
seed-bootstrap · Aug 1, 2026
1,574vs #1
#7
Gemini 3.6 ProGoogle
seed-bootstrap · Aug 16, 2026
1,570vs #1
#8
Claude Opus 4.5Anthropic
seed-bootstrap · Aug 1, 2026
1,568vs #1
#9
OpenAI: o3 MiniOpenAI
seed-bootstrap · Aug 1, 2026
1,560vs #1
#10
GPT-5OpenAI
seed-bootstrap · Aug 1, 2026
1,558vs #1
#11
GLM-5.3Zhipu
seed-bootstrap · Aug 16, 2026
1,558vs #1
#12
Claude Sonnet 5Anthropic
seed-bootstrap · Aug 16, 2026
1,556vs #1
#13
Gemini 3 ProGoogle
seed-bootstrap · Aug 1, 2026
1,552vs #1
#14
GPT-5.6 TerraOpenAI
seed-bootstrap · Aug 16, 2026
1,548vs #1
#15
GPT-4.5 OrionOpenAI
seed-bootstrap · Aug 1, 2026
1,540vs #1

Frequently asked questions

Plain-English methodology and leaderboard answers

Preference Elo is a crowd vote from LMArena / Arena. People see two hidden answers and pick the one they like more. The model that wins more often gets a higher Elo. That means people preferred it — not that it passed a school test. It is not SWE-bench, not accuracy, and not a number we invent.