CompareLLM
CompareLLM.ai
Live
LeaderboardModelsCompareStacksBest ofGuidesNewsMethod
…
CompareLLM
CompareLLM.ai
Precision Benchmarks

Programmatic, dated AI model benchmarks, head-to-head comparisons, and Stack Engine presets.

Daily ingest · 06:00 UTC

Analytics & Benchmarks

  • AI Model Leaderboard
  • Head-to-Head Compare Hub
  • Models Directory
  • Stack Engine Presets
  • Frontier Models
  • Open Weights Catalog

Guides & Intent Lists

  • Best LLM Lists (2026)
  • Best Coding LLM
  • Best Cheap LLM
  • Fastest Low-Latency LLM
  • Claude vs GPT Benchmark
  • What is Elo?
  • Methodology Guides
  • News & Dispatches

Transparency & API

  • Evaluation Methodology
  • Benchmark Changelog
  • Public JSON API
  • llms.txt Specification
  • Privacy Policy
  • Sign In / Account

© 2026 CompareLLM. Public benchmark data aggregated from Arena Elo, LiveBench, SWE-bench & OpenRouter.

Every score has a dated snapshot.

  1. Home
  2. Guides
  3. What is Elo on an AI leaderboard?
CompareLLM Practical Guide
Updated 2026-08-16

What is Elo on an AI leaderboard?

Elo is a crowd vote: people pick which hidden answer they like more. A higher Elo means more people preferred that model — not that it passed a test.

Quick answer

Imagine two AIs write an answer to the same question. A person sees both answers with the names hidden. They tap the one they like more. Nobody is grading math homework. They are just saying “this one felt better.”
1Section 1

In one minute

Imagine two AIs write an answer to the same question. A person sees both answers with the names hidden. They tap the one they like more. Nobody is grading math homework. They are just saying “this one felt better.”

After thousands of those votes, the AI that wins more often gets a higher Elo. A higher Elo means people preferred its answers more often. It does not mean the model is smarter, cheaper, faster, or better at coding.

2Section 2

Where the number comes from

Elo started as a chess rating: when two players compete, the winner takes points from the loser. On this site the “players” are language models. The votes come from public LMArena / Arena tables. We copy the dated snapshot. We do not run the voting booth, and we do not invent a missing cell.

The column is labeled Preference Elo so it is not confused with a school test or a 0–100 “intelligence index.”

3Section 3

What Elo is not

It is not a science exam. A friendly, long answer can beat a short, correct one if voters like the style.

It is not the coding test (SWE-bench: did an agent finish a real GitHub issue?). It is not LiveBench (objective tasks). It is not “did it tell the truth?” Coding Elo, when we have it, is a separate vote on coding prompts — do not mix it with general Elo.

4Section 4

How to use it here

Use Elo to shortlist a general chat assistant. Then open a vs page and check the coding test, speed, and price for the actual job. The highest Elo model is often the wrong pick for a cheap widget or a repo agent.

If two models are close (tens of points, not hundreds), treat the crown as noise. Ratings move as new votes arrive. Always read the as-of date.

5Section 5

Why other sites quote a different Elo

Different dumps, different dates, and different name matching. We refuse to scrape a third-party “intelligence index.” If our cell is blank, the public feed did not match a name yet.

Ready to evaluate your stack?

Calculate your optimal model weights with Stack Engine or compare top models head-to-head.

Stack EngineCompare Hub