How CompareLLM compares AI models
What Elo, LiveBench, SWE-bench, TTFT, and list price mean on this site — and what we refuse to invent.
One row is a snapshot, not a personality
Every cell on CompareLLM is a dated snapshot: a number, a unit, a source name, and an observed-at time. If we do not have a public machine-readable feed for a metric, the cell is blank. We do not paint radar axes with guessed 0–100 scores.
V1 is an aggregator plus optional first-party latency pings. We do not run SWE-bench or LiveBench ourselves. That is written on /methodology and on every compare page.
The five numbers people actually argue about
Preference Elo is crowd pairwise taste (LMArena / Arena). It is not a science exam.
LiveBench is a contamination-resistant objective suite. Only compare scores from the same release.
SWE-bench is “did this harness resolve a real GitHub issue?” Agent and split change the number.
TTFT and tok/s are latency and stream rate. List $/1M is not your invoice after cache and retries.
How a new vs page appears
We do not hand-write Claude vs GPT pages. If two models are indexable and share at least three metrics, /compare/{a}-vs-{b} exists (canonical A–Z slug). Frontier pairs are prerendered; the rest are created on first request.
New models arrive from the daily OpenRouter ingest as preview. They join the matrix when a second source matches or we promote the alias.
Ready to evaluate your stack?
Calculate your optimal model weights with Stack Engine or compare top models head-to-head.
