Compare AI Models on Verified Benchmarks, Speed & Real Cost
Independent crowd preference Elo, SWE-bench coding tests, and live $/1M API pricing with zero synthetic bias.
Find a high-scoring model that costs less
Top-Left = Best ValueVertical ↑
Crowd vote (Elo)
Two hidden answers. A person picks the one they like. Higher Elo means more wins — not a school test or a coding exam.
Horizontal →
Output (answer) price / 1 million tokens (USD)
Cost to generate the answer. Provider list prices are USD
Claude Opus 5
Elo
1,624
Quality score
Prompt / in
$10.00
per 1M in
Answer / out
$50.00
per 1M out
Scores 100% of Claude Fable 5's Elo at 0% lower price ($50.00 vs $50.00 / 1M)
Leaders Matchup
Multi-Dimensional Capability Radar
Each spoke is a skill. Farther from the center is better (0–100th percentile). Tap any dot to inspect details.
Live AI Model Leaderboard
Top 10 of 70 models.Top 30 of 70 models. Sortable rankings from dated snapshots.
Compare every row against
Open the Compare column to pick a different model for that row.
40 more not shown
Recent Benchmark Movements & Changelog
Live updates captured when scores shift across Chatbot Arena, LiveBench, SWE-bench, or when pricing and models change.
Recent AI News & Benchmark Briefings
Independent analysis on model promotions, price reductions, and verified score movements.
Shipped this month
New and Refreshed Models
Zhipu
GLM-5.3
Z.ai Aug 14 2026 post-train of the GLM-5.2 744B base. Coding-plan live; open weights promised after a two-week safety review (z.ai/blog/glm-5.3).
Alibaba
Qwen3.8 27B
Alibaba open-weight 27B drop dated Aug 14 2026. Dense enough to self-host; not a frontier MoE.
Gemini 3.7 Flash
Google Aug 13 2026 workhorse. Official intro price $0.75/$3.75 per 1M through Dec 31 2026; 1,048,576-token context (Google blog).
How to read this leaderboard
Plain-English methodology and leaderboard answers
- Preference Elo is a crowd vote from LMArena / Arena. People see two hidden answers and pick the one they like more. The model that wins more often gets a higher Elo. That means people preferred it — not that it passed a school test. It is not SWE-bench, not accuracy, and not a number we invent.

