Compare AI Models on Verified Benchmarks, Speed & Real Cost
Independent crowd preference Elo, SWE-bench coding tests, and live $/1M API pricing with zero synthetic bias.
Find a high-scoring model that costs less
Top-Left = Best ValueVertical β
Crowd vote (Elo)
Two hidden answers. A person picks the one they like. Higher Elo means more wins β not a school test or a coding exam.
Horizontal β
Output (answer) price / 1 million tokens (USD)
Cost to generate the answer. Provider list prices are USD
Claude Opus 5
Elo
1,624
Quality score
Input price
$10.00
per 1M tokens
Output price
$50.00
per 1M tokens
Scores 100% of Claude Fable 5's Elo at 0% lower price ($50.00 vs $50.00 / 1M)
Live AI Model Leaderboard
Top 10 of 26 models.Top 26 of 26 models. Sortable rankings from dated verified benchmarks.
Compare every row against
Open the Compare column to pick a different model for that row.
Leaders Matchup
Multi-Dimensional Capability Radar
Each spoke is a skill. Farther from the center is better (0β100th percentile). Tap any dot to inspect details.
Recent Benchmark Movements & Changelog
Live updates captured when scores shift across Chatbot Arena, LiveBench, SWE-bench, or when pricing and models change.
Recent AI News & Benchmark Briefings
Independent analysis on model promotions, price reductions, and verified score movements.
Shipped this month
New and Refreshed Models
Zhipu
GLM-5.3
Z.ai Aug 14 2026 post-train of the GLM-5.2 744B base. Coding-plan live; open weights promised after a two-week safety review (z.ai/blog/glm-5.3).
Alibaba
Qwen3.8 27B
Alibaba open-weight 27B drop dated Aug 14 2026. Dense enough to self-host; not a frontier MoE.
Gemini 3.7 Flash
Google Aug 13 2026 workhorse. Official intro price $0.75/$3.75 per 1M through Dec 31 2026; 1,048,576-token context (Google blog).
How to read this leaderboard
Plain-English methodology and leaderboard answers
- Preference Elo is a crowd vote from LMArena / Arena. People see two hidden answers and pick the one they like more. The model that wins more often gets a higher Elo. That means people preferred it β not that it passed a school test. It is not SWE-bench, not accuracy, and not a number we invent.

