CompareLLM
CompareLLM.ai
Live
Leaderboard
Quality & Reasoning
Overall Arena EloPrimary

LMSYS crowd human preference ranking

Coding Elo & SWE-benchCode

Real GitHub issue software solve rate

LiveBench Reasoning

Contamination-free automated tests

GPQA Diamond

PhD-level science & domain knowledge

Speed & Token Cost
Throughput (tok/s)Speed

Output generation token rate

Time to First Token (TTFT)

Response latency for voice & chat loops

Output Price ($/1M tokens)

Cost per million generated tokens

Live Pareto Frontier Scatter

Quality vs Cost efficiency boundary

Models
Model Classes
Frontier ModelsProprietary

Opus 4.5, GPT-5, Gemini 2.0 Pro

Open-Weight CatalogApache/MIT

Llama 3.3, DeepSeek, Qwen 2.5

🇨🇳 Chinese LLMsCN

DeepSeek V3, Qwen, GLM-5, MiniMax

Browse All 40+ Models
Top Providers
Anthropic

Claude Opus 4.5, Sonnet 4.5, Haiku

OpenAI

GPT-5, GPT-4.5, GPT-4o, o3

DeepSeek

DeepSeek V3, R1 Reasoning

Google

Gemini 2.0 Pro, Flash, Thinking

Compare
Popular Head-to-Head ShowdownsView all 48+ pairs →
🇨🇳 DeepSeek V3 vs 🇺🇸 GPT-5

East vs West frontier battle

Sonnet 4.5 vs 🇨🇳 DeepSeek V3

Everyday developer favorite

🇨🇳 Qwen 2.5 vs 🇺🇸 Llama 3.3

Open-weights value clash

Claude Opus 4.5 vs GPT-5

Flagship proprietary duel

Open Interactive Comparison Matrix
Best of & Stacks
Best LLM Lists (2026)
Best Coding LLM

SWE-bench verified repository tests

Best Cheap LLM

Sub-$1/1M token value powerhouses

Fastest Low-Latency LLM

Sub-200ms TTFT for voice & live chat

Claude vs GPT Benchmark

Anthropic vs OpenAI head-to-head

Stack Engine Presets
Cheap Coding Agents

Budget repo bots with high SWE-bench

Lowest-Latency Chat

Fast interactive support loops

Open-Weight Reasoning

Self-hostable reasoning power

Explore All 11 Presets
Research & News
Intelligence & Telemetry
News & Benchmark BriefingsDispatches

Verified model promotions & price shifts

Hourly Benchmark ChangelogLive

Dated ingest audit trail with exact diffs

Evaluation Methodology

Standardized scoring formulas & harnesses

What is Arena Elo?

Understanding blind pairwise human ratings

…
CompareLLM
CompareLLM.ai
Precision Benchmarks

Programmatic, dated AI model benchmarks, head-to-head comparisons, and Stack Engine presets.

Daily ingest · 06:00 UTC

Analytics & Benchmarks

  • AI Model Leaderboard
  • Head-to-Head Compare Hub
  • Models Directory
  • Stack Engine Presets
  • Frontier Models
  • Open Weights Catalog

Guides & Intent Lists

  • Best LLM Lists (2026)
  • Best Coding LLM
  • Best Cheap LLM
  • Fastest Low-Latency LLM
  • Claude vs GPT Benchmark
  • What is Elo?
  • Methodology Guides
  • News & Dispatches

Transparency & API

  • Evaluation Methodology
  • Benchmark Changelog
  • Public JSON API
  • llms.txt Specification
  • Privacy Policy
  • Sign In / Account

© 2026 CompareLLM. Public benchmark data aggregated from Arena Elo, LiveBench, SWE-bench & OpenRouter.

Every score has a dated snapshot.

  1. Home
  2. News Desk
Editorial Intelligence & Release Desk

AI Model News

34 dispatches published

Deep benchmark briefings, price reductions, and architecture analysis. Every dispatch links verified snapshot data and head-to-head showdowns.

AllLaunchesPricesVersusAnalysis
Claude Sonnet 4.5 is the previous workhorse — still a valid vs baselinelaunch
Feb 18, 2026·CompareLLM Intelligence Desk

Claude Sonnet 4.5 is the previous workhorse — still a valid vs baseline

Most of Opus 4.5 coding at a mid-tier price. Sonnet 5 replaced it as the buy; 4.5 remains a compare baseline.

claude-sonnet-4-5claude-sonnet-5
Read briefing
PrevPage 3 of 4Next
Previous
1234
Next
MiniMax M2.5: the Feb 2026 SWE-bench specialist still on the boardlaunch
Feb 12, 2026·CompareLLM Intelligence Desk

MiniMax M2.5: the Feb 2026 SWE-bench specialist still on the board

Tied near the top of official SWE-bench bash-only in Feb 2026. Coding lists should still see it.

minimax-m2-5
Read briefing
Mistral Large 3: the European flagship with function callinglaunch
Jan 18, 2026·CompareLLM Intelligence Desk

Mistral Large 3: the European flagship with function calling

Jan 2026. Not the Elo crown. Useful when residency and tools matter more than Arena.

mistral-large-3
Read briefing
Claude Haiku 4.5 at $1/$5: the cheap Claude people actually shiplaunch
Oct 22, 2025·CompareLLM Intelligence Desk

Claude Haiku 4.5 at $1/$5: the cheap Claude people actually ship

Official Anthropic list $1/$5. This is the SKU behind “cheapest Claude alternative” that is still Claude.

claude-haiku-4-5
Read briefing
Qwen3 235B: the open MoE Qwen, not the Max SKUlaunch
Sep 20, 2025·CompareLLM Intelligence Desk

Qwen3 235B: the open MoE Qwen, not the Max SKU

Sep 2025 open-weight mixture-of-experts. Self-host path. Don’t confuse it with Qwen 3 Max.

qwen-3-235bqwen-3-max
Read briefing
GPT-5 mini: the 2025 distill still useful as a cheap OpenAI baselinelaunch
Aug 8, 2025·CompareLLM Intelligence Desk

GPT-5 mini: the 2025 distill still useful as a cheap OpenAI baseline

High-volume agents that cannot pay Sol. Compare to Luna before you assume the new cheap SKU wins.

gpt-5-minigpt-5-6-luna
Read briefing
GPT-5 (2025) remains the legacy OpenAI flagship baselinelaunch
Aug 8, 2025·CompareLLM Intelligence Desk

GPT-5 (2025) remains the legacy OpenAI flagship baseline

Still a common compare target. 5.6 Sol replaced it as the buy; GPT-5 stays for old URLs.

gpt-5gpt-5-6-sol
Read briefing
Grok 4 remains the previous xAI flagship on this cataloglaunch
Jul 10, 2025·CompareLLM Intelligence Desk

Grok 4 remains the previous xAI flagship on this catalog

2025-07 row. 4.6 is the buy. We keep 4 so upgrade pairs and old inbound links resolve.

grok-4grok-4-6
Read briefing
Gemini 2.5 Flash: previous Google speed workhorse, still a baselinelaunch
Mar 25, 2025·CompareLLM Intelligence Desk

Gemini 2.5 Flash: previous Google speed workhorse, still a baseline

Useful as a cheap historical compare. 3.7 Flash is the 2026 volume SKU.

gemini-2-5-flashgemini-3-7-flash
Read briefing
Gemini 2.5 Pro remains the previous Google long-context flagshiplaunch
Mar 25, 2025·CompareLLM Intelligence Desk

Gemini 2.5 Pro remains the previous Google long-context flagship

2025-03 row. 3.x Pro replaced it. Kept for old “2.5 pro vs gpt-4o” inbound.

gemini-2-5-pro
Read briefing