CompareLLM
CompareLLM.ai
Live
LeaderboardModelsCompareStacksBest ofGuidesNewsMethod
…
CompareLLM
CompareLLM.ai
Precision Benchmarks

Programmatic, dated AI model benchmarks, head-to-head comparisons, and Stack Engine presets.

Daily ingest · 06:00 UTC

Analytics & Benchmarks

  • AI Model Leaderboard
  • Head-to-Head Compare Hub
  • Models Directory
  • Stack Engine Presets
  • Frontier Models
  • Open Weights Catalog

Guides & Intent Lists

  • Best LLM Lists (2026)
  • Best Coding LLM
  • Best Cheap LLM
  • Fastest Low-Latency LLM
  • Claude vs GPT Benchmark
  • What is Elo?
  • Methodology Guides
  • News & Dispatches

Transparency & API

  • Evaluation Methodology
  • Benchmark Changelog
  • Public JSON API
  • llms.txt Specification
  • Privacy Policy
  • Sign In / Account

© 2026 CompareLLM. Public benchmark data aggregated from Arena Elo, LiveBench, SWE-bench & OpenRouter.

Every score has a dated snapshot.

Workload Architecture Optimizer

AI Stack Engine Presets

Deterministic recommendations computed from verified benchmark percentiles. Objective weights published transparently on each page.

Workload Preset

cheap coding agents

SWE-bench first, then price. For CI bots and repo agents that cannot burn Opus prices.

DeepSeek V4 Flash
Score 76.2
Workload Preset

open-weight reasoning

Highest reasoning signal among models you can self-host or buy as open weights.

DeepSeek V4 Pro
Score 78.0
Workload Preset

lowest-latency chat

TTFT and tokens/sec first. For support widgets and voice-adjacent loops.

Gemini 3.7 Flash
Score 80.9
Workload Preset

long-context RAG

Context window and input price for stuffing large corpora.

GPT-5.6 Luna
Score 82.8
Workload Preset

frontier agents

Coding + preference Elo for computer-use and multi-step tools.

Gemini 3.7 Flash
Score 78.8
Workload Preset

vision and screenshots

Multimodal models only. Preference Elo and latency for UI-understanding jobs.

Gemini 3.7 Flash
Score 80.9
Workload Preset

cheapest hosted API

Output price first among models we still consider usable for chat.

Gemini 3.7 Flash
Score 80.9
Workload Preset

highest coding accuracy

SWE-bench first with no budget cap. For when the patch quality matters more than the invoice.

Claude Fable 5
Score 98.8
Workload Preset

writing and editing

Preference Elo first for long-form drafts. Price still counts if you generate all day.

Gemini 3.7 Flash
Score 80.9
Workload Preset

mid-tier workhorse

Sonnet / Terra / Flash class. Usable Elo without Opus or Sol prices.

Gemini 3.7 Flash
Score 80.9
Workload Preset

cheapest frontier

Models that still clear a high Elo bar, sorted so price hurts.

OpenAI: o3 Mini
Score 83.0