CompareLLM
CompareLLM.ai
Live
LeaderboardModelsCompareStacksBest ofGuidesNewsMethod
…
CompareLLM
CompareLLM.ai
Precision Benchmarks

Programmatic, dated AI model benchmarks, head-to-head comparisons, and Stack Engine presets.

Daily ingest · 06:00 UTC

Analytics & Benchmarks

  • AI Model Leaderboard
  • Head-to-Head Compare Hub
  • Models Directory
  • Stack Engine Presets
  • Frontier Models
  • Open Weights Catalog

Guides & Intent Lists

  • Best LLM Lists (2026)
  • Best Coding LLM
  • Best Cheap LLM
  • Fastest Low-Latency LLM
  • Claude vs GPT Benchmark
  • What is Elo?
  • Methodology Guides
  • News & Dispatches

Transparency & API

  • Evaluation Methodology
  • Benchmark Changelog
  • Public JSON API
  • llms.txt Specification
  • Privacy Policy
  • Sign In / Account

© 2026 CompareLLM. Public benchmark data aggregated from Arena Elo, LiveBench, SWE-bench & OpenRouter.

Every score has a dated snapshot.

Intent-Specific Benchmark Lists

Best AI Model Lists (2026)

Optimized rankings calculated transparently from verified benchmark snapshots.

query: best coding llm

Best coding LLM in 2026 (SWE-bench)

Models ranked for repo work using dated SWE-bench snapshots, then price and speed. Not a single lab score.

View ranked leaderboard
query: best cheap llm

Best cheap LLM in 2026

Lowest list-price models that still clear a usability bar. For high-volume chat and batch jobs.

View ranked leaderboard
query: fastest llm

Fastest LLM in 2026 (TTFT and tok/s)

Time-to-first-token and output speed from dated snapshots. For voice, widgets, and tight loops.

View ranked leaderboard
query: best open source llm

Best open-source LLM in 2026

Open-weight models ranked on preference Elo, LiveBench, and SWE-bench. Check the license before you ship.

View ranked leaderboard
query: best llm for rag

Best LLM for RAG in 2026

Context window and input list price first. For stuffing corpora, not for coding agents.

View ranked leaderboard
query: best llm for agents

Best LLM for agents in 2026

SWE-bench and preference Elo for tool-using and computer-use stacks.

View ranked leaderboard
query: best vision llm

Best vision LLM in 2026

Models marked multimodal, ranked on Elo and latency for screenshot and document jobs.

View ranked leaderboard
query: best llm for writing

Best LLM for writing in 2026

Preference Elo first for long-form writing. Price still matters if you generate all day.

View ranked leaderboard
query: best mid range llm

Best mid-range LLM in 2026

Workhorse models under $12/1M output. For daily coding and chat without Opus or Sol invoices.

View ranked leaderboard
query: cheapest frontier llm

Cheapest frontier LLM in 2026

Models that still clear a high preference-Elo bar, ranked so list price hurts.

View ranked leaderboard
query: best local llm

Best local / open-weight LLM in 2026

Open-weight rows you can self-host. Not a laptop VRAM guide — check the license and your box.

View ranked leaderboard
query: cheapest claude alternative

Cheapest Claude alternative in 2026

Low list-price hosted models if Opus/Sonnet is too expensive. Then open the vs page against Sonnet 5.

View ranked leaderboard
query: cheapest coding llm

Cheapest coding LLM in 2026

Repo agents ranked with SWE-bench first under a mid-tier budget. For CI bots that cannot burn Opus prices.

View ranked leaderboard
query: best llm for chat

Best LLM for chat in 2026

General assistants ranked for conversation quality. Elo first, then price if you chat all day.

View ranked leaderboard
query: claude vs gpt benchmark

Claude vs GPT benchmark (2026)

Current top Anthropic vs top OpenAI model: Elo, SWE-bench, speed, and price.

View matchup benchmark
query: gemini vs gpt benchmark

Gemini vs GPT benchmark (2026)

Current top Google vs top OpenAI model on preference Elo, SWE-bench, and list price.

View matchup benchmark
query: grok vs chatgpt

Grok vs ChatGPT benchmark (2026)

Current top xAI vs top OpenAI row. Price and Elo, not a personality contest.

View matchup benchmark
query: claude vs gemini

Claude vs Gemini benchmark (2026)

Current top Anthropic vs top Google model. Coding, Elo, and output price.

View matchup benchmark
query: deepseek vs claude

DeepSeek vs Claude benchmark (2026)

Current top DeepSeek vs top Anthropic. The usual cheap-vs-frontier question.

View matchup benchmark
query: llama vs claude

Llama vs Claude benchmark (2026)

Current top Meta open-weight row versus current top Anthropic row. Elo, SWE-bench, and price.

View matchup benchmark
query: kimi vs gpt

Kimi vs GPT benchmark (2026)

Current top Moonshot row versus current top OpenAI row. Open-weight-adjacent vs closed flagship.

View matchup benchmark
query: glm vs claude

GLM vs Claude benchmark (2026)

Current top Zhipu GLM row versus current top Anthropic row. Coding and list price.

View matchup benchmark