CompareLLM
CompareLLM.ai
Live
LeaderboardModelsCompareStacksBest ofGuidesNewsMethod
…
CompareLLM
CompareLLM.ai
Precision Benchmarks

Programmatic, dated AI model benchmarks, head-to-head comparisons, and Stack Engine presets.

Daily ingest · 06:00 UTC

Analytics & Benchmarks

  • AI Model Leaderboard
  • Head-to-Head Compare Hub
  • Models Directory
  • Stack Engine Presets
  • Frontier Models
  • Open Weights Catalog

Guides & Intent Lists

  • Best LLM Lists (2026)
  • Best Coding LLM
  • Best Cheap LLM
  • Fastest Low-Latency LLM
  • Claude vs GPT Benchmark
  • What is Elo?
  • Methodology Guides
  • News & Dispatches

Transparency & API

  • Evaluation Methodology
  • Benchmark Changelog
  • Public JSON API
  • llms.txt Specification
  • Privacy Policy
  • Sign In / Account

© 2026 CompareLLM. Public benchmark data aggregated from Arena Elo, LiveBench, SWE-bench & OpenRouter.

Every score has a dated snapshot.

Open Weights & Self-Hostable

Open-Weight AI Models

16 open-weights models ranked by preference Elo and SWE-bench. Read our open vs closed evaluation guide.

#1DeepSeek V4 ProDeepSeek
Compare vs all

Open-weight-adjacent DeepSeek flagship. High reasoning density per dollar.

Elo:1,536
SWE-bench:71.6%
Out:$0.87/1M
#2Qwen QwQ 32BAlibaba
Compare vs all

Alibaba specialized open reasoning model competing with frontier closed reasoning models.

Elo:1,495
SWE-bench:67.5%
Out:$1.2/1M
#3GLM-5.2Zhipu
Compare vs all

Zhipu flagship. Strong Chinese/English coding and agents.

Elo:1,492
SWE-bench:69.3%
Out:$0.9680000000000001/1M
#4DeepSeek V4 FlashDeepSeek
Compare vs all

DeepSeek Jul 31 2026 price-performance SKU. Public listings put it near $0.14/$0.28 per 1M tokens.

Elo:1,490
SWE-bench:71.2%
Out:$0.28/1M
#5MiniMax M2.5MiniMax
Compare vs all

MiniMax coding model. Tied near the top of official SWE-bench bash-only in Feb 2026.

Elo:1,478
SWE-bench:75.8%
Out:$0.8999999999999999/1M
#6Llama 4 MaverickMeta
Compare vs all

Meta natively multimodal open-weight flagship.

Elo:1,468
SWE-bench:57.1%
Out:$0.7999999999999999/1M
#7Qwen3 235BAlibaba
Compare vs all

Open-weight Qwen3 mixture-of-experts.

Elo:1,440
SWE-bench:59.4%
Out:$1.8199999999999998/1M
#8Qwen 2.5 Coder 32BAlibaba
Compare vs all

Alibaba dedicated open-weight code generation model with near-frontier SWE-bench Verified coding capability.

Elo:1,425
SWE-bench:65.2%
Out:$0.72/1M
#9Llama 4 ScoutMeta
Compare vs all

Meta open-weight Llama 4 long-context sibling of Maverick. Common public API lists sit near $0.08–$0.30 / $0.30–$0.70 per 1M; we store a conservative hosted list until OpenRouter overwrites.

Elo:1,412
SWE-bench:52.4%
Out:$0.3/1M
#10Qwen3.8 27BAlibaba
Compare vs all

Alibaba open-weight 27B drop dated Aug 14 2026. Dense enough to self-host; not a frontier MoE.

Elo:1,398
SWE-bench:58.8%
Out:$3.1999999999999997/1M
#11Llama 3.1 405BMeta
Compare vs all

Meta flagship open-weight 405B dense foundation model with 128k context window.

Elo:1,370
SWE-bench:58.4%
Out:$3.5/1M
#12DeepSeek Coder V2DeepSeek
Compare vs all

DeepSeek open-weight Mixture-of-Experts coding model supporting 338 programming languages and 128k context.

Elo:1,365
SWE-bench:60.5%
Out:$0.28/1M
#13DeepSeek R1DeepSeek
Compare vs all

Open-weights reasoning model trained with large-scale RL.

Elo:1,358
SWE-bench:65.2%
Out:$2.5/1M
#14DeepSeek V3DeepSeek
Compare vs all

Prior DeepSeek flagship. Baseline for v3 vs v4.

Elo:1,310
SWE-bench:48.6%
Out:$1.0287/1M
#15Llama 3.3 70BMeta
Compare vs all

Previous Meta 70B open-weight workhorse.

Elo:1,285
SWE-bench:45.1%
Out:$0.32/1M
#16Llama 3.1 70BMeta
Compare vs all

Llama 3.1 70B instruct. Predecessor to 3.3 70B and Llama 4.

Elo:1,240
SWE-bench:40.2%
Out:$0.39999999999999997/1M