CompareLLM
CompareLLM.ai
Live
LeaderboardModelsCompareStacksBest ofGuidesNewsMethod
…
CompareLLM
CompareLLM.ai
Precision Benchmarks

Programmatic, dated AI model benchmarks, head-to-head comparisons, and Stack Engine presets.

Daily ingest · 06:00 UTC

Analytics & Benchmarks

  • AI Model Leaderboard
  • Head-to-Head Compare Hub
  • Models Directory
  • Stack Engine Presets
  • Frontier Models
  • Open Weights Catalog

Guides & Intent Lists

  • Best LLM Lists (2026)
  • Best Coding LLM
  • Best Cheap LLM
  • Fastest Low-Latency LLM
  • Claude vs GPT Benchmark
  • What is Elo?
  • Methodology Guides
  • News & Dispatches

Transparency & API

  • Evaluation Methodology
  • Benchmark Changelog
  • Public JSON API
  • llms.txt Specification
  • Privacy Policy
  • Sign In / Account

© 2026 CompareLLM. Public benchmark data aggregated from Arena Elo, LiveBench, SWE-bench & OpenRouter.

Every score has a dated snapshot.

  1. Home
  2. Stack Engine
  3. open-weight reasoning
Deterministic Stack Engine

Best AI Stack for open-weight reasoning

Highest reasoning signal among models you can self-host or buy as open weights.

Objective Weight Distribution
Elo:30%LiveBench:30%GPQA:25%Out $:15%
Recommended PickScore 78.0 / 100

DeepSeek V4 Pro

DeepSeek · Open weights

elo p78livebench p82gpqa p81output price per m p65
Output: $0.87/1MSWE-bench: 71.6%
View full model fact sheet

Full Stack Leaderboard

Ranked alternatives optimized for open-weight reasoning based on objective benchmark weighting.

#2DeepSeek V4 Flash(DeepSeek)
Score 69.8
eloP61livebenchP71gpqaP71output price per mP83
Rank #2 in stackvs #1
#3Qwen QwQ 32B(Alibaba)
Score 69.2
eloP66livebenchP72gpqaP77output price per mP57
Rank #3 in stackvs #1
#4GLM-5.2(Zhipu)
Score 61.6
eloP64livebenchP59gpqaP61output price per mP63
Rank #4 in stackvs #1
#5MiniMax M2.5(MiniMax)
Score 53.1
eloP56livebenchP54gpqaP42output price per mP64
Rank #5 in stackvs #1
#6Llama 4 Maverick(Meta)
Score 47.1
eloP51livebenchP43gpqaP36output price per mP66
Rank #6 in stackvs #1
#7Qwen3 235B(Alibaba)
Score 45.4
eloP39livebenchP47gpqaP49output price per mP49
Rank #7 in stackvs #1
#8Qwen 2.5 Coder 32B(Alibaba)
Score 45.1
eloP35livebenchP56gpqaP30output price per mP69
Rank #8 in stackvs #1
#9DeepSeek R1(DeepSeek)
Score 36.0
eloP21livebenchP26gpqaP64output price per mP39
Rank #9 in stackvs #1
#10Llama 4 Scout(Meta)
Score 35.5
eloP32livebenchP28gpqaP21output price per mP82
Rank #10 in stackvs #1