CompareLLM.ai
Live
Leaderboard
Models
Compare
Best of & Stacks
Research & News
…
CompareLLM.ai
Precision Benchmarks

Programmatic, dated AI model benchmarks, head-to-head comparisons, and Stack Engine presets.

Daily ingest · 06:00 UTC

Analytics & Benchmarks

  • AI Model Leaderboard
  • Head-to-Head Compare Hub
  • Models Directory
  • Stack Engine Presets
  • Frontier Models
  • Open Weights Catalog

Guides & Intent Lists

  • Best LLM Lists (2026)
  • Best Coding LLM
  • Best Cheap LLM
  • Fastest Low-Latency LLM
  • Claude vs GPT Benchmark
  • What is Elo?
  • Methodology Guides
  • News & Dispatches

Transparency & API

  • Evaluation Methodology
  • Benchmark Changelog
  • Public JSON API
  • llms.txt Specification
  • Privacy Policy
  • Sign In / Account

© 2026 CompareLLM. Public benchmark data aggregated from Arena Elo, LiveBench, SWE-bench & OpenRouter.

Every score has a dated snapshot.

Theme:
Currency:
  1. Home
  2. News
  3. Claude 3.7 Sonnet stays as a historical hybrid-reasoning baseline
Claude 3.7 Sonnet stays as a historical hybrid-reasoning baseline
newsVerified Dispatch
CompareLLM Intelligence Desk·Feb 25, 2025

Claude 3.7 Sonnet stays as a historical hybrid-reasoning baseline

2025 hybrid-reasoning Sonnet. Deprecated for buying, kept for old compare URLs and “what did 3.7 score?”

Search intent: claude 3.7 sonnet benchmark

Evaluated Models (1):
Claude 3.7 SonnetAnthropic
Share Analysis:
WhatsAppTelegramXLinkedInReddit
Verified Benchmark Scorecard

Claude 3.7 Sonnet Performance & Efficiency Dashboard

Live benchmark scores, throughput speeds, and token pricing with dynamic peer comparison.

LMSYS Arena Elo1,362Claude 3.7 Sonnet
SWE-bench Verified70.3%Coding resolve %
Throughput Speed85 tok/sGeneration rate
Output Price / 1M$15.00List rate
Interactive Model Comparison:1 model selected
Claude 3.7 SonnetArticle
LMSYS Arena EloOverall human preference
quality
Claude 3.7 Sonnet
1,362
SWE-bench VerifiedGitHub issue resolve %
quality
Claude 3.7 Sonnet
70.3%
GPQA DiamondPhD-level science reasoning %
quality
Claude 3.7 Sonnet
78.5%
LiveBenchContamination-free reasoning
quality
Claude 3.7 Sonnet
58.4%
Throughput (Speed)Output tokens / second
speed
Claude 3.7 Sonnet
85 tok/s
Time-to-First-TokenInitial latency (ms)
speed
Claude 3.7 Sonnet
240 ms
Output Token Price$ per 1M reply tokens
price
Claude 3.7 Sonnet
$15/1M
Input Token Price$ per 1M prompt tokens
price
Claude 3.7 Sonnet
$3/1M
Context WindowMax token capacity
capacity
Claude 3.7 Sonnet
200k
Benchmark / Metric
Claude 3.7 SonnetAnthropic · Article Model
LMSYS Arena EloOverall human preference
1,362
SWE-bench VerifiedGitHub issue resolve %
70.3%
GPQA DiamondPhD-level science reasoning %
78.5%
LiveBenchContamination-free reasoning
58.4%
Throughput (Speed)Output tokens / second
85 tok/s
Time-to-First-TokenInitial latency (ms)
240 ms
Output Token Price$ per 1M reply tokens
$15/1M
Input Token Price$ per 1M prompt tokens
$3/1M
Context WindowMax token capacity
200k
Executive Key Takeaway
2025 hybrid-reasoning Sonnet. Deprecated for buying, kept for old compare URLs and “what did 3.7 score?”

Claude 3.7 Sonnet is marked deprecated in this catalog. It is the 2025 hybrid-reasoning Sonnet people still Google. Deleting it would 404 a year of vs pages. Buying it in 2026 is usually a mistake; citing it as a baseline is not. The model page says deprecated in the header on purpose.

Deprecated is not noindex-by-default

We still index historical rows that have enough dated metrics. We noindex thin pairs. 3.7 vs a modern Flash can stay visible if three metrics overlap; otherwise it is noindex HTML for humans with old links.

Empirical Evaluation & Architectural Analysis

Empirical evaluations recorded across the CompareLLM Leaderboard highlight how parameter scaling and inference optimization impact production throughput and cost-per-token economics. Reference the interactive scorecard above for verified snapshots or inspect the dedicated claude-3-7-sonnet dossier.

Do not put 3.7 in “best coding llm”

Stack Engine presets skip the wrong job. Best-coding will not crown a deprecated 2025 row unless the dated SWE-bench still wins — and it should not against 2026 SKUs.

Strategic Deployment Recommendation

  • Production Workloads: For enterprise workloads requiring strict reliability, cross-reference Top Coding LLMs and Cheapest High-Quality LLMs.
  • Direct Pair Comparison: Check model specifications at claude-3-7-sonnet specs.
  • Architectural Stacks: Recommended architectural configurations can be evaluated on the CompareLLM Stack Engine.

Frequently Asked Questions

Is this an CompareLLM lab test score? No. Linked model dossiers and compare showdowns reflect dated snapshots gathered from named public evaluation benchmarks. Seed catalog entries remain transparently timestamped until daily ingest updates them.

Where can I see live comparisons for this model? Explore verified model specs at claude-3-7-sonnet.

What primary search query does this briefing answer? claude 3.7 sonnet benchmark.

Share Analysis:
WhatsAppTelegramXLinkedInReddit

Head-to-Head Showdowns for Mentioned Models

Benchmark MatchupClaude 3.7 Sonnet vs GPT-5
Benchmark MatchupClaude 3.7 Sonnet vs GPT-5 mini
Back to All DispatchesExplore Comparison Matrix