Claude Opus 4.5 is the Feb 2026 Anthropic coding reference on our seed: official mini-SWE-agent harness, $15/$75 list on that generation. Later Opus 5 cut the list price and grew the window. We keep 4.5 so “what was the SWE-bench number in February” has a URL instead of a screenshot.
Harness names matter more than brand
SWE-bench without a harness is marketing. We keep the source URL and date. If ingest later writes a different split, the changelog will show 4.5 moving — that is a new snapshot, not time travel.
Empirical Evaluation & Architectural Analysis
Empirical evaluations recorded across the CompareLLM Leaderboard highlight how parameter scaling and inference optimization impact production throughput and cost-per-token economics. Reference the interactive scorecard above for verified snapshots or inspect the dedicated claude-opus-4-5 dossier.
Do not mix Feb and August Elo
Arena pools change. A Feb Elo and an August Elo are not the same exam. Always read observed-at on the cell.
Strategic Deployment Recommendation
- Production Workloads: For enterprise workloads requiring strict reliability, cross-reference Top Coding LLMs and Cheapest High-Quality LLMs.
- Direct Pair Comparison: Check model specifications at claude-opus-4-5 specs.
- Architectural Stacks: Recommended architectural configurations can be evaluated on the CompareLLM Stack Engine.
Frequently Asked Questions
Is this an CompareLLM lab test score? No. Linked model dossiers and compare showdowns reflect dated snapshots gathered from named public evaluation benchmarks. Seed catalog entries remain transparently timestamped until daily ingest updates them.
Where can I see live comparisons for this model? Explore verified model specs at claude-opus-4-5.
What primary search query does this briefing answer? claude opus 4.5 swe-bench.
