What is LiveBench?
Contamination-resistant objective tasks. Only compare scores from the same LiveBench release.
A benchmark that tries not to leak
LiveBench publishes new questions on a schedule so models cannot memorize the test from the public internet. That is the point: static exams leak into training data. Scores are percentages on that release’s suite.
Only compare two LiveBench numbers if they share a release. A 70 on last quarter’s set is not the same as a 70 on this quarter’s set. CompareLLM keeps the snapshot date so mixed-era cells are visible.
What it does not replace
LiveBench is not human preference (Elo) and not repo-level coding (SWE-bench). A model can lead LiveBench and still lose a chat vote or fail your private eval. Use it as one dated column.
Ready to evaluate your stack?
Calculate your optimal model weights with Stack Engine or compare top models head-to-head.
