CompareLLM Practical Guide
Updated 2026-08-16How to read LLM prices ($/1M tokens)
List price is not your invoice. Cache, retries, long context, and output tokens move the real bill.
1Section 1
Two numbers, not one
Input $/1M is what you pay to send the prompt (and usually the RAG dump). Output $/1M is what you pay for the reply. Coding agents spend more on output. RAG spends more on input.
CompareLLM pair pages sketch three list-price scenarios. They are not a finance model. Cache hits, batch APIs, and long-context surcharges are not in the cell.
2Section 2
Why a “cheap” model can cost more
If it retries three times or dumps a 100k context on every turn, the cheaper row loses. Compare the pair page, then measure on your own traffic.
Ready to evaluate your stack?
Calculate your optimal model weights with Stack Engine or compare top models head-to-head.
