LLM Economics
Compare hardware, models and inference setups by operating cost, token output and total cost of ownership.
Settings are stored in your browser and shared via the URL. Power draw is estimated from GPU TDP × count since no per-run measurement exists. Methodology →
Operating cost vs. TCO
Operating cost is electricity only. TCO adds acquisition (GPU or full system) and, if applicable, spreads it over the runtime.
Power draw (GPU TDP)
Since no wattage is measured per run, we use the GPU's manufacturer TDP × count as the estimated inference load – the GPU is the dominant consumer. CPU and mainboard are deliberately not added: otherwise a 450 W GPU would show a wrong system TDP, and for APUs (e.g. AMD Strix Halo / Radeon 8060S) the GPU and CPU of the same chip would be double-counted. Pure CPU nodes use the CPU TDP (+ board).
AVG, MIN, MAX, median
Economics metrics are computed per real benchmark run first and only then aggregated. Every cost metric stays bound to an actual run. For costs, MIN = cheapest run, MAX = most expensive.
Cost per 1M tokens
electricity_cost_per_1M = 1.000.000 × power_kw × price_per_kwh / tokens_per_second / 3600
Influence of price, utilization, acquisition
Higher electricity price and utilization raise running costs linearly. Acquisition spreads across more tokens as production rises, lowering total cost per token.
Performance vs. harness economics
Performance runs yield tokens/s → cost per token. Harness runs yield duration + points → cost per successful task / per point. Token-based harness costs only when real token metrics exist.
Why low token cost ≠ quality
A small, cheap model produces tokens cheaply but is not automatically best. Quality comes from the harness benchmark. Economics compares cost, not quality.
Data status & missing values
Every value is marked as measured, estimated (TDP), full price, partial price or missing. Missing mandatory values are never replaced by invented defaults.
