LLM Economics
Compare hardware, models and inference setups by operating cost, token output and total cost of ownership.
Operating cost vs. TCO
Operating cost is electricity only. TCO adds acquisition (GPU or full system) and, if applicable, spreads it over the runtime.
Power & acquisition
Since no wattage is measured per run, we estimate system power: GPU TDP × count (the GPU is the main load under inference) + CPU (during GPU runs the CPU mostly waits → idle watts + 15 %; for pure CPU runs the full CPU TDP) + mainboard. The GPU overview additionally shows the bare GPU TDP. For unified chips/APUs (e.g. AMD Strix Halo / Radeon 8060S, NVIDIA GB10) the GPU TDP already includes the CPU – it is not double-counted there. Acquisition covers the full system: GPU + CPU + mainboard + RAM + PSU.
AVG, MIN, MAX, median
Economics metrics are computed per real benchmark run first and only then aggregated. Every cost metric stays bound to an actual run. For costs, MIN = cheapest run, MAX = most expensive.
Cost per 1M tokens
electricity_cost_per_1M = 1.000.000 × power_kw × price_per_kwh / tokens_per_second / 3600
Influence of price, utilization, acquisition
Higher electricity price and utilization raise running costs linearly. Acquisition spreads across more tokens as production rises, lowering total cost per token.
Performance vs. harness economics
Performance runs yield tokens/s → cost per token. Harness runs yield duration + points → cost per successful task / per point. Token-based harness costs only when real token metrics exist.
Why low token cost ≠ quality
A small, cheap model produces tokens cheaply but is not automatically best. Quality comes from the harness benchmark. Economics compares cost, not quality.
Data status & missing values
Every value is marked as measured, estimated (TDP), full price, partial price or missing. Missing mandatory values are never replaced by invented defaults.
