Created bymario-alka.dePowered bygodcore.denoob2claw.detricoma.de
💰
LLM Economics

LLM Economics

Compare hardware, models and inference setups by operating cost, token output and total cost of ownership.

💸 Operating cost🖥️ TCO⚡ Tokens / euro🔋 Tokens / kWh
Reset

Settings are stored in your browser and shared via the URL. Power draw estimated from GPU TDP + CPU (idle + 15 %) + board. Methodology →

GPU economics

21 GPU configurations

Aggregated per configuration (incl. GPU count) – economics computed per run, then AVG/MIN/MAX. Number after the GPU = runs · models.

GPU configurations chart

Bars per configuration – switch between cost per 1M tokens, tokens per euro and throughput, plus best values (top model) and AVG. Below each GPU is the model that reaches the value. Cheaper/faster = longer bar.

List of graphics cards

#GPU configurationGen tok/s iCost/hEUR / 1M tok iTok/€Tok/kWhTCO i
1 NVIDIA GB10 (DGX Spark)i128 GB VRAM · EUR 4,000 · 140 W GPU-TDP
903.3 tok/s
gemma-4-E2B-it
EUR 0.042 EUR 0.013 77.43M 23.23M EUR 736
2 AMD Radeon PRO W7800 48GBi48 GB VRAM · EUR 2,808 · 260 W GPU-TDP
702.9 tok/s
gemma-4-E2B-it
EUR 0.081 EUR 0.032 31.24M 9.37M EUR 1,419
3 AMD Radeon 8060S GraphicsiEUR 3,000 · 120 W GPU-TDP
326.9 tok/s
Nemotron-Cascade-2-30B-A3B
EUR 0.036 EUR 0.031 32.69M 9.81M EUR 631
4 NVIDIA RTX A6000i48 GB VRAM · EUR 3,000 · 300 W GPU-TDP
1,122.5 tok/s
gemma-4-E2B-it
EUR 0.11 EUR 0.028 36.31M 10.89M EUR 1,950
5 3× NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Editioni288 GB VRAM · EUR 39,000 · 900 W GPU-TDP
2,143.6 tok/s
gemma-4-E2B-it
EUR 0.29 EUR 0.038 26.18M 7.85M EUR 5,033
6 2× NVIDIA RTX A6000i96 GB VRAM · EUR 6,000 · 600 W GPU-TDP
1,568.0 tok/s
gemma-4-E2B-it
EUR 0.20 EUR 0.036 28.04M 8.41M EUR 3,527
7 AMD Radeon PRO W7900 Dual Sloti48 GB VRAM · EUR 3,911 · 295 W GPU-TDP
343.9 tok/s
gpt-oss-20b
EUR 0.091 EUR 0.074 13.53M 4.06M EUR 1,603
8 2× AMD Radeon PRO W7800 48GBi96 GB VRAM · EUR 5,616 · 520 W GPU-TDP
323.1 tok/s
Meta-Llama-3.1-8B-Instruct
EUR 0.16 EUR 0.14 7.32M 2.19M EUR 2,786
9 NVIDIA Tesla P100 PCIe 16GBi16 GB VRAM · EUR 70 · 250 W GPU-TDP
111.1 tok/s
gpt-oss-20b
EUR 0.075 EUR 0.19 5.33M 1.60M EUR 1,314
10 2× AMD Radeon PRO W7900 Dual Sloti96 GB VRAM · EUR 7,822 · 590 W GPU-TDP
344.1 tok/s
gpt-oss-20b
EUR 0.18 EUR 0.15 6.88M 2.06M EUR 3,154
11 NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Editioni96 GB VRAM · EUR 13,000 · 300 W GPU-TDP
28.8 tok/s
Qwen3.6-27B
EUR 0.10 EUR 0.95 1.05M 315.38K EUR 1,728
12 NVIDIA GeForce RTX 5070 Tii16 GB VRAM · EUR 994 · 300 W GPU-TDP
1,744.0 tok/s
gemma-4-E2B-it
EUR 0.093 EUR 0.015 67.51M 20.25M EUR 1,577
13 4× AMD Radeon AI PRO R9700i128 GB VRAM · EUR 5,600 · 1,200 W GPU-TDP
290.3 tok/s
Nemotron-3.5-Lightning-30B-A3B
EUR 0.38 EUR 0.36 2.74M 822.17K EUR 6,680
14 AMD Radeon AI PRO R9700i32 GB VRAM · EUR 1,400 · 300 W GPU-TDP
5.6 tok/s
Qwen2.5-Coder-32B-Instruct
EUR 0.11 EUR 5.55 180.16K 54.05K EUR 1,950
15 Intel Arc Pro B70i32 GB VRAM · EUR 1,258 · 230 W GPU-TDP
826.2 tok/s
gemma-4-E2B-it
EUR 0.069 EUR 0.023 43.11M 12.93M EUR 1,209
16 NVIDIA GeForce RTX 3090 Tii24 GB VRAM · EUR 999 · 450 W GPU-TDP
1,492.8 tok/s
gemma-4-E2B-it
EUR 0.14 EUR 0.027 37.53M 11.26M EUR 2,365
17 2× NVIDIA GeForce RTX 2060i12 GB VRAM · EUR 220 · 320 W GPU-TDP
133.7 tok/s
gemma-4-E4B-it
EUR 0.10 EUR 0.20 5.01M 1.50M EUR 1,682
18 3× AMD Radeon AI PRO R9700i96 GB VRAM · EUR 4,200 · 900 W GPU-TDP
667.0 tok/s
Nemotron-3-Nano-4B
EUR 0.29 EUR 0.12 8.24M 2.47M EUR 5,104
19 NVIDIA GeForce RTX 5090i32 GB VRAM · EUR 3,300 · 575 W GPU-TDP
2,491.2 tok/s
gemma-4-E2B-it
EUR 0.19 EUR 0.021 48.10M 14.43M EUR 3,267
20 NVIDIA RTX PRO 6000 Blackwell Workstation Editioni96 GB VRAM · EUR 13,000 · 600 W GPU-TDP
2,414.4 tok/s
gemma-4-E2B-it
EUR 0.19 EUR 0.022 45.01M 13.50M EUR 3,384
21 Tesla V100-PCIE-32GBi32 GB VRAM · no price · –
361.8 tok/s
gemma-4-E2B-it

MIN = lowest cost (highest speed / lowest draw), MAX = highest cost. Every cost metric is bound to a real run (computed per run, then aggregated). Prices = GPU acquisition (unit price × count).

GPUby GPU

3.1052.3421.57981653,103.5237.04710.570Prefill (tok/s)Generation (tok/s)NVIDIA GeForce RTX 5090 - 2.491,2 tok/s Generation, 4.548 tok/s Prefill, TTFT 65.147 ms (218 Laufe)NVIDIA GeForce RTX 50...NVIDIA RTX PRO 6000 Blackwell Workstation Edition - 2.414,4 tok/s Generation, 6.972 tok/s Prefill, TTFT 27.182 ms (244 Laufe)NVIDIA RTX PRO 6000 B...NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition - 2.143,6 tok/s Generation, 6.219 tok/s Prefill, TTFT 9.221 ms (512 Laufe)NVIDIA RTX PRO 6000 B...NVIDIA GeForce RTX 5070 Ti - 1.744,0 tok/s Generation, 1.933 tok/s Prefill, TTFT 57.650 ms (222 Laufe)NVIDIA GeForce RTX 50...NVIDIA RTX A6000 - 1.568,0 tok/s Generation, 8.649 tok/s Prefill, TTFT 4.699 ms (992 Laufe)NVIDIA RTX A6000NVIDIA GeForce RTX 3090 Ti - 1.492,8 tok/s Generation, 3.396 tok/s Prefill, TTFT 81.183 ms (264 Laufe)NVIDIA GeForce RTX 30...NVIDIA GB10 (DGX Spark) - 903,3 tok/s Generation, 3.550 tok/s Prefill, TTFT 8.759 ms (82 Laufe)NVIDIA GB10 (DGX Spar...Intel Arc Pro B70 - 826,2 tok/s Generation, 858 tok/s Prefill, TTFT 45.853 ms (32 Laufe)Intel Arc Pro B70AMD Radeon PRO W7800 48GB - 702,9 tok/s Generation, 2.969 tok/s Prefill, TTFT 5.830 ms (149 Laufe)AMD Radeon PRO W7800 ...AMD Radeon AI PRO R9700 - 667,0 tok/s Generation, 1.970 tok/s Prefill, TTFT 65.510 ms (855 Laufe)AMD Radeon AI PRO R97...
NVIDIA GeForce RTX 5090 2.491,2 tok/sNVIDIA RTX PRO 6000 Blackwell Workstation Edition 2.414,4 tok/sNVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition 2.143,6 tok/sNVIDIA GeForce RTX 5070 Ti 1.744,0 tok/sNVIDIA RTX A6000 1.568,0 tok/sNVIDIA GeForce RTX 3090 Ti 1.492,8 tok/sNVIDIA GB10 (DGX Spark) 903,3 tok/sIntel Arc Pro B70 826,2 tok/sAMD Radeon PRO W7800 48GB 702,9 tok/sAMD Radeon AI PRO R9700 667,0 tok/s

Speed vs. tokens per euro

Speed vs. tokens per euro – bubble size = VRAM. Top-right = fast and economical.

Over 2 years projected

Projection per GPU configuration for the selected period: electricity cost and generated output tokens (per-run AVG × effective operating hours). Adjust period/utilization above.

Cost vs. generated tokens (2 years)

Cost over 2 years

Generated tokens over 2 years

Cost & tokens grow linearly with runtime. Power = GPU TDP + CPU (idle + 15 %) + board (estimated). Charts show the top 12 most economical configurations.