Created bymario-alka.dePowered bygodcore.denoob2claw.detricoma.de
Contributed byMarcel SommerOpenAI

gpt-oss-20b

Performance benchmark · measured on 04.08.2026 01:35

Benchmark-IDrun-20260804-052134-02d2cd
Timebench 3 - Kombi (Prefill + Generation)MoE20BRuntime: vLLM
Generation61,64tok/s
Prefill5.670,62tok/s
Time to First Token366,00ms
Total duration33,96s
Concurrency1parallel
Ranking in the field
40of 95 systems

Performance benchmark · Primary metric: Generation-Speed (tok/s) · 1× concurrent

This run is better than 59 % of all comparable systems.
Generation 61,6 tok/s
-29 % vs Ø 86,2
Prefill 5.670,6 tok/s
+113 % vs Ø 2.662,7
Time to First Token 366 ms
-99 % vs Ø 46.155
Distribution in the field0 – 246 tok/s
Ø 86 Median Dieser Lauf

Wie schlägt sich dieser Benchmark mit anderen Modellen?

How does this benchmark compare on other GPUs?

Same model on different hardware · 1× concurrent · Generation (tok/s)

Hardware

GPU: NVIDIA GeForce RTX 3090 Ti · 24 GB VRAM
CPU: AMD Ryzen 5 5600X 6-Core Processor
RAM: 30 GB
Mainboard: ASUSTeK COMPUTER INC. PRIME A520M-K

Setup

Runtime: vLLM
Quantization: -
Model: gpt-oss-20b

Configuration

benchmark-konfiguration — run-20260804-052134-02d2cd
# LLM-Benchmark Konfiguration # Modell : gpt-oss-20b # Engine : vLLM # Run-ID : run-20260804-052134-02d2cd # GPU : NVIDIA GeForce RTX 3090 Ti # CPU : AMD Ryzen 5 5600X 6-Core Processor # RAM : 30 GB bench@llm-benchmark:~$ /opt/vllm-gemma/venv/bin/python /opt/vllm-gemma/venv/bin/vllm serve openai/gpt-oss-20b \ --served-model-name gpt-oss-20b \ --host 192.168.41.116 \ --port 8000 \ --max-model-len 8192 \ --gpu-memory-utilization 0.95 \ --enforce-eager
Engine?Die Inferenz-Software, die das Modell ausliefert (z.B. vLLM oder llama.cpp). Sie bestimmt Geschwindigkeit, unterstuetzte Modellformate und welche Parameter ueberhaupt verfuegbar sind.vllm
Modellalias?Der Name, unter dem das Modell ueber die API angesprochen wird. Genau dieser Wert muss im Request-Feld 'model' stehen.gpt-oss-20b
Kontextlaenge?Maximale Anzahl Tokens (Eingabe + erzeugte Ausgabe zusammen), die das Modell pro Anfrage verarbeiten kann.8192
Alias?Anzeigename des Modells nach aussen (served model name), unabhaengig vom Dateinamen.gpt-oss-20b
Kontext?Groesse des Kontextfensters in Token. 0 = der beim Training verwendete Kontext des Modells.8192
GPU-Speicher?Anteil des GPU-Speichers (0 bis 1), den vLLM belegen darf. 0.92 = 92 %. Hoeher = mehr Platz fuer den KV-Cache (mehr/laengere parallele Anfragen), aber groesseres Risiko fuer 'Out of Memory'.0.95
Enforce-Eager?Schaltet die optimierte Graph-Ausfuehrung (CUDA-/HIP-Graphs) AB und rechnet Schritt fuer Schritt. Startet schneller und spart etwas VRAM, ist im laufenden Betrieb aber meist langsamer als mit Graphs.aktiv

All benchmarks of this model To leaderboard

Anzeige
Model comparison

gpt-oss-20b on various hardware

All published performance runs of this model – each bubble a variant: position = prefill (X) × generation (Y), bubble size = number of runs. Closer to the top right = faster. ★ Marked gold = this benchmark.

GPUby graphics card

2.5911.9431.2966480,008.07916.15724.236Prefill (tok/s)Generation (tok/s)NVIDIA RTX PRO 6000 Blackwell Workstation Edition - 1.995,4 tok/s Generation, 19.558 tok/s Prefill, TTFT 1.895 ms (8 Laufe)NVIDIA RTX PRO 6000 B...NVIDIA GeForce RTX 5090 - 1.890,7 tok/s Generation, 11.034 tok/s Prefill, TTFT 4.372 ms (18 Laufe)NVIDIA GeForce RTX 50...NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition - 1.835,7 tok/s Generation, 7.727 tok/s Prefill, TTFT 3.473 ms (32 Laufe)NVIDIA RTX PRO 6000 B...NVIDIA GeForce RTX 5070 Ti - 1.210,3 tok/s Generation, 4.544 tok/s Prefill, TTFT 11.765 ms (6 Laufe)NVIDIA GeForce RTX 50...AMD Radeon AI PRO R9700 - 340,5 tok/s Generation, 4.510 tok/s Prefill, TTFT 3.295 ms (2 Laufe)AMD Radeon AI PRO R97...AMD Radeon 8060S Graphics - 173,3 tok/s Generation, 1.028 tok/s Prefill, TTFT 16.849 ms (8 Laufe)AMD Radeon 8060S Grap...CPU-only - 14,5 tok/s Generation, 88 tok/s Prefill, TTFT 122.346 ms (3 Laufe)CPU-onlyNVIDIA GeForce RTX 3090 Ti - 1.113,5 tok/s Generation, 6.760 tok/s Prefill, TTFT 3.631 ms (22 Laufe) | DIESER LAUF★ NVIDIA GeForce RTX 30...
NVIDIA RTX PRO 6000 Blackwell Workstation Edition 1.995,4 tok/sNVIDIA GeForce RTX 5090 1.890,7 tok/sNVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition 1.835,7 tok/sNVIDIA GeForce RTX 5070 Ti 1.210,3 tok/s★ NVIDIA GeForce RTX 3090 Ti 1.113,5 tok/s this runAMD Radeon AI PRO R9700 340,5 tok/sAMD Radeon 8060S Graphics 173,3 tok/sCPU-only 14,5 tok/s

CPUby processor

2.5911.9431.2966480,008.07916.15724.236Prefill (tok/s)Generation (tok/s)AMD Ryzen 9 9950X 16-Core Processor - 1.995,4 tok/s Generation, 19.558 tok/s Prefill, TTFT 1.895 ms (8 Laufe)AMD Ryzen 9 9950X 16-...AMD Ryzen 7 5800X3D 8-Core Processor - 1.890,7 tok/s Generation, 11.034 tok/s Prefill, TTFT 4.372 ms (18 Laufe)AMD Ryzen 7 5800X3D 8...AMD Ryzen Threadripper PRO 9965WX 24-Cores - 1.835,7 tok/s Generation, 7.727 tok/s Prefill, TTFT 3.473 ms (32 Laufe)AMD Ryzen Threadrippe...AMD Ryzen Threadripper PRO 5975WX 32-Cores - 1.210,3 tok/s Generation, 4.544 tok/s Prefill, TTFT 11.765 ms (6 Laufe)AMD Ryzen Threadrippe...AMD Ryzen 9 8945HX with Radeon Graphics - 1.113,5 tok/s Generation, 6.303 tok/s Prefill, TTFT 4.194 ms (18 Laufe)AMD Ryzen 9 8945HX wi...AMD Ryzen Threadripper PRO 7955WX 16-Cores - 340,5 tok/s Generation, 4.510 tok/s Prefill, TTFT 3.295 ms (2 Laufe)AMD Ryzen Threadrippe...AMD RYZEN AI MAX+ 395 w/ Radeon 8060S - 173,3 tok/s Generation, 1.028 tok/s Prefill, TTFT 16.849 ms (8 Laufe)AMD RYZEN AI MAX+ 395...Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz - 14,5 tok/s Generation, 88 tok/s Prefill, TTFT 122.346 ms (3 Laufe)Intel(R) Xeon(R) CPU ...AMD Ryzen 5 5600X 6-Core Processor - 572,1 tok/s Generation, 8.817 tok/s Prefill, TTFT 1.101 ms (4 Laufe) | DIESER LAUF★ AMD Ryzen 5 5600X 6-C...
AMD Ryzen 9 9950X 16-Core Processor 1.995,4 tok/sAMD Ryzen 7 5800X3D 8-Core Processor 1.890,7 tok/sAMD Ryzen Threadripper PRO 9965WX 24-Cores 1.835,7 tok/sAMD Ryzen Threadripper PRO 5975WX 32-Cores 1.210,3 tok/sAMD Ryzen 9 8945HX with Radeon Graphics 1.113,5 tok/s★ AMD Ryzen 5 5600X 6-Core Processor 572,1 tok/s this runAMD Ryzen Threadripper PRO 7955WX 16-Cores 340,5 tok/sAMD RYZEN AI MAX+ 395 w/ Radeon 8060S 173,3 tok/sIntel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz 14,5 tok/s

MBby mainboard

2.5911.9431.2966480,008.07916.15724.236Prefill (tok/s)Generation (tok/s)ASUSTeK COMPUTER INC. ProArt X870E-CREATOR WIFI - 1.995,4 tok/s Generation, 19.558 tok/s Prefill, TTFT 1.895 ms (8 Laufe)ASUSTeK COMPUTER INC....ASUSTeK COMPUTER INC. ROG STRIX B550-A GAMING - 1.890,7 tok/s Generation, 11.034 tok/s Prefill, TTFT 4.372 ms (18 Laufe)ASUSTeK COMPUTER INC....ASUSTeK COMPUTER INC. Pro WS WRX90E-SAGE SE - 1.835,7 tok/s Generation, 7.538 tok/s Prefill, TTFT 3.463 ms (34 Laufe)ASUSTeK COMPUTER INC....ASUSTeK COMPUTER INC. Pro WS WRX80E-SAGE SE WIFI - 1.210,3 tok/s Generation, 4.544 tok/s Prefill, TTFT 11.765 ms (6 Laufe)ASUSTeK COMPUTER INC....Meigao Innovation Technology (Shen Zhen) Co., Ltd DRFXL (MotherBoard Series) - 1.113,5 tok/s Generation, 6.303 tok/s Prefill, TTFT 4.194 ms (18 Laufe)Meigao Innovation Tec...Bosgame AXB35-02 (BeyondMax Series) - 173,3 tok/s Generation, 1.028 tok/s Prefill, TTFT 16.849 ms (8 Laufe)Bosgame AXB35-02 (Bey...Dell Inc. PowerEdge R820 - 14,5 tok/s Generation, 88 tok/s Prefill, TTFT 122.346 ms (3 Laufe)Dell Inc. PowerEdge R...ASUSTeK COMPUTER INC. PRIME A520M-K - 572,1 tok/s Generation, 8.817 tok/s Prefill, TTFT 1.101 ms (4 Laufe) | DIESER LAUF★ ASUSTeK COMPUTER INC....
ASUSTeK COMPUTER INC. ProArt X870E-CREATOR WIFI 1.995,4 tok/sASUSTeK COMPUTER INC. ROG STRIX B550-A GAMING 1.890,7 tok/sASUSTeK COMPUTER INC. Pro WS WRX90E-SAGE SE 1.835,7 tok/sASUSTeK COMPUTER INC. Pro WS WRX80E-SAGE SE WIFI 1.210,3 tok/sMeigao Innovation Technology (Shen Zhen) Co., Ltd DRFXL (MotherBoard Series) 1.113,5 tok/s★ ASUSTeK COMPUTER INC. PRIME A520M-K 572,1 tok/s this runBosgame AXB35-02 (BeyondMax Series) 173,3 tok/sDell Inc. PowerEdge R820 14,5 tok/s

ENGby engine

2.3052.0121.7201.4271.1346.3277.6969.06610.435Prefill (tok/s)Generation (tok/s)llama.cpp - 1.995,4 tok/s Generation, 7.288 tok/s Prefill, TTFT 9.095 ms (65 Laufe)llama.cppvLLM - 1.443,9 tok/s Generation, 9.473 tok/s Prefill, TTFT 8.022 ms (34 Laufe) | DIESER LAUF★ vLLM
llama.cpp 1.995,4 tok/s★ vLLM 1.443,9 tok/s this run

DRVby driver

2.1952.0951.9951.8961.7967.5567.8788.2008.521Prefill (tok/s)Generation (tok/s)unbekannt - 1.995,4 tok/s Generation, 8.039 tok/s Prefill, TTFT 8.726 ms (99 Laufe)unbekannt
unbekannt 1.995,4 tok/s
💰 Economics

Economics of this run

Operating cost, TCO and comparison with the next-best runs of the same model at identical concurrency (1× concurrent). Methodology →

⚙️ ConfigurationAll metrics and charts below follow these settings – based on a 24-month runtime.Save to URLReset
⚡ Electricity price EUR/kWh
⚙️ System utilization 100 %
🖥️ Acquisition EUR
🔌 Idle 25 W
⚡ TDP 450 W
☁️ External LLM (API)
Electricity0.30 EUR/kWh
Avg power (incl. idle)450 W estimated (TDP)GPU 450 W full load
Avg cost / hourEUR 0.14
Electricity / 1M tokensEUR 0.61
Token / kWh493.12K
Acquisition (system)EUR 1,149 partial priceGPU EUR 999 · PSU EUR 150
Electricity (2 years)
TCO (2 years)EUR 3,514
Output tokens (2 years)3.89B
☁️ External LLM (API) – comparison
External LLM cost (2 years)
Savings vs. external (2 years)

All values above and the charts below take the configured system utilization into account: at X% the system generates only X% of the time, the rest it idles (25 W). Cost per hour drops (more idle), cost per token rises.

Cost over 2 years – electricity only

Cost over 2 years – incl. acquisition (TCO)

Speed vs. tokens per euro

Euro per 1M tokens

Comparison vs. API – economics per benchmark

gpt-oss-20bNVIDIA GeForce RTX 3090 Tigpt-oss-20bNVIDIA GeForce RTX 5090gpt-oss-20bNVIDIA RTX PRO 6000 Blackwell Workstation Editiongpt-oss-20bNVIDIA GeForce RTX 5090
Electricity cost (24 mo.)
Acquisition cost
Total cost (TCO)
Generated tokens (24 mo.)
Token price via API
Break-even point (days)
Result (savings / extra cost)

Comparison with up to 3 next-best runs of this model at the same concurrency (at least one on different hardware). Power = GPU TDP + CPU (idle + 15 %) + board (estimated), acquisition = full system (GPU + CPU + board + RAM + PSU), prices = stored market prices.

Contributed by

Marcel Sommer

@marcelsommer