Created bymario-alka.dePowered bygodcore.denoob2claw.detricoma.de
Contributed byMario AlkaQwen (Alibaba)

Qwen3.6-27B

Performance benchmark · measured on 20.07.2026 15:01

Benchmark-IDrun-20260722-165903-329c7f
Dense27BRuntime: godclaw
Generation28,80tok/s
Prefill599,42tok/s
Time to First Token121,70ms
Total duration107,14s
Concurrency1parallel
Ranking in the field
680of 1161 systems

Performance benchmark · Primary metric: Generation-Speed (tok/s) · 1× concurrent

This run is better than 41 % of all comparable systems.
Generation 28,8 tok/s
-61 % vs Ø 74,3
Prefill 599,4 tok/s
-74 % vs Ø 2.324,7
Time to First Token 122 ms
-100 % vs Ø 35.657
Distribution in the field0 – 405 tok/s
Ø 74 Median Dieser Lauf

Wie schlägt sich dieser Benchmark mit anderen Modellen?

How does this benchmark compare on other GPUs?

Same model on different hardware · 1× concurrent · Generation (tok/s)

Hardware

GPU: NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition · 96 GB VRAM
CPU: 32x AMD Ryzen 9 9950X 16-Core Processor
RAM: 92 GB

Setup

Runtime: godclaw
Quantization: -
Driver: NVIDIA 580.159.03 / CUDA 13.0
Model: Qwen3.6-27B

Configuration

benchmark-konfiguration — run-20260722-165903-329c7f
# LLM-Benchmark Konfiguration # Modell : Qwen3.6-27B # Run-ID : run-20260722-165903-329c7f # GPU : NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition # CPU : 32x AMD Ryzen 9 9950X 16-Core Processor # RAM : 92 GB bench@llm-benchmark:~$ cat benchmark.conf Konfigurationspfad Qwen/Qwen3.6-27B Engine vllm Kontextlaenge 65536 Python-Modul vllm.entrypoints.openai.api_server Host 0.0.0.0 Port 8000 GPU-Speicher 0.90 Nur-Sprachmodell aktiv Reasoning-Parser qwen3 Tool-Parser qwen3_coder
Engine?Die Inferenz-Software, die das Modell ausliefert (z.B. vLLM oder llama.cpp). Sie bestimmt Geschwindigkeit, unterstuetzte Modellformate und welche Parameter ueberhaupt verfuegbar sind.vllm
Modellalias?Der Name, unter dem das Modell ueber die API angesprochen wird. Genau dieser Wert muss im Request-Feld 'model' stehen.Qwen/Qwen3.6-27B
Kontextlaenge?Maximale Anzahl Tokens (Eingabe + erzeugte Ausgabe zusammen), die das Modell pro Anfrage verarbeiten kann.65536
Python-Modul?Startet vLLM als Python-Modul (python -m vllm.entrypoints.openai.api_server). Der folgende Wert ist der Modulname.vllm.entrypoints.openai.api_server
Modellpfad?Pfad oder HuggingFace-Repo des zu ladenden Modells (Positionsargument bei "vllm serve" oder --model).Qwen/Qwen3.6-27B
Alias?Name, unter dem vLLM das Modell in der API anbietet. Der Client muss im Request-Feld 'model' exakt diesen Namen senden.Qwen/Qwen3.6-27B
Host?Netzwerk-Adresse, an die der API-Server bindet. 0.0.0.0 = auf allen Netzwerk-Schnittstellen erreichbar.0.0.0.0
Port?TCP-Port, auf dem der API-Server lauscht (bei vLLM standardmaessig 8000).8000
GPU-Speicher?Anteil des GPU-Speichers (0 bis 1), den vLLM belegen darf. 0.92 = 92 %. Hoeher = mehr Platz fuer den KV-Cache (mehr/laengere parallele Anfragen), aber groesseres Risiko fuer 'Out of Memory'.0.90
Kontext?Maximale Kontextlaenge in Tokens (Eingabe + Ausgabe zusammen). Begrenzt, wie gross eine einzelne Anfrage sein darf.65536
Generation-Config?Quelle der Default-Sampling-Parameter (auto/vllm oder Pfad). Steuert die vom Modell vorgegebenen Defaults fuer temperature/top_p usw.vllm
Nur-Sprachmodell?Laedt bei multimodalen Modellen nur den Sprachteil. Spart VRAM, deaktiviert aber Bild-/Audio-Eingabe.aktiv
Reasoning-Parser?Trennt den Denk-/Reasoning-Teil der Antwort vom eigentlichen Inhalt (z.B. bei Modellen mit <think>-Bloecken).qwen3
Auto-Tool-Choice?Erlaubt dem Modell, Tools (Function-Calling) selbststaendig auszuwaehlen. MUSS aktiv sein, damit 'tools' im Request akzeptiert werden - sonst lehnt vLLM sie mit HTTP 400 ab.aktiv
Tool-Parser?Bestimmt, wie vLLM Tool-Aufrufe aus der Modellantwort ausliest (z.B. hermes, mistral, llama3_json). Muss zum Modell passen.qwen3_coder

All benchmarks of this model To leaderboard

Anzeige
Model comparison

Qwen3.6-27B on various hardware

All published performance runs of this model – each bubble a variant: position = prefill (X) × generation (Y), bubble size = number of runs. Closer to the top right = faster. ★ Marked gold = this benchmark.

GPUby graphics card

7045283521760,001.1052.2103.315Prefill (tok/s)Generation (tok/s)NVIDIA RTX PRO 6000 Blackwell Workstation Edition - 546,9 tok/s Generation, 2.739 tok/s Prefill, TTFT 9.126 ms (3 Laufe)NVIDIA RTX PRO 6000 B...NVIDIA GeForce RTX 5090 - 532,5 tok/s Generation, 2.539 tok/s Prefill, TTFT 9.812 ms (3 Laufe)NVIDIA GeForce RTX 50...NVIDIA GeForce RTX 3090 Ti - 264,2 tok/s Generation, 1.458 tok/s Prefill, TTFT 19.358 ms (3 Laufe)NVIDIA GeForce RTX 30...AMD Radeon AI PRO R9700 - 165,4 tok/s Generation, 1.380 tok/s Prefill, TTFT 28.271 ms (9 Laufe)AMD Radeon AI PRO R97...NVIDIA GeForce RTX 5070 Ti - 57,2 tok/s Generation, 447 tok/s Prefill, TTFT 80.318 ms (3 Laufe)NVIDIA GeForce RTX 50...AMD Radeon 8060S Graphics - 36,1 tok/s Generation, 536 tok/s Prefill, TTFT 41.917 ms (1 Lauf)AMD Radeon 8060S Grap...NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition - 464,1 tok/s Generation, 2.737 tok/s Prefill, TTFT 9.091 ms (11 Laufe) | DIESER LAUF★ NVIDIA RTX PRO 6000 B...
NVIDIA RTX PRO 6000 Blackwell Workstation Edition 546,9 tok/sNVIDIA GeForce RTX 5090 532,5 tok/s★ NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition 464,1 tok/s this runNVIDIA GeForce RTX 3090 Ti 264,2 tok/sAMD Radeon AI PRO R9700 165,4 tok/sNVIDIA GeForce RTX 5070 Ti 57,2 tok/sAMD Radeon 8060S Graphics 36,1 tok/s

CPUby processor

7045283521760,001.3012.6023.903Prefill (tok/s)Generation (tok/s)AMD Ryzen 7 5800X3D 8-Core Processor - 532,5 tok/s Generation, 2.539 tok/s Prefill, TTFT 9.812 ms (3 Laufe)AMD Ryzen 7 5800X3D 8...AMD Ryzen Threadripper PRO 9965WX 24-Cores - 464,1 tok/s Generation, 3.212 tok/s Prefill, TTFT 11.084 ms (9 Laufe)AMD Ryzen Threadrippe...AMD Ryzen 9 8945HX with Radeon Graphics - 264,2 tok/s Generation, 1.458 tok/s Prefill, TTFT 19.358 ms (3 Laufe)AMD Ryzen 9 8945HX wi...AMD Ryzen Threadripper PRO 7955WX 16-Cores - 165,4 tok/s Generation, 1.380 tok/s Prefill, TTFT 28.271 ms (9 Laufe)AMD Ryzen Threadrippe...AMD Ryzen Threadripper PRO 5975WX 32-Cores - 57,2 tok/s Generation, 447 tok/s Prefill, TTFT 80.318 ms (3 Laufe)AMD Ryzen Threadrippe...AMD RYZEN AI MAX+ 395 w/ Radeon 8060S - 36,1 tok/s Generation, 536 tok/s Prefill, TTFT 41.917 ms (1 Lauf)AMD RYZEN AI MAX+ 395...AMD Ryzen 9 9950X 16-Core Processor - 546,9 tok/s Generation, 1.883 tok/s Prefill, TTFT 5.524 ms (5 Laufe) | DIESER LAUF★ AMD Ryzen 9 9950X 16-...
★ AMD Ryzen 9 9950X 16-Core Processor 546,9 tok/s this runAMD Ryzen 7 5800X3D 8-Core Processor 532,5 tok/sAMD Ryzen Threadripper PRO 9965WX 24-Cores 464,1 tok/sAMD Ryzen 9 8945HX with Radeon Graphics 264,2 tok/sAMD Ryzen Threadripper PRO 7955WX 16-Cores 165,4 tok/sAMD Ryzen Threadripper PRO 5975WX 32-Cores 57,2 tok/sAMD RYZEN AI MAX+ 395 w/ Radeon 8060S 36,1 tok/s

MBby mainboard

7055293531760,001.0222.0453.067Prefill (tok/s)Generation (tok/s)ASUSTeK COMPUTER INC. ProArt X870E-CREATOR WIFI - 546,9 tok/s Generation, 2.204 tok/s Prefill, TTFT 6.875 ms (4 Laufe)ASUSTeK COMPUTER INC....ASUSTeK COMPUTER INC. ROG STRIX B550-A GAMING - 532,5 tok/s Generation, 2.539 tok/s Prefill, TTFT 9.812 ms (3 Laufe)ASUSTeK COMPUTER INC....ASUSTeK COMPUTER INC. Pro WS WRX90E-SAGE SE - 464,1 tok/s Generation, 2.296 tok/s Prefill, TTFT 19.677 ms (18 Laufe)ASUSTeK COMPUTER INC....Meigao Innovation Technology (Shen Zhen) Co., Ltd DRFXL (MotherBoard Series) - 264,2 tok/s Generation, 1.458 tok/s Prefill, TTFT 19.358 ms (3 Laufe)Meigao Innovation Tec...ASUSTeK COMPUTER INC. Pro WS WRX80E-SAGE SE WIFI - 57,2 tok/s Generation, 447 tok/s Prefill, TTFT 80.318 ms (3 Laufe)ASUSTeK COMPUTER INC....Bosgame AXB35-02 (BeyondMax Series) - 36,1 tok/s Generation, 536 tok/s Prefill, TTFT 41.917 ms (1 Lauf)Bosgame AXB35-02 (Bey...unbekannt - 28,8 tok/s Generation, 599 tok/s Prefill, TTFT 122 ms (1 Lauf)unbekannt
ASUSTeK COMPUTER INC. ProArt X870E-CREATOR WIFI 546,9 tok/sASUSTeK COMPUTER INC. ROG STRIX B550-A GAMING 532,5 tok/sASUSTeK COMPUTER INC. Pro WS WRX90E-SAGE SE 464,1 tok/sMeigao Innovation Technology (Shen Zhen) Co., Ltd DRFXL (MotherBoard Series) 264,2 tok/sASUSTeK COMPUTER INC. Pro WS WRX80E-SAGE SE WIFI 57,2 tok/sBosgame AXB35-02 (BeyondMax Series) 36,1 tok/sunbekannt 28,8 tok/s

ENGby engine

7055293531760,02169541.6912.428Prefill (tok/s)Generation (tok/s)llama.cpp - 546,9 tok/s Generation, 2.045 tok/s Prefill, TTFT 24.256 ms (31 Laufe)llama.cppunbekannt - 28,8 tok/s Generation, 599 tok/s Prefill, TTFT 122 ms (1 Lauf)unbekanntvLLM - 28,8 tok/s Generation, 599 tok/s Prefill, TTFT 122 ms (1 Lauf)vLLM
llama.cpp 546,9 tok/sunbekannt 28,8 tok/svLLM 28,8 tok/s

DRVby driver

7055293531760,02279421.6572.372Prefill (tok/s)Generation (tok/s)unbekannt - 546,9 tok/s Generation, 2.000 tok/s Prefill, TTFT 23.502 ms (32 Laufe)unbekanntNVIDIA 580.159.03 / CUDA 13.0 - 28,8 tok/s Generation, 599 tok/s Prefill, TTFT 122 ms (1 Lauf) | DIESER LAUF★ NVIDIA 580.159.03 / C...
unbekannt 546,9 tok/s★ NVIDIA 580.159.03 / CUDA 13.0 28,8 tok/s this run
💰 Economics

Economics of this run

Operating cost, TCO and comparison with the next-best runs of the same model at identical concurrency (1× concurrent). Methodology →

⚙️ ConfigurationAll metrics and charts below follow these settings – based on a 24-month runtime.Save to URLReset
⚡ Electricity price EUR/kWh
⚙️ System utilization 100 %
🖥️ Acquisition EUR
🔌 Idle 55 W
⚡ TDP 329 W
☁️ External LLM (API)
Electricity0.30 EUR/kWh
Avg power (incl. idle)329 W estimated (TDP)GPU 300 + CPU 29 W full load
Avg cost / hourEUR 0.099
Electricity / 1M tokensEUR 0.95
Token / kWh315.38K
Acquisition (system)EUR 13,799 partial priceGPU EUR 13,000 · CPU EUR 649 · PSU EUR 150
Electricity (2 years)
TCO (2 years)EUR 15,527
Output tokens (2 years)1.82B
☁️ External LLM (API) – comparison
External LLM cost (2 years)
Savings vs. external (2 years)

All values above and the charts below take the configured system utilization into account: at X% the system generates only X% of the time, the rest it idles (55 W). Cost per hour drops (more idle), cost per token rises.

Cost over 2 years – electricity only

Cost over 2 years – incl. acquisition (TCO)

Speed vs. tokens per euro

Euro per 1M tokens

Comparison vs. API – economics per benchmark

Qwen3.6-27BNVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionQwen3.6-27BNVIDIA GeForce RTX 5090Qwen3.6-27BNVIDIA RTX PRO 6000 Blackwell Workstation EditionQwen3.6-27B3x NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
Electricity cost (24 mo.)
Acquisition cost
Total cost (TCO)
Generated tokens (24 mo.)
Token price via API
Break-even point (days)
Result (savings / extra cost)

Comparison with up to 3 next-best runs of this model at the same concurrency (at least one on different hardware). Power = GPU TDP + CPU (idle + 15 %) + board (estimated), acquisition = full system (GPU + CPU + board + RAM + PSU), prices = stored market prices.

Contributed by

Mario Alka Administrator

@marioalka

Ich bin Unternehmer, Softwareentwickler und KI-Enthusiast. Seit vielen Jahren entwickle ich Unternehmenssoftware und beschäftige mich inzwischen fast täglich mit lokalen LLMs, KI-Agenten und leistungsfähiger KI-Hardware.

Mit LLM-Benchmark.de möchte ich eine Plattform schaffen, auf der Modelle, GPUs und Agenten objektiv und reproduzierbar miteinander verglichen werden.