Created bymario-alka.dePowered bygodcore.denoob2claw.detricoma.de
Contributed byMario AlkaQwen (Alibaba)

Qwen3-30B-A3B-Thinking-2507

Performance benchmark · measured on 24.07.2026 05:54

Benchmark-IDrun-20260724-082955-40e7fe
Timebench 3 - Kombi (Prefill + Generation)MoE30BRuntime: llama.cppQuantisierung: Q4_K_M
For context: Diese Plattform nutzt Unified Memory – die "VRAM" ist gemeinsamer System-RAM (APU/Superchip); das Modell teilt sich den Speicher mit dem System.
Generation90,65tok/s
Prefill1.230,29tok/s
Time to First Token16.519,50ms
Total duration101,11s
Concurrency10parallel
Ranking in the field
534of 997 systems

Performance benchmark · Primary metric: Generation-Speed (tok/s)

This run is better than 46 % of all comparable systems.
Generation 90,7 tok/s
-66 % vs Ø 270,4
Prefill 1.230,3 tok/s
-78 % vs Ø 5.568,2
Time to First Token 16.520 ms
+69 % vs Ø 9.777
Distribution in the field0 – 1.444 tok/s
Ø 270 Median Dieser Lauf

Wie schlägt sich dieser Benchmark?

gpt-oss-20bNVIDIA RTX PRO 6000 Blackwell Workstation Edition · run-20260723-194447-dc199d
1.443,9 tok/s
gpt-oss-20bNVIDIA GeForce RTX 5090 · run-20260724-004604-1c9023
1.386,8 tok/s
Qwen2.5-Coder-32B-InstructNVIDIA GeForce RTX 5090 · run-20260724-082948-5b32df
1.307,9 tok/s
gpt-oss-20bNVIDIA GeForce RTX 5090 · run-20260723-045447-edeed3
1.295,3 tok/s
gpt-oss-20b3× NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition · run-20260724-082940-10761f
1.238,3 tok/s
Qwen3-Coder-30B-A3B-InstructNVIDIA RTX PRO 6000 Blackwell Workstation Edition · run-20260724-004606-ebf474
1.070,6 tok/s
Qwen3.5-35B-A3BNVIDIA RTX PRO 6000 Blackwell Workstation Edition · run-20260724-082959-cd1403
1.028,6 tok/s
Qwen3-VL-30B-A3B-InstructNVIDIA RTX PRO 6000 Blackwell Workstation Edition · run-20260724-004607-c49303
1.028,5 tok/s
Ministral-3-14B-Reasoning-2512NVIDIA RTX PRO 6000 Blackwell Workstation Edition · run-20260723-195749-eebab2
1.015,7 tok/s
GLM-4.5-AirNVIDIA GeForce RTX 5090 · run-20260724-082959-2ea391
976,7 tok/s
Qwen3-Coder-30B-A3B-InstructNVIDIA GeForce RTX 5090 · run-20260723-045450-366dae
966,2 tok/s
Qwen3-30B-A3B-Instruct-2507NVIDIA RTX PRO 6000 Blackwell Workstation Edition · run-20260724-004604-860045
961,5 tok/s
Devstral-Small-2507NVIDIA GeForce RTX 5090 · run-20260724-082938-ad06f1
958,9 tok/s
Qwen3.6-35B-A3BNVIDIA RTX PRO 6000 Blackwell Workstation Edition · run-20260724-083000-8690d7
952,8 tok/s
Qwen3-30B-A3B-Thinking-2507 this runAMD Radeon 8060S Graphics · run-20260724-082955-40e7fe
90,7 tok/s

Configuration

benchmark-konfiguration — run-20260724-082955-40e7fe
# LLM-Benchmark Konfiguration # Modell : Qwen3-30B-A3B-Thinking-2507 # Engine : llama.cpp # Run-ID : run-20260724-082955-40e7fe # GPU : AMD Radeon 8060S Graphics # CPU : 32x AMD RYZEN AI MAX+ 395 w/ Radeon 8060S # RAM : 31 GB bench@llm-benchmark:~$ /home/godcore/llama.cpp/build-rocm/bin/llama-server \ -m /home/godcore/models/dl/gpt-oss-20b-Q4_K_M.gguf \ -a gpt-oss-20b \ --host 0.0.0.0 \ --port 8000 \ --gpu-layers 99 \ -cmoe \ --flash-attn on \ -c 26624 \ --parallel 12
Engine?Die Inferenz-Software, die das Modell ausliefert (z.B. vLLM oder llama.cpp). Sie bestimmt Geschwindigkeit, unterstuetzte Modellformate und welche Parameter ueberhaupt verfuegbar sind.llamacpp
Modellalias?Der Name, unter dem das Modell ueber die API angesprochen wird. Genau dieser Wert muss im Request-Feld 'model' stehen.gpt-oss-20b
Kontextlaenge?Maximale Anzahl Tokens (Eingabe + erzeugte Ausgabe zusammen), die das Modell pro Anfrage verarbeiten kann.26624
Modellpfad?Pfad zur GGUF-Modelldatei, die geladen und ausgeliefert wird./home/godcore/models/dl/gpt-oss-20b-Q4_K_M.gguf
Alias?Anzeigename des Modells nach aussen (served model name), unabhaengig vom Dateinamen.gpt-oss-20b
GPU-Layer?Anzahl der auf die GPU ausgelagerten Modell-Layer. Hoeher = mehr VRAM und schneller; der Rest laeuft auf der CPU. 999 = alles auf GPU.99
Flash Attention?FlashAttention fuer schnellere und speichersparende Attention. Wert on/off/auto je nach Build.on
Kontext?Groesse des Kontextfensters in Token. 0 = der beim Training verwendete Kontext des Modells.26624
Parallel?Anzahl paralleler Slots/Sequenzen, die der Server gleichzeitig bedient. Der Kontext wird auf die Slots aufgeteilt.12

All benchmarks of this model To leaderboard

Model comparison

Qwen3-30B-A3B-Thinking-2507 on various hardware

All published performance runs of this model – each bubble a variant: position = prefill (X) × generation (Y), bubble size = number of runs. Closer to the top right = faster. ★ Marked gold = this benchmark.

GPUby graphics card

1.1768825882940,0011.19822.39733.595Prefill (tok/s)Generation (tok/s)NVIDIA RTX PRO 6000 Blackwell Workstation Edition - 906,0 tok/s Generation, 27.103 tok/s Prefill, TTFT 608 ms (3 Laufe)NVIDIA RTX PRO 6000 B...NVIDIA GeForce RTX 3090 Ti - 883,7 tok/s Generation, 6.954 tok/s Prefill, TTFT 2.503 ms (15 Laufe)NVIDIA GeForce RTX 30...NVIDIA GeForce RTX 5090 - 873,7 tok/s Generation, 13.841 tok/s Prefill, TTFT 817 ms (15 Laufe)NVIDIA GeForce RTX 50...AMD Radeon AI PRO R9700 - 231,1 tok/s Generation, 4.030 tok/s Prefill, TTFT 1.973 ms (10 Laufe)AMD Radeon AI PRO R97...CPU-only - 8,5 tok/s Generation, 74 tok/s Prefill, TTFT 134.321 ms (3 Laufe)CPU-onlyAMD Radeon 8060S Graphics - 131,2 tok/s Generation, 1.195 tok/s Prefill, TTFT 8.927 ms (15 Laufe) | DIESER LAUF★ AMD Radeon 8060S Grap...
NVIDIA RTX PRO 6000 Blackwell Workstation Edition 906,0 tok/sNVIDIA GeForce RTX 3090 Ti 883,7 tok/sNVIDIA GeForce RTX 5090 873,7 tok/sAMD Radeon AI PRO R9700 231,1 tok/s★ AMD Radeon 8060S Graphics 131,2 tok/s this runCPU-only 8,5 tok/s

CPUby processor

1.1768825882940,0011.19822.39733.595Prefill (tok/s)Generation (tok/s)AMD Ryzen 9 9950X 16-Core Processor - 906,0 tok/s Generation, 27.103 tok/s Prefill, TTFT 608 ms (3 Laufe)AMD Ryzen 9 9950X 16-...AMD Ryzen 9 8945HX with Radeon Graphics - 883,7 tok/s Generation, 6.954 tok/s Prefill, TTFT 2.503 ms (15 Laufe)AMD Ryzen 9 8945HX wi...AMD Ryzen 7 5800X3D 8-Core Processor - 873,7 tok/s Generation, 13.841 tok/s Prefill, TTFT 817 ms (15 Laufe)AMD Ryzen 7 5800X3D 8...AMD Ryzen Threadripper PRO 7955WX 16-Cores - 231,1 tok/s Generation, 4.030 tok/s Prefill, TTFT 1.973 ms (10 Laufe)AMD Ryzen Threadrippe...Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz - 8,5 tok/s Generation, 74 tok/s Prefill, TTFT 134.321 ms (3 Laufe)Intel(R) Xeon(R) CPU ...AMD RYZEN AI MAX+ 395 w/ Radeon 8060S - 131,2 tok/s Generation, 1.195 tok/s Prefill, TTFT 8.927 ms (15 Laufe) | DIESER LAUF★ AMD RYZEN AI MAX+ 395...
AMD Ryzen 9 9950X 16-Core Processor 906,0 tok/sAMD Ryzen 9 8945HX with Radeon Graphics 883,7 tok/sAMD Ryzen 7 5800X3D 8-Core Processor 873,7 tok/sAMD Ryzen Threadripper PRO 7955WX 16-Cores 231,1 tok/s★ AMD RYZEN AI MAX+ 395 w/ Radeon 8060S 131,2 tok/s this runIntel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz 8,5 tok/s

MBby mainboard

1.1768825882940,004.9219.84214.763Prefill (tok/s)Generation (tok/s)unbekannt - 906,0 tok/s Generation, 11.916 tok/s Prefill, TTFT 1.564 ms (33 Laufe)unbekanntASUSTeK COMPUTER INC. Pro WS WRX90E-SAGE SE - 231,1 tok/s Generation, 4.030 tok/s Prefill, TTFT 1.973 ms (10 Laufe)ASUSTeK COMPUTER INC....Dell Inc. PowerEdge R820 - 8,5 tok/s Generation, 74 tok/s Prefill, TTFT 134.321 ms (3 Laufe)Dell Inc. PowerEdge R...Meigao Innovation Technology (Shen Zhen) Co., Ltd SHWSA (MS-S1 MAX) - 131,2 tok/s Generation, 1.195 tok/s Prefill, TTFT 8.927 ms (15 Laufe) | DIESER LAUF★ Meigao Innovation Tec...
unbekannt 906,0 tok/sASUSTeK COMPUTER INC. Pro WS WRX90E-SAGE SE 231,1 tok/s★ Meigao Innovation Technology (Shen Zhen) Co., Ltd SHWSA (MS-S1 MAX) 131,2 tok/s this runDell Inc. PowerEdge R820 8,5 tok/s

ENGby engine

1.0089438788137483.0367.50111.96716.433Prefill (tok/s)Generation (tok/s)vLLM - 906,0 tok/s Generation, 14.040 tok/s Prefill, TTFT 2.181 ms (14 Laufe)vLLMllama.cpp - 849,7 tok/s Generation, 5.428 tok/s Prefill, TTFT 12.291 ms (47 Laufe) | DIESER LAUF★ llama.cpp
vLLM 906,0 tok/s★ llama.cpp 849,7 tok/s this run

DRVby driver

1.1528645762880,005.64911.29916.948Prefill (tok/s)Generation (tok/s)unbekannt - 906,0 tok/s Generation, 8.850 tok/s Prefill, TTFT 21.063 ms (21 Laufe)unbekanntNVIDIA 580.159.03 / CUDA 13.0 - 873,7 tok/s Generation, 13.841 tok/s Prefill, TTFT 817 ms (15 Laufe)NVIDIA 580.159.03 / C...AMD 7.0.0-27-generic - 231,1 tok/s Generation, 4.030 tok/s Prefill, TTFT 1.973 ms (10 Laufe)AMD 7.0.0-27-genericROCm 7.2.0 - 131,2 tok/s Generation, 1.195 tok/s Prefill, TTFT 8.927 ms (15 Laufe) | DIESER LAUF★ ROCm 7.2.0
unbekannt 906,0 tok/sNVIDIA 580.159.03 / CUDA 13.0 873,7 tok/sAMD 7.0.0-27-generic 231,1 tok/s★ ROCm 7.2.0 131,2 tok/s this run
Contributed by

Mario Alka Administrator

@marioalka

Gruender von godcore.de