Created bymario-alka.dePowered bygodcore.denoob2claw.detricoma.de
Contributed byMario AlkaMistral AI

Magistral-Small-2509

Performance benchmark · measured on 23.07.2026 21:58

Benchmark-IDrun-20260724-004604-d065ac
Timebench 3 - Kombi (Prefill + Generation)Dense24BRuntime: vLLMQuantisierung: AWQ
Generation96,03tok/s
Prefill5.628,85tok/s
Time to First Token379,50ms
Total duration21,09s
Concurrency1parallel
Ranking in the field
47of 83 systems

Performance benchmark · Primary metric: Generation-Speed (tok/s) · 1× concurrent

This run is better than 44 % of all comparable systems.
Generation 96,0 tok/s
-34 % vs Ø 146,4
Prefill 5.628,9 tok/s
+43 % vs Ø 3.941,9
Time to First Token 380 ms
-98 % vs Ø 16.310
Distribution in the field0 – 395 tok/s
Ø 146 Median Dieser Lauf

Wie schlägt sich dieser Benchmark mit anderen Modellen?

gpt-oss-20bNVIDIA RTX PRO 6000 Blackwell Workstation Edition · run-20260730-035052-9f766a
395,2 tok/s
Nemotron-3-Nano-4BNVIDIA RTX PRO 6000 Blackwell Workstation Edition · run-20260727-181607-bffb11
393,5 tok/s
gemma-4-E2B-itNVIDIA RTX PRO 6000 Blackwell Workstation Edition · run-20260727-181606-6ea089
378,6 tok/s
Nemotron-3-Nano-Omni-30B-A3B-ReasoningNVIDIA RTX PRO 6000 Blackwell Workstation Edition · run-20260729-032119-4852bf
357,1 tok/s
Nemotron-3-Nano-30B-A3BNVIDIA RTX PRO 6000 Blackwell Workstation Edition · run-20260729-032121-643b00
357,0 tok/s
Nemotron-Cascade-2-30B-A3BNVIDIA RTX PRO 6000 Blackwell Workstation Edition · run-20260729-032120-4e3476
356,8 tok/s
Qwen3-30B-A3B-Thinking-2507NVIDIA RTX PRO 6000 Blackwell Workstation Edition · run-20260729-032123-b33938
312,1 tok/s
Qwen3-Coder-30B-A3B-InstructNVIDIA RTX PRO 6000 Blackwell Workstation Edition · run-20260729-032124-ac5167
311,3 tok/s
Laguna-XS-2.1NVIDIA RTX PRO 6000 Blackwell Workstation Edition · run-20260728-140954-cecd39
310,2 tok/s
Laguna-XS-2.1NVIDIA RTX PRO 6000 Blackwell Workstation Edition · run-20260728-140955-42efc5
309,9 tok/s
Laguna-XS-2.1NVIDIA RTX PRO 6000 Blackwell Workstation Edition · run-20260728-140955-abb113
309,5 tok/s
Qwen3-30B-A3B-Instruct-2507NVIDIA RTX PRO 6000 Blackwell Workstation Edition · run-20260729-032121-b5f531
304,6 tok/s
Qwen3-Omni-30B-A3B-ThinkingNVIDIA RTX PRO 6000 Blackwell Workstation Edition · run-20260729-032125-29dbef
304,4 tok/s
Qwen3-VL-30B-A3B-InstructNVIDIA RTX PRO 6000 Blackwell Workstation Edition · run-20260729-032125-021dfb
303,6 tok/s
Magistral-Small-2509 this runNVIDIA RTX PRO 6000 Blackwell Workstation Edition · run-20260724-004604-d065ac
96,0 tok/s

How does this benchmark compare on other GPUs?

Same model on different hardware · 1× concurrent · Generation (tok/s)

Configuration

benchmark-konfiguration — run-20260724-004604-d065ac
# LLM-Benchmark Konfiguration # Modell : Magistral-Small-2509 # Engine : vLLM # Run-ID : run-20260724-004604-d065ac # GPU : NVIDIA RTX PRO 6000 Blackwell Workstation Edition # CPU : AMD Ryzen 9 9950X 16-Core Processor # RAM : 92 GB bench@llm-benchmark:~$ /home/godcore/minimax-vllm-nightly/.venv/bin/python /home/godcore/minimax-vllm-nightly/.venv/bin/vllm serve cyankiwi/Magistral-Small-2509-AWQ-4bit \ --served-model-name Magistral-Small-2509 \ --dtype auto \ --max-model-len 8192 \ --gpu-memory-utilization 0.90 \ --trust-remote-code \ --host 0.0.0.0 \ --port 8000
Engine?Die Inferenz-Software, die das Modell ausliefert (z.B. vLLM oder llama.cpp). Sie bestimmt Geschwindigkeit, unterstuetzte Modellformate und welche Parameter ueberhaupt verfuegbar sind.vllm
Modellalias?Der Name, unter dem das Modell ueber die API angesprochen wird. Genau dieser Wert muss im Request-Feld 'model' stehen.Magistral-Small-2509
Kontextlaenge?Maximale Anzahl Tokens (Eingabe + erzeugte Ausgabe zusammen), die das Modell pro Anfrage verarbeiten kann.8192
Alias?Anzeigename des Modells nach aussen (served model name), unabhaengig vom Dateinamen.Magistral-Small-2509
Dtype?Zahlenformat der Modellgewichte bei der Berechnung (z.B. auto, float16, bfloat16). 'auto' waehlt automatisch das vom Modell empfohlene Format.auto
Kontext?Groesse des Kontextfensters in Token. 0 = der beim Training verwendete Kontext des Modells.8192
GPU-Speicher?Anteil des GPU-Speichers (0 bis 1), den vLLM belegen darf. 0.92 = 92 %. Hoeher = mehr Platz fuer den KV-Cache (mehr/laengere parallele Anfragen), aber groesseres Risiko fuer 'Out of Memory'.0.90

All benchmarks of this model To leaderboard

Anzeige
Model comparison

Magistral-Small-2509 on various hardware

All published performance runs of this model – each bubble a variant: position = prefill (X) × generation (Y), bubble size = number of runs. Closer to the top right = faster. ★ Marked gold = this benchmark.

GPUby graphics card

9537154772380,001.6063.2124.817Prefill (tok/s)Generation (tok/s)NVIDIA GeForce RTX 5090 - 737,0 tok/s Generation, 128 tok/s Prefill, TTFT 3.689 ms (3 Laufe)NVIDIA GeForce RTX 50...NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition - 637,2 tok/s Generation, 149 tok/s Prefill, TTFT 3.798 ms (12 Laufe)NVIDIA RTX PRO 6000 B...NVIDIA RTX A6000 - 534,8 tok/s Generation, 3.517 tok/s Prefill, TTFT 5.663 ms (27 Laufe)NVIDIA RTX A6000NVIDIA GeForce RTX 3090 Ti - 353,8 tok/s Generation, 1.856 tok/s Prefill, TTFT 4.508 ms (6 Laufe)NVIDIA GeForce RTX 30...AMD Radeon AI PRO R9700 - 230,0 tok/s Generation, 910 tok/s Prefill, TTFT 6.435 ms (15 Laufe)AMD Radeon AI PRO R97...NVIDIA GeForce RTX 5070 Ti - 159,0 tok/s Generation, 78 tok/s Prefill, TTFT 11.931 ms (3 Laufe)NVIDIA GeForce RTX 50...AMD Radeon PRO W7800 48GB - 153,0 tok/s Generation, 762 tok/s Prefill, TTFT 14.642 ms (3 Laufe)AMD Radeon PRO W7800 ...AMD Radeon PRO W7900 Dual Slot - 141,7 tok/s Generation, 1.590 tok/s Prefill, TTFT 8.535 ms (12 Laufe)AMD Radeon PRO W7900 ...Intel Arc Pro B70 - 37,9 tok/s Generation, 43 tok/s Prefill, TTFT 65.064 ms (3 Laufe)Intel Arc Pro B70NVIDIA Tesla P100 PCIe 16GB - 24,7 tok/s Generation, 189 tok/s Prefill, TTFT 35.758 ms (2 Laufe)NVIDIA Tesla P100 PCI...NVIDIA RTX PRO 6000 Blackwell Workstation Edition - 726,9 tok/s Generation, 3.891 tok/s Prefill, TTFT 2.194 ms (6 Laufe) | DIESER LAUF★ NVIDIA RTX PRO 6000 B...
NVIDIA GeForce RTX 5090 737,0 tok/s★ NVIDIA RTX PRO 6000 Blackwell Workstation Edition 726,9 tok/s this runNVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition 637,2 tok/sNVIDIA RTX A6000 534,8 tok/sNVIDIA GeForce RTX 3090 Ti 353,8 tok/sAMD Radeon AI PRO R9700 230,0 tok/sNVIDIA GeForce RTX 5070 Ti 159,0 tok/sAMD Radeon PRO W7800 48GB 153,0 tok/sAMD Radeon PRO W7900 Dual Slot 141,7 tok/sIntel Arc Pro B70 37,9 tok/sNVIDIA Tesla P100 PCIe 16GB 24,7 tok/s

CPUby processor

9507134752380,001.6023.2054.807Prefill (tok/s)Generation (tok/s)AMD Ryzen 7 5800X3D 8-Core Processor - 737,0 tok/s Generation, 128 tok/s Prefill, TTFT 3.689 ms (3 Laufe)AMD Ryzen 7 5800X3D 8...AMD Ryzen Threadripper PRO 9965WX 24-Cores - 637,2 tok/s Generation, 149 tok/s Prefill, TTFT 3.798 ms (12 Laufe)AMD Ryzen Threadrippe...AMD Ryzen Threadripper PRO 7955WX 16-Cores - 534,8 tok/s Generation, 2.586 tok/s Prefill, TTFT 5.939 ms (42 Laufe)AMD Ryzen Threadrippe...AMD Ryzen 9 8945HX with Radeon Graphics - 353,8 tok/s Generation, 1.856 tok/s Prefill, TTFT 4.508 ms (6 Laufe)AMD Ryzen 9 8945HX wi...AMD Ryzen Threadripper PRO 5975WX 32-Cores - 159,0 tok/s Generation, 1.200 tok/s Prefill, TTFT 10.119 ms (18 Laufe)AMD Ryzen Threadrippe...AMD Ryzen 9 7945HX with Radeon Graphics - 37,9 tok/s Generation, 102 tok/s Prefill, TTFT 53.342 ms (5 Laufe)AMD Ryzen 9 7945HX wi...AMD Ryzen 9 9950X 16-Core Processor - 726,9 tok/s Generation, 3.891 tok/s Prefill, TTFT 2.194 ms (6 Laufe) | DIESER LAUF★ AMD Ryzen 9 9950X 16-...
AMD Ryzen 7 5800X3D 8-Core Processor 737,0 tok/s★ AMD Ryzen 9 9950X 16-Core Processor 726,9 tok/s this runAMD Ryzen Threadripper PRO 9965WX 24-Cores 637,2 tok/sAMD Ryzen Threadripper PRO 7955WX 16-Cores 534,8 tok/sAMD Ryzen 9 8945HX with Radeon Graphics 353,8 tok/sAMD Ryzen Threadripper PRO 5975WX 32-Cores 159,0 tok/sAMD Ryzen 9 7945HX with Radeon Graphics 37,9 tok/s

MBby mainboard

9537154772380,001.6063.2124.817Prefill (tok/s)Generation (tok/s)ASUSTeK COMPUTER INC. ROG STRIX B550-A GAMING - 737,0 tok/s Generation, 128 tok/s Prefill, TTFT 3.689 ms (3 Laufe)ASUSTeK COMPUTER INC....ASUSTeK COMPUTER INC. Pro WS WRX90E-SAGE SE - 637,2 tok/s Generation, 2.044 tok/s Prefill, TTFT 5.463 ms (54 Laufe)ASUSTeK COMPUTER INC....Meigao Innovation Technology (Shen Zhen) Co., Ltd DRFXL (MotherBoard Series) - 353,8 tok/s Generation, 1.856 tok/s Prefill, TTFT 4.508 ms (6 Laufe)Meigao Innovation Tec...ASUSTeK COMPUTER INC. Pro WS WRX80E-SAGE SE WIFI - 159,0 tok/s Generation, 1.200 tok/s Prefill, TTFT 10.119 ms (18 Laufe)ASUSTeK COMPUTER INC....Shenzhen Meigao Electronic Equipment Co.,Ltd DRFXI (MotherBoard Series) - 37,9 tok/s Generation, 43 tok/s Prefill, TTFT 65.064 ms (3 Laufe)Shenzhen Meigao Elect...Shenzhen Meigao Electronic Equipment Co.,Ltd F1FXM (DeskMini Series) - 24,7 tok/s Generation, 189 tok/s Prefill, TTFT 35.758 ms (2 Laufe)Shenzhen Meigao Elect...ASUSTeK COMPUTER INC. ProArt X870E-CREATOR WIFI - 726,9 tok/s Generation, 3.891 tok/s Prefill, TTFT 2.194 ms (6 Laufe) | DIESER LAUF★ ASUSTeK COMPUTER INC....
ASUSTeK COMPUTER INC. ROG STRIX B550-A GAMING 737,0 tok/s★ ASUSTeK COMPUTER INC. ProArt X870E-CREATOR WIFI 726,9 tok/s this runASUSTeK COMPUTER INC. Pro WS WRX90E-SAGE SE 637,2 tok/sMeigao Innovation Technology (Shen Zhen) Co., Ltd DRFXL (MotherBoard Series) 353,8 tok/sASUSTeK COMPUTER INC. Pro WS WRX80E-SAGE SE WIFI 159,0 tok/sShenzhen Meigao Electronic Equipment Co.,Ltd DRFXI (MotherBoard Series) 37,9 tok/sShenzhen Meigao Electronic Equipment Co.,Ltd F1FXM (DeskMini Series) 24,7 tok/s

ENGby engine

9527144762380,002.3294.6586.986Prefill (tok/s)Generation (tok/s)llama.cpp - 737,0 tok/s Generation, 1.025 tok/s Prefill, TTFT 10.147 ms (74 Laufe)llama.cppunbekannt - 28,9 tok/s Generation, 1.586 tok/s Prefill, TTFT 1.672 ms (3 Laufe)unbekanntvLLM - 726,9 tok/s Generation, 5.783 tok/s Prefill, TTFT 2.618 ms (15 Laufe) | DIESER LAUF★ vLLM
llama.cpp 737,0 tok/s★ vLLM 726,9 tok/s this rununbekannt 28,9 tok/s

DRVby driver

9527144762380,007781.5572.335Prefill (tok/s)Generation (tok/s)unbekannt - 737,0 tok/s Generation, 1.889 tok/s Prefill, TTFT 6.918 ms (86 Laufe)unbekanntIntel 26.18.38308.4 - 37,9 tok/s Generation, 43 tok/s Prefill, TTFT 65.064 ms (3 Laufe)Intel 26.18.38308.4AMD 7.0.0-27-generic - 28,9 tok/s Generation, 1.586 tok/s Prefill, TTFT 1.672 ms (3 Laufe)AMD 7.0.0-27-generic
unbekannt 737,0 tok/sIntel 26.18.38308.4 37,9 tok/sAMD 7.0.0-27-generic 28,9 tok/s
💰 Economics

Economics of this run

Operating cost, TCO and comparison with the next-best runs of the same model at identical concurrency (1× concurrent). Methodology →

⚙️ ConfigurationAll metrics and charts below follow these settings – based on a 24-month runtime.Save to URLReset
⚡ Electricity price EUR/kWh
⚙️ System utilization 100 %
🖥️ Acquisition EUR
🔌 Idle 70 W
⚡ TDP 644 W
☁️ External LLM (API)
Electricity0.30 EUR/kWh
Avg power (incl. idle)644 W estimated (TDP)GPU 600 + CPU 29 + Board 15 W full load
Avg cost / hourEUR 0.19
Electricity / 1M tokensEUR 0.56
Token / kWh537.02K
Acquisition (system)EUR 15,248 full priceGPU EUR 13,000 · CPU EUR 649 · Board EUR 499 · RAM EUR 920 · PSU EUR 180
Electricity (2 years)
TCO (2 years)EUR 18,632
Output tokens (2 years)6.06B
☁️ External LLM (API) – comparison
External LLM cost (2 years)
Savings vs. external (2 years)

All values above and the charts below take the configured system utilization into account: at X% the system generates only X% of the time, the rest it idles (70 W). Cost per hour drops (more idle), cost per token rises.

Cost over 2 years – electricity only

Cost over 2 years – incl. acquisition (TCO)

Speed vs. tokens per euro

Euro per 1M tokens

Comparison vs. API – economics per benchmark

Magistral-Small-2509NVIDIA RTX PRO 6000 Blackwell Workstation EditionMagistral-Small-2509NVIDIA GeForce RTX 5090Magistral-Small-2509NVIDIA RTX PRO 6000 Blackwell Workstation EditionMagistral-Small-25093x NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
Electricity cost (24 mo.)
Acquisition cost
Total cost (TCO)
Generated tokens (24 mo.)
Token price via API
Break-even point (days)
Result (savings / extra cost)

Comparison with up to 3 next-best runs of this model at the same concurrency (at least one on different hardware). Power = GPU TDP + CPU (idle + 15 %) + board (estimated), acquisition = full system (GPU + CPU + board + RAM + PSU), prices = stored market prices.

Contributed by

Mario Alka Administrator

@marioalka

Ich bin Unternehmer, Softwareentwickler und KI-Enthusiast. Seit vielen Jahren entwickle ich Unternehmenssoftware und beschäftige mich inzwischen fast täglich mit lokalen LLMs, KI-Agenten und leistungsfähiger KI-Hardware.

Mit LLM-Benchmark.de möchte ich eine Plattform schaffen, auf der Modelle, GPUs und Agenten objektiv und reproduzierbar miteinander verglichen werden.