Created bymario-alka.dePowered bygodcore.denoob2claw.detricoma.de
Contributed byMario AlkaMistral AI

Mistral-Small-3.1-24B-Instruct-2503

Performance benchmark · measured on 23.07.2026 20:58

Benchmark-IDrun-20260724-004604-3c6d31
Timebench 3 - Kombi (Prefill + Generation)Dense24BRuntime: vLLMQuantisierung: AWQ
Generation8,99tok/s
Prefill4.686,67tok/s
Time to First Token494,50ms
Total duration228,88s
Concurrency1parallel
Ranking in the field
1287of 1549 systems

Performance benchmark · Primary metric: Generation-Speed (tok/s) · 1× concurrent

This run is better than 17 % of all comparable systems.
Generation 9,0 tok/s
-89 % vs Ø 79,0
Prefill 4.686,7 tok/s
+65 % vs Ø 2.837,1
Time to First Token 495 ms
-98 % vs Ø 26.963
Distribution in the field0 – 405 tok/s
Ø 79 Median Dieser Lauf

Wie schlägt sich dieser Benchmark mit anderen Modellen?

gpt-oss-20bNVIDIA GeForce RTX 5090 · run-20260724-004608-16af9b
404,6 tok/s
gpt-oss-20bNVIDIA RTX PRO 6000 Blackwell Workstation Edition · run-20260730-035052-9f766a
395,2 tok/s
Nemotron-3-Nano-4BNVIDIA GeForce RTX 5090 · run-20260728-194135-d456e7
393,5 tok/s
Nemotron-3-Nano-4BNVIDIA RTX PRO 6000 Blackwell Workstation Edition · run-20260727-181607-bffb11
393,5 tok/s
gpt-oss-20bNVIDIA GeForce RTX 5090 · run-20260724-004607-29bd16
388,9 tok/s
gpt-oss-20bNVIDIA GeForce RTX 5090 · run-20260729-032121-1c779f
388,2 tok/s
gemma-4-E2B-itNVIDIA RTX PRO 6000 Blackwell Workstation Edition · run-20260727-181606-6ea089
378,6 tok/s
gemma-4-E2B-itNVIDIA GeForce RTX 5090 · run-20260728-184455-e8b129
377,6 tok/s
Nemotron-3-Nano-4B3× NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition · run-20260728-194136-ac388e
362,3 tok/s
Nemotron-3-Nano-4B3× NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition · run-20260728-194136-5e5085
361,4 tok/s
Nemotron-3-Nano-30B-A3BNVIDIA GeForce RTX 5090 · run-20260729-105335-653788
359,4 tok/s
Nemotron-3-Nano-Omni-30B-A3B-ReasoningNVIDIA GeForce RTX 5090 · run-20260729-105334-2c4905
359,3 tok/s
Nemotron-Cascade-2-30B-A3BNVIDIA GeForce RTX 5090 · run-20260729-105334-44d1e7
359,0 tok/s
Nemotron-3-Nano-Omni-30B-A3B-ReasoningNVIDIA RTX PRO 6000 Blackwell Workstation Edition · run-20260729-032119-4852bf
357,1 tok/s
Mistral-Small-3.1-24B-Instruct-2503 this runNVIDIA RTX PRO 6000 Blackwell Workstation Edition · run-20260724-004604-3c6d31
9,0 tok/s

How does this benchmark compare on other GPUs?

Same model on different hardware · 1× concurrent · Generation (tok/s)

Configuration

benchmark-konfiguration — run-20260724-004604-3c6d31
# LLM-Benchmark Konfiguration # Modell : Mistral-Small-3.1-24B-Instruct-2503 # Engine : vLLM # Run-ID : run-20260724-004604-3c6d31 # GPU : NVIDIA RTX PRO 6000 Blackwell Workstation Edition # CPU : AMD Ryzen 9 9950X 16-Core Processor # RAM : 92 GB bench@llm-benchmark:~$ /home/godcore/minimax-vllm-nightly/.venv/bin/python /home/godcore/minimax-vllm-nightly/.venv/bin/vllm serve OPEA/Mistral-Small-3.1-24B-Instruct-2503-int4-AutoRound-awq-sym \ --served-model-name Mistral-Small-3.1-24B-Instruct-2503 \ --dtype auto \ --max-model-len 8192 \ --gpu-memory-utilization 0.90 \ --trust-remote-code \ --host 0.0.0.0 \ --port 8000
Engine?Die Inferenz-Software, die das Modell ausliefert (z.B. vLLM oder llama.cpp). Sie bestimmt Geschwindigkeit, unterstuetzte Modellformate und welche Parameter ueberhaupt verfuegbar sind.vllm
Modellalias?Der Name, unter dem das Modell ueber die API angesprochen wird. Genau dieser Wert muss im Request-Feld 'model' stehen.Mistral-Small-3.1-24B-Instruct-2503
Kontextlaenge?Maximale Anzahl Tokens (Eingabe + erzeugte Ausgabe zusammen), die das Modell pro Anfrage verarbeiten kann.8192
Alias?Anzeigename des Modells nach aussen (served model name), unabhaengig vom Dateinamen.Mistral-Small-3.1-24B-Instruct-2503
Dtype?Zahlenformat der Modellgewichte bei der Berechnung (z.B. auto, float16, bfloat16). 'auto' waehlt automatisch das vom Modell empfohlene Format.auto
Kontext?Groesse des Kontextfensters in Token. 0 = der beim Training verwendete Kontext des Modells.8192
GPU-Speicher?Anteil des GPU-Speichers (0 bis 1), den vLLM belegen darf. 0.92 = 92 %. Hoeher = mehr Platz fuer den KV-Cache (mehr/laengere parallele Anfragen), aber groesseres Risiko fuer 'Out of Memory'.0.90

All benchmarks of this model To leaderboard

Anzeige
Model comparison

Mistral-Small-3.1-24B-Instruct-2503 on various hardware

All published performance runs of this model – each bubble a variant: position = prefill (X) × generation (Y), bubble size = number of runs. Closer to the top right = faster. ★ Marked gold = this benchmark.

GPUby graphics card

27520613768,60,003.4356.87110.306Prefill (tok/s)Generation (tok/s)NVIDIA RTX A6000 - 214,9 tok/s Generation, 3.905 tok/s Prefill, TTFT 5.166 ms (27 Laufe)NVIDIA RTX A6000AMD Radeon PRO W7900 Dual Slot - 150,4 tok/s Generation, 1.883 tok/s Prefill, TTFT 8.173 ms (12 Laufe)AMD Radeon PRO W7900 ...AMD Radeon PRO W7800 48GB - 141,9 tok/s Generation, 1.851 tok/s Prefill, TTFT 7.766 ms (7 Laufe)AMD Radeon PRO W7800 ...AMD Radeon AI PRO R9700 - 33,1 tok/s Generation, 3.057 tok/s Prefill, TTFT 1.231 ms (16 Laufe)AMD Radeon AI PRO R97...AMD Radeon 8060S Graphics - 31,3 tok/s Generation, 1.069 tok/s Prefill, TTFT 21.414 ms (2 Laufe)AMD Radeon 8060S Grap...NVIDIA Tesla P100 PCIe 16GB - 23,8 tok/s Generation, 218 tok/s Prefill, TTFT 30.869 ms (2 Laufe)NVIDIA Tesla P100 PCI...NVIDIA RTX PRO 6000 Blackwell Workstation Edition - 88,3 tok/s Generation, 8.343 tok/s Prefill, TTFT 1.652 ms (3 Laufe) | DIESER LAUF★ NVIDIA RTX PRO 6000 B...
NVIDIA RTX A6000 214,9 tok/sAMD Radeon PRO W7900 Dual Slot 150,4 tok/sAMD Radeon PRO W7800 48GB 141,9 tok/s★ NVIDIA RTX PRO 6000 Blackwell Workstation Edition 88,3 tok/s this runAMD Radeon AI PRO R9700 33,1 tok/sAMD Radeon 8060S Graphics 31,3 tok/sNVIDIA Tesla P100 PCIe 16GB 23,8 tok/s

CPUby processor

27520613768,60,003.4356.87110.306Prefill (tok/s)Generation (tok/s)AMD Ryzen Threadripper PRO 7955WX 16-Cores - 214,9 tok/s Generation, 3.589 tok/s Prefill, TTFT 3.702 ms (43 Laufe)AMD Ryzen Threadrippe...AMD Ryzen Threadripper PRO 5975WX 32-Cores - 150,4 tok/s Generation, 1.871 tok/s Prefill, TTFT 8.023 ms (19 Laufe)AMD Ryzen Threadrippe...AMD RYZEN AI MAX+ 395 w/ Radeon 8060S - 31,3 tok/s Generation, 1.069 tok/s Prefill, TTFT 21.414 ms (2 Laufe)AMD RYZEN AI MAX+ 395...AMD Ryzen 9 7945HX with Radeon Graphics - 23,8 tok/s Generation, 218 tok/s Prefill, TTFT 30.869 ms (2 Laufe)AMD Ryzen 9 7945HX wi...AMD Ryzen 9 9950X 16-Core Processor - 88,3 tok/s Generation, 8.343 tok/s Prefill, TTFT 1.652 ms (3 Laufe) | DIESER LAUF★ AMD Ryzen 9 9950X 16-...
AMD Ryzen Threadripper PRO 7955WX 16-Cores 214,9 tok/sAMD Ryzen Threadripper PRO 5975WX 32-Cores 150,4 tok/s★ AMD Ryzen 9 9950X 16-Core Processor 88,3 tok/s this runAMD RYZEN AI MAX+ 395 w/ Radeon 8060S 31,3 tok/sAMD Ryzen 9 7945HX with Radeon Graphics 23,8 tok/s

MBby mainboard

27520613768,60,003.4356.87110.306Prefill (tok/s)Generation (tok/s)ASUSTeK COMPUTER INC. Pro WS WRX90E-SAGE SE - 214,9 tok/s Generation, 3.589 tok/s Prefill, TTFT 3.702 ms (43 Laufe)ASUSTeK COMPUTER INC....ASUSTeK COMPUTER INC. Pro WS WRX80E-SAGE SE WIFI - 150,4 tok/s Generation, 1.871 tok/s Prefill, TTFT 8.023 ms (19 Laufe)ASUSTeK COMPUTER INC....Bosgame AXB35-02 (BeyondMax Series) - 31,3 tok/s Generation, 1.069 tok/s Prefill, TTFT 21.414 ms (2 Laufe)Bosgame AXB35-02 (Bey...Shenzhen Meigao Electronic Equipment Co.,Ltd F1FXM (DeskMini Series) - 23,8 tok/s Generation, 218 tok/s Prefill, TTFT 30.869 ms (2 Laufe)Shenzhen Meigao Elect...ASUSTeK COMPUTER INC. ProArt X870E-CREATOR WIFI - 88,3 tok/s Generation, 8.343 tok/s Prefill, TTFT 1.652 ms (3 Laufe) | DIESER LAUF★ ASUSTeK COMPUTER INC....
ASUSTeK COMPUTER INC. Pro WS WRX90E-SAGE SE 214,9 tok/sASUSTeK COMPUTER INC. Pro WS WRX80E-SAGE SE WIFI 150,4 tok/s★ ASUSTeK COMPUTER INC. ProArt X870E-CREATOR WIFI 88,3 tok/s this runBosgame AXB35-02 (BeyondMax Series) 31,3 tok/sShenzhen Meigao Electronic Equipment Co.,Ltd F1FXM (DeskMini Series) 23,8 tok/s

ENGby engine

27320513668,20,01.0722.9704.8676.764Prefill (tok/s)Generation (tok/s)llama.cpp - 214,9 tok/s Generation, 2.080 tok/s Prefill, TTFT 8.497 ms (43 Laufe)llama.cppunbekannt - 33,1 tok/s Generation, 3.057 tok/s Prefill, TTFT 1.231 ms (8 Laufe)unbekanntvLLM - 99,3 tok/s Generation, 5.757 tok/s Prefill, TTFT 2.550 ms (18 Laufe) | DIESER LAUF★ vLLM
llama.cpp 214,9 tok/s★ vLLM 99,3 tok/s this rununbekannt 33,1 tok/s

DRVby driver

27320513668,20,02.8483.0233.1993.374Prefill (tok/s)Generation (tok/s)unbekannt - 214,9 tok/s Generation, 3.165 tok/s Prefill, TTFT 6.742 ms (61 Laufe)unbekanntAMD 7.0.0-27-generic - 33,1 tok/s Generation, 3.057 tok/s Prefill, TTFT 1.231 ms (8 Laufe)AMD 7.0.0-27-generic
unbekannt 214,9 tok/sAMD 7.0.0-27-generic 33,1 tok/s
💰 Economics

Economics of this run

Operating cost, TCO and comparison with the next-best runs of the same model at identical concurrency (1× concurrent). Methodology →

⚙️ ConfigurationAll metrics and charts below follow these settings – based on a 24-month runtime.Save to URLReset
⚡ Electricity price EUR/kWh
⚙️ System utilization 100 %
🖥️ Acquisition EUR
🔌 Idle 70 W
⚡ TDP 644 W
☁️ External LLM (API)
Electricity0.30 EUR/kWh
Avg power (incl. idle)644 W estimated (TDP)GPU 600 + CPU 29 + Board 15 W full load
Avg cost / hourEUR 0.19
Electricity / 1M tokensEUR 5.97
Token / kWh50.27K
Acquisition (system)EUR 15,248 full priceGPU EUR 13,000 · CPU EUR 649 · Board EUR 499 · RAM EUR 920 · PSU EUR 180
Electricity (2 years)
TCO (2 years)EUR 18,632
Output tokens (2 years)567.02M
☁️ External LLM (API) – comparison
External LLM cost (2 years)
Savings vs. external (2 years)

All values above and the charts below take the configured system utilization into account: at X% the system generates only X% of the time, the rest it idles (70 W). Cost per hour drops (more idle), cost per token rises.

Cost over 2 years – electricity only

Cost over 2 years – incl. acquisition (TCO)

Speed vs. tokens per euro

Euro per 1M tokens

Comparison vs. API – economics per benchmark

Mistral-Small-3.1-24B-Instruct-2503NVIDIA RTX PRO 6000 Blackwell Workstation EditionMistral-Small-3.1-24B-Instruct-2503AMD Radeon PRO W7900 Dual SlotMistral-Small-3.1-24B-Instruct-25032x NVIDIA RTX A6000Mistral-Small-3.1-24B-Instruct-25032x NVIDIA RTX A6000
Electricity cost (24 mo.)
Acquisition cost
Total cost (TCO)
Generated tokens (24 mo.)
Token price via API
Break-even point (days)
Result (savings / extra cost)

Comparison with up to 3 next-best runs of this model at the same concurrency (at least one on different hardware). Power = GPU TDP + CPU (idle + 15 %) + board (estimated), acquisition = full system (GPU + CPU + board + RAM + PSU), prices = stored market prices.

Contributed by

Mario Alka Administrator

@marioalka

Ich bin Unternehmer, Softwareentwickler und KI-Enthusiast. Seit vielen Jahren entwickle ich Unternehmenssoftware und beschäftige mich inzwischen fast täglich mit lokalen LLMs, KI-Agenten und leistungsfähiger KI-Hardware.

Mit LLM-Benchmark.de möchte ich eine Plattform schaffen, auf der Modelle, GPUs und Agenten objektiv und reproduzierbar miteinander verglichen werden.