Created bymario-alka.dePowered bygodcore.denoob2claw.detricoma.de
Contributed byMarcel SommerMistral AI

Devstral-Small-2-24B-Instruct-2512

Performance benchmark · measured on 03.08.2026 11:41

Benchmark-IDrun-20260804-052132-3e1a73
Timebench 3 - Kombi (Prefill + Generation)Dense24BRuntime: vLLMQuantisierung: AWQ
Generation182,39tok/s
Prefill3.280,60tok/s
Time to First Token4.105,00ms
Total duration65,74s
Concurrency5parallel
Ranking in the field
355of 604 systems

Performance benchmark · Primary metric: Generation-Speed (tok/s) · 5× concurrent

This run is better than 41 % of all comparable systems.
Generation 182,4 tok/s
-49 % vs Ø 361,1
Prefill 3.280,6 tok/s
-34 % vs Ø 4.933,5
Time to First Token 4.105 ms
-89 % vs Ø 38.932
Distribution in the field0 – 1.349 tok/s
Ø 355 Median Dieser Lauf

Wie schlägt sich dieser Benchmark mit anderen Modellen?

gemma-4-E2B-itNVIDIA RTX PRO 6000 Blackwell Workstation Edition · run-20260727-181606-0adf6d
1.349,3 tok/s
gemma-4-E2B-itNVIDIA GeForce RTX 5090 · run-20260728-184455-5b4937
1.301,9 tok/s
gemma-4-E2B-it3× NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition · run-20260728-184456-4111a9
1.195,7 tok/s
NVIDIA-Nemotron-3-Nano-4BNVIDIA RTX PRO 6000 Blackwell Workstation Edition · run-20260727-181607-30d864
1.188,7 tok/s
NVIDIA-Nemotron-3-Nano-4BNVIDIA GeForce RTX 5090 · run-20260728-194135-5432cd
1.164,8 tok/s
gpt-oss-20bNVIDIA RTX PRO 6000 Blackwell Workstation Edition · run-20260728-140954-9e81f9
1.136,3 tok/s
gpt-oss-20bNVIDIA RTX PRO 6000 Blackwell Workstation Edition · run-20260730-035052-09627f
1.135,0 tok/s
gpt-oss-20bNVIDIA GeForce RTX 5090 · run-20260729-032121-058a31
1.113,5 tok/s
NVIDIA-Nemotron-3-Nano-4B3× NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition · run-20260728-194136-836a91
1.090,6 tok/s
NVIDIA-Nemotron-3-Nano-4B3× NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition · run-20260727-181606-8b381e
1.077,4 tok/s
NVIDIA-Nemotron-3-Nano-4B3× NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition · run-20260728-194136-b8a818
1.071,0 tok/s
NVIDIA-Nemotron-3-Nano-4B3× NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition · run-20260728-194136-038e92
1.051,1 tok/s
gpt-oss-20b3× NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition · run-20260729-032121-ffb18a
1.015,5 tok/s
NVIDIA-Nemotron-3-Nano-30B-A3BNVIDIA RTX PRO 6000 Blackwell Workstation Edition · run-20260729-032121-305ad6
1.007,8 tok/s
Devstral-Small-2-24B-Instruct-2512 this runNVIDIA GeForce RTX 3090 Ti · run-20260804-052132-3e1a73
182,4 tok/s

How does this benchmark compare on other GPUs?

Same model on different hardware · 5× concurrent · Generation (tok/s)

Hardware

GPU: NVIDIA GeForce RTX 3090 Ti · 24 GB VRAM
CPU: AMD Ryzen 5 5600X 6-Core Processor
RAM: 30 GB
Mainboard: ASUSTeK COMPUTER INC. PRIME A520M-K

Setup

Runtime: vLLM
Quantization: AWQ
Model: Devstral-Small-2-24B-Instruct-2512

Configuration

benchmark-konfiguration — run-20260804-052132-3e1a73
# LLM-Benchmark Konfiguration # Modell : Devstral-Small-2-24B-Instruct-2512 # Engine : vLLM # Run-ID : run-20260804-052132-3e1a73 # GPU : NVIDIA GeForce RTX 3090 Ti # CPU : AMD Ryzen 5 5600X 6-Core Processor # RAM : 30 GB bench@llm-benchmark:~$ /opt/vllm-gemma/venv/bin/python /opt/vllm-gemma/venv/bin/vllm serve /home/godcore/models/devstral24-awq \ --served-model-name Devstral-Small-2-24B-Instruct-2512 \ --host 192.168.41.116 \ --port 8000 \ --dtype half \ --max-model-len 32768 \ --gpu-memory-utilization 0.92 \ --enforce-eager \ --limit-mm-per-prompt '{"image":0}'
Engine?Die Inferenz-Software, die das Modell ausliefert (z.B. vLLM oder llama.cpp). Sie bestimmt Geschwindigkeit, unterstuetzte Modellformate und welche Parameter ueberhaupt verfuegbar sind.vllm
Modellalias?Der Name, unter dem das Modell ueber die API angesprochen wird. Genau dieser Wert muss im Request-Feld 'model' stehen.Devstral-Small-2-24B-Instruct-2512
Kontextlaenge?Maximale Anzahl Tokens (Eingabe + erzeugte Ausgabe zusammen), die das Modell pro Anfrage verarbeiten kann.32768
Alias?Anzeigename des Modells nach aussen (served model name), unabhaengig vom Dateinamen.Devstral-Small-2-24B-Instruct-2512
Dtype?Zahlenformat der Modellgewichte bei der Berechnung (z.B. auto, float16, bfloat16). 'auto' waehlt automatisch das vom Modell empfohlene Format.half
Kontext?Groesse des Kontextfensters in Token. 0 = der beim Training verwendete Kontext des Modells.32768
GPU-Speicher?Anteil des GPU-Speichers (0 bis 1), den vLLM belegen darf. 0.92 = 92 %. Hoeher = mehr Platz fuer den KV-Cache (mehr/laengere parallele Anfragen), aber groesseres Risiko fuer 'Out of Memory'.0.92
Enforce-Eager?Schaltet die optimierte Graph-Ausfuehrung (CUDA-/HIP-Graphs) AB und rechnet Schritt fuer Schritt. Startet schneller und spart etwas VRAM, ist im laufenden Betrieb aber meist langsamer als mit Graphs.aktiv
limit-mm-per-prompt?Obergrenze fuer multimodale Eingaben pro Prompt (z.B. Anzahl Bilder). Nur fuer multimodale Modelle relevant.{"image":0}

All benchmarks of this model To leaderboard

Anzeige
Model comparison

Devstral-Small-2-24B-Instruct-2512 on various hardware

All published performance runs of this model – each bubble a variant: position = prefill (X) × generation (Y), bubble size = number of runs. Closer to the top right = faster. ★ Marked gold = this benchmark.

GPUby graphics card

9226914612300,003.3026.6049.906Prefill (tok/s)Generation (tok/s)NVIDIA RTX PRO 6000 Blackwell Workstation Edition - 725,6 tok/s Generation, 8.070 tok/s Prefill, TTFT 3.939 ms (6 Laufe)NVIDIA RTX PRO 6000 B...NVIDIA GeForce RTX 5090 - 620,3 tok/s Generation, 6.308 tok/s Prefill, TTFT 6.898 ms (3 Laufe)NVIDIA GeForce RTX 50...NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition - 569,7 tok/s Generation, 8.224 tok/s Prefill, TTFT 7.901 ms (9 Laufe)NVIDIA RTX PRO 6000 B...NVIDIA GeForce RTX 5070 Ti - 108,5 tok/s Generation, 1.619 tok/s Prefill, TTFT 37.385 ms (3 Laufe)NVIDIA GeForce RTX 50...NVIDIA GeForce RTX 3090 Ti - 355,4 tok/s Generation, 3.516 tok/s Prefill, TTFT 7.254 ms (9 Laufe) | DIESER LAUF★ NVIDIA GeForce RTX 30...
NVIDIA RTX PRO 6000 Blackwell Workstation Edition 725,6 tok/sNVIDIA GeForce RTX 5090 620,3 tok/sNVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition 569,7 tok/s★ NVIDIA GeForce RTX 3090 Ti 355,4 tok/s this runNVIDIA GeForce RTX 5070 Ti 108,5 tok/s

CPUby processor

9226914612300,003.3026.6049.906Prefill (tok/s)Generation (tok/s)AMD Ryzen 9 9950X 16-Core Processor - 725,6 tok/s Generation, 8.070 tok/s Prefill, TTFT 3.939 ms (6 Laufe)AMD Ryzen 9 9950X 16-...AMD Ryzen 7 5800X3D 8-Core Processor - 620,3 tok/s Generation, 6.308 tok/s Prefill, TTFT 6.898 ms (3 Laufe)AMD Ryzen 7 5800X3D 8...AMD Ryzen Threadripper PRO 9965WX 24-Cores - 569,7 tok/s Generation, 8.224 tok/s Prefill, TTFT 7.901 ms (9 Laufe)AMD Ryzen Threadrippe...AMD Ryzen 9 8945HX with Radeon Graphics - 331,6 tok/s Generation, 3.668 tok/s Prefill, TTFT 14.010 ms (3 Laufe)AMD Ryzen 9 8945HX wi...AMD Ryzen Threadripper PRO 5975WX 32-Cores - 108,5 tok/s Generation, 1.619 tok/s Prefill, TTFT 37.385 ms (3 Laufe)AMD Ryzen Threadrippe...AMD Ryzen 5 5600X 6-Core Processor - 355,4 tok/s Generation, 3.440 tok/s Prefill, TTFT 3.876 ms (6 Laufe) | DIESER LAUF★ AMD Ryzen 5 5600X 6-C...
AMD Ryzen 9 9950X 16-Core Processor 725,6 tok/sAMD Ryzen 7 5800X3D 8-Core Processor 620,3 tok/sAMD Ryzen Threadripper PRO 9965WX 24-Cores 569,7 tok/s★ AMD Ryzen 5 5600X 6-Core Processor 355,4 tok/s this runAMD Ryzen 9 8945HX with Radeon Graphics 331,6 tok/sAMD Ryzen Threadripper PRO 5975WX 32-Cores 108,5 tok/s

MBby mainboard

9226914612300,003.3026.6049.906Prefill (tok/s)Generation (tok/s)ASUSTeK COMPUTER INC. ProArt X870E-CREATOR WIFI - 725,6 tok/s Generation, 8.070 tok/s Prefill, TTFT 3.939 ms (6 Laufe)ASUSTeK COMPUTER INC....ASUSTeK COMPUTER INC. ROG STRIX B550-A GAMING - 620,3 tok/s Generation, 6.308 tok/s Prefill, TTFT 6.898 ms (3 Laufe)ASUSTeK COMPUTER INC....ASUSTeK COMPUTER INC. Pro WS WRX90E-SAGE SE - 569,7 tok/s Generation, 8.224 tok/s Prefill, TTFT 7.901 ms (9 Laufe)ASUSTeK COMPUTER INC....Meigao Innovation Technology (Shen Zhen) Co., Ltd DRFXL (MotherBoard Series) - 331,6 tok/s Generation, 3.668 tok/s Prefill, TTFT 14.010 ms (3 Laufe)Meigao Innovation Tec...ASUSTeK COMPUTER INC. Pro WS WRX80E-SAGE SE WIFI - 108,5 tok/s Generation, 1.619 tok/s Prefill, TTFT 37.385 ms (3 Laufe)ASUSTeK COMPUTER INC....ASUSTeK COMPUTER INC. PRIME A520M-K - 355,4 tok/s Generation, 3.440 tok/s Prefill, TTFT 3.876 ms (6 Laufe) | DIESER LAUF★ ASUSTeK COMPUTER INC....
ASUSTeK COMPUTER INC. ProArt X870E-CREATOR WIFI 725,6 tok/sASUSTeK COMPUTER INC. ROG STRIX B550-A GAMING 620,3 tok/sASUSTeK COMPUTER INC. Pro WS WRX90E-SAGE SE 569,7 tok/s★ ASUSTeK COMPUTER INC. PRIME A520M-K 355,4 tok/s this runMeigao Innovation Technology (Shen Zhen) Co., Ltd DRFXL (MotherBoard Series) 331,6 tok/sASUSTeK COMPUTER INC. Pro WS WRX80E-SAGE SE WIFI 108,5 tok/s

ENGby engine

89469048728379,63.8464.9215.9967.071Prefill (tok/s)Generation (tok/s)llama.cpp - 645,1 tok/s Generation, 6.363 tok/s Prefill, TTFT 12.610 ms (21 Laufe)llama.cppunbekannt - 247,7 tok/s Generation, 4.553 tok/s Prefill, TTFT 4.826 ms (2 Laufe)unbekanntvLLM - 725,6 tok/s Generation, 5.018 tok/s Prefill, TTFT 2.630 ms (7 Laufe) | DIESER LAUF★ vLLM
★ vLLM 725,6 tok/s this runllama.cpp 645,1 tok/sunbekannt 247,7 tok/s

DRVby driver

7987627266896535.5735.8106.0476.284Prefill (tok/s)Generation (tok/s)unbekannt - 725,6 tok/s Generation, 5.929 tok/s Prefill, TTFT 9.762 ms (30 Laufe)unbekannt
unbekannt 725,6 tok/s
💰 Economics

Economics of this run

Operating cost, TCO and comparison with the next-best runs of the same model at identical concurrency (5× concurrent). Methodology →

⚙️ ConfigurationAll metrics and charts below follow these settings – based on a 24-month runtime.Save to URLReset
⚡ Electricity price EUR/kWh
⚙️ System utilization 100 %
🖥️ Acquisition EUR
🔌 Idle 25 W
⚡ TDP 450 W
☁️ External LLM (API)
Electricity0.30 EUR/kWh
Avg power (incl. idle)450 W estimated (TDP)GPU 450 W full load
Avg cost / hourEUR 0.14
Electricity / 1M tokensEUR 0.21
Token / kWh1.46M
Acquisition (system)EUR 1,149 partial priceGPU EUR 999 · PSU EUR 150
Electricity (2 years)
TCO (2 years)EUR 3,514
Output tokens (2 years)11.50B
☁️ External LLM (API) – comparison
External LLM cost (2 years)
Savings vs. external (2 years)

All values above and the charts below take the configured system utilization into account: at X% the system generates only X% of the time, the rest it idles (25 W). Cost per hour drops (more idle), cost per token rises.

Cost over 2 years – electricity only

Cost over 2 years – incl. acquisition (TCO)

Speed vs. tokens per euro

Euro per 1M tokens

Comparison vs. API – economics per benchmark

Devstral-Small-2-24B-Instruct-2512NVIDIA GeForce RTX 3090 TiDevstral-Small-2-24B-Instruct-2512NVIDIA RTX PRO 6000 Blackwell Workstation EditionDevstral-Small-2-24B-Instruct-2512NVIDIA RTX PRO 6000 Blackwell Workstation EditionDevstral-Small-2-24B-Instruct-2512NVIDIA GeForce RTX 5090
Electricity cost (24 mo.)
Acquisition cost
Total cost (TCO)
Generated tokens (24 mo.)
Token price via API
Break-even point (days)
Result (savings / extra cost)

Comparison with up to 3 next-best runs of this model at the same concurrency (at least one on different hardware). Power = GPU TDP + CPU (idle + 15 %) + board (estimated), acquisition = full system (GPU + CPU + board + RAM + PSU), prices = stored market prices.

Contributed by

Marcel Sommer

@marcelsommer