Mistral AI
Devstral-Small-2-24B-Instruct-2512
Harnessbenchmark · gemessen am 16.07.2026 07:54
Benchmark-ID
run-20260722-161611-5f2f86Tool Usage Standard 1.0Dense24BRuntime: godclawQuantisierung: Q8_0
Punkte2.597,00von 2.980,00
Punktequote87,15%
Aufgaben bestanden16von 16
Gesamtdauer165,00s
Einordnung im Feld
2von 30 Systemen
Harnessbenchmark · Leitmetrik: Punktequote (%)
Dieser Lauf ist besser als 97 % aller vergleichbaren Systeme.
Punktequote
87 %
+60 % vs Ø 54,4
Aufgaben bestanden
16
-14 % vs Ø 18,6
Gesamtdauer
165,0 s
-93 % vs Ø 2.395,9
Verteilung im Feld10 – 91 %
Wie schlägt sich dieser Benchmark?
Gemma-4-26B-A4B-itNVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition · run-20260722-161611-05d866
91 %
Devstral-Small-2-24B-Instruct-2512 dieser LaufAMD GPU · run-20260722-161611-5f2f86
87 %
Gemma-4-26B-A4B-itAMD GPU · run-20260722-161611-12a767
82 %
MiniMax-M2.53× NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition · run-20260722-161611-f6cfe1
81 %
Gemma-4-26B-A4B-itAMD GPU · run-20260722-161611-9b7325
80 %
MiniMax-M2.53× NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition · run-20260722-161611-54019f
79 %
Gemma-4-26B-A4B-itAMD GPU · run-20260722-161611-0064a9
76 %
Devstral-Small-2-24B-Instruct-2512 gleiches ModellAMD GPU · run-20260722-161611-798dcc
76 %
Gemma-4-26B-A4B-itAMD GPU · run-20260722-161611-78ff78
74 %
MiniMax-M2.5-BF16-INT4-AWQNVIDIA GB10 · run-20260722-165444-51f2b3
71 %
Devstral-Small-2-24B-Instruct-2512 gleiches ModellAMD GPU · run-20260722-161611-23baf9
69 %
MiniMax-M2.53× NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition · run-20260722-161611-a58abe
68 %
MiniMax-M2.53× NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition · run-20260722-161611-2a2c91
65 %
MiniMax-M2.53× NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition · run-20260722-161610-99f36d
61 %
Hardware
GPU: AMD GPU · 0 GB VRAM
CPU: 32x AMD RYZEN AI MAX+ 395 w/ Radeon 8060S
RAM: 31 GB
Setup
Runtime: godclaw
Quantisierung: Q8_0
Modell: Devstral-Small-2-24B-Instruct-2512
Aufgaben (16)
| Aufgabe | Status | Punkte | Dauer |
|---|---|---|---|
| Antwortzeit Test | bestanden | 38 / 50 | 4,00 s |
| Format Test | bestanden | 73 / 80 | 3,00 s |
| Dateisystem Setup | bestanden | 113 / 120 | 3,00 s |
| Daten Generierung | bestanden | 98 / 150 | 23,00 s |
| Parsing + Transformation | bestanden | 146 / 200 | 27,00 s |
| Fehlerbehandlung JSON | bestanden | 107 / 150 | 19,00 s |
| Python Tool Usage | bestanden | 216 / 220 | 2,00 s |
| PHP Tool Usage | bestanden | 143 / 220 | 35,00 s |
| Export + Reimport | bestanden | 233 / 240 | 4,00 s |
| Archivierung | bestanden | 172 / 180 | 3,00 s |
| Hash Berechnung | bestanden | 193 / 200 | 3,00 s |
| Textsuche | bestanden | 172 / 180 | 3,00 s |
| Idempotenz Test | bestanden | 211 / 220 | 4,00 s |
| Rechte Problem | bestanden | 183 / 220 | 17,00 s |
| Final JSON | bestanden | 223 / 250 | 11,00 s |
| Finale Ausgabe | bestanden | 276 / 300 | 4,00 s |
Konfiguration
# LLM-Benchmark Konfiguration
# Modell : Devstral-Small-2-24B-Instruct-2512
# Engine : llama.cpp
# Run-ID : run-20260722-161611-5f2f86
# GPU : AMD GPU
# CPU : 32x AMD RYZEN AI MAX+ 395 w/ Radeon 8060S
# RAM : 31 GB
bench@llm-benchmark:~$ /home/godcore/llama.cpp/build-rocm/bin/llama-server \
--model /home/godcore/models/devstral-small-2-24b-q8_0-gguf/Devstral-Small-2-24B-Instruct-2512-Q8_0.gguf \
--alias mistralai/Devstral-Small-2-24B-Instruct-2512 \
--host 0.0.0.0 \
--port 8000 \
--ctx-size 131072 \
--parallel 1 \
--temp 0.2 \
--repeat-penalty 1.05 \
--repeat-last-n 256 \
--gpu-layers 99 \
--flash-attn on \
--jinja \
--chat-template-file /home/godcore/models/devstral-small-2-24b-q8_0-gguf/chat_template.jinja \
--threads 8 \
--threads-batch 16
| Engine | llamacpp |
| Modellalias | mistralai/Devstral-Small-2-24B-Instruct-2512 |
| Kontextlaenge | 131072 |
| Modellpfad | /home/godcore/models/devstral-small-2-24b-q8_0-gguf/Devstral-Small-2-24B-Instruct-2512-Q8_0.gguf |
| Parallel | 1 |
| Temperatur | 0.2 |
| Repeat-Penalty | 1.05 |
| Repeat-Last-N | 256 |
| GPU-Layer | 99 |
| Flash Attention | on |
| Jinja | aktiv |
| Chat-Template | /home/godcore/models/devstral-small-2-24b-q8_0-gguf/chat_template.jinja |
| Threads | 8 |
| Batch-Threads | 16 |
