Mario Alka
@marioalka · Administrator · Mitglied seit 07/2026
Gruender von godcore.de und Betreiber von llm-benchmark.de. Ich teste KI-Sprachmodelle auf echter Hardware, messe Tempo (Token/s) und Aufgaben-Qualitaet und teile die Ergebnisse offen und reproduzierbar.
627Benchmarks beigesteuert
597Performance-Läufe
30Harness-Läufe
50Getestete Modelle
8Genutzte GPUs
1.444Bester Durchsatz · tok/s
90,6%Beste Erfolgsquote
Genutzte Hardware
AMD Radeon AI PRO R9700199× · 341 tok/sNVIDIA GeForce RTX 3090 Ti184× · 902 tok/sAMD Radeon 8060S Graphics78× · 173 tok/sNVIDIA GeForce RTX 509065× · 1.387 tok/sNVIDIA RTX PRO 6000 Blackwell Workstation Edition30× · 1.444 tok/sNVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition14× · 63 tok/sAMD GPU12×NVIDIA GB109× · 174 tok/s
Getestete Modelle
gpt-oss-20b47×Qwen3-30B-A3B-Instruct-250738×Qwen3-Coder-30B-A3B-Instruct38×Ministral-3-14B-Reasoning-251238×Qwen3-30B-A3B-Thinking-250737×Devstral-Small-2-24B-Instruct-251235×Qwen2.5-Coder-32B-Instruct29×Mistral-Small-3.1-24B-Instruct-250329×Magistral-Small-250929×Codestral-22B-v0.128×Qwen3-VL-30B-A3B-Instruct27×Qwen3.5-35B-A3B24×Qwen3.6-35B-A3B22×Devstral-Small-250718×
Top Performance-Läufe
| Modell | Hardware | tok/s | TTFT | Datum |
|---|---|---|---|---|
| gpt-oss-20b 10× | NVIDIA RTX PRO 6000 Blackwell Workstation Edition vLLM | 1.444 | 560 ms | 23.07.2026 |
| gpt-oss-20b 10× | NVIDIA GeForce RTX 5090 vLLM | 1.387 | 1.205 ms | 23.07.2026 |
| gpt-oss-20b 10× | NVIDIA GeForce RTX 5090 vLLM | 1.295 | 808 ms | 23.07.2026 |
| Qwen3-Coder-30B-A3B-Instruct 10× | NVIDIA RTX PRO 6000 Blackwell Workstation Edition vLLM | 1.071 | 976 ms | 23.07.2026 |
| Qwen3-VL-30B-A3B-Instruct 10× | NVIDIA RTX PRO 6000 Blackwell Workstation Edition vLLM | 1.029 | 916 ms | 23.07.2026 |
| Ministral-3-14B-Reasoning-2512 10× | NVIDIA RTX PRO 6000 Blackwell Workstation Edition vLLM | 1.016 | 1.289 ms | 23.07.2026 |
| Qwen3-Coder-30B-A3B-Instruct 10× | NVIDIA GeForce RTX 5090 vLLM | 966 | 1.793 ms | 23.07.2026 |
| Qwen3-30B-A3B-Instruct-2507 10× | NVIDIA RTX PRO 6000 Blackwell Workstation Edition vLLM | 962 | 1.633 ms | 23.07.2026 |
| Codestral-22B-v0.1 5× | NVIDIA GeForce RTX 5090 llama.cpp | 945 | 1.034 ms | 23.07.2026 |
| Devstral-Small-2-24B-Instruct-2512 10× | NVIDIA GeForce RTX 5090 llama.cpp | 925 | 1.518 ms | 24.07.2026 |
Harness-Ergebnisse
| Modell | Hardware | Erfolgsquote | Punkte | Datum |
|---|---|---|---|---|
| Gemma-4-26B-A4B-it | NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition | 90,6% | 2.701/2.980 | 16.04.2026 |
| Devstral-Small-2-24B-Instruct-2512 | AMD GPU | 87,1% | 2.597/2.980 | 16.07.2026 |
| Gemma-4-26B-A4B-it | AMD GPU | 81,9% | 2.440/2.980 | 08.04.2026 |
| MiniMax-M2.5 | NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition | 81,4% | 2.427/2.980 | 13.07.2026 |
| Gemma-4-26B-A4B-it | AMD GPU | 80,0% | 2.383/2.980 | 08.04.2026 |
| MiniMax-M2.5 | NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition | 78,5% | 2.340/2.980 | 06.04.2026 |
| Gemma-4-26B-A4B-it | AMD GPU | 76,1% | 2.268/2.980 | 16.07.2026 |
| Devstral-Small-2-24B-Instruct-2512 | AMD GPU | 75,7% | 5.738/7.580 | 17.07.2026 |
| Gemma-4-26B-A4B-it | AMD GPU | 73,7% | 2.197/2.980 | 13.04.2026 |
| MiniMax-M2.5-BF16-INT4-AWQ | NVIDIA GB10 | 70,7% | 1.245/1.760 | 04.04.2026 |
