🏆
Ranglisten
LLM Leaderboard
Reale Messergebnisse aller getesteten Sprachmodelle auf echter, benannter Hardware – bewertet nach roher Geschwindigkeit (Performance in Token pro Sekunde, Prefill und Zeit bis zum ersten Token) und nach praktischer Aufgaben-Qualität in vollständigen Agenten- und Chat-Abläufen (Harness). Wähle unten einen Benchmark-Typ oder filtere nach Modell, Hersteller und Hardware, um genau zu sehen, was jeweils geprüft wird und wie die Ergebnisse zustande kommen.
⚡ Performance (tok/s)🤖 Harness-Qualität👥 Concurrency🖥️ echte Hardware
| # | Modell / Hersteller | Messwerte | Parallel | GPU / CPU / RAM | Runtime | ||
|---|---|---|---|---|---|---|---|
| 1 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 647,58 tok/s TG Prefill 10.804 · TTFT 10.971 ms | 10× | NVIDIA RTX PRO 6000 Blackwell Workstation EditionAMD Ryzen 9 9950X 16-Core Processor · 92 GB RAM | llama.cppgodclaw | Details → | |
| 2 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 625,05 tok/s TG Prefill 9.619 · TTFT 11.866 ms | 10× | NVIDIA GeForce RTX 5090AMD Ryzen 7 5800X3D 8-Core Processor · 126 GB RAM | llama.cppcodex_cli | Details → | |
| 3 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 572,20 tok/s TG Prefill 16.777 · TTFT 11.592 ms | 10× | 3x NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen Threadripper PRO 9965WX 24-Cores · 125 GB RAM | llama.cppgodclaw | Details → | |
| 4 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 565,47 tok/s TG Prefill 14.569 · TTFT 11.929 ms | 10× | 3x NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen Threadripper PRO 9965WX 24-Cores · 125 GB RAM | llama.cppgodclaw | Details → | |
| 5 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 555,42 tok/s TG Prefill 13.977 · TTFT 12.231 ms | 10× | 3x NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen Threadripper PRO 9965WX 24-Cores · 125 GB RAM | llama.cppgodclaw | Details → | |
| 6 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 462,21 tok/s TG Prefill 9.382 · TTFT 15.360 ms | 10× | 3x NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen Threadripper PRO 9965WX 24-Cores · 125 GB RAM | llama.cppgodclaw | Details → | |
| 7 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 350,70 tok/s TG Prefill 11.051 · TTFT 3.677 ms | 5× | NVIDIA RTX PRO 6000 Blackwell Workstation EditionAMD Ryzen 9 9950X 16-Core Processor · 92 GB RAM | llama.cppgodclaw | Details → | |
| 8 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 344,54 tok/s TG Prefill 9.890 · TTFT 3.957 ms | 5× | NVIDIA GeForce RTX 5090AMD Ryzen 7 5800X3D 8-Core Processor · 126 GB RAM | llama.cppcodex_cli | Details → | |
| 9 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 317,44 tok/s TG Prefill 6.072 · TTFT 23.382 ms | 10× | NVIDIA GeForce RTX 3090 TiAMD Ryzen 9 8945HX with Radeon Graphics · 92 GB RAM | llama.cppwebsocket | Details → | |
| 10 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 308,08 tok/s TG Prefill 16.739 · TTFT 3.491 ms | 5× | 3x NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen Threadripper PRO 9965WX 24-Cores · 125 GB RAM | llama.cppgodclaw | Details → | |
| 11 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 305,62 tok/s TG Prefill 13.287 · TTFT 3.629 ms | 5× | 3x NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen Threadripper PRO 9965WX 24-Cores · 125 GB RAM | llama.cppgodclaw | Details → | |
| 12 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 301,47 tok/s TG Prefill 11.995 · TTFT 3.913 ms | 5× | 3x NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen Threadripper PRO 9965WX 24-Cores · 125 GB RAM | llama.cppgodclaw | Details → | |
| 13 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 257,36 tok/s TG Prefill 9.035 · TTFT 4.979 ms | 5× | 3x NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen Threadripper PRO 9965WX 24-Cores · 125 GB RAM | llama.cppgodclaw | Details → | |
| 14 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 205,80 tok/s TG Prefill 7.519 · TTFT 31.654 ms | 10× | 3x AMD Radeon AI PRO R9700AMD Ryzen Threadripper PRO 7955WX 16-Cores · 184 GB RAM | llama.cppgodclaw | Details → | |
| 15 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 205,12 tok/s TG Prefill 6.044 · TTFT 32.830 ms | 10× | 3x AMD Radeon AI PRO R9700AMD Ryzen Threadripper PRO 7955WX 16-Cores · 184 GB RAM | llama.cppgodclaw | Details → | |
| 16 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 192,69 tok/s TG Prefill 3.898 · TTFT 37.605 ms | 10× | 3x AMD Radeon AI PRO R9700AMD Ryzen Threadripper PRO 7955WX 16-Cores · 184 GB RAM | llama.cppgodclaw | Details → | |
| 17 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 173,35 tok/s TG Prefill 4.438 · TTFT 8.093 ms | 5× | NVIDIA GeForce RTX 3090 TiAMD Ryzen 9 8945HX with Radeon Graphics · 92 GB RAM | llama.cppwebsocket | Details → | |
| 18 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 137,47 tok/s TG Prefill 5.577 · TTFT 8.649 ms | 10× | 2x AMD Radeon PRO W7900 Dual SlotAMD Ryzen Threadripper PRO 5975WX 32-Cores · 247 GB RAM | llama.cppgodclaw | Details → | |
| 19 | Devstral-Small-250724BQ8_0Mistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 130,71 tok/s TG Prefill 3.496 · TTFT 13.343 ms | 10× | 2x AMD Radeon PRO W7900 Dual SlotAMD Ryzen Threadripper PRO 5975WX 32-Cores · 247 GB RAM | llama.cppgodclaw | Details → | |
| 20 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 130,59 tok/s TG Prefill 3.587 · TTFT 13.505 ms | 10× | AMD Radeon PRO W7900 Dual SlotAMD Ryzen Threadripper PRO 5975WX 32-Cores · 247 GB RAM | llama.cppgodclaw | Details → | |
| 21 | Devstral-Small-250724BQ8_0Mistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 120,47 tok/s TG Prefill 2.594 · TTFT 19.870 ms | 10× | AMD Radeon PRO W7900 Dual SlotAMD Ryzen Threadripper PRO 5975WX 32-Cores · 247 GB RAM | llama.cppgodclaw | Details → | |
| 22 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 109,40 tok/s TG Prefill 6.025 · TTFT 9.558 ms | 5× | 3x AMD Radeon AI PRO R9700AMD Ryzen Threadripper PRO 7955WX 16-Cores · 184 GB RAM | llama.cppgodclaw | Details → | |
| 23 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 109,11 tok/s TG Prefill 4.704 · TTFT 10.335 ms | 5× | 3x AMD Radeon AI PRO R9700AMD Ryzen Threadripper PRO 7955WX 16-Cores · 184 GB RAM | llama.cppgodclaw | Details → | |
| 24 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 103,94 tok/s TG Prefill 2.848 · TTFT 64.620 ms | 10× | NVIDIA GeForce RTX 5070 TiAMD Ryzen Threadripper PRO 5975WX 32-Cores · 247 GB RAM | llama.cppgodclaw | Details → | |
| 25 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 103,69 tok/s TG Prefill 2.847 · TTFT 13.024 ms | 5× | 3x AMD Radeon AI PRO R9700AMD Ryzen Threadripper PRO 7955WX 16-Cores · 184 GB RAM | llama.cppgodclaw | Details → | |
| 26 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 103,26 tok/s TG Prefill 2.824 · TTFT 64.279 ms | 10× | NVIDIA GeForce RTX 5070 TiAMD Ryzen Threadripper PRO 5975WX 32-Cores · 247 GB RAM | llama.cppgodclaw | Details → | |
| 27 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 98,13 tok/s TG Prefill 2.634 · TTFT 1.583 ms | 1× | NVIDIA GeForce RTX 5090AMD Ryzen 7 5800X3D 8-Core Processor · 126 GB RAM | llama.cppcodex_cli | Details → | |
| 28 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 96,56 tok/s TG Prefill 3.644 · TTFT 1.108 ms | 1× | NVIDIA RTX PRO 6000 Blackwell Workstation EditionAMD Ryzen 9 9950X 16-Core Processor · 92 GB RAM | llama.cppgodclaw | Details → | |
| 29 | Devstral-Small-250724BQ8_0Mistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 94,89 tok/s TG Prefill 2.986 · TTFT 7.982 ms | 5× | 2x AMD Radeon PRO W7900 Dual SlotAMD Ryzen Threadripper PRO 5975WX 32-Cores · 247 GB RAM | llama.cppgodclaw | Details → | |
| 30 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 92,66 tok/s TG Prefill 7.341 · TTFT 475 ms | 1× | 3x NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen Threadripper PRO 9965WX 24-Cores · 125 GB RAM | llama.cppgodclaw | Details → | |
| 31 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 92,18 tok/s TG Prefill 7.259 · TTFT 472 ms | 1× | 3x NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen Threadripper PRO 9965WX 24-Cores · 125 GB RAM | llama.cppgodclaw | Details → | |
| 32 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 92,06 tok/s TG Prefill 8.093 · TTFT 424 ms | 1× | 3x NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen Threadripper PRO 9965WX 24-Cores · 125 GB RAM | llama.cppgodclaw | Details → | |
| 33 | Devstral-Small-250724BQ8_0Mistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 87,84 tok/s TG Prefill 1.877 · TTFT 12.604 ms | 5× | AMD Radeon PRO W7900 Dual SlotAMD Ryzen Threadripper PRO 5975WX 32-Cores · 247 GB RAM | llama.cppgodclaw | Details → | |
| 34 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 82,97 tok/s TG Prefill 5.467 · TTFT 629 ms | 1× | 3x NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen Threadripper PRO 9965WX 24-Cores · 125 GB RAM | llama.cppgodclaw | Details → | |
| 35 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 77,16 tok/s TG Prefill 3.656 · TTFT 5.517 ms | 5× | 2x AMD Radeon PRO W7900 Dual SlotAMD Ryzen Threadripper PRO 5975WX 32-Cores · 247 GB RAM | llama.cppgodclaw | Details → | |
| 36 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 73,53 tok/s TG Prefill 2.273 · TTFT 8.522 ms | 5× | AMD Radeon PRO W7900 Dual SlotAMD Ryzen Threadripper PRO 5975WX 32-Cores · 247 GB RAM | llama.cppgodclaw | Details → | |
| 37 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 57,81 tok/s TG Prefill 2.629 · TTFT 20.620 ms | 5× | NVIDIA GeForce RTX 5070 TiAMD Ryzen Threadripper PRO 5975WX 32-Cores · 247 GB RAM | llama.cppgodclaw | Details → | |
| 38 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 57,59 tok/s TG Prefill 2.617 · TTFT 21.158 ms | 5× | NVIDIA GeForce RTX 5070 TiAMD Ryzen Threadripper PRO 5975WX 32-Cores · 247 GB RAM | llama.cppgodclaw | Details → | |
| 39 | Devstral-Small-250724BMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 57,12 tok/s TG Prefill 3.392 · TTFT 922 ms | 1× | NVIDIA GeForce RTX 3090 TiAMD Ryzen 5 5600X 6-Core Processor · 30 GB RAM | llamacpp | Details → | |
| 40 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 56,94 tok/s TG Prefill 3.347 · TTFT 1.030 ms | 1× | NVIDIA GeForce RTX 3090 TiAMD Ryzen 9 8945HX with Radeon Graphics · 92 GB RAM | llama.cppwebsocket | Details → | |
| 41 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 38,93 tok/s TG Prefill 1.644 · TTFT 2.060 ms | 1× | AMD Radeon PRO W7900 Dual SlotAMD Ryzen Threadripper PRO 5975WX 32-Cores · 247 GB RAM | llama.cppgodclaw | Details → | |
| 42 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 38,81 tok/s TG Prefill 2.539 · TTFT 1.355 ms | 1× | 2x AMD Radeon PRO W7900 Dual SlotAMD Ryzen Threadripper PRO 5975WX 32-Cores · 247 GB RAM | llama.cppgodclaw | Details → | |
| 43 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 32,95 tok/s TG Prefill 2.127 · TTFT 1.623 ms | 1× | 3x AMD Radeon AI PRO R9700AMD Ryzen Threadripper PRO 7955WX 16-Cores · 184 GB RAM | llama.cppgodclaw | Details → | |
| 44 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 32,89 tok/s TG Prefill 3.402 · TTFT 1.011 ms | 1× | 3x AMD Radeon AI PRO R9700AMD Ryzen Threadripper PRO 7955WX 16-Cores · 184 GB RAM | llama.cppgodclaw | Details → | |
| 45 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 32,59 tok/s TG Prefill 4.220 · TTFT 814 ms | 1× | 3x AMD Radeon AI PRO R9700AMD Ryzen Threadripper PRO 7955WX 16-Cores · 184 GB RAM | llama.cppgodclaw | Details → | |
| 46 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) i | 31,41 tok/s TG Prefill 1.961 · TTFT 1.394 ms | 1× | 3x AMD Radeon AI PRO R9700AMD Ryzen Threadripper PRO 7955WX 16-Cores · 184 GB RAM | llama.cppgodclaw | Details → | |
| 47 | Devstral-Small-250724BQ4_K_MMistral AI Performancebenchmark i | 31,41 tok/s TG Prefill 1.961 · TTFT 1.394 ms | 1× | 3x AMD Radeon AI PRO R970032x AMD Ryzen Threadripper PRO 7955WX 16-Cores · 184 GB RAM | godclaw | Details → | |
| 48 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) i | 30,97 tok/s TG Prefill 1.167 · TTFT 1.890 ms | 1× | 3x AMD Radeon AI PRO R9700AMD Ryzen Threadripper PRO 7955WX 16-Cores · 184 GB RAM | llama.cppgodclaw | Details → | |
| 49 | Devstral-Small-250724BQ4_K_MMistral AI Performancebenchmark i | 30,97 tok/s TG Prefill 1.167 · TTFT 1.890 ms | 1× | 3x AMD Radeon AI PRO R970032x AMD Ryzen Threadripper PRO 7955WX 16-Cores · 184 GB RAM | godclaw | Details → | |
| 50 | Devstral-Small-250724BQ8_0Mistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 28,95 tok/s TG Prefill 1.131 · TTFT 3.039 ms | 1× | AMD Radeon PRO W7900 Dual SlotAMD Ryzen Threadripper PRO 5975WX 32-Cores · 247 GB RAM | llama.cppgodclaw | Details → | |
| 51 | Devstral-Small-250724BQ8_0Mistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 28,65 tok/s TG Prefill 1.841 · TTFT 1.870 ms | 1× | 2x AMD Radeon PRO W7900 Dual SlotAMD Ryzen Threadripper PRO 5975WX 32-Cores · 247 GB RAM | llama.cppgodclaw | Details → | |
| 52 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) i | 28,45 tok/s TG Prefill 890 · TTFT 3.580 ms | 1× | 3x AMD Radeon AI PRO R9700AMD Ryzen Threadripper PRO 7955WX 16-Cores · 184 GB RAM | llama.cppgodclaw | Details → | |
| 53 | Devstral-Small-250724BQ4_K_MMistral AI Performancebenchmark i | 28,45 tok/s TG Prefill 890 · TTFT 3.580 ms | 1× | 3x AMD Radeon AI PRO R970032x AMD Ryzen Threadripper PRO 7955WX 16-Cores · 184 GB RAM | godclaw | Details → | |
| 54 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 23,15 tok/s TG Prefill 308 · TTFT 64.768 ms | 5× | NVIDIA Tesla P100 PCIe 16GBAMD Ryzen 9 7945HX with Radeon Graphics · 60 GB RAM | llama.cppopenclaw_cli | Details → | |
| 55 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 19,14 tok/s TG Prefill 1.510 · TTFT 2.317 ms | 1× | NVIDIA GeForce RTX 5070 TiAMD Ryzen Threadripper PRO 5975WX 32-Cores · 247 GB RAM | llama.cppgodclaw | Details → | |
| 56 | Devstral-Small-250724BQ4_K_MMistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 13,34 tok/s TG Prefill 243 · TTFT 13.913 ms | 1× | NVIDIA Tesla P100 PCIe 16GBAMD Ryzen 9 7945HX with Radeon Graphics · 60 GB RAM | llama.cppopenclaw_cli | Details → |
