⚡
Performance-Benchmark
Wie schnell ist das Modell?
Roher Durchsatz auf echter Hardware – wie viele Token ein Modell pro Sekunde erzeugt, wie schnell es den Prompt verarbeitet und wie kurz die Zeit bis zum ersten Token ausfällt. Je höher, desto besser.
⚡ Generation (tok/s)🔄 Prefill⏱️ TTFT👥 Concurrency
| # | Modell / Hersteller | Messwerte | Parallel | GPU / CPU / RAM | Runtime | ||
|---|---|---|---|---|---|---|---|
| 1 | gpt-oss-20bOpenAI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 796,82 tok/s TG Prefill 11.892 · TTFT 1.781 ms | 10× | NVIDIA GeForce RTX 3090 Ti32x AMD Ryzen 9 8945HX with Radeon Graphics · 92 GB RAM | vLLMwebsocket | Details → | |
| 2 | Ministral-3-14B-Reasoning-2512Mistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 524,37 tok/s TG Prefill 7.751 · TTFT 4.130 ms | 10× | NVIDIA GeForce RTX 3090 Ti32x AMD Ryzen 9 8945HX with Radeon Graphics · 92 GB RAM | vLLMwebsocketAWQ | Details → | |
| 3 | gpt-oss-20bOpenAI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 483,69 tok/s TG Prefill 9.270 · TTFT 1.070 ms | 5× | NVIDIA GeForce RTX 3090 Ti32x AMD Ryzen 9 8945HX with Radeon Graphics · 92 GB RAM | vLLMwebsocket | Details → | |
| 4 | Ministral-3-14B-Reasoning-2512Mistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 338,30 tok/s TG Prefill 5.943 · TTFT 2.299 ms | 5× | NVIDIA GeForce RTX 3090 Ti32x AMD Ryzen 9 8945HX with Radeon Graphics · 92 GB RAM | vLLMwebsocketAWQ | Details → | |
| 5 | gpt-oss-20bOpenAI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 176,18 tok/s TG Prefill 8.074 · TTFT 245 ms | 1× | NVIDIA GeForce RTX 3090 Ti32x AMD Ryzen 9 8945HX with Radeon Graphics · 92 GB RAM | vLLMwebsocket | Details → | |
| 6 | Ministral-3-14B-Reasoning-2512Mistral AI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 89,62 tok/s TG Prefill 2.949 · TTFT 726 ms | 1× | NVIDIA GeForce RTX 3090 Ti32x AMD Ryzen 9 8945HX with Radeon Graphics · 92 GB RAM | vLLMwebsocketAWQ | Details → |
