⚡
Performance benchmark
How fast is the model?
Raw throughput on real hardware – how many tokens a model generates per second, how quickly it processes the prompt and how short the time to first token is. Higher is better.
⚡ Generation (tok/s)🔄 Prefill⏱️ TTFT👥 Concurrency
| # | Model / Maker | Metrics | Parallel | GPU / CPU / RAM | Runtime | ||
|---|---|---|---|---|---|---|---|
| 1 | gpt-oss-20b20BOpenAI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 111,12 tok/s TG Prefill 616 · TTFT 16.066 ms | 5× | NVIDIA Tesla P100 PCIe 16GBAMD Ryzen 9 7945HX with Radeon Graphics · 60 GB RAM | llama.cppopenclaw_cliMXFP4 | Details → | |
| 2 | gpt-oss-20b20BOpenAI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 105,30 tok/s TG Prefill 694 · TTFT 29.007 ms | 10× | NVIDIA Tesla P100 PCIe 16GBAMD Ryzen 9 7945HX with Radeon Graphics · 60 GB RAM | llama.cppopenclaw_cliMXFP4 | Details → | |
| 3 | gemma-4-E4B-it8BGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 97,08 tok/s TG Prefill 757 · TTFT 13.077 ms | 5× | NVIDIA Tesla P100 PCIe 16GBAMD Ryzen 9 7945HX with Radeon Graphics · 60 GB RAM | llama.cppopenclaw_cliQ4_K_M | Details → | |
| 4 | Meta-Llama-3.1-8B-Instruct8BMeta Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 67,58 tok/s TG Prefill 888 · TTFT 15.224 ms | 5× | NVIDIA Tesla P100 PCIe 16GBAMD Ryzen 9 7945HX with Radeon Graphics · 60 GB RAM | llama.cppopenclaw_cliQ4_K_M | Details → | |
| 5 | Ornith-1.0-9B9BDeepReinforce Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 64,46 tok/s TG Prefill 592 · TTFT 16.869 ms | 5× | NVIDIA Tesla P100 PCIe 16GBAMD Ryzen 9 7945HX with Radeon Graphics · 60 GB RAM | llama.cppopenclaw_cliQ4_K_M | Details → | |
| 6 | gpt-oss-20b20BOpenAI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 63,18 tok/s TG Prefill 627 · TTFT 3.155 ms | 1× | NVIDIA Tesla P100 PCIe 16GBAMD Ryzen 9 7945HX with Radeon Graphics · 60 GB RAM | llama.cppopenclaw_cliMXFP4 | Details → | |
| 7 | gemma-4-E4B-it8BGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 58,96 tok/s TG Prefill 818 · TTFT 24.766 ms | 10× | NVIDIA Tesla P100 PCIe 16GBAMD Ryzen 9 7945HX with Radeon Graphics · 60 GB RAM | llama.cppopenclaw_cliQ4_K_M | Details → | |
| 8 | gemma-4-12B-it12BGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 45,75 tok/s TG Prefill 305 · TTFT 32.520 ms | 5× | NVIDIA Tesla P100 PCIe 16GBAMD Ryzen 9 7945HX with Radeon Graphics · 60 GB RAM | llama.cppopenclaw_cliQ4_K_M | Details → | |
| 9 | gemma-4-E4B-it8BGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 43,23 tok/s TG Prefill 758 · TTFT 2.613 ms | 1× | NVIDIA Tesla P100 PCIe 16GBAMD Ryzen 9 7945HX with Radeon Graphics · 60 GB RAM | llama.cppopenclaw_cliQ4_K_M | Details → | |
| 10 | Ministral-3-14B-Reasoning-251214BMistral AI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 41,46 tok/s TG Prefill 476 · TTFT 25.832 ms | 5× | NVIDIA Tesla P100 PCIe 16GBAMD Ryzen 9 7945HX with Radeon Graphics · 60 GB RAM | llama.cppopenclaw_cliQ4_K_M | Details → | |
| 11 | Ornith-1.0-9B9BDeepReinforce Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 37,65 tok/s TG Prefill 776 · TTFT 29.085 ms | 10× | NVIDIA Tesla P100 PCIe 16GBAMD Ryzen 9 7945HX with Radeon Graphics · 60 GB RAM | llama.cppopenclaw_cliQ4_K_M | Details → | |
| 12 | Meta-Llama-3.1-8B-Instruct8BMeta Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 37,29 tok/s TG Prefill 1.115 · TTFT 27.826 ms | 10× | NVIDIA Tesla P100 PCIe 16GBAMD Ryzen 9 7945HX with Radeon Graphics · 60 GB RAM | llama.cppopenclaw_cliQ4_K_M | Details → | |
| 13 | Meta-Llama-3.1-8B-Instruct8BMeta Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 35,46 tok/s TG Prefill 589 · TTFT 4.189 ms | 1× | NVIDIA Tesla P100 PCIe 16GBAMD Ryzen 9 7945HX with Radeon Graphics · 60 GB RAM | llama.cppopenclaw_cliQ4_K_M | Details → | |
| 14 | Ornith-1.0-9B9BDeepReinforce Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 33,54 tok/s TG Prefill 547 · TTFT 3.535 ms | 1× | NVIDIA Tesla P100 PCIe 16GBAMD Ryzen 9 7945HX with Radeon Graphics · 60 GB RAM | llama.cppopenclaw_cliQ4_K_M | Details → | |
| 15 | gemma-4-12B-it12BGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 25,04 tok/s TG Prefill 333 · TTFT 61.247 ms | 10× | NVIDIA Tesla P100 PCIe 16GBAMD Ryzen 9 7945HX with Radeon Graphics · 60 GB RAM | llama.cppopenclaw_cliQ4_K_M | Details → | |
| 16 | Magistral-Small-250924BMistral AI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 24,74 tok/s TG Prefill 212 · TTFT 57.889 ms | 5× | NVIDIA Tesla P100 PCIe 16GBAMD Ryzen 9 7945HX with Radeon Graphics · 60 GB RAM | llama.cppopenclaw_cliQ4_K_M | Details → | |
| 17 | Mistral-Small-3.1-24B-Instruct-250324BMistral AI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 23,84 tok/s TG Prefill 277 · TTFT 48.247 ms | 5× | NVIDIA Tesla P100 PCIe 16GBAMD Ryzen 9 7945HX with Radeon Graphics · 60 GB RAM | llama.cppopenclaw_cliQ4_K_M | Details → | |
| 18 | Devstral-Small-2-24B-Instruct-251224BMistral AI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 23,29 tok/s TG Prefill 273 · TTFT 55.104 ms | 5× | NVIDIA Tesla P100 PCIe 16GBAMD Ryzen 9 7945HX with Radeon Graphics · 60 GB RAM | llama.cppopenclaw_cliQ4_K_M | Details → | |
| 19 | Devstral-Small-250724BMistral AI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 23,15 tok/s TG Prefill 308 · TTFT 64.768 ms | 5× | NVIDIA Tesla P100 PCIe 16GBAMD Ryzen 9 7945HX with Radeon Graphics · 60 GB RAM | llama.cppopenclaw_cliQ4_K_M | Details → | |
| 20 | Ministral-3-14B-Reasoning-251214BMistral AI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 23,11 tok/s TG Prefill 639 · TTFT 44.373 ms | 10× | NVIDIA Tesla P100 PCIe 16GBAMD Ryzen 9 7945HX with Radeon Graphics · 60 GB RAM | llama.cppopenclaw_cliQ4_K_M | Details → | |
| 21 | Ministral-3-14B-Reasoning-251214BMistral AI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 22,41 tok/s TG Prefill 388 · TTFT 5.855 ms | 1× | NVIDIA Tesla P100 PCIe 16GBAMD Ryzen 9 7945HX with Radeon Graphics · 60 GB RAM | llama.cppopenclaw_cliQ4_K_M | Details → | |
| 22 | gemma-4-12B-it12BGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 22,30 tok/s TG Prefill 298 · TTFT 6.721 ms | 1× | NVIDIA Tesla P100 PCIe 16GBAMD Ryzen 9 7945HX with Radeon Graphics · 60 GB RAM | llama.cppopenclaw_cliQ4_K_M | Details → | |
| 23 | Codestral-22B-v0.122BMistral AI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 22,24 tok/s TG Prefill 263 · TTFT 57.267 ms | 5× | NVIDIA Tesla P100 PCIe 16GBAMD Ryzen 9 7945HX with Radeon Graphics · 60 GB RAM | llama.cppopenclaw_cliQ4_K_M | Details → | |
| 24 | Mistral-Small-3.1-24B-Instruct-250324BMistral AI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 13,57 tok/s TG Prefill 159 · TTFT 13.491 ms | 1× | NVIDIA Tesla P100 PCIe 16GBAMD Ryzen 9 7945HX with Radeon Graphics · 60 GB RAM | llama.cppopenclaw_cliQ4_K_M | Details → | |
| 25 | Magistral-Small-250924BMistral AI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 13,55 tok/s TG Prefill 166 · TTFT 13.627 ms | 1× | NVIDIA Tesla P100 PCIe 16GBAMD Ryzen 9 7945HX with Radeon Graphics · 60 GB RAM | llama.cppopenclaw_cliQ4_K_M | Details → | |
| 26 | Codestral-22B-v0.122BMistral AI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 13,51 tok/s TG Prefill 185 · TTFT 14.723 ms | 1× | NVIDIA Tesla P100 PCIe 16GBAMD Ryzen 9 7945HX with Radeon Graphics · 60 GB RAM | llama.cppopenclaw_cliQ4_K_M | Details → | |
| 27 | Devstral-Small-2-24B-Instruct-251224BMistral AI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 13,47 tok/s TG Prefill 197 · TTFT 13.734 ms | 1× | NVIDIA Tesla P100 PCIe 16GBAMD Ryzen 9 7945HX with Radeon Graphics · 60 GB RAM | llama.cppopenclaw_cliQ4_K_M | Details → | |
| 28 | Devstral-Small-250724BMistral AI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 13,34 tok/s TG Prefill 243 · TTFT 13.913 ms | 1× | NVIDIA Tesla P100 PCIe 16GBAMD Ryzen 9 7945HX with Radeon Graphics · 60 GB RAM | llama.cppopenclaw_cliQ4_K_M | Details → |
