⚡
Performance benchmark
How fast is the model?
Raw throughput on real hardware – how many tokens a model generates per second, how quickly it processes the prompt and how short the time to first token is. Higher is better.
⚡ Generation (tok/s)🔄 Prefill⏱️ TTFT👥 Concurrency
| # | Model / Maker | Metrics | Parallel | GPU / CPU / RAM | Runtime | ||
|---|---|---|---|---|---|---|---|
| 1 | gemma-4-E2B-it5BGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 826,24 tok/s TG Prefill 3.229 · TTFT 15.223 ms | 10× | 2x AMD Radeon GraphicsAMD Ryzen 9 7945HX with Radeon Graphics · 29 GB RAM | llama.cppopenclaw_cliQ4_K_M | Details → | |
| 2 | gemma-4-E4B-it8BGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 474,76 tok/s TG Prefill 1.760 · TTFT 26.540 ms | 10× | 2x AMD Radeon GraphicsAMD Ryzen 9 7945HX with Radeon Graphics · 29 GB RAM | llama.cppopenclaw_cliQ4_K_M | Details → | |
| 3 | gemma-4-E2B-it5BGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 436,25 tok/s TG Prefill 3.526 · TTFT 5.370 ms | 5× | 2x AMD Radeon GraphicsAMD Ryzen 9 7945HX with Radeon Graphics · 29 GB RAM | llama.cppopenclaw_cliQ4_K_M | Details → | |
| 4 | gemma-4-E4B-it8BGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 267,65 tok/s TG Prefill 1.544 · TTFT 9.962 ms | 5× | 2x AMD Radeon GraphicsAMD Ryzen 9 7945HX with Radeon Graphics · 29 GB RAM | llama.cppopenclaw_cliQ4_K_M | Details → | |
| 5 | gemma-4-E2B-it5BGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 134,26 tok/s TG Prefill 1.113 · TTFT 2.031 ms | 1× | 2x AMD Radeon GraphicsAMD Ryzen 9 7945HX with Radeon Graphics · 29 GB RAM | llama.cppopenclaw_cliQ4_K_M | Details → | |
| 6 | gemma-4-E4B-it8BGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 84,70 tok/s TG Prefill 1.022 · TTFT 2.264 ms | 1× | 2x AMD Radeon GraphicsAMD Ryzen 9 7945HX with Radeon Graphics · 29 GB RAM | llama.cppopenclaw_cliQ4_K_M | Details → |
