Real performance
TTFT, prefill and generation tokens/s instead of theoretical peak numbers.
We benchmark AI language models on real hardware: speed in tokens per second and response time – plus quality on real agent tasks. Transparent, reproducible, directly comparable.
TTFT, prefill and generation tokens/s instead of theoretical peak numbers.
Points and success rates from complete agent and chat runs.
Model, quantization, runtime and hardware stay visible on every run.
Speed (tokens/s) and quality (harness) of the most-tested models – directly comparable.
| # | Model / Maker | Metrics | Parallel | GPU / CPU / RAM | Runtime | ||
|---|---|---|---|---|---|---|---|
| 1 | gemma-4-E2B-it5BQ4_K_MGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 2.491,20 tok/s TG Prefill 15.814 · TTFT 4.382 ms | 10× | NVIDIA GeForce RTX 5090AMD Ryzen 7 5800X3D 8-Core Processor · 126 GB RAM | llama.cppcodex_cli | Details → | |
| 2 | gemma-4-E2B-it5BQ4_K_MGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 2.414,37 tok/s TG Prefill 20.277 · TTFT 4.107 ms | 10× | NVIDIA RTX PRO 6000 Blackwell Workstation EditionAMD Ryzen 9 9950X 16-Core Processor · 92 GB RAM | llama.cppgodclaw | Details → | |
| 3 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 2.182,67 tok/s TG Prefill 10.599 · TTFT 5.598 ms | 10× | NVIDIA RTX PRO 6000 Blackwell Workstation EditionAMD Ryzen 9 9950X 16-Core Processor · 92 GB RAM | llama.cppgodclaw | Details → | |
| 4 | gemma-4-E2B-it5BQ4_K_MGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 2.143,55 tok/s TG Prefill 13.632 · TTFT 4.819 ms | 10× | 3x NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen Threadripper PRO 9965WX 24-Cores · 125 GB RAM | llama.cppgodclaw | Details → | |
| 5 | gpt-oss-20b20BQ4_K_MOpenAI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 1.995,37 tok/s TG Prefill 16.158 · TTFT 5.309 ms | 10× | NVIDIA RTX PRO 6000 Blackwell Workstation EditionAMD Ryzen 9 9950X 16-Core Processor · 92 GB RAM | llama.cppgodclaw | Details → |
Generation throughput (tokens/s) per GPU – minimum, average and maximum across all performance runs.
Generation throughput (tokens/s) per CPU – minimum, average and maximum across all performance runs.
We use technically necessary cookies. External content (e.g. YouTube videos) is only loaded with your consent and then sets cookies from YouTube/Google. More in our privacy policy.