🏆
Rankings
LLM Leaderboard
Real measurements of every tested language model on real, named hardware – rated by raw speed (performance in tokens per second, prefill and time to first token) and by practical task quality in complete agent and chat runs (harness). Pick a benchmark type below or filter by model, maker and hardware to see exactly what is tested and how the results are produced.
⚡ Performance (tok/s)🤖 Harness quality👥 Concurrency🖥️ real hardware
| # | Model / Maker | Metrics | Parallel | GPU / CPU / RAM | Runtime | ||
|---|---|---|---|---|---|---|---|
| 1 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 2.182,67 tok/s TG Prefill 10.599 · TTFT 5.598 ms | 10× | NVIDIA RTX PRO 6000 Blackwell Workstation EditionAMD Ryzen 9 9950X 16-Core Processor · 92 GB RAM | llama.cppgodclaw | Details → | |
| 2 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 1.937,67 tok/s TG Prefill 9.452 · TTFT 6.288 ms | 10× | 3x NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen Threadripper PRO 9965WX 24-Cores · 125 GB RAM | llama.cppgodclaw | Details → | |
| 3 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 1.911,59 tok/s TG Prefill 10.116 · TTFT 6.245 ms | 10× | 3x NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen Threadripper PRO 9965WX 24-Cores · 125 GB RAM | llama.cppgodclaw | Details → | |
| 4 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 1.896,98 tok/s TG Prefill 10.984 · TTFT 6.227 ms | 10× | 3x NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen Threadripper PRO 9965WX 24-Cores · 125 GB RAM | llama.cppgodclaw | Details → | |
| 5 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 1.878,73 tok/s TG Prefill 7.717 · TTFT 6.690 ms | 10× | NVIDIA GeForce RTX 5090AMD Ryzen 7 5800X3D 8-Core Processor · 126 GB RAM | llama.cppcodex_cli | Details → | |
| 6 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 1.830,31 tok/s TG Prefill 9.399 · TTFT 6.598 ms | 10× | 3x NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen Threadripper PRO 9965WX 24-Cores · 125 GB RAM | llama.cppgodclaw | Details → | |
| 7 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 1.196,98 tok/s TG Prefill 6.074 · TTFT 10.280 ms | 10× | NVIDIA GeForce RTX 5070 TiAMD Ryzen Threadripper PRO 5975WX 32-Cores · 247 GB RAM | llama.cppgodclaw | Details → | |
| 8 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 1.188,73 tok/s TG Prefill 9.032 · TTFT 2.065 ms | 5× | NVIDIA RTX PRO 6000 Blackwell Workstation EditionAMD Ryzen 9 9950X 16-Core Processor · 92 GB RAM | llama.cppgodclaw | Details → | |
| 9 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 1.164,80 tok/s TG Prefill 6.428 · TTFT 2.540 ms | 5× | NVIDIA GeForce RTX 5090AMD Ryzen 7 5800X3D 8-Core Processor · 126 GB RAM | llama.cppcodex_cli | Details → | |
| 10 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 1.090,59 tok/s TG Prefill 9.551 · TTFT 2.101 ms | 5× | 3x NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen Threadripper PRO 9965WX 24-Cores · 125 GB RAM | llama.cppgodclaw | Details → | |
| 11 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 1.077,40 tok/s TG Prefill 9.899 · TTFT 2.078 ms | 5× | 3x NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen Threadripper PRO 9965WX 24-Cores · 125 GB RAM | llama.cppgodclaw | Details → | |
| 12 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 1.071,04 tok/s TG Prefill 6.552 · TTFT 2.620 ms | 5× | 3x NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen Threadripper PRO 9965WX 24-Cores · 125 GB RAM | llama.cppgodclaw | Details → | |
| 13 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 1.051,12 tok/s TG Prefill 7.358 · TTFT 2.444 ms | 5× | 3x NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen Threadripper PRO 9965WX 24-Cores · 125 GB RAM | llama.cppgodclaw | Details → | |
| 14 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 992,11 tok/s TG Prefill 5.920 · TTFT 12.083 ms | 10× | NVIDIA GeForce RTX 3090 TiAMD Ryzen 9 8945HX with Radeon Graphics · 92 GB RAM | llama.cppwebsocket | Details → | |
| 15 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 676,42 tok/s TG Prefill 5.367 · TTFT 3.579 ms | 5× | NVIDIA GeForce RTX 5070 TiAMD Ryzen Threadripper PRO 5975WX 32-Cores · 247 GB RAM | llama.cppgodclaw | Details → | |
| 16 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 667,02 tok/s TG Prefill 4.019 · TTFT 17.735 ms | 10× | 3x AMD Radeon AI PRO R9700AMD Ryzen Threadripper PRO 7955WX 16-Cores · 184 GB RAM | llama.cppgodclaw | Details → | |
| 17 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 655,34 tok/s TG Prefill 4.936 · TTFT 17.123 ms | 10× | 3x AMD Radeon AI PRO R9700AMD Ryzen Threadripper PRO 7955WX 16-Cores · 184 GB RAM | llama.cppgodclaw | Details → | |
| 18 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 629,37 tok/s TG Prefill 4.530 · TTFT 17.991 ms | 10× | 3x AMD Radeon AI PRO R9700AMD Ryzen Threadripper PRO 7955WX 16-Cores · 184 GB RAM | llama.cppgodclaw | Details → | |
| 19 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 596,31 tok/s TG Prefill 5.389 · TTFT 4.001 ms | 5× | NVIDIA GeForce RTX 3090 TiAMD Ryzen 9 8945HX with Radeon Graphics · 92 GB RAM | llama.cppwebsocket | Details → | |
| 20 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 393,52 tok/s TG Prefill 1.898 · TTFT 1.302 ms | 1× | NVIDIA GeForce RTX 5090AMD Ryzen 7 5800X3D 8-Core Processor · 126 GB RAM | llama.cppcodex_cli | Details → | |
| 21 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 393,46 tok/s TG Prefill 4.408 · TTFT 489 ms | 1× | NVIDIA RTX PRO 6000 Blackwell Workstation EditionAMD Ryzen 9 9950X 16-Core Processor · 92 GB RAM | llama.cppgodclaw | Details → | |
| 22 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 382,25 tok/s TG Prefill 3.656 · TTFT 5.903 ms | 5× | 3x AMD Radeon AI PRO R9700AMD Ryzen Threadripper PRO 7955WX 16-Cores · 184 GB RAM | llama.cppgodclaw | Details → | |
| 23 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 375,87 tok/s TG Prefill 4.619 · TTFT 5.434 ms | 5× | 3x AMD Radeon AI PRO R9700AMD Ryzen Threadripper PRO 7955WX 16-Cores · 184 GB RAM | llama.cppgodclaw | Details → | |
| 24 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 367,75 tok/s TG Prefill 3.397 · TTFT 29.121 ms | 10× | NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAM | llama.cppgodclaw | Details → | |
| 25 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 362,28 tok/s TG Prefill 4.209 · TTFT 548 ms | 1× | 3x NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen Threadripper PRO 9965WX 24-Cores · 125 GB RAM | llama.cppgodclaw | Details → | |
| 26 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 361,36 tok/s TG Prefill 5.030 · TTFT 428 ms | 1× | 3x NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen Threadripper PRO 9965WX 24-Cores · 125 GB RAM | llama.cppgodclaw | Details → | |
| 27 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 358,48 tok/s TG Prefill 4.148 · TTFT 5.814 ms | 5× | 3x AMD Radeon AI PRO R9700AMD Ryzen Threadripper PRO 7955WX 16-Cores · 184 GB RAM | llama.cppgodclaw | Details → | |
| 28 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 356,73 tok/s TG Prefill 6.221 · TTFT 346 ms | 1× | 3x NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen Threadripper PRO 9965WX 24-Cores · 125 GB RAM | llama.cppgodclaw | Details → | |
| 29 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 355,57 tok/s TG Prefill 2.753 · TTFT 1.003 ms | 1× | 3x NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen Threadripper PRO 9965WX 24-Cores · 125 GB RAM | llama.cppgodclaw | Details → | |
| 30 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 354,29 tok/s TG Prefill 3.544 · TTFT 30.045 ms | 10× | NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAM | llama.cppgodclaw | Details → | |
| 31 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 236,91 tok/s TG Prefill 4.079 · TTFT 528 ms | 1× | NVIDIA GeForce RTX 3090 TiAMD Ryzen 9 8945HX with Radeon Graphics · 92 GB RAM | llama.cppwebsocket | Details → | |
| 32 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 233,69 tok/s TG Prefill 2.376 · TTFT 950 ms | 1× | NVIDIA GeForce RTX 5070 TiAMD Ryzen Threadripper PRO 5975WX 32-Cores · 247 GB RAM | llama.cppgodclaw | Details → | |
| 33 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 210,25 tok/s TG Prefill 3.483 · TTFT 8.798 ms | 5× | NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAM | llama.cppgodclaw | Details → | |
| 34 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 202,30 tok/s TG Prefill 3.149 · TTFT 9.380 ms | 5× | NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAM | llama.cppgodclaw | Details → | |
| 35 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 139,75 tok/s TG Prefill 3.025 · TTFT 712 ms | 1× | 3x AMD Radeon AI PRO R9700AMD Ryzen Threadripper PRO 7955WX 16-Cores · 184 GB RAM | llama.cppgodclaw | Details → | |
| 36 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 136,24 tok/s TG Prefill 3.620 · TTFT 595 ms | 1× | 3x AMD Radeon AI PRO R9700AMD Ryzen Threadripper PRO 7955WX 16-Cores · 184 GB RAM | llama.cppgodclaw | Details → | |
| 37 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 129,52 tok/s TG Prefill 3.211 · TTFT 671 ms | 1× | 3x AMD Radeon AI PRO R9700AMD Ryzen Threadripper PRO 7955WX 16-Cores · 184 GB RAM | llama.cppgodclaw | Details → | |
| 38 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) i | 125,68 tok/s TG Prefill 2.208 · TTFT 11.151 ms | 10× | 2x NVIDIA GeForce RTX 2060Intel(R) Core(TM) i5-7400 CPU @ 3.00GHz · 7 GB RAM | llama.cppopenclaw_cli | Details → | |
| 39 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) i | 116,24 tok/s TG Prefill 1.891 · TTFT 5.975 ms | 5× | 2x NVIDIA GeForce RTX 2060Intel(R) Core(TM) i5-7400 CPU @ 3.00GHz · 7 GB RAM | llama.cppopenclaw_cli | Details → | |
| 40 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) i | 106,92 tok/s TG Prefill 1.319 · TTFT 18.921 ms | 10× | 2x NVIDIA GeForce RTX 2060Intel(R) Core(TM) i5-7400 CPU @ 3.00GHz · 7 GB RAM | llama.cppopenclaw_cli | Details → | |
| 41 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 106,11 tok/s TG Prefill 1.440 · TTFT 17.007 ms | 10× | 2x NVIDIA GeForce RTX 2060Intel(R) Core(TM) i5-7400 CPU @ 3.00GHz · 7 GB RAM | llama.cppopenclaw_cli | Details → | |
| 42 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) i | 100,31 tok/s TG Prefill 1.092 · TTFT 10.399 ms | 5× | 2x NVIDIA GeForce RTX 2060Intel(R) Core(TM) i5-7400 CPU @ 3.00GHz · 7 GB RAM | llama.cppopenclaw_cli | Details → | |
| 43 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 100,04 tok/s TG Prefill 1.231 · TTFT 9.195 ms | 5× | 2x NVIDIA GeForce RTX 2060Intel(R) Core(TM) i5-7400 CPU @ 3.00GHz · 7 GB RAM | llama.cppopenclaw_cli | Details → | |
| 44 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) i | 93,42 tok/s TG Prefill 2.630 · TTFT 822 ms | 1× | 3x AMD Radeon AI PRO R9700AMD Ryzen Threadripper PRO 7955WX 16-Cores · 184 GB RAM | llama.cppgodclaw | Details → | |
| 45 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) i | 80,94 tok/s TG Prefill 1.564 · TTFT 1.376 ms | 1× | 2x NVIDIA GeForce RTX 2060Intel(R) Core(TM) i5-7400 CPU @ 3.00GHz · 7 GB RAM | llama.cppopenclaw_cli | Details → | |
| 46 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 76,87 tok/s TG Prefill 2.223 · TTFT 1.006 ms | 1× | NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAM | llama.cppgodclaw | Details → | |
| 47 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 72,70 tok/s TG Prefill 1.798 · TTFT 1.470 ms | 1× | NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAM | llama.cppgodclaw | Details → | |
| 48 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) i | 72,37 tok/s TG Prefill 865 · TTFT 2.489 ms | 1× | 2x NVIDIA GeForce RTX 2060Intel(R) Core(TM) i5-7400 CPU @ 3.00GHz · 7 GB RAM | llama.cppopenclaw_cli | Details → | |
| 49 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 71,70 tok/s TG Prefill 945 · TTFT 2.277 ms | 1× | 2x NVIDIA GeForce RTX 2060Intel(R) Core(TM) i5-7400 CPU @ 3.00GHz · 7 GB RAM | llama.cppopenclaw_cli | Details → | |
| 50 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 7,97 tok/s TG Prefill 72 · TTFT 153.653 ms | 5× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclaw | Details → | |
| 51 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 4,94 tok/s TG Prefill 41 · TTFT 62.079 ms | 1× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclaw | Details → | |
| 52 | Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 0,25 tok/s TG Prefill 75 · TTFT 202.135 ms | 10× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclaw | Details → |
