N
Manufacturer
NVIDIA
Nemotron-Modellfamilie von NVIDIA.
Best token generation
2.182,7 tok/s
Nemotron-3-Nano-4B
Best prefill
44.171,1 tok/s
Nemotron-3-Nano-4B
∅ Token generation
244,5 tok/s
Average of 567 runs
∅ Prefill
3.504,1 tok/s
Average of 567 runs
Best TTFT
155 ms
Nemotron-3.5-Lightning-30B-A3B
Best harness rate
0%
no harness run yet
🏆
Nemotron-3-Nano-4B · fastest model at 2.182,7 tok/s token generation
Best prefill of this manufacturer: 44.171,1 tok/s (Nemotron-3-Nano-4B)
Best prefill of this manufacturer: 44.171,1 tok/s (Nemotron-3-Nano-4B)
Top models by token generation
Best measured generation throughput per model (tok/s), published performance benchmarks.
Models
All benchmarks →12 assigned
Llama-3.1-Nemotron-Ultra-253B-v1253B53 runs
4,3tok/s Ø
Min 0,1Max 62,1
Llama-3.3-Nemotron-Super-49B-v1.549B45 runs
77,6tok/s Ø
Min 2,5Max 364,1
Nemotron-3-Nano-30B-A3B30B83 runs
359,4tok/s Ø
Min 0,3Max 1.773,5
Nemotron-3-Nano-4B4B64 runs
552,8tok/s Ø
Min 0,3Max 2.182,7
Nemotron-3-Nano-Omni-30B-A3B-Reasoning30B61 runs
407,2tok/s Ø
Min 31,4Max 1.791,5
Nemotron-3-Super-120B-A12B120B56 runs
85,5tok/s Ø
Min 1,9Max 600,9
Nemotron-3-Super-120B-A12B-Q8120B0 runs
No performance data yet
Nemotron-3.5-Lightning-30B-A3B30B27 runs
186,2tok/s Ø
Min 50,7Max 290,3
Nemotron-Cascade-2-30B-A3B30B74 runs
365,5tok/s Ø
Min 0,3Max 1.768,1
Nemotron-H-47B-Reasoning-128K47B59 runs
45,9tok/s Ø
Min 3,2Max 203,5
Nemotron-Labs-3-Puzzle-75B-A9B75B0 runs
No performance data yet
OpenReasoning-Nemotron-32B32B45 runs
117,4tok/s Ø
Min 1,1Max 481,0
