Created bymario-alka.dePowered bygodcore.denoob2claw.detricoma.de
Manufacturer

NVIDIA

Nemotron-Modellfamilie von NVIDIA.

12Models
567Benchmarks
567Performance
0Harness
Best token generation
2.182,7 tok/s
Nemotron-3-Nano-4B
Best prefill
44.171,1 tok/s
Nemotron-3-Nano-4B
∅ Token generation
244,5 tok/s
Average of 567 runs
∅ Prefill
3.504,1 tok/s
Average of 567 runs
Best TTFT
155 ms
Nemotron-3.5-Lightning-30B-A3B
Best harness rate
0%
no harness run yet
🏆
Nemotron-3-Nano-4B · fastest model at 2.182,7 tok/s token generation
Best prefill of this manufacturer: 44.171,1 tok/s (Nemotron-3-Nano-4B)

Top models by token generation

Best measured generation throughput per model (tok/s), published performance benchmarks.

Nemotron-3-Nano-4B
2.182,7 tok/s
Nemotron-3-Nano-Omni-30B-A3B-Reasoning
1.791,5 tok/s
Nemotron-3-Nano-30B-A3B
1.773,5 tok/s
Nemotron-Cascade-2-30B-A3B
1.768,1 tok/s
Nemotron-3-Super-120B-A12B
600,9 tok/s
OpenReasoning-Nemotron-32B
481,0 tok/s
Llama-3.3-Nemotron-Super-49B-v1.5
364,1 tok/s
Nemotron-3.5-Lightning-30B-A3B
290,3 tok/s