N
Hersteller
NVIDIA
Nemotron-Modellfamilie von NVIDIA.
Beste Token-Generierung
2.182,7 tok/s
Nemotron-3-Nano-4B
Bester Prefill
44.171,1 tok/s
Nemotron-3-Nano-4B
∅ Token-Generierung
244,5 tok/s
Schnitt aus 567 Läufen
∅ Prefill
3.504,1 tok/s
Schnitt aus 567 Läufen
Bester TTFT
155 ms
Nemotron-3.5-Lightning-30B-A3B
Beste Harness-Quote
0%
noch kein Harness-Lauf
🏆
Nemotron-3-Nano-4B · schnellstes Modell mit 2.182,7 tok/s Token-Generierung
Bester Prefill des Herstellers: 44.171,1 tok/s (Nemotron-3-Nano-4B)
Bester Prefill des Herstellers: 44.171,1 tok/s (Nemotron-3-Nano-4B)
Top-Modelle nach Token-Generierung
Bester gemessener Generierungs-Durchsatz je Modell (tok/s), veröffentlichte Performancebenchmarks.
Modelle
Alle Benchmarks →12 zugeordnet
Llama-3.1-Nemotron-Ultra-253B-v1253B53 Läufe
4,3tok/s Ø
Min 0,1Max 62,1
Llama-3.3-Nemotron-Super-49B-v1.549B45 Läufe
77,6tok/s Ø
Min 2,5Max 364,1
Nemotron-3-Nano-30B-A3B30B83 Läufe
359,4tok/s Ø
Min 0,3Max 1.773,5
Nemotron-3-Nano-4B4B64 Läufe
552,8tok/s Ø
Min 0,3Max 2.182,7
Nemotron-3-Nano-Omni-30B-A3B-Reasoning30B61 Läufe
407,2tok/s Ø
Min 31,4Max 1.791,5
Nemotron-3-Super-120B-A12B120B56 Läufe
85,5tok/s Ø
Min 1,9Max 600,9
Nemotron-3-Super-120B-A12B-Q8120B0 Läufe
Noch keine Performancedaten
Nemotron-3.5-Lightning-30B-A3B30B27 Läufe
186,2tok/s Ø
Min 50,7Max 290,3
Nemotron-Cascade-2-30B-A3B30B74 Läufe
365,5tok/s Ø
Min 0,3Max 1.768,1
Nemotron-H-47B-Reasoning-128K47B59 Läufe
45,9tok/s Ø
Min 3,2Max 203,5
Nemotron-Labs-3-Puzzle-75B-A9B75B0 Läufe
Noch keine Performancedaten
OpenReasoning-Nemotron-32B32B45 Läufe
117,4tok/s Ø
Min 1,1Max 481,0
