Created bymario-alka.dePowered bygodcore.denoob2claw.detricoma.de
🧠
Prozessor

NVIDIA Grace

NVIDIA
86Benchmarks
17Modelle
81Performance
5Harness
Beste Generation
903,3 tok/s
gemma-4-E2B-it
∅ Generation
149,8 tok/s
Schnitt aus 81 Läufen
Bester Prefill
7.470 tok/s
Prompt-Verarbeitung
Bester TTFT
105 ms
Time to First Token
Beste Harness-Quote
71%
5 Harness-Läufe
🏆
gemma-4-E2B-it läuft am schnellsten auf dieser CPU · 903,3 tok/s
Bester Prefill: 7.470 tok/s · 17 Modelle getestet

Top-Modelle

Bester gemessener Durchsatz je Modell auf NVIDIA Grace – umschaltbar nach Generation, Prefill oder Kombiwert.

gemma-4-E2B-it5B
903,3 tok/s
gemma-4-E4B-it8B
488,2 tok/s
Nemotron-3-Nano-4B4B
367,8 tok/s
Ling-3.0-flash127.5B
253,7 tok/s
gpt-oss-20b20B
202,1 tok/s
MiniMax-M2.7230B
174,2 tok/s
Qwen3.6-35B-A3B35B
167,1 tok/s
Performanceprofil

Durchsatz auf NVIDIA Grace

Jede Blase steht für ein Modell, eine Engine bzw. eine GPU – Position: Prompt-Verarbeitung (X) × Ausgabe (Y), Blasengröße: Anzahl Messläufe.

MODnach Modell

1.1678755832920,002.5775.1537.730Prefill (tok/s)Generation (tok/s)gemma-4-E2B-it - 903,3 tok/s Generation, 6.421 tok/s Prefill, TTFT 7.548 ms (4 Laufe)gemma-4-E2B-itgemma-4-E4B-it - 488,2 tok/s Generation, 3.420 tok/s Prefill, TTFT 10.757 ms (8 Laufe)gemma-4-E4B-itNemotron-3-Nano-4B - 367,8 tok/s Generation, 2.933 tok/s Prefill, TTFT 13.303 ms (6 Laufe)Nemotron-3-Nano-4BLing-3.0-flash - 253,7 tok/s Generation, 1.287 tok/s Prefill, TTFT 24.114 ms (14 Laufe)Ling-3.0-flashgpt-oss-20b - 202,1 tok/s Generation, 4.104 tok/s Prefill, TTFT 2.448 ms (3 Laufe)gpt-oss-20bQwen3-Coder-30B-A3B-Instruct - 182,9 tok/s Generation, 4.998 tok/s Prefill, TTFT 2.913 ms (3 Laufe)Qwen3-Coder-30B-A3B-I...MiniMax-M2.7 - 174,2 tok/s Generation, 4.952 tok/s Prefill, TTFT 9.535 ms (10 Laufe)MiniMax-M2.7Qwen3-30B-A3B-Instruct-2507 - 169,9 tok/s Generation, 5.183 tok/s Prefill, TTFT 2.923 ms (3 Laufe)Qwen3-30B-A3B-Instruc...Qwen3.6-35B-A3B - 167,1 tok/s Generation, 4.210 tok/s Prefill, TTFT 1.687 ms (8 Laufe)Qwen3.6-35B-A3BNemotron-Cascade-2-30B-A3B - 148,8 tok/s Generation, 2.840 tok/s Prefill, TTFT 3.808 ms (6 Laufe)Nemotron-Cascade-2-30...MiniMax-M2.5 - 105,6 tok/s Generation, 4.884 tok/s Prefill, TTFT 3.004 ms (6 Laufe)MiniMax-M2.5Nemotron-3-Nano-30B-A3B - 58,3 tok/s Generation, 5.466 tok/s Prefill, TTFT 395 ms (1 Lauf)Nemotron-3-Nano-30B-A...glm-4.7-flash - 48,9 tok/s Generation, 4.055 tok/s Prefill, TTFT 536 ms (1 Lauf)glm-4.7-flashGemma-4-26B-A4B - 47,1 tok/s Generation, 4.378 tok/s Prefill, TTFT 451 ms (1 Lauf)Gemma-4-26B-A4BQwen3.5-35B-A3B - 46,9 tok/s Generation, 3.713 tok/s Prefill, TTFT 537 ms (1 Lauf)Qwen3.5-35B-A3BQwen3-Coder-Next - 38,5 tok/s Generation, 2.124 tok/s Prefill, TTFT 1.146 ms (1 Lauf)Qwen3-Coder-Next
gemma-4-E2B-it 903,3 tok/sgemma-4-E4B-it 488,2 tok/sNemotron-3-Nano-4B 367,8 tok/sLing-3.0-flash 253,7 tok/sgpt-oss-20b 202,1 tok/sQwen3-Coder-30B-A3B-Instruct 182,9 tok/sMiniMax-M2.7 174,2 tok/sQwen3-30B-A3B-Instruct-2507 169,9 tok/sQwen3.6-35B-A3B 167,1 tok/sNemotron-Cascade-2-30B-A3B 148,8 tok/sMiniMax-M2.5 105,6 tok/sNemotron-3-Nano-30B-A3B 58,3 tok/sglm-4.7-flash 48,9 tok/sGemma-4-26B-A4B 47,1 tok/sQwen3.5-35B-A3B 46,9 tok/sQwen3-Coder-Next 38,5 tok/s

ENGnach Engine

1.1398555702850,02.5383.4844.4315.377Prefill (tok/s)Generation (tok/s)llama.cpp - 903,3 tok/s Generation, 3.125 tok/s Prefill, TTFT 11.857 ms (50 Laufe)llama.cppvLLM - 174,2 tok/s Generation, 4.790 tok/s Prefill, TTFT 4.555 ms (26 Laufe)vLLM
llama.cpp 903,3 tok/svLLM 174,2 tok/s

GPUnach Grafikkarte

9949489038588133.4733.6213.7683.916Prefill (tok/s)Generation (tok/s)NVIDIA GB10 (DGX Spark) - 903,3 tok/s Generation, 3.695 tok/s Prefill, TTFT 9.359 ms (76 Laufe)NVIDIA GB10 (DGX Spar...
NVIDIA GB10 (DGX Spark) 903,3 tok/s
Plattform

Mainboards mit dieser CPU

Auf diesen Boards wurde NVIDIA Grace gemessen – anklicken für Board-Details mit CPUs und Modellen.

Durchsatz & Latenz

Performancebenchmark

76 veroeffentlichte Performance-Läufe auf NVIDIA Grace.

Im Leaderboard →
Messwert:
Größe
#Modell / HerstellerMesswerteParallelGPU / CPU / RAMRuntime
1gemma-4-E2B-it5BGoogle PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
903,30 tok/s TG
Prefill 6.960 · TTFT 11.382 ms
10×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppgodclawQ4_K_M
2gemma-4-E2B-it5BGoogle PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
860,88 tok/s TG
Prefill 6.641 · TTFT 11.413 ms
10×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppgodclawQ4_K_M
3gemma-4-E4B-it8BGoogle PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
488,18 tok/s TG
Prefill 4.229 · TTFT 20.915 ms
10×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppgodclawQ4_K_M
4gemma-4-E4B-it8BGoogle PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
484,77 tok/s TG
Prefill 4.124 · TTFT 21.088 ms
10×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppgodclawQ4_K_M
5gemma-4-E2B-it5BGoogle PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
478,55 tok/s TG
Prefill 6.243 · TTFT 3.601 ms
5×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppgodclawQ4_K_M
6gemma-4-E2B-it5BGoogle PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
461,54 tok/s TG
Prefill 5.838 · TTFT 3.795 ms
5×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppgodclawQ4_K_M
7gemma-4-E4B-it8BGoogle PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
455,06 tok/s TG
Prefill 3.954 · TTFT 21.269 ms
10×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppgodclawQ4_K_M
8Nemotron-3-Nano-4B4BNVIDIA PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
367,75 tok/s TG
Prefill 3.397 · TTFT 29.121 ms
10×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppgodclawQ4_K_M
9Nemotron-3-Nano-4B4BNVIDIA PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
354,29 tok/s TG
Prefill 3.544 · TTFT 30.045 ms
10×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppgodclawQ4_K_M
10Ling-3.0-flash127.5BinclusionAI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
253,67 tok/s TG
Prefill 1.397 · TTFT 48.511 ms
10×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppgodclawQ4_K_M
11gemma-4-E4B-it8BGoogle PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
251,39 tok/s TG
Prefill 3.708 · TTFT 6.599 ms
5×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppgodclawQ4_K_M
12Ling-3.0-flash127.5BinclusionAI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
250,66 tok/s TG
Prefill 1.396 · TTFT 48.986 ms
10×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppgodclawQ4_K_M
13Ling-3.0-flash127.5BinclusionAI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
247,76 tok/s TG
Prefill 1.399 · TTFT 49.349 ms
10×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppgodclawQ4_K_M
14gemma-4-E4B-it8BGoogle PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
246,68 tok/s TG
Prefill 3.165 · TTFT 7.126 ms
5×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppgodclawQ4_K_M
15gemma-4-E4B-it8BGoogle PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
245,62 tok/s TG
Prefill 3.740 · TTFT 6.690 ms
5×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppgodclawQ4_K_M
16Nemotron-3-Nano-4B4BNVIDIA PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
210,25 tok/s TG
Prefill 3.483 · TTFT 8.798 ms
5×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppgodclawQ4_K_M
17Nemotron-3-Nano-4B4BNVIDIA PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
202,30 tok/s TG
Prefill 3.149 · TTFT 9.380 ms
5×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppgodclawQ4_K_M
18gpt-oss-20b20BOpenAI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
202,12 tok/s TG
Prefill 4.412 · TTFT 4.498 ms
10×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppopenclaw_cliMXFP4
19Qwen3-Coder-30B-A3B-Instruct30BQwen (Alibaba) PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
182,87 tok/s TG
Prefill 7.295 · TTFT 5.014 ms
10×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
20MiniMax-M2.7230BMiniMax PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
174,20 tok/s TG
Prefill 5.042 · TTFT 47.620 ms
10×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMvLLMgodclawAWQ
21Qwen3-30B-A3B-Instruct-250730BQwen (Alibaba) PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
169,89 tok/s TG
Prefill 7.470 · TTFT 5.007 ms
10×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
22Qwen3.6-35B-A3B35BQwen (Alibaba) PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
167,11 tok/s TG
Prefill 2.814 · TTFT 6.937 ms
10×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
23gpt-oss-20b20BOpenAI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
148,88 tok/s TG
Prefill 4.312 · TTFT 2.296 ms
5×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppopenclaw_cliMXFP4
24Nemotron-Cascade-2-30B-A3B30BNVIDIA PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
148,79 tok/s TG
Prefill 3.362 · TTFT 6.588 ms
10×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
25Qwen3-Coder-30B-A3B-Instruct30BQwen (Alibaba) PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
146,91 tok/s TG
Prefill 4.736 · TTFT 2.900 ms
5×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
26Ling-3.0-flash127.5BinclusionAI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
142,76 tok/s TG
Prefill 1.244 · TTFT 17.133 ms
5×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppgodclawQ4_K_M
27Ling-3.0-flash127.5BinclusionAI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
140,35 tok/s TG
Prefill 1.237 · TTFT 17.315 ms
5×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppgodclawQ4_K_M
28Ling-3.0-flash127.5BinclusionAI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
139,82 tok/s TG
Prefill 1.244 · TTFT 17.279 ms
5×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppgodclawQ4_K_M
29Qwen3-30B-A3B-Instruct-250730BQwen (Alibaba) PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
139,24 tok/s TG
Prefill 5.091 · TTFT 2.946 ms
5×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
30Qwen3.6-35B-A3B35BQwen (Alibaba) PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
135,61 tok/s TG
Prefill 2.590 · TTFT 3.731 ms
5×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
31Nemotron-Cascade-2-30B-A3B30BNVIDIA PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
123,54 tok/s TG
Prefill 2.897 · TTFT 3.743 ms
5×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
32Ling-3.0-flash127.5BinclusionAI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
108,98 tok/s TG
Prefill 1.446 · TTFT 66.475 ms
10×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
33MiniMax-M2.7230BMiniMax PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
108,63 tok/s TG
Prefill 4.568 · TTFT 13.673 ms
5×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMvLLMgodclawAWQ
34MiniMax-M2.7230BMiniMax PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
106,12 tok/s TG
Prefill 5.634 · TTFT 12.854 ms
5×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMvLLMgodclawAWQ
35MiniMax-M2.5230BMiniMax PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
105,57 tok/s TG
Prefill 6.997 · TTFT 5.319 ms
10×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMvLLMgodclawAWQ
36MiniMax-M2.5230BMiniMax PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
103,49 tok/s TG
Prefill 7.016 · TTFT 5.294 ms
10×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMvLLMgodclawAWQ
37MiniMax-M2.7230BMiniMax PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
103,25 tok/s TG
Prefill 6.620 · TTFT 5.066 ms
10×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMvLLMgodclawAWQ
38Nemotron-Cascade-2-30B-A3B30BNVIDIA PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
101,36 tok/s TG
Prefill 3.237 · TTFT 6.850 ms
10×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppopenclaw_cliQ8_0
39Qwen3.6-35B-A3B35BQwen (Alibaba) PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
101,18 tok/s TG
Prefill 6.144 · TTFT 314 ms
1×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMvLLMgodclawNVFP4
40Qwen3.6-35B-A3B35BQwen (Alibaba) PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
100,59 tok/s TG
Prefill 6.798 · TTFT 285 ms
1×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMvLLMgodclawNVFP4
41Qwen3.6-35B-A3B35BQwen (Alibaba) PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
92,96 tok/s TG
Prefill 4.802 · TTFT 404 ms
1×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMvLLMgodclawNVFP4
42Ling-3.0-flash127.5BinclusionAI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
90,05 tok/s TG
Prefill 1.600 · TTFT 43.239 ms
10×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
43Qwen3-Coder-30B-A3B-Instruct30BQwen (Alibaba) PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
86,23 tok/s TG
Prefill 2.963 · TTFT 824 ms
1×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
44Nemotron-Cascade-2-30B-A3B30BNVIDIA PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
84,61 tok/s TG
Prefill 2.849 · TTFT 3.806 ms
5×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppopenclaw_cliQ8_0
45Qwen3-30B-A3B-Instruct-250730BQwen (Alibaba) PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
84,18 tok/s TG
Prefill 2.987 · TTFT 816 ms
1×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
46gpt-oss-20b20BOpenAI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
82,16 tok/s TG
Prefill 3.589 · TTFT 550 ms
1×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppopenclaw_cliMXFP4
47MiniMax-M2.5230BMiniMax PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
77,74 tok/s TG
Prefill 5.460 · TTFT 2.641 ms
5×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMvLLMgodclawAWQ
48Nemotron-Cascade-2-30B-A3B30BNVIDIA PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
76,95 tok/s TG
Prefill 2.356 · TTFT 929 ms
1×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
49Nemotron-3-Nano-4B4BNVIDIA PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
76,87 tok/s TG
Prefill 2.223 · TTFT 1.006 ms
1×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppgodclawQ4_K_M
50Qwen3.6-35B-A3B35BQwen (Alibaba) PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
74,19 tok/s TG
Prefill 4.923 · TTFT 393 ms
1×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMvLLMgodclawNVFP4
51MiniMax-M2.5230BMiniMax PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
73,44 tok/s TG
Prefill 5.550 · TTFT 2.628 ms
5×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMvLLMgodclawAWQ
52Nemotron-3-Nano-4B4BNVIDIA PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
72,70 tok/s TG
Prefill 1.798 · TTFT 1.470 ms
1×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppgodclawQ4_K_M
53Qwen3.6-35B-A3B35BQwen (Alibaba) PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
68,40 tok/s TG
Prefill 2.292 · TTFT 842 ms
1×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
54MiniMax-M2.7230BMiniMax PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
66,75 tok/s TG
Prefill 6.946 · TTFT 4.901 ms
10×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMvLLMgodclawAWQ
55MiniMax-M2.7230BMiniMax PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
64,45 tok/s TG
Prefill 7.167 · TTFT 4.743 ms
10×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMvLLMgodclawAWQ
56gemma-4-E4B-it8BGoogle PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
63,50 tok/s TG
Prefill 2.576 · TTFT 769 ms
1×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppgodclawQ4_K_M
57Nemotron-Cascade-2-30B-A3B30BNVIDIA PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
62,04 tok/s TG
Prefill 2.341 · TTFT 935 ms
1×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppopenclaw_cliQ8_0
58gemma-4-E4B-it8BGoogle PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
60,50 tok/s TG
Prefill 1.862 · TTFT 1.599 ms
1×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppgodclawQ4_K_M
59Nemotron-3-Nano-30B-A3B30BNVIDIA PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
58,25 tok/s TG
Prefill 5.466 · TTFT 395 ms
1×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMvLLMgodclawNVFP4
60MiniMax-M2.7230BMiniMax PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
52,46 tok/s TG
Prefill 5.268 · TTFT 2.720 ms
5×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMvLLMgodclawAWQ
61MiniMax-M2.7230BMiniMax PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
51,61 tok/s TG
Prefill 5.448 · TTFT 2.676 ms
5×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMvLLMgodclawAWQ
62Ling-3.0-flash127.5BinclusionAI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
49,10 tok/s TG
Prefill 1.157 · TTFT 10.523 ms
5×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
63glm-4.7-flashZ.ai (Zhipu) PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
48,92 tok/s TG
Prefill 4.055 · TTFT 536 ms
1×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMvLLMgodclawAWQ
64Ling-3.0-flash127.5BinclusionAI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
48,48 tok/s TG
Prefill 1.201 · TTFT 2.031 ms
1×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppgodclawQ4_K_M
65Ling-3.0-flash127.5BinclusionAI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
48,44 tok/s TG
Prefill 1.189 · TTFT 2.051 ms
1×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppgodclawQ4_K_M
66Ling-3.0-flash127.5BinclusionAI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
47,20 tok/s TG
Prefill 1.176 · TTFT 2.071 ms
1×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppgodclawQ4_K_M
67Gemma-4-26B-A4B26BGoogle PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
47,06 tok/s TG
Prefill 4.378 · TTFT 451 ms
1×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMvLLMgodclawNVFP4
68Qwen3.5-35B-A3B35BQwen (Alibaba) PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
46,91 tok/s TG
Prefill 3.713 · TTFT 537 ms
1×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMvLLMgodclawFP8
69Qwen3.6-35B-A3B35BQwen (Alibaba) PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
46,79 tok/s TG
Prefill 3.315 · TTFT 589 ms
1×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMvLLMgodclawFP8
70Ling-3.0-flash127.5BinclusionAI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
46,50 tok/s TG
Prefill 1.156 · TTFT 10.550 ms
5×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
71Qwen3-Coder-NextQwen (Alibaba) PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
38,54 tok/s TG
Prefill 2.124 · TTFT 1.146 ms
1×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMvLLMgodclawFP8
72MiniMax-M2.7230BMiniMax PerformancebenchmarkPerformancetest Small 1.0
34,37 tok/s TG
Prefill 503 · TTFT 105 ms
1×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMvLLMgodclawAWQ
73Ling-3.0-flash127.5BinclusionAI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
33,12 tok/s TG
Prefill 1.171 · TTFT 2.084 ms
1×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
74MiniMax-M2.5230BMiniMax PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
32,23 tok/s TG
Prefill 2.184 · TTFT 1.050 ms
1×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMvLLMgodclawAWQ
75MiniMax-M2.5230BMiniMax PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
31,93 tok/s TG
Prefill 2.096 · TTFT 1.093 ms
1×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMvLLMgodclawAWQ
76MiniMax-M2.7230BMiniMax PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation)
31,52 tok/s TG
Prefill 2.328 · TTFT 991 ms
1×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMvLLMgodclawAWQ
Agent- & Chat-Bewertung

Harnessbenchmark

4 veroeffentlichte Harness-Läufe auf NVIDIA Grace.

Im Leaderboard →
Messwert:
#Modell / HerstellerMesswerteGPU / CPU / RAMRuntime
1MiniMax-M2.5230BMiniMax HarnessbenchmarkTool Usage V 1.0
1.245 Pkt
70,7% · 15/15 Aufg. · 301,0 s
NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMopenclaw_cli
2MiniMax-M2.7230BMiniMax HarnessbenchmarkTool-Parcours Dossier 40 V1.0
3.641 Pkt
61,1% · 40/40 Aufg. · 8.209,0 s
NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMvLLMgodclawAWQ
3MiniMax-M2.5230BMiniMax HarnessbenchmarkTool Usage Standard 1.0
1.762 Pkt
59,1% · 16/16 Aufg. · 526,0 s
NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMopenclaw_cli
4MiniMax-M2.5230BMiniMax HarnessbenchmarkTextgenerierung V 1.0
790 Pkt
52,0% · 15/15 Aufg. · 261,0 s
NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMopenclaw_cli

← Alle Hardware

Technische Daten

Technische Daten – NVIDIA Grace

HerstellerNVIDIA
Beste Generation903,3 tok/s
Bester Prefill7.470 tok/s
Modelle17
Messläufe (Performance)81