Created bymario-alka.dePowered bygodcore.denoob2claw.detricoma.de
🧠
Processor

10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85

Unbekannt
72Benchmarks
17Models
68Performance
4Harness
Best generation
903,3 tok/s
gemma-4-E2B-it
∅ Generation
167,1 tok/s
Schnitt aus 68 Läufen
Best prefill
7.470 tok/s
Prompt processing
Best TTFT
285 ms
Time to First Token
Best harness rate
71%
4 Harness-Läufe
🏆
gemma-4-E2B-it läuft am schnellsten auf dieser CPU · 903,3 tok/s
Bester Prefill: 7.470 tok/s · 17 Modelle getestet

Top models

Bester gemessener Durchsatz je Modell auf 10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 – umschaltbar nach Generation, Prefill oder Kombiwert.

gemma-4-E2B-it5B
903,3 tok/s
gemma-4-E4B-it8B
488,2 tok/s
Nemotron-3-Nano-4B4B
367,8 tok/s
Ling-3.0-flash127.5B
253,7 tok/s
gpt-oss-20b20B
202,1 tok/s
MiniMax-M2.7230B
174,2 tok/s
Qwen3.6-35B-A3B35B
167,1 tok/s
Performance profile

Durchsatz auf 10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85

Jede Blase steht für ein Modell, eine Engine bzw. eine GPU – Position: Prompt-Verarbeitung (X) × Ausgabe (Y), Blasengröße: Anzahl Messläufe.

MODnach Modell

1.1728795862930,002.5775.1557.732Prefill (tok/s)Generation (tok/s)gemma-4-E2B-it - 903,3 tok/s Generation, 6.421 tok/s Prefill, TTFT 7.548 ms (4 Laufe)gemma-4-E2B-itgemma-4-E4B-it - 488,2 tok/s Generation, 3.420 tok/s Prefill, TTFT 10.757 ms (8 Laufe)gemma-4-E4B-itNemotron-3-Nano-4B - 367,8 tok/s Generation, 2.933 tok/s Prefill, TTFT 13.303 ms (6 Laufe)Nemotron-3-Nano-4BLing-3.0-flash - 253,7 tok/s Generation, 1.278 tok/s Prefill, TTFT 22.647 ms (15 Laufe)Ling-3.0-flashgpt-oss-20b - 202,1 tok/s Generation, 4.104 tok/s Prefill, TTFT 2.448 ms (3 Laufe)gpt-oss-20bQwen3-Coder-30B-A3B-Instruct - 182,9 tok/s Generation, 4.998 tok/s Prefill, TTFT 2.913 ms (3 Laufe)Qwen3-Coder-30B-A3B-I...MiniMax-M2.7 - 174,2 tok/s Generation, 4.335 tok/s Prefill, TTFT 20.488 ms (3 Laufe)MiniMax-M2.7Qwen3-30B-A3B-Instruct-2507 - 169,9 tok/s Generation, 5.183 tok/s Prefill, TTFT 2.923 ms (3 Laufe)Qwen3-30B-A3B-Instruc...Qwen3.6-35B-A3B - 167,1 tok/s Generation, 4.210 tok/s Prefill, TTFT 1.687 ms (8 Laufe)Qwen3.6-35B-A3BNemotron-Cascade-2-30B-A3B - 148,8 tok/s Generation, 2.840 tok/s Prefill, TTFT 3.808 ms (6 Laufe)Nemotron-Cascade-2-30...MiniMax-M2.5 - 103,5 tok/s Generation, 4.887 tok/s Prefill, TTFT 2.995 ms (3 Laufe)MiniMax-M2.5Nemotron-3-Nano-30B-A3B - 58,3 tok/s Generation, 5.466 tok/s Prefill, TTFT 395 ms (1 Lauf)Nemotron-3-Nano-30B-A...glm-4.7-flash - 48,9 tok/s Generation, 4.055 tok/s Prefill, TTFT 536 ms (1 Lauf)glm-4.7-flashGemma-4-26B-A4B - 47,1 tok/s Generation, 4.378 tok/s Prefill, TTFT 451 ms (1 Lauf)Gemma-4-26B-A4BQwen3.5-35B-A3B - 46,9 tok/s Generation, 3.713 tok/s Prefill, TTFT 537 ms (1 Lauf)Qwen3.5-35B-A3BQwen3-Coder-Next - 38,5 tok/s Generation, 2.124 tok/s Prefill, TTFT 1.146 ms (1 Lauf)Qwen3-Coder-NextNemotron-3-Super-120B-A12B - 11,4 tok/s Generation, 1.495 tok/s Prefill, TTFT 1.449 ms (1 Lauf)Nemotron-3-Super-120B...
gemma-4-E2B-it 903,3 tok/sgemma-4-E4B-it 488,2 tok/sNemotron-3-Nano-4B 367,8 tok/sLing-3.0-flash 253,7 tok/sgpt-oss-20b 202,1 tok/sQwen3-Coder-30B-A3B-Instruct 182,9 tok/sMiniMax-M2.7 174,2 tok/sQwen3-30B-A3B-Instruct-2507 169,9 tok/sQwen3.6-35B-A3B 167,1 tok/sNemotron-Cascade-2-30B-A3B 148,8 tok/sMiniMax-M2.5 103,5 tok/sNemotron-3-Nano-30B-A3B 58,3 tok/sglm-4.7-flash 48,9 tok/sGemma-4-26B-A4B 47,1 tok/sQwen3.5-35B-A3B 46,9 tok/sQwen3-Coder-Next 38,5 tok/sNemotron-3-Super-120B-A12B 11,4 tok/s

ENGnach Engine

1.1398555702850,02.5853.3584.1324.906Prefill (tok/s)Generation (tok/s)llama.cpp - 903,3 tok/s Generation, 3.086 tok/s Prefill, TTFT 11.666 ms (51 Laufe)llama.cppvLLM - 174,2 tok/s Generation, 4.404 tok/s Prefill, TTFT 4.526 ms (17 Laufe)vLLM
llama.cpp 903,3 tok/svLLM 174,2 tok/s

GPUnach Grafikkarte

9949489038588133.2113.3473.4843.621Prefill (tok/s)Generation (tok/s)NVIDIA GB10 (DGX Spark) - 903,3 tok/s Generation, 3.416 tok/s Prefill, TTFT 9.881 ms (68 Laufe)NVIDIA GB10 (DGX Spar...
NVIDIA GB10 (DGX Spark) 903,3 tok/s
Platform

Mainboards with this CPU

Auf diesen Boards wurde 10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 gemessen – anklicken für Board-Details mit CPUs und Modellen.

Throughput & latency

Performancebenchmark

68 veroeffentlichte Performance-Läufe auf 10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85.

In the leaderboard →
Metric:
Size
#Model / MakerMetricsParallelGPU / CPU / RAMRuntime
1gemma-4-E2B-it5BGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
903,30 tok/s TG
Prefill 6.960 · TTFT 11.382 ms
10×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppgodclawQ4_K_M
2gemma-4-E2B-it5BGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
860,88 tok/s TG
Prefill 6.641 · TTFT 11.413 ms
10×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppgodclawQ4_K_M
3gemma-4-E4B-it8BGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
488,18 tok/s TG
Prefill 4.229 · TTFT 20.915 ms
10×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppgodclawQ4_K_M
4gemma-4-E4B-it8BGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
484,77 tok/s TG
Prefill 4.124 · TTFT 21.088 ms
10×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppgodclawQ4_K_M
5gemma-4-E2B-it5BGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
478,55 tok/s TG
Prefill 6.243 · TTFT 3.601 ms
5×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppgodclawQ4_K_M
6gemma-4-E2B-it5BGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
461,54 tok/s TG
Prefill 5.838 · TTFT 3.795 ms
5×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppgodclawQ4_K_M
7gemma-4-E4B-it8BGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
455,06 tok/s TG
Prefill 3.954 · TTFT 21.269 ms
10×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppgodclawQ4_K_M
8Nemotron-3-Nano-4B4BNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
367,75 tok/s TG
Prefill 3.397 · TTFT 29.121 ms
10×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppgodclawQ4_K_M
9Nemotron-3-Nano-4B4BNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
354,29 tok/s TG
Prefill 3.544 · TTFT 30.045 ms
10×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppgodclawQ4_K_M
10Ling-3.0-flash127.5BinclusionAI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
253,67 tok/s TG
Prefill 1.397 · TTFT 48.511 ms
10×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppgodclawQ4_K_M
11gemma-4-E4B-it8BGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
251,39 tok/s TG
Prefill 3.708 · TTFT 6.599 ms
5×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppgodclawQ4_K_M
12Ling-3.0-flash127.5BinclusionAI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
250,66 tok/s TG
Prefill 1.396 · TTFT 48.986 ms
10×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppgodclawQ4_K_M
13Ling-3.0-flash127.5BinclusionAI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
247,76 tok/s TG
Prefill 1.399 · TTFT 49.349 ms
10×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppgodclawQ4_K_M
14gemma-4-E4B-it8BGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
246,68 tok/s TG
Prefill 3.165 · TTFT 7.126 ms
5×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppgodclawQ4_K_M
15gemma-4-E4B-it8BGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
245,62 tok/s TG
Prefill 3.740 · TTFT 6.690 ms
5×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppgodclawQ4_K_M
16Nemotron-3-Nano-4B4BNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
210,25 tok/s TG
Prefill 3.483 · TTFT 8.798 ms
5×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppgodclawQ4_K_M
17Nemotron-3-Nano-4B4BNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
202,30 tok/s TG
Prefill 3.149 · TTFT 9.380 ms
5×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppgodclawQ4_K_M
18gpt-oss-20b20BOpenAI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
202,12 tok/s TG
Prefill 4.412 · TTFT 4.498 ms
10×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppopenclaw_cliMXFP4
19Qwen3-Coder-30B-A3B-Instruct30BQwen (Alibaba) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
182,87 tok/s TG
Prefill 7.295 · TTFT 5.014 ms
10×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
20MiniMax-M2.7230BMiniMax Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
174,20 tok/s TG
Prefill 5.042 · TTFT 47.620 ms
10×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMvLLMgodclawAWQ
21Qwen3-30B-A3B-Instruct-250730BQwen (Alibaba) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
169,89 tok/s TG
Prefill 7.470 · TTFT 5.007 ms
10×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
22Qwen3.6-35B-A3B35BQwen (Alibaba) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
167,11 tok/s TG
Prefill 2.814 · TTFT 6.937 ms
10×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
23gpt-oss-20b20BOpenAI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
148,88 tok/s TG
Prefill 4.312 · TTFT 2.296 ms
5×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppopenclaw_cliMXFP4
24Nemotron-Cascade-2-30B-A3B30BNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
148,79 tok/s TG
Prefill 3.362 · TTFT 6.588 ms
10×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
25Qwen3-Coder-30B-A3B-Instruct30BQwen (Alibaba) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
146,91 tok/s TG
Prefill 4.736 · TTFT 2.900 ms
5×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
26Ling-3.0-flash127.5BinclusionAI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
142,76 tok/s TG
Prefill 1.244 · TTFT 17.133 ms
5×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppgodclawQ4_K_M
27Ling-3.0-flash127.5BinclusionAI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
140,35 tok/s TG
Prefill 1.237 · TTFT 17.315 ms
5×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppgodclawQ4_K_M
28Ling-3.0-flash127.5BinclusionAI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
139,82 tok/s TG
Prefill 1.244 · TTFT 17.279 ms
5×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppgodclawQ4_K_M
29Qwen3-30B-A3B-Instruct-250730BQwen (Alibaba) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
139,24 tok/s TG
Prefill 5.091 · TTFT 2.946 ms
5×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
30Qwen3.6-35B-A3B35BQwen (Alibaba) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
135,61 tok/s TG
Prefill 2.590 · TTFT 3.731 ms
5×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
31Nemotron-Cascade-2-30B-A3B30BNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
123,54 tok/s TG
Prefill 2.897 · TTFT 3.743 ms
5×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
32Ling-3.0-flash127.5BinclusionAI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
108,98 tok/s TG
Prefill 1.446 · TTFT 66.475 ms
10×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
33MiniMax-M2.7230BMiniMax Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
106,12 tok/s TG
Prefill 5.634 · TTFT 12.854 ms
5×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMvLLMgodclawAWQ
34MiniMax-M2.5230BMiniMax Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
103,49 tok/s TG
Prefill 7.016 · TTFT 5.294 ms
10×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMvLLMgodclawAWQ
35Nemotron-Cascade-2-30B-A3B30BNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
101,36 tok/s TG
Prefill 3.237 · TTFT 6.850 ms
10×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppopenclaw_cliQ8_0
36Qwen3.6-35B-A3B35BQwen (Alibaba) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
101,18 tok/s TG
Prefill 6.144 · TTFT 314 ms
1×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMvLLMgodclawNVFP4
37Qwen3.6-35B-A3B35BQwen (Alibaba) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
100,59 tok/s TG
Prefill 6.798 · TTFT 285 ms
1×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMvLLMgodclawNVFP4
38Qwen3.6-35B-A3B35BQwen (Alibaba) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
92,96 tok/s TG
Prefill 4.802 · TTFT 404 ms
1×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMvLLMgodclawNVFP4
39Ling-3.0-flash127.5BinclusionAI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
90,05 tok/s TG
Prefill 1.600 · TTFT 43.239 ms
10×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
40Qwen3-Coder-30B-A3B-Instruct30BQwen (Alibaba) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
86,23 tok/s TG
Prefill 2.963 · TTFT 824 ms
1×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
41Nemotron-Cascade-2-30B-A3B30BNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
84,61 tok/s TG
Prefill 2.849 · TTFT 3.806 ms
5×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppopenclaw_cliQ8_0
42Qwen3-30B-A3B-Instruct-250730BQwen (Alibaba) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
84,18 tok/s TG
Prefill 2.987 · TTFT 816 ms
1×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
43gpt-oss-20b20BOpenAI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
82,16 tok/s TG
Prefill 3.589 · TTFT 550 ms
1×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppopenclaw_cliMXFP4
44MiniMax-M2.5230BMiniMax Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
77,74 tok/s TG
Prefill 5.460 · TTFT 2.641 ms
5×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMvLLMgodclawAWQ
45Nemotron-Cascade-2-30B-A3B30BNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
76,95 tok/s TG
Prefill 2.356 · TTFT 929 ms
1×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
46Nemotron-3-Nano-4B4BNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
76,87 tok/s TG
Prefill 2.223 · TTFT 1.006 ms
1×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppgodclawQ4_K_M
47Qwen3.6-35B-A3B35BQwen (Alibaba) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
74,19 tok/s TG
Prefill 4.923 · TTFT 393 ms
1×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMvLLMgodclawNVFP4
48Nemotron-3-Nano-4B4BNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
72,70 tok/s TG
Prefill 1.798 · TTFT 1.470 ms
1×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppgodclawQ4_K_M
49Qwen3.6-35B-A3B35BQwen (Alibaba) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
68,40 tok/s TG
Prefill 2.292 · TTFT 842 ms
1×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
50gemma-4-E4B-it8BGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
63,50 tok/s TG
Prefill 2.576 · TTFT 769 ms
1×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppgodclawQ4_K_M
51Nemotron-Cascade-2-30B-A3B30BNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
62,04 tok/s TG
Prefill 2.341 · TTFT 935 ms
1×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppopenclaw_cliQ8_0
52gemma-4-E4B-it8BGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
60,50 tok/s TG
Prefill 1.862 · TTFT 1.599 ms
1×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppgodclawQ4_K_M
53Nemotron-3-Nano-30B-A3B30BNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
58,25 tok/s TG
Prefill 5.466 · TTFT 395 ms
1×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMvLLMgodclawNVFP4
54Ling-3.0-flash127.5BinclusionAI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
49,10 tok/s TG
Prefill 1.157 · TTFT 10.523 ms
5×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
55glm-4.7-flashZ.ai (Zhipu) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
48,92 tok/s TG
Prefill 4.055 · TTFT 536 ms
1×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMvLLMgodclawAWQ
56Ling-3.0-flash127.5BinclusionAI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
48,48 tok/s TG
Prefill 1.201 · TTFT 2.031 ms
1×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppgodclawQ4_K_M
57Ling-3.0-flash127.5BinclusionAI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
48,44 tok/s TG
Prefill 1.189 · TTFT 2.051 ms
1×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppgodclawQ4_K_M
58Ling-3.0-flash127.5BinclusionAI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
47,20 tok/s TG
Prefill 1.176 · TTFT 2.071 ms
1×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppgodclawQ4_K_M
59Gemma-4-26B-A4B26BGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
47,06 tok/s TG
Prefill 4.378 · TTFT 451 ms
1×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMvLLMgodclawNVFP4
60Qwen3.5-35B-A3B35BQwen (Alibaba) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
46,91 tok/s TG
Prefill 3.713 · TTFT 537 ms
1×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMvLLMgodclawFP8
61Qwen3.6-35B-A3B35BQwen (Alibaba) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
46,79 tok/s TG
Prefill 3.315 · TTFT 589 ms
1×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMvLLMgodclawFP8
62Ling-3.0-flash127.5BinclusionAI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
46,50 tok/s TG
Prefill 1.156 · TTFT 10.550 ms
5×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
63Qwen3-Coder-NextQwen (Alibaba) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
38,54 tok/s TG
Prefill 2.124 · TTFT 1.146 ms
1×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMvLLMgodclawFP8
64Ling-3.0-flash127.5BinclusionAI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
33,12 tok/s TG
Prefill 1.171 · TTFT 2.084 ms
1×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
65MiniMax-M2.5230BMiniMax Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
32,23 tok/s TG
Prefill 2.184 · TTFT 1.050 ms
1×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMvLLMgodclawAWQ
66MiniMax-M2.7230BMiniMax Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
31,52 tok/s TG
Prefill 2.328 · TTFT 991 ms
1×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMvLLMgodclawAWQ
67Ling-3.0-flash127.5BinclusionAI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
28,27 tok/s TG
Prefill 1.151 · TTFT 2.118 ms
1×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMllama.cppopenclaw_cliQ4_K_M
68Nemotron-3-Super-120B-A12B120BNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
11,39 tok/s TG
Prefill 1.495 · TTFT 1.449 ms
1×NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMvLLMgodclawNVFP4
Agent & chat rating

Harnessbenchmark

4 veroeffentlichte Harness-Läufe auf 10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85.

In the leaderboard →
Metric:
#Model / MakerMetricsGPU / CPU / RAMRuntime
1MiniMax-M2.5230BMiniMax Harness benchmarkTool Usage V 1.0
1.245 Pkt
70,7% · 15/15 Aufg. · 301,0 s
NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMopenclaw_cli
2MiniMax-M2.5230BMiniMax Harness benchmarkTool Usage Standard 1.0
1.762 Pkt
59,1% · 16/16 Aufg. · 526,0 s
NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMopenclaw_cli
3MiniMax-M2.5230BMiniMax Harness benchmarkTextgenerierung V 1.0
790 Pkt
52,0% · 15/15 Aufg. · 261,0 s
NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMopenclaw_cli
4MiniMax-M2.7230BMiniMax Harness benchmarkTool Usage Standard 1.0
858 Pkt
28,8% · 16/16 Aufg. · 1.222,0 s
NVIDIA GB10 (DGX Spark)10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85 · 120 GB RAMgodclaw

← All hardware

Technische Daten

Technische Daten – 10 x ARM implementer 0x41, part 0xd87 + 10 x ARM implementer 0x41, part 0xd85

HerstellerUnbekannt
Best generation903,3 tok/s
Best prefill7.470 tok/s
Models17
Messläufe (Performance)68