Created bymario-alka.dePowered bygodcore.denoob2claw.detricoma.de
Google

Gemma-4-26B-A4B26B

Benchmark profile and published results.

MoE26B
Position in the field

Best values compared

Best metrics of this model against the minimum, average and maximum of all published systems.

Generation47,1 tok/s
Min 0,0Ø 253,2Max 2.491,2
Unter dem Durchschnitt · 2823 Systeme im Feld
Prefill4.378 tok/s
Min 4Ø 3.470Max 39.029
Ueber dem Durchschnitt · 2823 Systeme im Feld
Time to First Token451 ms
Min 20Ø 45.698Max 535.235
Ueber dem Durchschnitt · 2821 Systeme im Feld
Performance profile

Throughput by hardware & engine

Each bubble is a GPU, CPU or engine – position shows prompt processing (X) and output speed (Y), bubble size the number of runs.

GPUby graphics card

51,849,447,144,742,44.1154.2904.4664.641Prefill (tok/s)Generation (tok/s)NVIDIA GB10 - 47,1 tok/s Generation, 4.378 tok/s Prefill, TTFT 451 ms (1 Lauf)NVIDIA GB10
NVIDIA GB10 47,1 tok/s

CPUby processor

51,849,447,144,742,44.1154.2904.4664.641Prefill (tok/s)Generation (tok/s)NVIDIA Grace - 47,1 tok/s Generation, 4.378 tok/s Prefill, TTFT 451 ms (1 Lauf)NVIDIA Grace
NVIDIA Grace 47,1 tok/s

ENGby engine

51,849,447,144,742,44.1154.2904.4664.641Prefill (tok/s)Generation (tok/s)vLLM - 47,1 tok/s Generation, 4.378 tok/s Prefill, TTFT 451 ms (1 Lauf)vLLM
vLLM 47,1 tok/s
Throughput & latency

Performance benchmark

Metric:
#Model / MakerMetricsParallelGPU / CPU / RAMRuntime
1Gemma-4-26B-A4B26BNVFP4Google Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
47,06 tok/s TG
Prefill 4.378 · TTFT 451 ms
1×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMvLLMgodclaw
Agent & chat rating

Harness benchmark

Noch keine Harnessbenchmarks fuer dieses Modell.