Created bymario-alka.dePowered bygodcore.denoob2claw.detricoma.de
Google

Gemma-4-26B-A4B26B

Benchmark profile and published results.

MoE26B
Position in the field

Best values compared

Best metrics of this model against the minimum, average and maximum of all published systems.

Generation47,1 tok/s
Min 11,4Ø 148,9Max 903,3
Unter dem Durchschnitt · 82 Systeme im Feld
Prefill4.378 tok/s
Min 503Ø 3.550Max 7.470
Ueber dem Durchschnitt · 82 Systeme im Feld
Time to First Token451 ms
Min 105Ø 8.759Max 66.475
Ueber dem Durchschnitt · 82 Systeme im Feld
Performance profile

Throughput by hardware & engine

Each bubble is a GPU, CPU or engine – position shows prompt processing (X) and output speed (Y), bubble size the number of runs.

GPUby graphics card

51,849,447,144,742,44.1154.2904.4664.641Prefill (tok/s)Generation (tok/s)NVIDIA GB10 - 47,1 tok/s Generation, 4.378 tok/s Prefill, TTFT 451 ms (1 Lauf)NVIDIA GB10
NVIDIA GB10 47,1 tok/s

CPUby processor

51,849,447,144,742,44.1154.2904.4664.641Prefill (tok/s)Generation (tok/s)NVIDIA Grace - 47,1 tok/s Generation, 4.378 tok/s Prefill, TTFT 451 ms (1 Lauf)NVIDIA Grace
NVIDIA Grace 47,1 tok/s

ENGby engine

51,849,447,144,742,44.1154.2904.4664.641Prefill (tok/s)Generation (tok/s)vLLM - 47,1 tok/s Generation, 4.378 tok/s Prefill, TTFT 451 ms (1 Lauf)vLLM
vLLM 47,1 tok/s
Throughput & latency

Performance benchmark

Metric:
#Model / MakerMetricsParallelGPU / CPU / RAMRuntime
1Gemma-4-26B-A4B26BNVFP4Google Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
47,06 tok/s TG
Prefill 4.378 · TTFT 451 ms
1×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMvLLMgodclaw
Agent & chat rating

Harness benchmark

Noch keine Harnessbenchmarks fuer dieses Modell.