Google
gemma-4-26B-A4B-it-assistant26B
Benchmark profile and published results.
MoE26B
Position in the field
Best values compared
Best metrics of this model against the minimum, average and maximum of all published systems.
Generation74,5 tok/s
Min 7,0Ø 61,0Max 326,9
Ueber dem Durchschnitt · 47 Systeme im Feld
Prefill867 tok/s
Min 22Ø 830Max 3.208
Ueber dem Durchschnitt · 47 Systeme im Feld
Time to First Token2.545 ms
Min 194Ø 16.838Max 210.255
Ueber dem Durchschnitt · 47 Systeme im Feld
Performance profile
Throughput by hardware & engine
Each bubble is a GPU, CPU or engine – position shows prompt processing (X) and output speed (Y), bubble size the number of runs.
Throughput & latency
Performance benchmark
| # | Model / Maker | Metrics | Parallel | GPU / CPU / RAM | Runtime | ||
|---|---|---|---|---|---|---|---|
| 1 | gemma-4-26B-A4B-it-assistant26BGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 74,51 tok/s TG Prefill 867 · TTFT 23.582 ms | 10× | AMD Radeon 8060S GraphicsAMD RYZEN AI MAX+ 395 w/ Radeon 8060S · 31 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 2 | gemma-4-26B-A4B-it-assistant26BGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 67,34 tok/s TG Prefill 795 · TTFT 12.430 ms | 5× | AMD Radeon 8060S GraphicsAMD RYZEN AI MAX+ 395 w/ Radeon 8060S · 31 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 3 | gemma-4-26B-A4B-it-assistant26BGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 32,97 tok/s TG Prefill 777 · TTFT 2.545 ms | 1× | AMD Radeon 8060S GraphicsAMD RYZEN AI MAX+ 395 w/ Radeon 8060S · 31 GB RAM | llama.cppgodclawQ4_K_M | Details → |
Agent & chat rating
Harness benchmark
Noch keine Harnessbenchmarks fuer dieses Modell.
