Created bymario-alka.dePowered bygodcore.denoob2claw.detricoma.de
Z.ai (Zhipu)

glm-4.7-flash

Benchmark profile and published results.

Position in the field

Best values compared

Best metrics of this model against the minimum, average and maximum of all published systems.

Generation48,9 tok/s
Min 11,4Ø 148,9Max 903,3
Unter dem Durchschnitt · 82 Systeme im Feld
Prefill4.055 tok/s
Min 503Ø 3.550Max 7.470
Ueber dem Durchschnitt · 82 Systeme im Feld
Time to First Token536 ms
Min 105Ø 8.759Max 66.475
Ueber dem Durchschnitt · 82 Systeme im Feld
Performance profile

Throughput by hardware & engine

Each bubble is a GPU, CPU or engine – position shows prompt processing (X) and output speed (Y), bubble size the number of runs.

GPUby graphics card

53,851,448,946,544,03.8123.9744.1364.298Prefill (tok/s)Generation (tok/s)NVIDIA GB10 - 48,9 tok/s Generation, 4.055 tok/s Prefill, TTFT 536 ms (1 Lauf)NVIDIA GB10
NVIDIA GB10 48,9 tok/s

CPUby processor

53,851,448,946,544,03.8123.9744.1364.298Prefill (tok/s)Generation (tok/s)NVIDIA Grace - 48,9 tok/s Generation, 4.055 tok/s Prefill, TTFT 536 ms (1 Lauf)NVIDIA Grace
NVIDIA Grace 48,9 tok/s

ENGby engine

53,851,448,946,544,03.8123.9744.1364.298Prefill (tok/s)Generation (tok/s)vLLM - 48,9 tok/s Generation, 4.055 tok/s Prefill, TTFT 536 ms (1 Lauf)vLLM
vLLM 48,9 tok/s
Throughput & latency

Performance benchmark

Metric:
#Model / MakerMetricsParallelGPU / CPU / RAMRuntime
1glm-4.7-flashAWQZ.ai (Zhipu) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
48,92 tok/s TG
Prefill 4.055 · TTFT 536 ms
1×NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAMvLLMgodclaw
Agent & chat rating

Harness benchmark

Noch keine Harnessbenchmarks fuer dieses Modell.