Created bymario-alka.dePowered bygodcore.denoob2claw.detricoma.de
MiniMax

MiniMax-M2.7-AWQ-4bit

Performancebenchmark · gemessen am 23.07.2026 05:33

Benchmark-IDrun-20260723-055735-60de93
Timebench 3 - Kombi (Prefill + Generation)MoE230BRuntime: vLLMQuantisierung: AWQ
Zur Einordnung: Diese Plattform nutzt Unified Memory – die "VRAM" ist gemeinsamer System-RAM (APU/Superchip); das Modell teilt sich den Speicher mit dem System.
Generation31,52tok/s
Prefill2.328,05tok/s
Time to First Token991,00ms
Gesamtdauer66,97s
Concurrency1parallel
Einordnung im Feld
158von 246 Systemen

Performancebenchmark · Leitmetrik: Generation-Speed (tok/s)

Dieser Lauf ist besser als 36 % aller vergleichbaren Systeme.
Generation 31,5 tok/s
-68 % vs Ø 97,1
Prefill 2.328,1 tok/s
-33 % vs Ø 3.494,7
Time to First Token 991 ms
-94 % vs Ø 17.581
Verteilung im Feld0 – 1.295 tok/s
Ø 97 Median Dieser Lauf

Wie schlägt sich dieser Benchmark?

Hardware

GPU: NVIDIA GB10 · 128 GB VRAM
CPU: NVIDIA Grace
RAM: 120 GB

Setup

Runtime: vLLM
Quantisierung: AWQ
Modell: MiniMax-M2.7-AWQ-4bit

Konfiguration

benchmark-konfiguration — run-20260723-055735-60de93
# LLM-Benchmark Konfiguration # Modell : MiniMax-M2.7-AWQ-4bit # Engine : vLLM # Run-ID : run-20260723-055735-60de93 # GPU : NVIDIA GB10 # CPU : NVIDIA Grace # RAM : 120 GB bench@llm-benchmark:~$ /usr/bin/python /usr/local/bin/vllm serve cyankiwi/MiniMax-M2.7-AWQ-4bit \ --host 0.0.0.0 \ --port 8000 \ --tensor-parallel-size 2 \ --distributed-executor-backend ray \ --gpu-memory-utilization 0.8 \ --max-model-len 196608 \ --max-num-seqs 4 \ --load-format fastsafetensors \ --tool-call-parser minimax_m2 \ --reasoning-parser minimax_m2 \ --enable-auto-tool-choice \ --kv-cache-dtype fp8_e4m3 \ --enable-prefix-caching \ --trust-remote-code \ --enforce-eager \ --served-model-name cyankiwi/MiniMax-M2.7-AWQ-4bit
Enginevllm
Modellaliascyankiwi/MiniMax-M2.7-AWQ-4bit
Kontextlaenge196608
Tensor-Parallel2
Executor-Backendray
GPU-Speicher0.8
Kontext196608
Max-Sequenzen4
Load-Formatfastsafetensors
Tool-Parserminimax_m2
Reasoning-Parserminimax_m2
Auto-Tool-Choiceaktiv
KV-Cache-Dtypefp8_e4m3
Prefix-Cachingaktiv
Trust-Remote-Codeaktiv
Enforce-Eageraktiv
Aliascyankiwi/MiniMax-M2.7-AWQ-4bit

Alle Benchmarks dieses Modells Zum Leaderboard