⚡
Performance benchmark
How fast is the model?
Raw throughput on real hardware – how many tokens a model generates per second, how quickly it processes the prompt and how short the time to first token is. Higher is better.
⚡ Generation (tok/s)🔄 Prefill⏱️ TTFT👥 Concurrency
| # | Model / Maker | Metrics | Parallel | GPU / CPU / RAM | Runtime | ||
|---|---|---|---|---|---|---|---|
| 1 | Ling-3.0-flash127.5BinclusionAI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 92,62 tok/s TG Prefill 1.185 · TTFT 24.873 ms | 10× | 2x NVIDIA RTX A6000AMD EPYC 7203P 8-Core Processor | llama.cppgodclawQ4_K_M | Details → | |
| 2 | Ling-3.0-flash127.5BinclusionAI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 73,50 tok/s TG Prefill 900 · TTFT 14.075 ms | 5× | 2x NVIDIA RTX A6000AMD EPYC 7203P 8-Core Processor | llama.cppgodclawQ4_K_M | Details → | |
| 3 | Ling-3.0-flash127.5BinclusionAI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 30,41 tok/s TG Prefill 713 · TTFT 3.416 ms | 1× | 2x NVIDIA RTX A6000AMD EPYC 7203P 8-Core Processor | llama.cppgodclawQ4_K_M | Details → | |
| 4 | Qwen3.8-Flash-Next180BQwen (Alibaba) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 25,97 tok/s TG Prefill 204 · TTFT 110.341 ms | 10× | 2x NVIDIA RTX A6000AMD EPYC 7203P 8-Core Processor | llama.cppgodclawQ4_K_XL | Details → | |
| 5 | Qwen3.8-Flash-Next180BQwen (Alibaba) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 19,36 tok/s TG Prefill 158 · TTFT 64.302 ms | 5× | 2x NVIDIA RTX A6000AMD EPYC 7203P 8-Core Processor | llama.cppgodclawQ4_K_XL | Details → | |
| 6 | Qwen3.8-Flash-Next180BQwen (Alibaba) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 8,95 tok/s TG Prefill 94 · TTFT 24.951 ms | 1× | 2x NVIDIA RTX A6000AMD EPYC 7203P 8-Core Processor | llama.cppgodclawQ4_K_XL | Details → |
