Qwen (Alibaba)
Qwen2.5-7B-Instruct7B
Benchmark profile and published results.
Dense7B
Position in the field
Best values compared
Best metrics of this model against the minimum, average and maximum of all published systems.
Generation160,3 tok/s
Min 1,1Ø 71,9Max 259,9
Ueber dem Durchschnitt · 78 Systeme im Feld
Prefill4.979 tok/s
Min 8Ø 2.203Max 9.775
Ueber dem Durchschnitt · 78 Systeme im Feld
Time to First Token1.142 ms
Min 569Ø 46.548Max 382.793
Ueber dem Durchschnitt · 78 Systeme im Feld
Performance profile
Throughput by hardware & engine
Each bubble is a GPU, CPU or engine – position shows prompt processing (X) and output speed (Y), bubble size the number of runs.
Throughput & latency
Performance benchmark
| # | Model / Maker | Metrics | Parallel | GPU / CPU / RAM | Runtime | ||
|---|---|---|---|---|---|---|---|
| 1 | Qwen2.5-7B-Instruct7BQwen (Alibaba) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 160,26 tok/s TG Prefill 4.979 · TTFT 6.510 ms | 10× | 2x NVIDIA GeForce RTX 20604x Intel(R) Xeon(R) CPU E5-4657L v2 @ 2.40GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 2 | Qwen2.5-7B-Instruct7BQwen (Alibaba) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 108,79 tok/s TG Prefill 3.816 · TTFT 3.796 ms | 5× | 2x NVIDIA GeForce RTX 20604x Intel(R) Xeon(R) CPU E5-4657L v2 @ 2.40GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 3 | Qwen2.5-7B-Instruct7BQwen (Alibaba) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 56,92 tok/s TG Prefill 2.151 · TTFT 1.142 ms | 1× | 2x NVIDIA GeForce RTX 20604x Intel(R) Xeon(R) CPU E5-4657L v2 @ 2.40GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → |
Agent & chat rating
Harness benchmark
Noch keine Harnessbenchmarks fuer dieses Modell.
