Qwen (Alibaba)
Qwen2.5-7B-Instruct7B
Benchmark profile and published results.
Dense7B
Position in the field
Best values compared
Best metrics of this model against the minimum, average and maximum of all published systems.
Generation160,3 tok/s
Min 0,0Ø 245,9Max 2.491,2
Unter dem Durchschnitt · 4037 Systeme im Feld
Prefill4.979 tok/s
Min 4Ø 4.711Max 55.260
Ueber dem Durchschnitt · 4037 Systeme im Feld
Time to First Token1.142 ms
Min 20Ø 34.263Max 535.235
Ueber dem Durchschnitt · 4035 Systeme im Feld
Performance profile
Throughput by hardware & engine
Each bubble is a GPU, CPU or engine – position shows prompt processing (X) and output speed (Y), bubble size the number of runs.
Throughput & latency
Performance benchmark
| # | Model / Maker | Metrics | Parallel | GPU / CPU / RAM | Runtime | ||
|---|---|---|---|---|---|---|---|
| 1 | Qwen2.5-7B-Instruct7BQwen (Alibaba) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 160,26 tok/s TG Prefill 4.979 · TTFT 6.510 ms | 10× | 2x NVIDIA GeForce RTX 20604x Intel(R) Xeon(R) CPU E5-4657L v2 @ 2.40GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 2 | Qwen2.5-7B-Instruct7BQwen (Alibaba) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 108,79 tok/s TG Prefill 3.816 · TTFT 3.796 ms | 5× | 2x NVIDIA GeForce RTX 20604x Intel(R) Xeon(R) CPU E5-4657L v2 @ 2.40GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 3 | Qwen2.5-7B-Instruct7BQwen (Alibaba) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 56,92 tok/s TG Prefill 2.151 · TTFT 1.142 ms | 1× | 2x NVIDIA GeForce RTX 20604x Intel(R) Xeon(R) CPU E5-4657L v2 @ 2.40GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → |
Agent & chat rating
Harness benchmark
Noch keine Harnessbenchmarks fuer dieses Modell.
