Qwen (Alibaba)
Qwen2.5-3B-Instruct3B
Benchmark profile and published results.
Dense3B
Position in the field
Best values compared
Best metrics of this model against the minimum, average and maximum of all published systems.
Generation259,9 tok/s
Min 1,1Ø 71,9Max 259,9
Spitzenwert im Feld · 78 Systeme im Feld
Prefill9.739 tok/s
Min 8Ø 2.203Max 9.775
Spitzenwert im Feld · 78 Systeme im Feld
Time to First Token596 ms
Min 569Ø 46.548Max 382.793
Ueber dem Durchschnitt · 78 Systeme im Feld
Performance profile
Throughput by hardware & engine
Each bubble is a GPU, CPU or engine – position shows prompt processing (X) and output speed (Y), bubble size the number of runs.
Throughput & latency
Performance benchmark
| # | Model / Maker | Metrics | Parallel | GPU / CPU / RAM | Runtime | ||
|---|---|---|---|---|---|---|---|
| 1 | Qwen2.5-3B-Instruct3BQwen (Alibaba) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 259,86 tok/s TG Prefill 9.253 · TTFT 3.361 ms | 10× | 2x NVIDIA GeForce RTX 20604x Intel(R) Xeon(R) CPU E5-4657L v2 @ 2.40GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 2 | Qwen2.5-3B-Instruct3BQwen (Alibaba) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 225,82 tok/s TG Prefill 9.739 · TTFT 3.277 ms | 10× | 2x NVIDIA GeForce RTX 20604x Intel(R) Xeon(R) CPU E5-4657L v2 @ 2.40GHz · 504 GB RAM | llama.cppgodclawQ8_0 | Details → | |
| 3 | Qwen2.5-3B-Instruct3BQwen (Alibaba) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 187,50 tok/s TG Prefill 6.698 · TTFT 2.090 ms | 5× | 2x NVIDIA GeForce RTX 20604x Intel(R) Xeon(R) CPU E5-4657L v2 @ 2.40GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 4 | Qwen2.5-3B-Instruct3BQwen (Alibaba) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 168,19 tok/s TG Prefill 7.879 · TTFT 1.850 ms | 5× | 2x NVIDIA GeForce RTX 20604x Intel(R) Xeon(R) CPU E5-4657L v2 @ 2.40GHz · 504 GB RAM | llama.cppgodclawQ8_0 | Details → | |
| 5 | Qwen2.5-3B-Instruct3BQwen (Alibaba) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 106,40 tok/s TG Prefill 3.776 · TTFT 651 ms | 1× | 2x NVIDIA GeForce RTX 20604x Intel(R) Xeon(R) CPU E5-4657L v2 @ 2.40GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 6 | Qwen2.5-3B-Instruct3BQwen (Alibaba) Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 73,93 tok/s TG Prefill 4.119 · TTFT 596 ms | 1× | 2x NVIDIA GeForce RTX 20604x Intel(R) Xeon(R) CPU E5-4657L v2 @ 2.40GHz · 504 GB RAM | llama.cppgodclawQ8_0 | Details → |
Agent & chat rating
Harness benchmark
Noch keine Harnessbenchmarks fuer dieses Modell.
