Meta
Llama-3.2-3B-Instruct3B
Benchmark profile and published results.
Dense3B
Position in the field
Best values compared
Best metrics of this model against the minimum, average and maximum of all published systems.
Generation226,6 tok/s
Min 1,1Ø 71,9Max 259,9
Ueber dem Durchschnitt · 78 Systeme im Feld
Prefill9.775 tok/s
Min 8Ø 2.203Max 9.775
Spitzenwert im Feld · 78 Systeme im Feld
Time to First Token645 ms
Min 569Ø 46.548Max 382.793
Ueber dem Durchschnitt · 78 Systeme im Feld
Performance profile
Throughput by hardware & engine
Each bubble is a GPU, CPU or engine – position shows prompt processing (X) and output speed (Y), bubble size the number of runs.
Throughput & latency
Performance benchmark
| # | Model / Maker | Metrics | Parallel | GPU / CPU / RAM | Runtime | ||
|---|---|---|---|---|---|---|---|
| 1 | Llama-3.2-3B-Instruct3BMeta Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 226,59 tok/s TG Prefill 9.526 · TTFT 3.488 ms | 10× | 2x NVIDIA GeForce RTX 20604x Intel(R) Xeon(R) CPU E5-4657L v2 @ 2.40GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 2 | Llama-3.2-3B-Instruct3BMeta Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 215,97 tok/s TG Prefill 9.775 · TTFT 3.283 ms | 10× | 2x NVIDIA GeForce RTX 20604x Intel(R) Xeon(R) CPU E5-4657L v2 @ 2.40GHz · 504 GB RAM | llama.cppgodclawQ8_0 | Details → | |
| 3 | Llama-3.2-3B-Instruct3BMeta Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 173,08 tok/s TG Prefill 7.158 · TTFT 2.024 ms | 5× | 2x NVIDIA GeForce RTX 20604x Intel(R) Xeon(R) CPU E5-4657L v2 @ 2.40GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 4 | Llama-3.2-3B-Instruct3BMeta Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 158,90 tok/s TG Prefill 7.616 · TTFT 1.908 ms | 5× | 2x NVIDIA GeForce RTX 20604x Intel(R) Xeon(R) CPU E5-4657L v2 @ 2.40GHz · 504 GB RAM | llama.cppgodclawQ8_0 | Details → | |
| 5 | Llama-3.2-3B-Instruct3BMeta Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 100,62 tok/s TG Prefill 3.646 · TTFT 677 ms | 1× | 2x NVIDIA GeForce RTX 20604x Intel(R) Xeon(R) CPU E5-4657L v2 @ 2.40GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 6 | Llama-3.2-3B-Instruct3BMeta Performance benchmarkTimebench 3 - Kombi (Prefill + Generation) | 70,69 tok/s TG Prefill 3.840 · TTFT 645 ms | 1× | 2x NVIDIA GeForce RTX 20604x Intel(R) Xeon(R) CPU E5-4657L v2 @ 2.40GHz · 504 GB RAM | llama.cppgodclawQ8_0 | Details → |
Agent & chat rating
Harness benchmark
Noch keine Harnessbenchmarks fuer dieses Modell.
