Created bymario-alka.dePowered bygodcore.denoob2claw.detricoma.de
Meta

Llama-3.2-3B-Instruct3B

Benchmark profile and published results.

Dense3B
Position in the field

Best values compared

Best metrics of this model against the minimum, average and maximum of all published systems.

Generation226,6 tok/s
Min 0,0Ø 245,9Max 2.491,2
Unter dem Durchschnitt · 4037 Systeme im Feld
Prefill9.775 tok/s
Min 4Ø 4.711Max 55.260
Ueber dem Durchschnitt · 4037 Systeme im Feld
Time to First Token645 ms
Min 20Ø 34.263Max 535.235
Ueber dem Durchschnitt · 4035 Systeme im Feld
Performance profile

Throughput by hardware & engine

Each bubble is a GPU, CPU or engine – position shows prompt processing (X) and output speed (Y), bubble size the number of runs.

GPUby graphics card

2492382272152046.5116.7887.0657.342Prefill (tok/s)Generation (tok/s)NVIDIA GeForce RTX 2060 - 226,6 tok/s Generation, 6.927 tok/s Prefill, TTFT 2.004 ms (6 Laufe)NVIDIA GeForce RTX 20...
NVIDIA GeForce RTX 2060 226,6 tok/s

CPUby processor

2492382272152046.5116.7887.0657.342Prefill (tok/s)Generation (tok/s)Intel(R) Xeon(R) CPU E5-4657L v2 @ 2.40GHz - 226,6 tok/s Generation, 6.927 tok/s Prefill, TTFT 2.004 ms (6 Laufe)Intel(R) Xeon(R) CPU ...
Intel(R) Xeon(R) CPU E5-4657L v2 @ 2.40GHz 226,6 tok/s

ENGby engine

2492382272152046.5116.7887.0657.342Prefill (tok/s)Generation (tok/s)llama.cpp - 226,6 tok/s Generation, 6.927 tok/s Prefill, TTFT 2.004 ms (6 Laufe)llama.cpp
llama.cpp 226,6 tok/s
Throughput & latency

Performance benchmark

Metric:
#Model / MakerMetricsParallelGPU / CPU / RAMRuntime
1Llama-3.2-3B-Instruct3BMeta Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
226,59 tok/s TG
Prefill 9.526 · TTFT 3.488 ms
10×2x NVIDIA GeForce RTX 20604x Intel(R) Xeon(R) CPU E5-4657L v2 @ 2.40GHz · 504 GB RAMllama.cppgodclawQ4_K_M
2Llama-3.2-3B-Instruct3BMeta Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
215,97 tok/s TG
Prefill 9.775 · TTFT 3.283 ms
10×2x NVIDIA GeForce RTX 20604x Intel(R) Xeon(R) CPU E5-4657L v2 @ 2.40GHz · 504 GB RAMllama.cppgodclawQ8_0
3Llama-3.2-3B-Instruct3BMeta Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
173,08 tok/s TG
Prefill 7.158 · TTFT 2.024 ms
5×2x NVIDIA GeForce RTX 20604x Intel(R) Xeon(R) CPU E5-4657L v2 @ 2.40GHz · 504 GB RAMllama.cppgodclawQ4_K_M
4Llama-3.2-3B-Instruct3BMeta Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
158,90 tok/s TG
Prefill 7.616 · TTFT 1.908 ms
5×2x NVIDIA GeForce RTX 20604x Intel(R) Xeon(R) CPU E5-4657L v2 @ 2.40GHz · 504 GB RAMllama.cppgodclawQ8_0
5Llama-3.2-3B-Instruct3BMeta Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
100,62 tok/s TG
Prefill 3.646 · TTFT 677 ms
1×2x NVIDIA GeForce RTX 20604x Intel(R) Xeon(R) CPU E5-4657L v2 @ 2.40GHz · 504 GB RAMllama.cppgodclawQ4_K_M
6Llama-3.2-3B-Instruct3BMeta Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
70,69 tok/s TG
Prefill 3.840 · TTFT 645 ms
1×2x NVIDIA GeForce RTX 20604x Intel(R) Xeon(R) CPU E5-4657L v2 @ 2.40GHz · 504 GB RAMllama.cppgodclawQ8_0
Agent & chat rating

Harness benchmark

Noch keine Harnessbenchmarks fuer dieses Modell.