Created bymario-alka.dePowered bygodcore.denoob2claw.detricoma.de
Open data · reproducible tests

Which model is fast & smart?

We benchmark AI language models on real hardware: speed in tokens per second and response time – plus quality on real agent tasks. Transparent, reproducible, directly comparable.

3863 benchmarks69 models2 test types
Live benchmark
Who generates text the fastest? Real measurements, real hardware.
1
gemma-4-E2B-itGeForce RTX 5090
2.491,2tok/s
2
NVIDIA-Nemotron-3-Nano-4BRTX PRO 6000 Blackwe
2.182,7tok/s
3
gpt-oss-20bRTX PRO 6000 Blackwe
1.995,4tok/s
4
Ornith-1.0-35BRTX PRO 6000 Blackwe
1.799,4tok/s
5
Nemotron-3-Nano-Omni-30B-A3B-ReasoningRTX PRO 6000 Blackwe
1.791,5tok/s
Peak2.491tok/s
View ranking →
01

Real performance

TTFT, prefill and generation tokens/s instead of theoretical peak numbers.

02

Harness quality

Points and success rates from complete agent and chat runs.

03

Full transparency

Model, quantization, runtime and hardware stay visible on every run.

Community favorites

Top models by benchmarks

Speed (tokens/s) and quality (harness) of the most-tested models – directly comparable.

All models →
Current measurements

Leaderboard Snapshot

Alle Ergebnisse →
Metric:
Size
#Model / MakerMetricsParallelGPU / CPU / RAMRuntime
1gemma-4-E2B-it5BGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
2.491,20 tok/s TG
Prefill 15.814 · TTFT 4.382 ms
10×NVIDIA GeForce RTX 5090AMD Ryzen 7 5800X3D 8-Core Processor · 126 GB RAMllama.cppcodex_cliQ4_K_M
2gemma-4-E2B-it5BGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
2.414,37 tok/s TG
Prefill 20.277 · TTFT 4.107 ms
10×NVIDIA RTX PRO 6000 Blackwell Workstation EditionAMD Ryzen 9 9950X 16-Core Processor · 92 GB RAMllama.cppgodclawQ4_K_M
3NVIDIA-Nemotron-3-Nano-4B4BNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
2.182,67 tok/s TG
Prefill 10.599 · TTFT 5.598 ms
10×NVIDIA RTX PRO 6000 Blackwell Workstation EditionAMD Ryzen 9 9950X 16-Core Processor · 92 GB RAMllama.cppgodclawQ4_K_M
4gemma-4-E2B-it5BGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
2.143,55 tok/s TG
Prefill 13.632 · TTFT 4.819 ms
10×3x NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen Threadripper PRO 9965WX 24-Cores · 125 GB RAMllama.cppgodclawQ4_K_M
5gpt-oss-20b20BOpenAI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
1.995,37 tok/s TG
Prefill 16.158 · TTFT 5.309 ms
10×NVIDIA RTX PRO 6000 Blackwell Workstation EditionAMD Ryzen 9 9950X 16-Core Processor · 92 GB RAMllama.cppgodclawQ4_K_M
Frisch gemessen

Neuste Benchmarks

Alle Ergebnisse →
Performance benchmark11.08.2026
Qwen3.5-122B-A10B
Qwen (Alibaba)
24,12 tok/s
Mario Alka
Performance benchmark11.08.2026
Qwen3.5-122B-A10B
Qwen (Alibaba)
17,52 tok/s
Mario Alka
Performance benchmark11.08.2026
Qwen3.5-122B-A10B
Qwen (Alibaba)
7,93 tok/s
Mario Alka
Performance benchmark11.08.2026
Laguna-M.1
Poolside
2,95 tok/s
Mario Alka
Performance benchmark10.08.2026
Laguna-M.1
Poolside
1,57 tok/s
Mario Alka
Performance benchmark10.08.2026
Laguna-M.1
Poolside
2,91 tok/s
Mario Alka
Performance benchmark10.08.2026
Qwen3-30B-A3B-Instruct-2507
Qwen (Alibaba)
143,56 tok/s
Mario Alka
Performance benchmark10.08.2026
Qwen3-30B-A3B-Instruct-2507
Qwen (Alibaba)
140,06 tok/s
Mario Alka
Performance benchmark10.08.2026
Qwen3-30B-A3B-Instruct-2507
Qwen (Alibaba)
78,10 tok/s
Mario Alka
Performance benchmark10.08.2026
Qwen3.5-122B-A10B
Qwen (Alibaba)
29,08 tok/s
Mario Alka
Performance benchmark10.08.2026
Qwen3.5-122B-A10B
Qwen (Alibaba)
23,94 tok/s
Mario Alka
Performance benchmark10.08.2026
Qwen3.5-122B-A10B
Qwen (Alibaba)
10,46 tok/s
Mario Alka
Performance benchmark10.08.2026
Tencent-Hy3-295B-A21B
Tencent
1,57 tok/s
Mario Alka
Performance benchmark10.08.2026
Laguna-M.1
Poolside
2,84 tok/s
Mario Alka
Performance benchmark10.08.2026
Llama-3_1-Nemotron-Ultra-253B-v1
NVIDIA
0,00 tok/s
Mario Alka
Performance benchmark10.08.2026
Tencent-Hy3-295B-A21B
Tencent
0,98 tok/s
Mario Alka
Performance benchmark10.08.2026
Laguna-M.1
Poolside
2,90 tok/s
Mario Alka
Performance benchmark10.08.2026
Llama-3_1-Nemotron-Ultra-253B-v1
NVIDIA
0,00 tok/s
Mario Alka
Performance benchmark09.08.2026
Llama-3_1-Nemotron-Ultra-253B-v1
NVIDIA
0,09 tok/s
Mario Alka
Performance benchmark09.08.2026
Llama-3_1-Nemotron-Ultra-253B-v1
NVIDIA
0,00 tok/s
Mario Alka
Performance benchmark09.08.2026
Tencent-Hy3-295B-A21B
Tencent
0,00 tok/s
Mario Alka
Performance benchmark09.08.2026
Llama-3_1-Nemotron-Ultra-253B-v1
NVIDIA
0,00 tok/s
Mario Alka
Performance benchmark09.08.2026
Laguna-M.1
Poolside
2,85 tok/s
Mario Alka
Performance benchmark09.08.2026
Tencent-Hy3-295B-A21B
Tencent
0,00 tok/s
Mario Alka
Performance benchmark09.08.2026
Laguna-S-2.1-INT4
Poolside
774,90 tok/s
Mario Alka