Created bymario-alka.dePowered bygodcore.denoob2claw.detricoma.de
Open data · reproducible tests

Which model is fast & smart?

We benchmark AI language models on real hardware: speed in tokens per second and response time – plus quality on real agent tasks. Transparent, reproducible, directly comparable.

2853 benchmarks74 models2 test types
Live benchmark
Who generates text the fastest? Real measurements, real hardware.
1
gemma-4-E2B-itGeForce RTX 5090
2.491,2tok/s
2
Nemotron-3-Nano-4BRTX PRO 6000 Blackwe
2.182,7tok/s
3
gpt-oss-20bRTX PRO 6000 Blackwe
1.995,4tok/s
4
Ornith-1.0-35BRTX PRO 6000 Blackwe
1.799,4tok/s
5
Nemotron-3-Nano-Omni-30B-A3B-ReasoningRTX PRO 6000 Blackwe
1.791,5tok/s
Peak2.491tok/s
View ranking →
01

Real performance

TTFT, prefill and generation tokens/s instead of theoretical peak numbers.

02

Harness quality

Points and success rates from complete agent and chat runs.

03

Full transparency

Model, quantization, runtime and hardware stay visible on every run.

Community favorites

Top models by benchmarks

Speed (tokens/s) and quality (harness) of the most-tested models – directly comparable.

All models →
Current measurements

Leaderboard Snapshot

Alle Ergebnisse →
Metric:
Size
#Model / MakerMetricsParallelGPU / CPU / RAMRuntime
1gemma-4-E2B-it5BQ4_K_MGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
2.491,20 tok/s TG
Prefill 15.814 · TTFT 4.382 ms
10×NVIDIA GeForce RTX 5090AMD Ryzen 7 5800X3D 8-Core Processor · 126 GB RAMllama.cppcodex_cli
2gemma-4-E2B-it5BQ4_K_MGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
2.414,37 tok/s TG
Prefill 20.277 · TTFT 4.107 ms
10×NVIDIA RTX PRO 6000 Blackwell Workstation EditionAMD Ryzen 9 9950X 16-Core Processor · 92 GB RAMllama.cppgodclaw
3Nemotron-3-Nano-4B4BQ4_K_MNVIDIA Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
2.182,67 tok/s TG
Prefill 10.599 · TTFT 5.598 ms
10×NVIDIA RTX PRO 6000 Blackwell Workstation EditionAMD Ryzen 9 9950X 16-Core Processor · 92 GB RAMllama.cppgodclaw
4gemma-4-E2B-it5BQ4_K_MGoogle Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
2.143,55 tok/s TG
Prefill 13.632 · TTFT 4.819 ms
10×3x NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen Threadripper PRO 9965WX 24-Cores · 125 GB RAMllama.cppgodclaw
5gpt-oss-20b20BQ4_K_MOpenAI Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
1.995,37 tok/s TG
Prefill 16.158 · TTFT 5.309 ms
10×NVIDIA RTX PRO 6000 Blackwell Workstation EditionAMD Ryzen 9 9950X 16-Core Processor · 92 GB RAMllama.cppgodclaw
Frisch gemessen

Neuste Benchmarks

Alle Ergebnisse →
Performance benchmark22.08.2026
Meta-Llama-3.1-8B-Instruct
Meta
45,73 tok/s
Mario Alka
Performance benchmark22.08.2026
Nemotron-3-Nano-4B
NVIDIA
106,92 tok/s
Mario Alka
Performance benchmark22.08.2026
Nemotron-3-Nano-4B
NVIDIA
100,31 tok/s
Mario Alka
Performance benchmark22.08.2026
Nemotron-3-Nano-4B
NVIDIA
72,37 tok/s
Mario Alka
Performance benchmark22.08.2026
Meta-Llama-3.1-8B-Instruct
Meta
48,95 tok/s
Mario Alka
Performance benchmark22.08.2026
Nemotron-3-Nano-4B
NVIDIA
125,68 tok/s
Mario Alka
Performance benchmark22.08.2026
Nemotron-3-Nano-4B
NVIDIA
116,24 tok/s
Mario Alka
Performance benchmark22.08.2026
Nemotron-3-Nano-4B
NVIDIA
80,94 tok/s
Mario Alka
Performance benchmark21.08.2026
Devstral-Small-2-24B-Instruct-2512
Mistral AI
23,29 tok/s
Mario Alka
Performance benchmark21.08.2026
Devstral-Small-2-24B-Instruct-2512
Mistral AI
13,47 tok/s
Mario Alka
Performance benchmark21.08.2026
Devstral-Small-2507
Mistral AI
23,15 tok/s
Mario Alka
Performance benchmark21.08.2026
Devstral-Small-2507
Mistral AI
13,34 tok/s
Mario Alka
Performance benchmark21.08.2026
Magistral-Small-2509
Mistral AI
24,74 tok/s
Mario Alka
Performance benchmark21.08.2026
Magistral-Small-2509
Mistral AI
13,55 tok/s
Mario Alka
Performance benchmark21.08.2026
Mistral-Small-3.1-24B-Instruct-2503
Mistral AI
23,84 tok/s
Mario Alka
Performance benchmark21.08.2026
Mistral-Small-3.1-24B-Instruct-2503
Mistral AI
13,57 tok/s
Mario Alka
Performance benchmark21.08.2026
Codestral-22B-v0.1
Mistral AI
22,24 tok/s
Mario Alka
Performance benchmark21.08.2026
Codestral-22B-v0.1
Mistral AI
13,51 tok/s
Mario Alka
Performance benchmark21.08.2026
gpt-oss-20b
OpenAI
105,30 tok/s
Mario Alka
Performance benchmark21.08.2026
gpt-oss-20b
OpenAI
111,12 tok/s
Mario Alka
Performance benchmark21.08.2026
gpt-oss-20b
OpenAI
63,18 tok/s
Mario Alka
Performance benchmark21.08.2026
Ministral-3-14B-Reasoning-2512
Mistral AI
23,11 tok/s
Mario Alka
Performance benchmark21.08.2026
Ministral-3-14B-Reasoning-2512
Mistral AI
41,46 tok/s
Mario Alka
Performance benchmark21.08.2026
Ministral-3-14B-Reasoning-2512
Mistral AI
22,41 tok/s
Mario Alka
Performance benchmark21.08.2026
gemma-4-12B-it
Google
25,04 tok/s
Mario Alka