Created bymario-alka.dePowered bygodcore.denoob2claw.detricoma.de
Reproducibility

Benchmark methodology

Two test types answer two different questions.

Performance benchmark

Direct call to an OpenAI-compatible chat completions API. Time to first token, prefill tokens/s, generation tokens/s and total runtime are recorded. Model, runtime, quantization and system data are part of the result.

Harness benchmark

The model runs through a complete agent or chat harness. Points achieved, maximum possible points and success rate are rated. This makes practical task completion visible alongside speed.

Comparability

Results are only directly reliable under identical benchmarks and comparable settings. Verified runs are marked in the leaderboard.