Performance benchmark
Direct call to an OpenAI-compatible chat completions API. Time to first token, prefill tokens/s, generation tokens/s and total runtime are recorded. Model, runtime, quantization and system data are part of the result.
Harness benchmark
The model runs through a complete agent or chat harness. Points achieved, maximum possible points and success rate are rated. This makes practical task completion visible alongside speed.
Comparability
Results are only directly reliable under identical benchmarks and comparable settings. Verified runs are marked in the leaderboard.
