🤖
Harness benchmark
How well does the model solve real tasks?
Quality on real agent and chat tasks – the model must handle real tasks with tools and multiple steps. Scored by success rate and task points achieved.
🎯 Success rate🏆 Task points🔧 Tool use🧠 real tasks
| # | Model / Maker | Metrics | GPU / CPU / RAM | Runtime | ||
|---|---|---|---|---|---|---|
| 1 | Gemma-4-26B-A4B-it26BGoogle Harness benchmarkTool Usage Standard 1.0 | 2.701 Pkt 90,6% · 16/16 Aufg. · 118,0 s | NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen 9 9950X 16-Core Processor · 92 GB RAM | openclaw_cli | Details → | |
| 2 | Qwen3.6-27B27BQwen (Alibaba) Harness benchmarkTool Usage Standard 1.0 | 1.220 Pkt 40,9% · 16/16 Aufg. · 1.036,0 s | NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen 9 9950X 16-Core Processor · 92 GB RAM | vLLMgodclaw | Details → | |
| 3 | Qwen3.6-27B27BQwen (Alibaba) Harness benchmarkTool-Parcours Dossier 40 V1.0 | 2.267 Pkt 38,0% · 38/40 Aufg. · 10.331,0 s | NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen 9 9950X 16-Core Processor · 92 GB RAM | vLLMgodclaw | Details → |
