🤖
Harness benchmark
How well does the model solve real tasks?
Quality on real agent and chat tasks – the model must handle real tasks with tools and multiple steps. Scored by success rate and task points achieved.
🎯 Success rate🏆 Task points🔧 Tool use🧠 real tasks
| # | Model / Maker | Metrics | GPU / CPU / RAM | Runtime | ||
|---|---|---|---|---|---|---|
| 1 | Devstral-Small-2-24B-Instruct-251224BMistral AI Harness benchmarkTool Usage Standard 1.0 | 2.597 Pkt 87,1% · 16/16 Aufg. · 165,0 s | AMD Radeon 8060S GraphicsAMD RYZEN AI MAX+ 395 w/ Radeon 8060S · 31 GB RAM | llama.cppgodclawQ8_0 | Details → | |
| 2 | Gemma-4-26B-A4B-it26BGoogle Harness benchmarkTool Usage Standard 1.0 | 2.440 Pkt 81,9% · 16/16 Aufg. · 218,0 s | AMD Radeon 8060S GraphicsAMD RYZEN AI MAX+ 395 w/ Radeon 8060S · 123 GB RAM | openclaw_cli | Details → | |
| 3 | Gemma-4-26B-A4B-it26BGoogle Harness benchmarkTool Usage Standard 1.0 | 2.383 Pkt 80,0% · 16/16 Aufg. · 245,0 s | AMD Radeon 8060S GraphicsAMD RYZEN AI MAX+ 395 w/ Radeon 8060S · 123 GB RAM | openclaw_cli | Details → | |
| 4 | Gemma-4-26B-A4B-it26BGoogle Harness benchmarkTool Usage Standard 1.0 | 2.268 Pkt 76,1% · 16/16 Aufg. · 347,0 s | AMD Radeon 8060S GraphicsAMD RYZEN AI MAX+ 395 w/ Radeon 8060S · 31 GB RAM | llama.cppgodclawQ8_0 | Details → | |
| 5 | Devstral-Small-2-24B-Instruct-251224BMistral AI Harness benchmarkGodBrain MCP Deep 40 - Wissensbasis-Lebenszyklus 1.0 | 5.738 Pkt 75,7% · 40/40 Aufg. · 1.879,0 s | AMD Radeon 8060S GraphicsAMD RYZEN AI MAX+ 395 w/ Radeon 8060S · 31 GB RAM | llama.cppgodclawQ8_0 | Details → | |
| 6 | Gemma-4-26B-A4B-it26BGoogle Harness benchmarkTool Usage Standard 1.0 | 2.197 Pkt 73,7% · 16/16 Aufg. · 375,0 s | AMD Radeon 8060S GraphicsAMD RYZEN AI MAX+ 395 w/ Radeon 8060S · 31 GB RAM | openclaw_cli | Details → | |
| 7 | Devstral-Small-2-24B-Instruct-251224BMistral AI Harness benchmarkTool Usage Extrem 40 - Tech-Radar Mission 1.0 | 5.629 Pkt 68,6% · 40/40 Aufg. · 2.018,0 s | AMD Radeon 8060S GraphicsAMD RYZEN AI MAX+ 395 w/ Radeon 8060S · 31 GB RAM | llama.cppgodclawQ8_0 | Details → | |
| 8 | Gemma-4-26B-A4B-it26BGoogle Harness benchmarkGodBrain MCP Deep 40 - Wissensbasis-Lebenszyklus 1.0 | 2.271 Pkt 30,0% · 18/40 Aufg. · 982,0 s | AMD Radeon 8060S GraphicsAMD RYZEN AI MAX+ 395 w/ Radeon 8060S · 31 GB RAM | llama.cppgodclawQ8_0 | Details → | |
| 9 | Gemma-4-26B-A4B-it26BGoogle Harness benchmarkTool-Parcours Dossier 40 V1.0 | 1.489 Pkt 25,0% · 22/40 Aufg. · 8.783,0 s | AMD Radeon 8060S GraphicsAMD RYZEN AI MAX+ 395 w/ Radeon 8060S · 31 GB RAM | llama.cppgodclawQ8_0 | Details → | |
| 10 | Devstral-Small-2-24B-Instruct-251224BMistral AI Harness benchmarkTool-Parcours Dossier 40 V1.0 | 1.436 Pkt 24,1% · 21/40 Aufg. · 8.887,0 s | AMD Radeon 8060S GraphicsAMD RYZEN AI MAX+ 395 w/ Radeon 8060S · 31 GB RAM | llama.cppgodclawQ8_0 | Details → | |
| 11 | Qwen2.5-32B-Instruct-AWQ32BQwen (Alibaba) Harness benchmarkTool Usage V 1.0 | 235 Pkt 13,4% · 15/15 Aufg. · 1.783,0 s | AMD Radeon 8060S GraphicsAMD RYZEN AI MAX+ 395 w/ Radeon 8060S · 31 GB RAM | openclaw_cli | Details → | |
| 12 | gemma-4-31B-it31BGoogle Harness benchmarkTool Usage Standard 1.0 | 298 Pkt 10,0% · 16/16 Aufg. · 9.625,0 s | AMD Radeon 8060S GraphicsAMD RYZEN AI MAX+ 395 w/ Radeon 8060S · 123 GB RAM | openclaw_cli | Details → | |
| 13 | Qwen2.5-32B-Instruct-AWQ32BQwen (Alibaba) Harness benchmarkTool Usage Standard 1.0 | 298 Pkt 10,0% · 16/16 Aufg. · 4.687,0 s | AMD Radeon 8060S GraphicsAMD RYZEN AI MAX+ 395 w/ Radeon 8060S · 123 GB RAM | openclaw_cli | Details → |
