🤖
Harness benchmark
How well does the model solve real tasks?
Quality on real agent and chat tasks – the model must handle real tasks with tools and multiple steps. Scored by success rate and task points achieved.
🎯 Success rate🏆 Task points🔧 Tool use🧠 real tasks
| # | Model / Maker | Metrics | GPU / CPU / RAM | Runtime | ||
|---|---|---|---|---|---|---|
| 1 | MiniMax-M2.5230BMiniMax Harness benchmarkTool Usage V 1.0 | 1.245 Pkt 70,7% · 15/15 Aufg. · 301,0 s | NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAM | openclaw_cli | Details → | |
| 2 | MiniMax-M2.7230BMiniMax Harness benchmarkTool-Parcours Dossier 40 V1.0 | 3.641 Pkt 61,1% · 40/40 Aufg. · 8.209,0 s | NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAM | vLLMgodclawAWQ | Details → | |
| 3 | MiniMax-M2.5230BMiniMax Harness benchmarkTool Usage Standard 1.0 | 1.762 Pkt 59,1% · 16/16 Aufg. · 526,0 s | NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAM | openclaw_cli | Details → | |
| 4 | MiniMax-M2.5230BMiniMax Harness benchmarkTextgenerierung V 1.0 | 790 Pkt 52,0% · 15/15 Aufg. · 261,0 s | NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAM | openclaw_cli | Details → | |
| 5 | MiniMax-M2.7230BMiniMax Harness benchmarkTool Usage Standard 1.0 | 858 Pkt 28,8% · 16/16 Aufg. · 1.222,0 s | NVIDIA GB10 (DGX Spark)NVIDIA Grace · 120 GB RAM | godclaw | Details → |
