Q
Manufacturer
Qwen (Alibaba)
Qwen-Modellfamilie von Alibaba.
Best token generation
1.716,6 tok/s
Qwen3-Coder-30B-A3B-Instruct
Best prefill
39.028,9 tok/s
Qwen3-30B-A3B-Thinking-2507
∅ Token generation
257,2 tok/s
Average of 1222 runs
∅ Prefill
6.070,2 tok/s
Average of 1222 runs
Best TTFT
122 ms
Qwen3.6-27B
Best harness rate
41%
Qwen3.6-27B
🏆
Qwen3-Coder-30B-A3B-Instruct · fastest model at 1.716,6 tok/s token generation
Best prefill of this manufacturer: 39.028,9 tok/s (Qwen3-30B-A3B-Thinking-2507)
Best prefill of this manufacturer: 39.028,9 tok/s (Qwen3-30B-A3B-Thinking-2507)
Top models by token generation
Best measured generation throughput per model (tok/s), published performance benchmarks.
Models
All benchmarks →19 assigned
Huihui-Qwen3-VL-30B-A3B-Instruct-abliterated30B0 runs
No performance data yet
Qwen-AgentWorld-35B-A3B35B69 runs
341,7tok/s Ø
Min 21,8Max 1.481,2
Qwen2.5-32B-Instruct-AWQ32B54 runs
139,2tok/s Ø
Min 4,4Max 510,3
Qwen2.5-72B-Instruct72B32 runs
43,9tok/s Ø
Min 0,8Max 239,2
Qwen2.5-Coder-32B-Instruct32B93 runs
99,7tok/s Ø
Min 4,4Max 500,6
Qwen3-30B-A3B-Instruct-250730B121 runs
304,4tok/s Ø
Min 1,0Max 1.671,0
Qwen3-30B-A3B-Thinking-250730B106 runs
329,2tok/s Ø
Min 0,6Max 1.683,8
Qwen3-Coder-30B-A3B-Instruct30B112 runs
325,4tok/s Ø
Min 1,0Max 1.716,6
Qwen3-Coder-Next54 runs
292,0tok/s Ø
Min 18,1Max 1.253,3
Qwen3-Omni-30B-A3B-Thinking30B68 runs
383,1tok/s Ø
Min 26,2Max 1.673,6
Qwen3-VL-30B-A3B-Instruct30B93 runs
334,3tok/s Ø
Min 16,3Max 1.655,9
Qwen3.5-122B-A10B122B54 runs
92,7tok/s Ø
Min 3,4Max 564,6
Qwen3.5-122B-A10B-Q8122B0 runs
No performance data yet
Qwen3.5-35B-A3B35B102 runs
327,9tok/s Ø
Min 5,6Max 1.520,4
Qwen3.6-27B27B66 runs
139,6tok/s Ø
Min 8,8Max 546,9
Qwen3.6-35B-A3B35B151 runs
259,1tok/s Ø
Min 4,3Max 1.489,1
Qwen3.8-27B27B51 runs
103,5tok/s Ø
Min 6,9Max 426,5
Qwen3.8-Flash-Next180B0 runs
No performance data yet
Qwen3.8-Flash-Next-Uncensored-NVFP4180B0 runs
No performance data yet
