M
Manufacturer
Mistral AI
Franzoesisches KI-Labor; Mistral, Mixtral, Codestral, Devstral, Ministral, Magistral.
Best token generation
1.060,0 tok/s
Ministral-3-14B-Reasoning-2512
Best prefill
20.061,1 tok/s
Ministral-3-14B-Reasoning-2512
∅ Token generation
140,5 tok/s
Average of 636 runs
∅ Prefill
3.123,1 tok/s
Average of 636 runs
Best TTFT
20 ms
Magistral-Small-2509
Best harness rate
87%
Devstral-Small-2-24B-Instruct-2512
🏆
Ministral-3-14B-Reasoning-2512 · fastest model at 1.060,0 tok/s token generation
Best prefill of this manufacturer: 20.061,1 tok/s (Ministral-3-14B-Reasoning-2512)
Best prefill of this manufacturer: 20.061,1 tok/s (Ministral-3-14B-Reasoning-2512)
Top models by token generation
Best measured generation throughput per model (tok/s), published performance benchmarks.
Models
All benchmarks →10 assigned
Codestral-22B-v0.122B89 runs
129,0tok/s Ø
Min 13,0Max 641,4
Devstral-2-123B-Instruct-2512123B50 runs
22,1tok/s Ø
Min 0,0Max 143,1
Devstral-Small-2-24B-Instruct-251224B94 runs
156,1tok/s Ø
Min 8,9Max 725,6
Devstral-Small-250724B86 runs
150,3tok/s Ø
Min 13,3Max 647,6
Magistral-Small-250924B92 runs
172,9tok/s Ø
Min 4,0Max 737,0
Mamba-Codestral-7B-v0.17B0 runs
No performance data yet
Ministral-3-14B-Reasoning-251214B106 runs
253,1tok/s Ø
Min 22,4Max 1.060,0
Mistral-Medium-3.5-128B128B28 runs
10,2tok/s Ø
Min 0,6Max 141,6
Mistral-Small-3.1-24B-Instruct-250324B69 runs
63,2tok/s Ø
Min 5,0Max 214,9
Mistral-Small-4-119B-2603119B26 runs
93,1tok/s Ø
Min 5,8Max 959,6
