What hardware do you need to run MiniCPM5 2B?
MiniCPM5 2B at Q4_K_M is 1.6 GB of weights, with a context ceiling of 128k tokens. Not yet rated: released after our last ratings pass. The best score per gigabyte on this list: 13 on the index at 1.6 GB.
How good is it, really?
On the Artificial Analysis Intelligence Index v4.3 it scores 13, which puts it in the Below every hosted tier band. Fine for simple, well-specified tasks. Noticeably less capable than anything the big labs sell today. Score source. See the whole table.
- Summarising — not rated
- Translation — not rated
- Everyday coding — not rated
- Reasoning & maths — not rated
- Agentic work — not rated
What it costs either way
Nobody rents MiniCPM5 2B by the token. The closest hosted match, Granite 4.2 8B, costs $0.06 per million input tokens and $0.25 per million output (OpenRouter, cheapest active endpoint, checked 2026-09-09). Buying a machine only beats that if you use it hard enough, for long enough, that the hardware price divides down below the rental bill.
Machines that run it
| Machine | Price | Speed at 32k | Pay-back | |
|---|---|---|---|---|
| Mac mini M6, 16GB | $899 | 39 tok/s estimated | Pays back in 74 years | Run the numbers |
| Strix Halo Framework Desktop, 32GB | $1,269 | 65 tok/s estimated | Pays back in 106 years | Run the numbers |
| MacBook Air M5 (13-inch), 16GB | $1,299 | 39 tok/s estimated | Pays back in 106 years | Run the numbers |
| MacBook Pro M5 (14-inch), 16GB | $1,999 | 39 tok/s estimated | Pays back in 164 years | Run the numbers |
| Mac Studio M5 Max, 36GB | $2,499 | 116 tok/s estimated | Pays back in 201 years | Run the numbers |
| DGX Spark GB10 Grace Blackwell, 128GB | $4,699 | 69 tok/s estimated | Pays back in 393 years | Run the numbers |
One machine per family, cheapest first. Speeds are measured where a public benchmark exists and estimated from memory bandwidth otherwise; the calculator says which for any configuration.
The specifics
- Parameters
- 2.5B
- Quantisation
- Q4_K_M
- Weights on disk
- 1.6 GB
- KV cache
- 1.4 GB at 32k context — Standard grouped-query attention, 16 query heads to 2 KV heads.
- Maximum context
- 128k tokens (128k)
- Licence
- Apache 2.0
- Sources
- source 1, source 2