What hardware do you need to run Granite 4.2 30B?
Granite 4.2 30B at Q4_K_M is 18 GB of weights, with a context ceiling of 128k tokens. Not yet rated: released after our last ratings pass. Dense 30B; the context costs far more memory than the hybrid models of the same size.
How good is it, really?
On the Artificial Analysis Intelligence Index v4.3 it scores 15, which puts it in the Haiku-class band. In the same band as Anthropic's cheap, fast tier. Every current OpenAI model scores above this band. Score source. See the whole table.
- Summarising — not rated
- Translation — not rated
- Everyday coding — not rated
- Reasoning & maths — not rated
- Agentic work — not rated
What it costs either way
Nobody rents Granite 4.2 30B by the token. The closest hosted match, GLM-4.7-Flash, costs $0.061 per million input tokens and $0.4 per million output (OpenRouter, cheapest active endpoint, checked 2026-09-09). Buying a machine only beats that if you use it hard enough, for long enough, that the hardware price divides down below the rental bill.
Machines that run it
| Machine | Price | Speed at 32k | Pay-back | |
|---|---|---|---|---|
| Strix Halo Framework Desktop, 64GB | $1,959 | 7.3 tok/s estimated | Pays back in 378 years | Run the numbers |
| Mac mini M5 Pro, 48GB | $2,299 | 8.8 tok/s estimated | Pays back in 360 years | Run the numbers |
| Mac Studio M5 Max, 36GB | $2,499 | 13 tok/s estimated | Pays back in 276 years | Run the numbers |
| MacBook Pro M5 Pro (16-inch), 48GB | $3,599 | 8.8 tok/s estimated | Pays back in 563 years | Run the numbers |
| DGX Spark GB10 Grace Blackwell, 128GB | $4,699 | 7.8 tok/s estimated | Pays back in 1,017 years | Run the numbers |
One machine per family, cheapest first. Speeds are measured where a public benchmark exists and estimated from memory bandwidth otherwise; the calculator says which for any configuration.
The specifics
- Parameters
- 29.3B
- Quantisation
- Q4_K_M
- Weights on disk
- 18 GB
- KV cache
- 8.6 GB at 32k context — Plain grouped-query attention on all 64 layers, so long contexts are expensive here. head_dim derived from hidden_size ÷ heads.
- Maximum context
- 128k tokens (128k)
- Licence
- Apache 2.0
- Sources
- source 1