Sunk Cost sunkcost.ai Data checked 2026-09-03

What hardware do you need to run Granite 4.2 8B?

Granite 4.2 8B at Q4_K_M is 5.3 GB of weights, with a context ceiling of 128k tokens. Not yet rated: released after our last ratings pass. IBM's small model, aimed at enterprises that care about provenance.

Cheapest machine that runs itMac mini M6, 24GB at $1,099
Fastest of the ones listedMac Studio M5 Ultra, 96GB — 84 tok/s at 32k context
Honest answer on costPays back in 108 years at 500k tokens a day.

How good is it, really?

On the Artificial Analysis Intelligence Index v4.3 it scores 12, which puts it in the Below every hosted tier band. Fine for simple, well-specified tasks. Noticeably less capable than anything the big labs sell today. Score source. See the whole table.

What it costs either way

Renting the same model costs $0.06 per million input tokens and $0.25 per million output (OpenRouter, cheapest active endpoint, checked 2026-09-09). Buying a machine only beats that if you use it hard enough, for long enough, that the hardware price divides down below the rental bill.

Machines that run it

MachinePriceSpeed at 32kPay-back
Mac mini M6, 24GB $1,099 12 tok/s estimated Pays back in 108 years Run the numbers
Strix Halo Framework Desktop, 32GB $1,269 18 tok/s estimated Pays back in 139 years Run the numbers
MacBook Pro M5 (14-inch), 32GB $2,399 11 tok/s estimated Pays back in 243 years Run the numbers
Mac Studio M5 Max, 36GB $2,499 32 tok/s estimated Pays back in 234 years Run the numbers
DGX Spark GB10 Grace Blackwell, 128GB $4,699 19 tok/s estimated Pays back in 528 years Run the numbers

One machine per family, cheapest first. Speeds are measured where a public benchmark exists and estimated from memory bandwidth otherwise; the calculator says which for any configuration.

The specifics

Parameters
8B
Quantisation
Q4_K_M
Weights on disk
5.3 GB
KV cache
5.4 GB at 32k context — Plain grouped-query attention on all 40 layers, so the context costs real memory. head_dim is absent from the config and derived as hidden_size ÷ heads.
Maximum context
128k tokens (128k)
Licence
Apache 2.0
Sources
source 1