Sunk Cost sunkcost.ai Data checked 2026-09-03

What hardware do you need to run Gemma 4 12B?

Gemma 4 12B at Q4_K_M is 7.1 GB of weights, with a context ceiling of 256k tokens. Not yet rated: released after our last ratings pass. Google's mid-size Gemma 4. The windowed layers keep the context nearly free.

Cheapest machine that runs itMac mini M6, 16GB at $899
Fastest of the ones listedMac Studio M5 Ultra, 96GB — 113 tok/s at 32k context
Honest answer on costPays back in 71 years at 500k tokens a day.

How good is it, really?

On the Artificial Analysis Intelligence Index v4.3 it scores 14 (reasoning mode; 9 without), which puts it in the Below every hosted tier band. Fine for simple, well-specified tasks. Noticeably less capable than anything the big labs sell today. Score source. See the whole table.

What it costs either way

Nobody rents Gemma 4 12B by the token. The closest hosted match, Qwen3.5 9B, costs $0.08 per million input tokens and $0.13 per million output (OpenRouter, cheapest active endpoint, checked 2026-09-15). Buying a machine only beats that if you use it hard enough, for long enough, that the hardware price divides down below the rental bill.

Machines that run it

MachinePriceSpeed at 32kPay-back
Mac mini M6, 16GB $899 14 tok/s estimated Pays back in 71 years Run the numbers
Strix Halo Framework Desktop, 32GB $1,269 24 tok/s estimated Pays back in 104 years Run the numbers
MacBook Air M5 (13-inch), 16GB $1,299 14 tok/s estimated Pays back in 102 years Run the numbers
MacBook Pro M5 (14-inch), 16GB $1,999 14 tok/s estimated Pays back in 157 years Run the numbers
Mac Studio M5 Max, 36GB $2,499 43 tok/s estimated Pays back in 187 years Run the numbers
DGX Spark GB10 Grace Blackwell, 128GB $4,699 26 tok/s estimated Pays back in 391 years Run the numbers

One machine per family, cheapest first. Speeds are measured where a public benchmark exists and estimated from memory bandwidth otherwise; the calculator says which for any configuration.

The specifics

Parameters
12B
Quantisation
Q4_K_M
Weights on disk
7.1 GB
KV cache
0.9 GB at 32k context — 8 of 48 layers grow with the context, and those use a different geometry from the other 40: one KV head at 512 wide (2,048 B/token) against eight heads at 256 wide on the windowed layers, which stop growing at 1,024 tokens.
Maximum context
256k tokens (256k)
Licence
Apache 2.0
Sources
source 1, source 2