Sunk Cost sunkcost.ai Data checked 2026-09-03

What hardware do you need to run Ling 3.0 tiny?

Ling 3.0 tiny at Q4_K_M is 4.8 GB of weights, with a context ceiling of 128k tokens. Not yet rated: released after our last ratings pass. 1.3B active in a 7.9B mixture, so it runs quickly on very little memory.

Cheapest machine that runs itMac mini M6, 16GB at $899
Fastest of the ones listedMac Studio M5 Ultra, 96GB — 150 tok/s at 32k context
Honest answer on costPays back in 80 years at 500k tokens a day.

How good is it, really?

On the Artificial Analysis Intelligence Index v4.3 it scores 12, which puts it in the Below every hosted tier band. Fine for simple, well-specified tasks. Noticeably less capable than anything the big labs sell today. Score source. See the whole table.

What it costs either way

Nobody rents Ling 3.0 tiny by the token. The closest hosted match, Granite 4.2 8B, costs $0.06 per million input tokens and $0.25 per million output (OpenRouter, cheapest active endpoint, checked 2026-09-09). Buying a machine only beats that if you use it hard enough, for long enough, that the hardware price divides down below the rental bill.

Machines that run it

MachinePriceSpeed at 32kPay-back
Mac mini M6, 16GB $899 19 tok/s estimated Pays back in 80 years Run the numbers
Strix Halo Framework Desktop, 32GB $1,269 59 tok/s estimated Pays back in 107 years Run the numbers
MacBook Air M5 (13-inch), 16GB $1,299 19 tok/s estimated Pays back in 115 years Run the numbers
MacBook Pro M5 (14-inch), 16GB $1,999 19 tok/s estimated Pays back in 177 years Run the numbers
Mac Studio M5 Max, 36GB $2,499 57 tok/s estimated Pays back in 212 years Run the numbers
DGX Spark GB10 Grace Blackwell, 128GB $4,699 74 tok/s estimated Pays back in 391 years Run the numbers

One machine per family, cheapest first. Speeds are measured where a public benchmark exists and estimated from memory bandwidth otherwise; the calculator says which for any configuration.

The specifics

Parameters
7.9B, of which 1.3B are active per token
Quantisation
Q4_K_M
Weights on disk
4.8 GB
KV cache
1.6 GB at 32k context — Hybrid: three linear-attention layers per full-attention layer. Those layers use compressed latent attention, which this figure does NOT model, so the KV cache shown is an over-estimate.
Maximum context
128k tokens (128k)
Licence
MIT
Sources
source 1, source 2