Sunk Cost sunkcost.ai Data checked 2026-09-03

What hardware do you need to run Ling 3.0 flash?

Ling 3.0 flash at Q4_K_M is 78 GB of weights, with a context ceiling of 256k tokens. Not yet rated: released after our last ratings pass. MIT-licensed 124B with only 5.1B active, for machines with a lot of memory.

Cheapest machine that runs itStrix Halo Framework Desktop, 128GB at $3,449
Shortest pay-backMac Studio M5 Max, 128GB — Pays back in 3,743 years
Fastest of the ones listedMac Studio M5 Ultra, 256GB — 52 tok/s at 32k context
Honest answer on costPays back in 4,468 years at 500k tokens a day.

How good is it, really?

On the Artificial Analysis Intelligence Index v4.3 it scores 25, which puts it in the Haiku-class band. In the same band as Anthropic's cheap, fast tier. Every current OpenAI model scores above this band. Score source. See the whole table.

What it costs either way

Renting the same model costs $0.021 per million input tokens and $0.063 per million output (OpenRouter, cheapest active endpoint, checked 2026-09-09). Buying a machine only beats that if you use it hard enough, for long enough, that the hardware price divides down below the rental bill.

Machines that run it

MachinePriceSpeed at 32kPay-back
Strix Halo Framework Desktop, 128GB $3,449 20 tok/s estimated Pays back in 4,468 years Run the numbers
DGX Spark GB10 Grace Blackwell, 128GB $4,699 26 tok/s estimated Pays back in 4,106 years Run the numbers
Mac Studio M5 Max, 128GB $5,099 26 tok/s estimated Pays back in 3,743 years Run the numbers
MacBook Pro M5 Max (16-inch), 128GB $6,999 26 tok/s estimated Pays back in 5,138 years Run the numbers

One machine per family, cheapest first. Speeds are measured where a public benchmark exists and estimated from memory bandwidth otherwise; the calculator says which for any configuration.

The specifics

Parameters
124B, of which 5.1B are active per token
Quantisation
Q4_K_M
Weights on disk
78 GB
KV cache
3.8 GB at 32k context — Hybrid: five linear-attention layers per full-attention layer. Those seven layers use compressed latent attention, which this figure does NOT model, so the KV cache shown is an over-estimate.
Maximum context
256k tokens (256k)
Licence
MIT
Sources
source 1, source 2