Sunk Cost sunkcost.ai Data checked 2026-09-03

What hardware do you need to run Qwen3.8 27B?

Qwen3.8 27B at Q4_K_M is 16 GB of weights, with a context ceiling of 256k tokens. Not yet rated: released after our last ratings pass. Sizes and prices are current.

Cheapest machine that runs itStrix Halo Framework Desktop, 32GB at $1,269
Shortest pay-backMac mini M6, 32GB — Pays back in 17 years
Fastest of the ones listedMac Studio M5 Ultra, 96GB — 48 tok/s at 32k context
Honest answer on costPays back in 17 years at 500k tokens a day.

How good is it, really?

On the Artificial Analysis Intelligence Index v4.3 it scores 34 (xhigh reasoning effort; 22 without reasoning), which puts it in the Sonnet-class band. In the same league as the labs' mainstream models on this index. Score source. See the whole table.

What it costs either way

Renting the same model costs $0.32 per million input tokens and $2.5 per million output (OpenRouter, cheapest active endpoint, checked 2026-09-03). Buying a machine only beats that if you use it hard enough, for long enough, that the hardware price divides down below the rental bill.

Machines that run it

MachinePriceSpeed at 32kPay-back
Strix Halo Framework Desktop, 32GB $1,269 10 tok/s estimated Pays back in 17 years Run the numbers
Mac mini M6, 32GB $1,299 6.9 tok/s estimated Pays back in 17 years Run the numbers
MacBook Pro M5 (14-inch), 32GB $2,399 6.2 tok/s estimated Pays back in 31 years Run the numbers
Mac Studio M5 Max, 36GB $2,499 19 tok/s estimated Pays back in 32 years Run the numbers
DGX Spark GB10 Grace Blackwell, 128GB $4,699 11 tok/s estimated Pays back in 62 years Run the numbers

One machine per family, cheapest first. Speeds are measured where a public benchmark exists and estimated from memory bandwidth otherwise; the calculator says which for any configuration.

The specifics

Parameters
27.8B
Quantisation
Q4_K_M
Weights on disk
16 GB
KV cache
2.1 GB at 32k context — Hybrid: 48 linear-attention layers; only 16 full-attention layers hold a growing KV cache.
Maximum context
256k tokens (256k natively; up to 1M with YaRN)
Licence
Apache 2.0
Sources
source 1, source 2