Sunk Cost sunkcost.ai Data checked 2026-09-03

What hardware do you need to run Nemotron 3.5 Lightning 30B-A3B?

Nemotron 3.5 Lightning 30B-A3B at Q4_K_M is 25 GB of weights, with a context ceiling of 256k tokens. Not yet rated: released after our last ratings pass. NVIDIA's Mamba hybrid: almost no KV cache, so long contexts cost nothing in memory.

Cheapest machine that runs itStrix Halo Framework Desktop, 64GB at $1,959
Fastest of the ones listedMac Studio M5 Ultra, 96GB — 137 tok/s at 32k context
Honest answer on costPays back in 134 years at 500k tokens a day.

How good is it, really?

On the Artificial Analysis Intelligence Index v4.3 it scores 14, which puts it in the Below every hosted tier band. Fine for simple, well-specified tasks. Noticeably less capable than anything the big labs sell today. Score source. See the whole table.

What it costs either way

Renting the same model costs $0.08 per million input tokens and $0.2 per million output (OpenRouter, cheapest active endpoint, checked 2026-09-09). Buying a machine only beats that if you use it hard enough, for long enough, that the hardware price divides down below the rental bill.

Machines that run it

MachinePriceSpeed at 32kPay-back
Strix Halo Framework Desktop, 64GB $1,959 54 tok/s estimated Pays back in 134 years Run the numbers
Mac mini M5 Pro, 48GB $2,299 35 tok/s estimated Pays back in 166 years Run the numbers
Mac Studio M5 Max, 36GB $2,499 53 tok/s estimated Pays back in 172 years Run the numbers
MacBook Pro M5 Pro (16-inch), 48GB $3,599 35 tok/s estimated Pays back in 260 years Run the numbers
DGX Spark GB10 Grace Blackwell, 128GB $4,699 68 tok/s estimated Pays back in 318 years Run the numbers

One machine per family, cheapest first. Speeds are measured where a public benchmark exists and estimated from memory bandwidth otherwise; the calculator says which for any configuration.

The specifics

Parameters
31.6B, of which 3B are active per token
Quantisation
Q4_K_M
Weights on disk
25 GB
KV cache
0.2 GB at 32k context — Only 6 of 52 layers are attention at all; 23 are Mamba blocks with a fixed-size state and 23 are feed-forward only. The context is close to free.
Maximum context
256k tokens (256k)
Licence
OpenMDW 1.1
Sources
source 1