What hardware do you need to run Qwen3 235B-A22B Instruct 2507?
Qwen3 235B-A22B Instruct 2507 at Q4_K_M is 142 GB of weights, with a context ceiling of 256k tokens. The closest thing to a frontier model that fits in 256 GB. Needs an Ultra-class Mac.
How good is it, really?
On the Artificial Analysis Intelligence Index v4.3 it scores 13 (reasoning mode; 12 without), which puts it in the Below every hosted tier band. Fine for simple, well-specified tasks. Noticeably less capable than anything the big labs sell today. Score source. See the whole table.
- Summarising — good
- Translation — good
- Everyday coding — good
- Reasoning & maths — good
- Agentic work — usable
What it costs either way
Renting the same model costs $0.0875 per million input tokens and $0.35 per million output (OpenRouter, cheapest active endpoint, checked 2026-09-03). Buying a machine only beats that if you use it hard enough, for long enough, that the hardware price divides down below the rental bill.
Machines that run it
| Machine | Price | Speed at 32k | Pay-back | |
|---|---|---|---|---|
| Mac Studio M5 Ultra, 256GB | $10,799 | 18 tok/s estimated | Pays back in 977 years | Run the numbers |
One machine per family, cheapest first. Speeds are measured where a public benchmark exists and estimated from memory bandwidth otherwise; the calculator says which for any configuration.
The specifics
- Parameters
- 235.1B, of which 22B are active per token
- Quantisation
- Q4_K_M
- Weights on disk
- 142 GB
- KV cache
- 6.3 GB at 32k context
- Maximum context
- 256k tokens (256k natively)
- Licence
- Apache 2.0
- Sources
- source 1, source 2