What hardware do you need to run GLM-4.7-Flash?
GLM-4.7-Flash at Q4_K_M is 18 GB of weights, with a context ceiling of 198k tokens. Not yet rated: released after our last ratings pass. Replaces GLM-4.5-Air at a quarter of the size and a higher score.
How good is it, really?
On the Artificial Analysis Intelligence Index v4.3 it scores 15 (reasoning mode; 11 without), which puts it in the Haiku-class band. In the same band as Anthropic's cheap, fast tier. Every current OpenAI model scores above this band. Score source. See the whole table.
- Summarising — not rated
- Translation — not rated
- Everyday coding — not rated
- Reasoning & maths — not rated
- Agentic work — not rated
What it costs either way
Renting the same model costs $0.061 per million input tokens and $0.4 per million output (OpenRouter, cheapest active endpoint, checked 2026-09-09). Buying a machine only beats that if you use it hard enough, for long enough, that the hardware price divides down below the rental bill.
Machines that run it
| Machine | Price | Speed at 32k | Pay-back | |
|---|---|---|---|---|
| Strix Halo Framework Desktop, 32GB | $1,269 | 40 tok/s estimated | Pays back in 96 years | Run the numbers |
| Mac mini M6, 32GB | $1,299 | 14 tok/s estimated | Pays back in 103 years | Run the numbers |
| MacBook Pro M5 (14-inch), 32GB | $2,399 | 13 tok/s estimated | Pays back in 195 years | Run the numbers |
| Mac Studio M5 Max, 36GB | $2,499 | 39 tok/s estimated | Pays back in 192 years | Run the numbers |
| DGX Spark GB10 Grace Blackwell, 128GB | $4,699 | 50 tok/s estimated | Pays back in 351 years | Run the numbers |
One machine per family, cheapest first. Speeds are measured where a public benchmark exists and estimated from memory bandwidth otherwise; the calculator says which for any configuration.
The specifics
- Parameters
- 31.2B, of which 3B are active per token
- Quantisation
- Q4_K_M
- Weights on disk
- 18 GB
- KV cache
- 1.8 GB at 32k context — Latent attention: the cache holds one 512-wide compressed vector plus 64 RoPE dimensions per layer, 1,152 bytes a token, not the 20 KB a head count would suggest. The config states no head_dim and it cannot be derived.
- Maximum context
- 198k tokens (198k)
- Licence
- MIT
- Sources
- source 1