What hardware do you need to run Gemma 4 E4B?
Gemma 4 E4B at QAT Q4_0 is 5.2 GB of weights, with a context ceiling of 128k tokens. Not yet rated: released after our last ratings pass. Google's quantisation-aware build for small machines; the published weights are already 4-bit.
How good is it, really?
On the Artificial Analysis Intelligence Index v4.3 it scores 9 (reasoning mode; 7 without), which puts it in the Below every hosted tier band. Fine for simple, well-specified tasks. Noticeably less capable than anything the big labs sell today. Score source. See the whole table.
- Summarising — not rated
- Translation — not rated
- Everyday coding — not rated
- Reasoning & maths — not rated
- Agentic work — not rated
What it costs either way
Nobody rents Gemma 4 E4B by the token. The closest hosted match, gpt-oss-20b, costs $0.02 per million input tokens and $0.1 per million output (OpenRouter, cheapest active endpoint, checked 2026-09-03). Buying a machine only beats that if you use it hard enough, for long enough, that the hardware price divides down below the rental bill.
Machines that run it
| Machine | Price | Speed at 32k | Pay-back | |
|---|---|---|---|---|
| Mac mini M6, 16GB | $899 | 13 tok/s estimated | Pays back in 466 years | Run the numbers |
| Strix Halo Framework Desktop, 32GB | $1,269 | 41 tok/s estimated | Pays back in 452 years | Run the numbers |
| MacBook Air M5 (13-inch), 16GB | $1,299 | 13 tok/s estimated | Pays back in 674 years | Run the numbers |
| MacBook Pro M5 (14-inch), 16GB | $1,999 | 13 tok/s estimated | Pays back in 1,036 years | Run the numbers |
| Mac Studio M5 Max, 36GB | $2,499 | 40 tok/s estimated | Pays back in 958 years | Run the numbers |
| DGX Spark GB10 Grace Blackwell, 128GB | $4,699 | 51 tok/s estimated | Pays back in 1,571 years | Run the numbers |
One machine per family, cheapest first. Speeds are measured where a public benchmark exists and estimated from memory bandwidth otherwise; the calculator says which for any configuration.
The specifics
- Parameters
- 8B, of which 4.5B are active per token
- Quantisation
- QAT Q4_0
- Weights on disk
- 5.2 GB
- KV cache
- 0.6 GB at 32k context — Only the first 24 layers hold a cache at all: the config shares KV across the last 18, which have no key/value projections. Of those 24, four grow with the context and twenty stop at a 512-token window.
- Maximum context
- 128k tokens (128k)
- Licence
- Apache 2.0
- Sources
- source 1