In short
Budget local-model memory for weights, context cache and runtime, not just the download. 32GB is a useful smaller-model starting point; larger dense models can require 64GB or 128GB depending on quantisation and context.
| Dense model | Approximate Q4 weights | System-memory planning range |
|---|---|---|
| 7B | 4–5GB | 16–32GB |
| 14B | 8–10GB | 24–32GB |
| 32B | 18–22GB | 32–64GB |
| 70B | 40–48GB | 64–128GB |
Editorial estimates for typical four-bit formats and moderate context. Architecture, cache precision and runtime change the requirement.
How much RAM for 7B, 14B, 32B and 70B models?
Use the table as a planning range, not a guaranteed minimum. Four-bit storage is roughly half a byte per parameter before overhead. Longer context, multimodal inputs and concurrent sessions can exceed these suggested ranges.
What quantisation does to memory
Quantisation reduces weight precision and file size, with possible quality and speed trade-offs. A theoretical four-bit 70B payload is about 35GB before overhead. Actual formats also store scales, metadata and higher-precision tensors; runtime and context require more capacity.
Which mini PC tier fits which model?
We prefer Geekom A9 Max for smaller-model desktop work and A9 Mega for larger shared-memory workloads. Verify accelerator-visible capacity. Fitting weights in system RAM is different from keeping the intended workload on the GPU at useful speed.
Frequently asked questions
Can 32GB of RAM run a 70B model?
A conventional four-bit dense 70B workload exceeds that capacity after weights and working memory are included. More aggressive formats change the trade-offs.
Is 64GB enough for local AI?
It is enough for many quantised workloads, but not every model or context. Leave room for system work and check GPU-visible memory.