Comfortable on the Q2_0_g128 GGUF + Q4 KV profile.
- Model size
- 27B · 1.71-bit · 262K max
- Model format
- Q2_0_g128 GGUF + Q4 KV
- Runs with
- llama.cpp / MLX
The published full-context peak is about 12.8 GB with a Q4 KV cache, so 16 GB leaves practical runtime headroom.