Comfortable on the SHQ8 MTP GGUF profile.
Edition: Quantized edition by Wepiqx
- Model size
- 9B · up to 1M context
- Model format
- SHQ8 MTP GGUF
- Runs with
- llama.cpp
The 5.8–6.4 GB artifacts fit, but the advertised 1M context requires far more memory than this profile.