This is a future cluster target. Even 4×80 GB is tight once runtime state is included; validate 8×80 GB manually.
Edition: Abliterated edition by Huihui AI
- Model size
- 753B MoE · 1M max
- Model format
- Q2 GGUF · split files
- Runs with
- llama.cpp · multi-GPU
This is a future cluster target. Even 4×80 GB is tight once runtime state is included; validate 8×80 GB manually.