12 GB accelerator memory
Q2_0_g128 GGUF · reduced context · llama.cpp / MLX
The model card reports roughly 8.4–8.7 GB at 4K–10K context; 12 GB is the safer reduced-context boundary.
Prism ML
Text model
A 27B multimodal reasoning model compressed into a laptop-sized ternary build.
Capabilities
The most useful reasons to choose this model, without making you read through its repository first.
Local reasoning assistants
Tool-using workflows
Multimodal analysis
For providers
The provider requirements live here, separate from the information you need to choose and use the model.
12 GB accelerator memory
Q2_0_g128 GGUF · reduced context · llama.cpp / MLX
The model card reports roughly 8.4–8.7 GB at 4K–10K context; 12 GB is the safer reduced-context boundary.
CPU + 16 GB RAM
Q2_0_g128 GGUF · reduced context · llama.cpp
The 7.2 GB packed model supports CPU inference; throughput depends heavily on memory bandwidth and context length.
16 GB accelerator memory
Q2_0_g128 GGUF + Q4 KV · llama.cpp / MLX
The published full-context peak is about 12.8 GB with a Q4 KV cache, so 16 GB leaves practical runtime headroom.
Source and license
OpenMayhem links back to the source repository so you can inspect the model card, files, license, limitations, and creator guidance directly.