The memory is plausible, but this artifact needs an OpenMayhem benchmark first.
- Model size
- 20B · 50-step reference
- Model format
- FP8 / INT8 + offload
- Runs with
- vLLM-Omni
Quantization and layerwise CPU offload may work, but OpenMayhem must benchmark the exact artifact first.