About This Resource

An inference project that streams mixture-of-experts model weights from storage to work within limited memory. It explores running large models on existing hardware with a small native implementation.