About This Resource

A platform for hosting and scaling language-model inference on an organization's own infrastructure. It combines inference services with request routing, buffering, model management, and an administration interface.