A self-hosted service for running language, vision, speech, and other AI models through compatible APIs.
Models & Inference
Model releases, local runtimes, serving systems, and tools for running model predictions or generation.
Saved Links in Models & Inference
An inference project that streams mixture-of-experts model weights from storage to work within limited memory.
An experimental inference tool for running mixture-of-experts language models on Macs by streaming expert weights from SSD storage.
A C/C++ speech-recognition inference library built on ggml.
An AI gateway that exposes multiple model providers through a shared API.
Microsoft's research project for generating conversational speech from text.
A platform for hosting and scaling language-model inference on an organization's own infrastructure.
A desktop AI assistant that can run downloaded models locally or connect to hosted model services.
A service for adjusting language-model serving capacity to demand.
The implementation of a document image parsing model from ByteDance.
A library for extracting structured data from language model responses.
A Kubernetes platform for deploying and serving machine learning models.