A Kubernetes-native framework for distributed language model inference.
Models & Inference
Model releases, local runtimes, serving systems, and tools for running model predictions or generation.
Saved Links in Models & Inference
A text-to-speech model from Nari Labs for generating dialogue.
An illustrated guide to understanding and applying language models.
The code and notebooks accompanying Hands-On Large Language Models.
A site focused on reviews of generative AI models.
A presentation surveying ways to improve language model application performance.
Simon Willison's retrospective on language model releases and developments in December 2024.
A DigitalOcean tutorial on deploying Hugging Face inference services on GPU Droplets.
A collection of notebooks and examples for using the Gemini API.
A Nous Research instruction-tuned model based on Llama 3.2 3B.
Cohere's announcement of the Command R7B language model.
An audio-language model from Nexa AI for processing audio and producing text responses.