A benchmark for testing whether language models can generate working Svelte components.
AI Evaluation
Benchmarks, testing, guardrails, and observation of model or agent behavior.
Saved Links in AI Evaluation
The Svelte team's repository for AI experiments, including its language model benchmark.
Anthropic describes the design of its multi-agent research system.
Damien Charlotin's database of court decisions involving AI-generated false or misleading material.
An observability service for inspecting application behavior through traces, logs, and metrics.
An Anthropic study of the values expressed by language models in real conversations.
Documentation for testing and evaluating prompts, models, and AI applications.
A site focused on reviews of generative AI models.
A presentation surveying ways to improve language model application performance.
A platform associated with monitoring AI systems and applying guardrails to model interactions.
An essay examining claims about the pace of AI progress.
ARC Prize analyzes OpenAI o3's results on the ARC-AGI benchmark at its December 2024 announcement.