Weights & Biases

Weights & Biases

MLOps platform for tracking AI experiments, comparing models, and managing the full LLM application lifecycle — the standard for AI teams.

Data & Analytics New Freemium ★ 4.7
Disclosure: Some links on this page are affiliate links. If you sign up through them, we may earn a commission at no extra cost to you. This helps support NewAIAppsList and lets us keep the directory free. We only recommend tools we believe add real value.

Weights & Biases (W&B) is the industry-standard platform for ML experiment tracking, model evaluation, and LLM application management. Used by over a million developers including researchers at OpenAI, DeepMind, and NVIDIA.

Free for unlimited public projects and limited private projects. Team ($50/user/mo) gives unlimited private projects. Enterprise pricing adds compliance and SSO.

Core features: Experiment tracking (log every run’s metrics, hyperparameters, and outputs), Weave (LLM observability — trace every LLM call, evaluate outputs, monitor production), Artifacts (version control for datasets and model checkpoints), Reports (collaborative analysis documents), and Sweeps (automated hyperparameter optimisation).

W&B Weave is the go-to tool for teams building with LLMs: traces every API call, logs inputs and outputs, measures latency and cost, and enables A/B testing of prompts and models in production.

Pros: Industry standard, Weave is excellent for LLM observability, comprehensive experiment tracking, integrates with all major ML frameworks, strong documentation.
Cons: Complex for beginners, Team plan is expensive, some features require custom setup.

Best for: ML engineers, data scientists, AI researchers, and teams building production LLM applications who need rigorous experiment tracking and observability.