You shipped a new version, swapped a model, or changed a prompt. Did it actually make things better, or just different? Experiments answer that from your production traffic, with no A/B test to instrument.
An experiment compares variants, and each variant is a slice of your traffic defined by a search query, filters, and a time range. Latitude computes the metrics for each variant and shows you the deltas side by side, across sessions, users, tools, signals, and behaviours.
Pick a preset and you have a comparison in seconds:
- A/B Test splits traffic by an A tag and a B tag
- Versions puts your current version against the last three tagged versions
- Failures compares successful sessions against sessions with errors
- Outliers puts normal sessions against the longest, costliest, and slowest to first token
- Seasonal spreads four variants across the year, or start from Custom with two empty variants
Once it runs, each variant becomes a card with headline panels for sessions, users, total cost, and median duration, and a full metric breakdown below. Every delta is measured against your baseline variant and colored by direction, so a rise in cost or error count lights up while a genuine improvement reads as a win.
Available now in the API and MCP too, so you can spin up variant comparisons from your own tooling.

