Decision replay
Replay runs the current decision policy over historical requests. Scope it by task, time window, and a sample limit of up to 500. Results show requested, evaluated, and unroutable counts, projected model distribution, actual cost, projected cost, and projected savings. Replay predicts decisions; it does not call the candidate or prove response quality.Cost and latency reports
Saved cost reports preserve a replay projection with its time window, model distribution, and actual versus projected cost. Latency projections indicate whether they are estimated and identify their basis. Treat an unavailable latency basis as unknown, not zero. Report jobs can apply structured traffic filters or an interpreted natural-language question. Review the interpreted filters before relying on the result.Live-shadow comparison
Shadow sends a sampled copy of an eligible request to a candidate while the live response remains authoritative. Paired results compare cost, latency, output tokens, and shadow errors. Set a target, sample rate, and monthly shadow budget. A shadow result is evidence about the sampled traffic, not an automatic production promotion.Quality evaluations
An evaluation replays sampled traffic against a candidate and grades primary and candidate responses. Runs progress through queued, running, complete, failed, or budget-exceeded states. A completed verdict includes:- sample count
- primary and candidate averages
- per-dimension scores
- threshold margin and minimum sample requirement
- a readiness result
- measured latency averages when available
Control spend
Understand which limits apply to live traffic, shadow calls, evaluations, and custom routers.