Run a golden dataset across prompt variants — see which wins on pass-rate, cost and latency, then promote it.
Share of rows each variant passed. Winner in accent.
Notional ₹ per row (usage-value, not a bill).
p50 / p95 response time (ms).
One row per variant. Promote repoints the winning prompt's prod tag. Owner/admin only.
| Variant | Prompt / model | Passed | Pass-rate | Mean ₹ | Total ₹ | p50 ms | p95 ms | Errors | |
|---|---|---|---|---|---|---|---|---|---|
| No run selected. | |||||||||
Notional ₹ is usage-value, not a charge. Experiment traffic rides the loopback self-call so routing, prompt-render and residency pinning apply exactly as production.