08. Multiple Metrics and Guardrail Verdicts

A real launch decision is rarely one t-test. It's a goal metric that needs to move, plus a handful of guardrail metrics that must not get meaningfully worse — and checking each one at the ordinary 5% significance level, independently, quietly inflates the odds that at least one guardrail trips by chance alone. This chapter builds a single SQL query that computes every metric's test statistic and reduces them to one ship/hold verdict, with the guardrail threshold corrected for how many guardrails there are.

Why Checking Guardrails One at a Time Is a Trap

Suppose an experiment tracks one goal metric and two guardrail metrics...

💎

Premium membership required

Upgrade to premium to access the full chapter.