AB Testing

Practical deep-dives into ab testing, experiment design, and statistics related. The core topics every data analyst needs to master.

Sequential A/B Testing: The Core Concept

Why peeking at results mid-experiment inflates false positives — and how sequential testing fixes it

A/B TestingSequential TestingStatistics
Read article →

The Peeking Problem: Why Every Team Keeps Refreshing the Dashboard

Every team wants their experiment to win, so everyone keeps refreshing the dashboard until it looks good — and until the guardrail metric doesn't look obviously bad. Here's why that behavior quietly inflates your false positive rate.

A/B TestingPeekingBeginnerStatistics
Read article →

Why a Fixed 7-Day Rule Doesn't Actually Solve Peeking

A fixed check-in day fixes optional stopping on the primary metric — but guardrail metrics like revenue still only get one noisy snapshot, with no way to catch harm early.

A/B TestingPeekingGuardrail MetricsStatistics
Read article →

Group Sequential Testing: A Definition

Sign in required. What Group Sequential Testing actually is, when it's worth the complexity, and how it solves both the peeking problem and the fixed-window problem at once.

A/B TestingSequential TestingGSTStatistics
Read article →

A Worked Example: Spending Your Alpha Across Three Interim Looks

Premium. A concrete day 1 / day 3 / day 7 threshold schedule, the simplest way to sanity-check it, and why that simple check is only an approximation once you account for correlated looks.

A/B TestingAlpha SpendingSequential TestingStatistics
Read article →

GST Alpha Spending Playbook: A 12-Case Boundary Reference for Practical Interim Schedules

Premium. Precomputed, Monte Carlo-verified boundary tables for the four most common interim schedules across three spending function styles — twelve cases in one reference.

A/B TestingAlpha SpendingGSTReference
Read article →

Goal Metrics vs. Guardrail Metrics: Why They Need Different Tests

Every experiment has metrics you're trying to move and metrics you're trying to protect. Testing both with the same superiority test is why guardrail failures slip through unnoticed.

A/B TestingGuardrail MetricsBeginnerStatistics
Read article →

"Flat" Is Not Evidence: The Silent Failure Mode of Guardrail Analysis

"Not significant, so we're good" is the most common misreading in experimentation. A walkthrough of how an underpowered guardrail lets a real regression ship, quietly, again and again.

A/B TestingGuardrail MetricsStatisticsBeginner
Read article →

Non-Inferiority Testing: A Definition

Sign in required. Flipping the null hypothesis to ask the question a guardrail actually needs answered — what the non-inferiority margin is, and how to read a one-sided confidence bound.

A/B TestingNon-InferiorityGuardrail MetricsStatistics
Read article →

Why Non-Inferiority Tests Need Bigger Samples (and How Much)

Ruling out harm down to a specific margin takes more data than detecting any difference at all — and the gap grows quadratically as the margin shrinks.

A/B TestingNon-InferioritySample SizeStatistics
Read article →

How Many Guardrails Is Too Many? The Other Multiplicity Problem

Watching several guardrails at once doesn't inflate false positives — it quietly erodes power instead. Why the fix looks like Bonferroni, applied to β instead of α.

A/B TestingGuardrail MetricsMultiple TestingStatistics
Read article →

Setting the Non-Inferiority Margin: A Worked Example

Premium. Deriving a real revenue guardrail's margin from business impact, discovering it's statistically unaffordable, and the two honest levers that fix it.

A/B TestingNon-InferiorityGuardrail MetricsStatistics
Read article →

Guardrail Testing Playbook: Margins, Sample Sizes, and Corrected Boundaries

Premium. Precomputed sample sizes across metric types, margin levels, and guardrail counts, plus a ship-decision matrix for turning goal and guardrail results into a launch call.

A/B TestingNon-InferiorityGuardrail MetricsReference
Read article →

An ML Engineer Asked Two Questions About a GIF Experiment — One Was Survivorship Bias, One Was Encouragement Design

Two questions that sounded like the same instinct — narrow the analysis down to the users who matter. One breaks randomization outright. The other is encouragement design, and the fix keeps every user in the analysis instead of filtering any of them out.

A/B TestingEncouragement DesignSurvivorship BiasCausal Inference
Read article →

Checking LATE = ITT / r Against the Literature: What Holds, and What Needs a Caveat

Cross-checking the simple ITT-to-LATE division against Bloom (1984), Angrist & Pischke, an HHS policy brief, and the randomization-inference literature — what carries over exactly, and where the standard error needs 2SLS instead.

A/B TestingEncouragement DesignCausal InferenceStatistics
Read article →

Encouragement Design in Practice: An ML Engineer's Question About a Diluted Experiment

Premium. A GIF-thumbnail removal experiment, an ML engineer's question about dilution, and the encouragement-design math that turns a diluted ITT estimate into the effect that actually matters.

A/B TestingEncouragement DesignCausal InferenceStatistics
Read article →

Sample Ratio Mismatch and Survivorship Bias: A Post-Mortem on a GIF-Removal Experiment

Premium. Why an activation subgroup defined on 'saw the treatment-relevant item' produced a 67:33 sample ratio mismatch, and how that mismatch manufactured a guardrail lift that was never real.

A/B TestingSRMSurvivorship BiasStatistics
Read article →

The Math Behind Encouragement Design: A Sensitivity Table for Converting ITT to LATE

Premium. How the estimated treatment effect changes as the exposure rate shifts, why a lower exposure rate implies a larger true effect, and how this connects to the instrumental-variables literature.

A/B TestingEncouragement DesignCausal InferenceReference
Read article →

New Products vs. Established Products: A Factorial Design for Separating Two Competing Effects

Premium. When new, cold-start products and established products compete for the same feed slots, a simple two-arm test can't tell a direct effect from a reallocation effect. A factorial design that separates them, plus a proportionality check for spotting a hidden quality problem.

A/B TestingCold StartExperiment DesignStatistics
Read article →

What Does Arm D Actually Measure? Interaction Effects, Not Quality Transfer

Premium. Arm D doesn't test whether a quality effect transfers to new products — that's already confounded into arm B. D exists for exactly one reason: measuring whether stronger established-product ranking quietly claws back the exposure a cold-start fix was built to win.

A/B TestingCold StartExperiment DesignInteraction Effects
Read article →