Experiment analysis

Analyzing A/B Test Results

Evaluate randomized experiments using assignment integrity, predefined outcomes, effect sizes, uncertainty, guardrails, and practical decision thresholds.

How this page is maintained

Written for learners, checked against the sources below, and reviewed every quarter. Last reviewed July 27, 2026.

Short answer

Analyze an A/B test by preserving randomized assignment, checking exposure and sample integrity, estimating the treatment-control difference on a predefined primary outcome, and reporting uncertainty with practical magnitude. Review guardrails and implementation failures before deciding. A small p-value does not prove importance, permanence, or the mechanism behind the effect.

Who this is for: Product and growth analysts interpreting controlled experiments for business decisions without overstating statistical evidence.

  • Analyze participants according to randomized assignment unless the experimental design specifies another valid estimand.
  • Report effect size and uncertainty for the primary metric alongside guardrails and sample counts.
  • Use a prewritten decision rule and investigate integrity before interpreting surprising results.

Start from the experiment plan

Confirm unit of randomization, eligible population, variants, allocation, start and stop rules, primary metric, guardrails, minimum detectable effect, and analysis window. These choices should exist before results are examined. Selecting the best-looking metric or segment afterward inflates false discoveries and turns exploration into unacknowledged hypothesis generation.

Match analysis unit to assignment. If stores were randomized, treating individual purchases as independent overstates information. If users can appear on several devices, identity rules must prevent cross-variant contamination. Preserve original assignment even when exposure fails; this intention-to-treat view estimates the effect of offering the treatment under actual delivery conditions.

Check experimental integrity

Compare assigned counts with expected allocation and investigate sample-ratio mismatch. Check eligibility, duplicate assignments, bot or employee exclusions defined in advance, logging coverage, variant exposure, and pre-experiment balance on important characteristics. Do not remove inconvenient observations after seeing their outcomes.

Plot daily assignment and metric data to locate outages, ramp changes, novelty periods, and interference. Ensure the full outcome window has elapsed for all included units. If treatment changes tracking rather than behavior, an apparent lift may be measurement bias. Document deviations and decide whether the planned estimate remains interpretable.

Estimate effect and uncertainty

Calculate treatment and control outcomes, absolute difference, relative difference when meaningful, and a confidence interval using a method appropriate to the metric and randomization. Show sample sizes and raw counts. For heavy-tailed revenue, inspect distributions and use a planned robust method rather than quietly deleting high-value observations.

A confidence interval communicates estimates compatible with the data under the method's assumptions. It does not assign probability that this fixed interval contains the effect after calculation. Avoid reducing the result to significant or not significant. Compare the interval with a practical decision threshold and include cost, implementation risk, and reversibility.

Decide with guardrails and context

Review guardrail outcomes such as refunds, latency, complaints, retention, and distributional harms. Correctly account for multiple primary claims if the plan includes them. Segment findings should be labeled exploratory unless powered and specified in advance. A positive average can conceal damage to a consequential population.

State ship, do not ship, continue, or run a targeted follow-up with reasons. Include expected business impact and uncertainty. Monitor after rollout because experimental duration may not capture learning, seasonality, or operational scaling. Preserve the analysis code, assignment snapshot, metric version, and decision so another analyst can reproduce it.

Evaluate a faster checkout experiment

An ecommerce team tests a shorter checkout and sees conversion increase from 10.0 to 10.4 percent after the planned two-week run.

  1. Confirm user-level randomization, expected allocation, complete seven-day conversion windows, and no cross-device reassignment.
  2. Check exposure logs, daily sample ratios, payment errors, and whether the new variant records completion differently.
  3. Estimate the 0.4 percentage-point effect with a confidence interval and convert plausible values to incremental orders and margin.
  4. Review refunds, failed payments, support contacts, page latency, and prespecified customer guardrails.
  5. Apply the planned minimum worthwhile effect and document whether to ship, extend for integrity reasons, or stop.
Result: The recommendation reflects assignment quality, business magnitude, uncertainty, and customer costs rather than celebrating a percentage change alone.

A/B test analysis record

Complete this record from the approved experiment plan and final data.

  • Design: hypothesis, eligibility, randomization unit, variants, allocation, duration, and interference risk.
  • Measures: primary outcome, guardrails, window, minimum worthwhile effect, exclusions, and metric versions.
  • Integrity: assigned counts, sample ratio, exposure, contamination, missingness, balance, and deviations.
  • Estimate: group values, absolute and relative effect, interval, method, sample size, and practical impact.
  • Decision: threshold, guardrail result, segment status, action, owner, follow-up, and reproducibility links.

Common mistakes

  • Stopping when the result first becomes statistically significant despite a fixed-duration plan.
  • Dropping assigned users who did not receive the treatment and thereby measuring a self-selected group.
  • Testing many metrics and segments, then presenting the strongest chance result as the original hypothesis.

Try one

An experiment's primary metric improves, but the confidence interval includes effects smaller than the company's minimum worthwhile change. What should the report say?

The report should give the point estimate, interval, raw group values, sample sizes, and practical threshold, then state that the data remain compatible with effects too small to justify rollout. It should review integrity and guardrails and follow the planned rule. A strong answer may recommend no change or a justified additional experiment, but it must not equate a favorable point estimate with a settled business win.

Sources

Learn this with a tutor

Tell LearnLive what you already know and what you need to do with analyze an a/b test.

Build this course