A product experiment deliberately changes an intervention and compares outcomes to learn about cause. Start with a decision and falsifiable hypothesis, define population, assignment, primary measure, guardrails, duration, and analysis before launch, then verify execution. Not every question needs an A/B test; prototypes, observational research, or staged pilots may fit better.
Who this is for: Product managers, analysts, designers, and engineers evaluating whether a product change caused a meaningful effect.
- Tie the hypothesis to a mechanism and decision rather than testing whether any metric moves.
- Predefine population, comparison, primary measure, guardrails, sample logic, and stopping rules.
- Interpret practical size, uncertainty, execution quality, and segment effects before choosing an action.
Choose an experimentable question
Clarify the decision and causal claim. 'Showing delivery estimates before checkout will increase completed orders because shoppers can judge arrival suitability' names intervention, outcome, population, and mechanism. A broad request to test a redesign invites selective interpretation. Confirm that the result would change whether the team ships, revises, or stops the idea.
Use an experiment when assignment to conditions is feasible, ethical, and informative. Interviews are better for understanding motives; usability tests reveal interaction failures; historical analysis can diagnose patterns. Do not randomize access to legally required, safety-critical, or clearly beneficial protections merely to obtain a cleaner comparison.
Design the comparison
Define eligible users, unit of assignment, control experience, treatment, allocation, and contamination risks. Assigning individual members may fail when teammates share the changed workflow; the workspace may be the correct unit. Keep other differences minimal so the observed effect can be attributed to the intended intervention.
Estimate the sample and duration based on baseline rate, smallest effect worth acting on, acceptable error, traffic, and full behavior cycle. Avoid ending when a dashboard first looks favorable. Include weekly patterns and delayed outcomes where relevant. Analysts should review the design before implementation, not only after ambiguous data arrives.
Define measurement and safeguards
Choose one primary outcome closely linked to the hypothesis. Add diagnostic measures for the mechanism and guardrails for cancellation, errors, latency, complaints, accessibility, or unequal impact. Document event definitions, exclusions, identity handling, and attribution windows. Validate instrumentation before exposing the full sample.
Write analysis and stopping rules in advance. Decide how missing data, repeated exposure, multiple devices, outliers, and multiple comparisons will be handled. Set operational stop conditions for severe harm even if statistical targets are unmet. Ensure consent, privacy, and review practices fit the intervention and affected population.
Read results as evidence
First verify allocation, exposure, instrumentation, sample composition, and guardrails. Then report effect size and uncertainty, not only a binary significance label. Compare the estimate with the smallest useful effect. A precise tiny increase may not justify engineering, support, or customer complexity.
Inspect prespecified segments and mechanism measures without mining endless cuts for a positive story. Consider novelty, seasonality, concurrent launches, and spillover. Record the decision and rationale, including inconclusive outcomes. A failed hypothesis can improve the opportunity model; it is not a reason to quietly redefine success after seeing results.
Test a checkout delivery estimate
Research suggests shoppers abandon purchases when they cannot tell whether an order will arrive before an event.
- Define eligible shoppers and hypothesize that a pre-checkout date estimate increases completed orders through reduced timing uncertainty.
- Randomize by shopper, keep price and checkout unchanged, and prevent treatment information from leaking into control sessions.
- Set completed order as primary, estimate interaction as mechanism, and cancellations, late deliveries, and support contacts as guardrails.
- Validate date accuracy and event tracking, then run for the planned sample and complete weekly cycle.
- Compare practical effect and guardrails with the shipping threshold, then document ship, revise, or stop.
Experiment protocol
Approve this protocol before exposing customers to an experimental product condition.
- Decision and hypothesis: intervention, population, mechanism, primary outcome, and action each result supports.
- Design: eligibility, assignment unit, control, treatment, allocation, contamination, sample, duration, and ramp.
- Measurement: primary metric, diagnostics, guardrails, event definitions, attribution, and instrumentation check.
- Analysis: smallest useful effect, uncertainty method, exclusions, segments, missing data, stopping, and validity risks.
- Operations: ethics, privacy, approvals, monitoring, rollback, result owner, decision record, and communication.
Common mistakes
- Launching with many possible success metrics and choosing the favorable one after results appear.
- Randomizing individuals when the treatment changes a shared team experience and contaminates conditions.
- Shipping a statistically detectable effect without checking practical size, guardrails, or operating cost.
Try one
An experiment increases notification clicks but also increases opt-outs. Explain what should determine the product decision.
A strong answer returns to the predefined primary outcome, customer-value mechanism, opt-out guardrail, effect sizes, and decision thresholds. Clicks alone may represent curiosity or pressure rather than durable value. The team should verify execution, examine relevant prespecified segments, consider longer-term behavior, and revise or reject the treatment if the guardrail breach outweighs the intended benefit.
Sources
- Atlassian product discovery guideOfficial guidance on examining customer problems and testing product ideas.
- Atlassian product metrics guideOfficial guidance on product metrics for engagement, retention, and business performance.