ATLASResearch
methods
Quantitative/ Research design

A/B testing

Randomise a product change in live use

A/B testing randomly assigns eligible units to product or service variants and compares prespecified outcomes. The unit may be a user, account, session or cluster; it must match exposure and dependence. Reliable tests require instrumentation checks, an appropriate duration and a defined analysis plan. Guardrail metrics examine harms or trade-offs alongside the main outcome.

WHEN IT FITS

Choose it when a reversible product change can be randomly exposed to sufficient eligible traffic and meaningful outcomes can be measured without contamination, serious interference or unresolved consent and governance issues.

Strengths

  • Estimates causal effects in a live operating context
  • Can test decision-relevant behaviour rather than stated intention

Limitations

  • Low traffic or rare outcomes can make tests impractical
  • Interference, novelty and metric gaming complicate interpretation

Know the boundary

Repeatedly checking a conventional p-value and stopping at significance invalidates its usual interpretation.

USED ACROSS
Computer scienceBusiness & MBAEducation