I knew the general idea, effect size, significance level, power, but when I tried to walk through it out loud I got tangled up in the direction of the relationship between sample size and power.
Start by explaining the key inputs: significance level (α), power (1-β), effect size, and variance. Then describe how to use the appropriate formula or simulation to compute sample size, and emphasize the importance of estimating effect size from historical data or a pilot. Finally, discuss practical considerations like traffic constraints and business impact.
Pro tip: At Amazon, always tie sample size to business impact: a smaller effect size may require an infeasibly large sample, so discuss trade-offs between power, effect size, and duration. Also, mention sequential testing or Bayesian methods if applicable, as Amazon often uses them.
Specify significance level (α), desired power (1-β), minimum detectable effect (MDE), and variance (or baseline conversion rate).
Select formula-based (e.g., for proportions or means) or simulation-based approach, depending on metric and complexity.
Use historical data, pilot studies, or domain knowledge to estimate MDE and variance; be conservative if uncertain.
Apply the formula or run simulations to calculate required sample size per variant, adjusting for unequal allocation if needed.
Check assumptions (e.g., normality), consider practical constraints (traffic, duration), and possibly adjust for multiple comparisons or sequential testing.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.