Structure your answer by first outlining the experiment design (metrics, statistical test, MDE, sample size/duration) and then addressing the post-rollout discrepancy. Emphasize the importance of guardrail metrics, power analysis, and potential biases like novelty effects or selection bias. For the diagnosis, systematically consider internal and external validity threats and propose concrete checks.
Pro tip: When explaining the discrepancy, highlight that the original test may have suffered from novelty effect or selection bias, and suggest running a holdback experiment to isolate the true effect. Also, mention that segment-level analysis can reveal heterogeneous treatment effects that explain the diluted lift.
Choose primary metric: purchase conversion rate. Select guardrail metrics: revenue per user, unsubscribe rate, complaint rate, and engagement metrics to ensure no negative impact.
Use a two-sample proportion test (e.g., Z-test) for conversion rate. Consider sequential testing or Bayesian methods if peeking. Set significance level (α=0.05) and power (1-β=0.8).
Set MDE based on business relevance (e.g., 5% relative lift). Calculate sample size using power analysis formula for proportions. Estimate duration based on daily traffic and allocation.
Investigate potential causes: novelty effect, selection bias, external validity, implementation differences, or segment heterogeneity. Propose checks: holdback experiment, segment analysis, and comparison of test and rollout populations.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.