This is basically a full DS case study crammed into one question.
Structure your answer as a clear experiment design narrative: start with a testable hypothesis tied to a business goal, then detail the experimental design (randomization, metrics, sample size/power), and finish with pitfalls and a decision framework. Emphasize how you would balance statistical rigor with practical constraints at PayPal, such as user experience and revenue impact.
Pro tip: Show maturity by discussing guardrail metrics (e.g., page load time, customer support contacts) and the importance of pre-registering your analysis plan to avoid p-hacking. Also, mention that you would run a pre-experiment power analysis and consider sequential testing if early stopping is needed.
State a clear, falsifiable hypothesis (e.g., 'The new feature will increase user engagement by X%') and define primary, secondary, and guardrail metrics. Tie them to PayPal's business objectives like conversion, retention, or revenue.
Specify randomization unit (e.g., user-level), control/treatment groups, and duration. Address potential interference, novelty effects, and ensure proper exposure logging.
Calculate required sample size using baseline metric, minimum detectable effect (MDE), significance level (α), and power (1-β). Discuss trade-offs between MDE and runtime.
List common pitfalls (e.g., peeking, multiple comparisons, SRM, novelty/primacy effects) and how you would mitigate them (e.g., sequential testing, Bonferroni correction, sample ratio mismatch checks).
Outline analysis approach (e.g., intention-to-treat, CUPED variance reduction) and decision criteria: ship if primary metric improves significantly without harming guardrails, iterate if inconclusive, or kill if negative.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.