This was basically five questions in a trenchcoat.
Start by framing the experiment around the core hypothesis that a stricter spam filter reduces spam friend requests without harming legitimate friend request dynamics. Then walk through the design choices—randomization unit, metrics, power analysis—and explicitly address how you'd handle incomplete spam labels using proxy metrics and sensitivity analyses.
Pro tip: Emphasize that you would pre-register your analysis plan and define guardrail metrics upfront to avoid p-hacking, and mention that you'd use a holdout group to measure long-term effects on user engagement.
State a clear null and alternative hypothesis: the stricter filter reduces spam friend requests (primary) without decreasing legitimate friend requests or overall engagement (guardrails). Define primary, secondary, and guardrail metrics with precise definitions.
Decide whether to randomize at the user level (e.g., by recipient or sender) or at the request level, considering network effects and interference. Discuss trade-offs and potential for spillover, and propose a cluster-randomized design if needed.
Estimate baseline rates for key metrics, specify minimum detectable effect (MDE), and calculate required sample size and duration. Account for multiple comparisons and consider sequential testing if peeking is a concern.
Acknowledge that spam labels are often incomplete or noisy. Propose using proxy metrics (e.g., user reports, block rates) and conducting sensitivity analyses to bound the treatment effect under different labeling assumptions.
Outline the statistical tests (e.g., t-test, CUPED for variance reduction), subgroup analyses, and how you'll interpret results. Define decision criteria and next steps based on outcomes.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.