This is a lot to hold in your head at once and I think I fumbled the ordering.
Start by framing the experiment around a clear product hypothesis and the decision it will inform, then walk through the metrics hierarchy (primary, secondary, guardrails), randomization unit, power analysis, and monitoring plan in a logical sequence. Emphasize trade-offs and practical constraints at Airbnb's scale, such as network effects and seasonality.
Pro tip: Show you understand that experiment design is not just statistical but also operational: mention how you'd pre-register the analysis plan and set up automated alerts for guardrail metrics to catch issues early without p-hacking.
Articulate the product change, the expected impact, and the decision the experiment will inform (e.g., launch, iterate, or kill). This anchors the metrics and design choices.
Select a primary success metric tied to the hypothesis, secondary metrics for deeper insight, and guardrail metrics to ensure no harm to user experience or business health.
Determine the randomization unit (e.g., user, listing, city) based on interference risks, then calculate required sample size using power analysis, accounting for baseline rates and minimum detectable effect.
Set up real-time dashboards for key metrics, define stopping rules, and pre-register the analysis plan including segmentation and multiple testing corrections.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.