This thing had seven sub-parts and I definitely didn't pace myself well.
Structure your answer as a full experiment lifecycle: start with business goals and a falsifiable hypothesis, define metrics and randomization, then cover power, instrumentation, analysis, and decision rules. Emphasize how you'd handle selection bias and heterogeneous effects, since dashers self-select into bike mode. End with a clear scale/rollback framework tied to guardrails.
Pro tip: Treat the opt-in nature as a feature, not a bug: propose an encouragement design or intent-to-treat analysis to estimate causal effects while acknowledging self-selection. This shows you understand real-world experimentation constraints beyond textbook A/B tests.
State the business objective (e.g., increase delivery efficiency and dasher satisfaction while maintaining reliability) and a falsifiable primary hypothesis, such as: 'Enabling bike mode for car dashers will increase deliveries per active hour by at least 5% without degrading on-time delivery rate.'
Define primary success metrics (e.g., deliveries per active hour, dasher retention) and guardrails (e.g., on-time delivery rate, customer rating, cancellation rate, cost per delivery) with precise formulas, time windows, and data sources.
Randomize at the dasher level (or city-level if interference is a concern) because treatment is applied to individuals and spillover between dashers is limited; justify why this unit balances bias and power.
Specify inputs: baseline metric values, minimum detectable effect (MDE), significance level (α=0.05), power (1-β=0.80), and expected variance; calculate required sample size and duration. Pre-rollout, ensure logging of mode selection, delivery events, and guardrail metrics, and run A/A tests.
Analyze intent-to-treat and per-protocol effects, test heterogeneous effects by dasher tenure, city density, and vehicle type, and correct for selection bias using propensity score matching or instrumental variables. Decide scale-up if primary metric improves and guardrails are not violated; rollback if guardrails degrade significantly.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.