Start by defining success metrics for both riders and drivers, then outline a randomized controlled experiment (A/B test) comparing the new model to the current one, and finally propose methods to disentangle model accuracy from behavioral changes, such as using a holdout group or instrumental variables.
Pro tip: Emphasize the importance of guardrail metrics to ensure the new model doesn't negatively impact overall marketplace health, and suggest using a switchback or staggered rollout to account for time-based confounders.
Identify key metrics for riders (e.g., wait time, walk time, cancellation rate, satisfaction) and drivers (e.g., idle time, trip completion rate, earnings per hour) that reflect the impact of the new model.
Propose a randomized controlled trial where riders are randomly assigned to either the new model or the current model, ensuring proper randomization and sample size calculation.
Use a holdout group that receives no ETA information or a fixed ETA to measure pure model accuracy, or employ causal inference methods like instrumental variables to separate behavioral effects.
Compare metrics between groups, check for statistical significance, and consider segment-level analysis to understand heterogeneous effects.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.