This thing had five sub-parts and I felt the walls closing in around part (c).
Start by clarifying the business goal and defining the causal question: does providing thermal bags reduce cold-food refunds? Then outline a randomized controlled trial with courier-level randomization, pre-registered metrics, and a detailed analysis plan that accounts for noncompliance and stratification. Emphasize practical constraints like interference and cost, and propose a robust analysis that estimates the intent-to-treat (ITT) effect and possibly the complier average causal effect (CACE).
Pro tip: Mention that you would stratify by historical cold-food refund rate and market to improve power, and that you'd use a CACE analysis to estimate the effect for couriers who actually use the bags, since noncompliance is likely.
Clarify the treatment (providing thermal bags) and outcome (cold-food refunds). Choose courier-level randomization because the intervention is at the courier level, but discuss potential interference if couriers share bags or if customers order from multiple couriers.
Primary metric: cold-food refund rate per delivery (or per courier). Guardrails: total delivery time, customer ratings, courier satisfaction, and cost of bags. Also consider secondary metrics like overall refund rate and repeat order rate.
Stratify by market, courier tenure, and historical cold-food refund rate to reduce variance and ensure balance. Calculate sample size based on expected effect size, power, and intra-courier correlation (if multiple deliveries per courier).
Since not all couriers given bags will use them, plan an ITT analysis as the primary analysis, and a CACE analysis using randomization as an instrument to estimate the effect for compliers. Track bag usage via surveys or app prompts.
Pre-register the analysis plan including primary and secondary metrics, subgroup analyses, and handling of missing data. Use regression adjustment for stratification variables and cluster-robust standard errors if needed. Monitor for novelty effects and ensure blinding where possible.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.