This was basically seven questions stapled together and handed to me like it was one thing.
Frame the staggered rollout as a quasi-experiment using a modern difference-in-differences design with not-yet-treated units as controls, explicitly addressing immortal-time bias via a risk-set definition. Then layer on robustness checks: pre-trends, negative controls, spillover adjustments, and a matching/weighting backup, while quantifying power for the smallest meaningful effect.
Pro tip: Lead with the identification threat—staggered adoption with heterogeneous effects—and name-drop Callaway & Sant'Anna or Sun & Abraham to show you know why TWFE fails. Then tie every design choice back to the two outcomes (CSAT and 4-week retention) and the household spillover, showing you can balance rigor with product relevance.
Specify treatment as market-channel-time adoption of reminders, and define the risk set as users eligible at each time (avoiding immortal-time bias by aligning time zero with eligibility). State the estimand: ATT of reminders on CSAT and 4-week retention, allowing for household spillovers.
Use a modern DID estimator (e.g., Callaway & Sant'Anna) with not-yet-treated as controls, including unit and time fixed effects and clustering SEs at the market-channel level. For spillovers, consider a household-level exposure model or spatial/network DID.
Test pre-trends via event-study plots and joint F-tests; include two negative controls (e.g., unrelated health outcome, pre-period placebo treatment) and a falsification test (e.g., randomize treatment timing in placebo).
If parallel trends fail, use propensity score matching or inverse probability weighting on pre-treatment covariates (demographics, channel usage, baseline health) to construct a comparable control group, then re-estimate DID.
Address missing CSAT via multiple imputation or IPW; model channel self-selection with a Heckman selection or instrumental variable; outline power/MDE using simulation or formulas for clustered designs, targeting the smallest detectable effect on retention.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.