This was basically six questions stapled together and I definitely felt the seams.
Start by framing the staggered rollout as a DiD design, then explain why TWFE fails under heterogeneous effects and what it actually identifies. Walk through implementing Callaway-Sant'Anna or Sun-Abraham, including group-time ATTs, aggregation, event-study, pre-trend testing, and inference. Finally, discuss reconciling with PSM and practical considerations for TikTok's 50 regions.
Pro tip: Emphasize that with 50 regions, cluster-robust SEs may be unreliable; use wild cluster bootstrap or randomization inference. Also, highlight that PSM and DiD answer different questions—PSM estimates ATT for treated regions, while DiD estimates ATT under parallel trends—so reconciliation requires careful interpretation.
Describe how TWFE with staggered adoption and heterogeneous effects uses already-treated units as controls, leading to biased estimates. Clarify that TWFE identifies a variance-weighted average of treatment effects, which can be negative even if all ATTs are positive.
Outline the steps: define groups by treatment timing, estimate group-time ATTs using not-yet-treated or never-treated as controls, then aggregate with appropriate weights (e.g., group-size weights) to get overall ATT and event-study estimates.
Create event-time indicators relative to treatment, plot coefficients, and test joint significance of pre-treatment coefficients using an F-test. Discuss sensitivity to binning and anticipation effects.
With 50 clusters, use wild cluster bootstrap or randomization inference to obtain valid p-values. Mention that standard cluster-robust SEs may over-reject.
Explain that PSM can be used to create a matched control group, then run DiD on matched sample. Discuss how PSM and DiD address different biases and how to interpret discrepancies.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.