← Tubi Interview Insights

Tubi·Data Scientist·Technical Phone Screen·Senior

SeniorPrefer not to say
Jun 2026Remote

Summary

Tubi data scientist interview with a single very involved measurement question about Super Bowl ad effectiveness. The whole thing was basically one giant case study broken into seven sub-parts, which felt more like a take-home exam than a conversation.

Questions Asked (1)

Q1

Your app ran a 30-second national Super Bowl ad. Design a full measurement plan to estimate the incremental impact on installs and revenue, covering KPIs and measurement windows, identification strategies without a clean control group, a difference-in-differences setup and its assumptions, confounders, uncertainty estimation, falsification checks, and a power calculation for detecting a 4% lift in daily installs.

A/B Testing & ExperimentationProduct Analytics & MetricsData Modeling
Author's notes

This was one question with seven sub-parts and I genuinely did not expect the power calculation at the end.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by framing the measurement problem as a causal inference challenge where the ad creates a national treatment with no clean control, then propose a multi-pronged identification strategy combining synthetic control, difference-in-differences with matched markets, and time-series methods. Structure your answer around defining KPIs and windows, building a credible counterfactual, quantifying uncertainty, and validating with falsification and power analyses.

Pro tip: Acknowledge that a 30-second national ad is a one-time, non-randomized shock, so no single method is definitive; instead, triangulate estimates from multiple designs and explicitly state which assumptions are most fragile. Also, mention that Tubi's ad likely drove both installs and engagement, so revenue should be modeled as a function of installs and retention, not just a direct lift.

1. Define KPIs and measurement windows

Specify primary metrics (daily installs, revenue) and secondary metrics (sign-ups, sessions, retention), along with pre-period (e.g., 4-8 weeks before), event window (game day ± 1 day), and post-period (2-4 weeks after) to capture immediate and carryover effects.

2. Choose identification strategies for causal impact

Since there is no clean control, propose synthetic control using similar apps/markets, difference-in-differences with matched markets (e.g., DMAs with similar pre-trends), and interrupted time series with counterfactual forecasting; discuss assumptions like parallel trends and no spillovers.

3. Address confounders and falsification checks

Identify confounders (seasonality, concurrent campaigns, competitor ads, macro trends) and propose falsification tests: placebo tests on pre-period, checking unaffected metrics (e.g., web traffic), and verifying no pre-trend differences.

4. Estimate uncertainty and conduct power analysis

Use bootstrap or Bayesian methods for credible intervals, and perform a power calculation for detecting a 4% lift in daily installs given historical variance, sample size, and desired power (80%) and significance (5%).

Key Points to Mention

  • Difference-in-differences setup: treatment group (national) vs. control (matched markets or synthetic control), pre/post periods, and the parallel trends assumption.
  • Confounders: Super Bowl Sunday effects, other advertisers, promotions, seasonality, and organic trends.
  • Uncertainty estimation: bootstrapping, Bayesian structural time series, or permutation tests to quantify confidence intervals.
  • Falsification checks: placebo tests on pre-period, testing on metrics that shouldn't be affected (e.g., app crashes), and checking for pre-trends.
  • Power calculation: formula for detecting 4% lift, requiring historical standard deviation of daily installs, sample size (days), and effect size; mention minimum detectable effect.
  • Measurement windows: immediate (0-24h), short-term (1-7 days), and long-term (2-4 weeks) to capture installs and revenue carryover.

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.