← Uber Interview Insights

Uber·Data Scientist·Technical Phone Screen·Senior

SeniorPrefer not to say
May 2026Remote

Summary

Uber DS interview that was basically one massive causal inference case study. The whole thing centered on rider incentives and marketplace effects, and it went deep fast. Walked out not totally sure how I did.

Questions Asked (7)

Q1

You're designing a rider-side discount offer targeted by a propensity model. What marketplace health metrics would you define to evaluate the program, and what are the formulas?

Product Analytics & MetricsA/B Testing & Experimentation
Author's notes

I started rattling off conversion rate and GMV like a reflex, which was the wrong move.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by framing the program's goal: to increase rider engagement and marketplace liquidity without eroding profitability. Then define metrics across rider behavior, marketplace efficiency, and unit economics, providing clear formulas for each. Finally, emphasize the importance of A/B testing to measure incremental impact and guard against cannibalization.

Pro tip: Always distinguish between correlation and causation by highlighting the need for a control group; propensity models can create selection bias, so measure incremental lift, not just raw differences.

1. Define program objectives

Clarify the primary goal (e.g., increase ride frequency, improve retention) and secondary goals (e.g., maintain driver utilization, protect margins). This ensures metrics align with business strategy.

2. Select rider behavior metrics

Choose metrics that capture changes in rider behavior, such as redemption rate, incremental rides per redeemed offer, and retention lift. Provide formulas like Redemption Rate = (# Offers Redeemed) / (# Offers Sent).

3. Select marketplace health metrics

Include metrics that reflect overall marketplace efficiency, such as driver utilization, ETA, and match rate. For example, Driver Utilization = (Total Driver Hours on Trip) / (Total Driver Hours Online).

4. Select unit economics metrics

Define profitability metrics like incremental gross bookings, cost per incremental ride, and ROI. For instance, ROI = (Incremental Gross Profit - Program Cost) / Program Cost.

5. Incorporate experimentation design

Explain how to measure incrementality via A/B testing, including holdout groups and statistical significance. Mention metrics like lift and p-value to validate impact.

Key Points to Mention

  • Incremental lift vs. absolute metrics: Use control groups to isolate the program's effect.
  • Redemption rate and its impact on ride frequency.
  • Driver utilization and ETA as indicators of marketplace balance.
  • Cost per incremental ride and ROI to ensure financial sustainability.
  • Cannibalization: Check if discounts are subsidizing rides that would have happened anyway.
  • Long-term retention and lifetime value (LTV) effects, not just short-term gains.

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q2

How would you estimate the causal effect of this incentive program under three different identification scenarios: randomized geo holdouts, a regression discontinuity design using offer score thresholds, and an instrumental variable approach using operational friction?

A/B Testing & ExperimentationProduct Analytics & Metrics
Author's notes

This was the part I found most interesting and also most stressful.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Structure your answer by first clarifying the causal estimand (e.g., ATE or LATE) and the assumptions required for each identification strategy. Then, for each scenario, explain how you would implement the method, check assumptions, and interpret the estimate in the context of Uber's incentive program. Emphasize the trade-offs between internal validity and generalizability across the three designs.

Pro tip: Show awareness that geo holdouts often suffer from interference and spillover effects, so you'd complement them with switchback or cluster randomization if possible. Also, mention that IV estimates a LATE, which may not generalize to all drivers, and that RD requires a sharp threshold and no manipulation of the running variable.

1. Define the causal question and estimand

Clarify what causal effect you want to estimate (e.g., effect of incentive on driver hours or completions) and for which population (e.g., all drivers or those near the threshold). Specify whether you're targeting ATE, ATT, or LATE.

2. Randomized geo holdouts

Explain that randomizing at the geo level allows estimating the ATE if geos are independent and no spillover. Discuss power, balance checks, and potential interference (e.g., drivers crossing geos). Suggest using cluster-robust standard errors and possibly a difference-in-differences if pre-period data available.

3. Regression discontinuity (RD)

Describe using offer score thresholds to compare drivers just above and below the cutoff. Check for manipulation of the running variable (e.g., McCrary test), continuity of covariates, and bandwidth selection. Estimate LATE at the threshold, and discuss external validity.

4. Instrumental variable (IV) with operational friction

Propose operational friction (e.g., app glitches, payment delays) as an instrument for treatment take-up. Argue relevance and exclusion restriction: friction affects participation but not outcomes except through the program. Estimate LATE for compliers, and discuss potential violations (e.g., friction correlated with driver quality).

5. Compare and synthesize

Discuss how the three estimates might differ due to different complier populations and assumptions. Recommend triangulation and sensitivity analyses (e.g., placebo tests, alternative specifications) to strengthen causal claims.

Key Points to Mention

  • Causal estimand: ATE vs. LATE and implications for generalizability
  • Assumptions: SUTVA, ignorability, exclusion restriction, continuity in RD
  • Interference and spillover in geo experiments; mitigation via switchback or cluster randomization
  • Manipulation checks and bandwidth sensitivity in RD
  • Instrument relevance and exclusion restriction in IV; weak instrument diagnostics
  • Triangulation across methods and sensitivity analysis to assess robustness

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q3

Write out the ROI formula for this incentive, accounting for cannibalization, subsidy burn, surge and ETA externalities, and long-term habit formation. How do you estimate LTV lift and amortize acquisition cost, and how do you put confidence intervals on the final number?

Pricing & MonetizationA/B Testing & ExperimentationProduct Analytics & Metrics
Author's notes

Honestly the externalities piece is where I added the most value.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by defining the ROI formula as incremental profit divided by total investment, explicitly subtracting cannibalization, subsidy burn, and externality costs from gross incremental revenue. Then explain how to estimate LTV lift using holdout groups and cohort analysis, amortize acquisition cost over expected lifetime, and quantify uncertainty via bootstrapping or Bayesian methods to produce confidence intervals.

Pro tip: Emphasize that cannibalization and externalities are often the largest hidden costs; propose using a geo-based holdout or switchback experiment to isolate true incrementality, and always report ROI as a distribution (e.g., 90% credible interval) rather than a point estimate to reflect decision risk.

1. Define the ROI formula with all adjustments

Write ROI = (Incremental Gross Profit - Cannibalization Cost - Subsidy Burn - Surge/ETA Externality Cost + Long-term Habit Value) / Total Investment. Clearly label each component and explain how it is measured.

2. Estimate incremental lift and cannibalization

Use a randomized controlled experiment (e.g., switchback or geo holdout) to measure the treatment effect on orders and revenue. Cannibalization is the negative spillover to non-incentivized products or segments, estimated via difference-in-differences or causal inference methods.

3. Quantify subsidy burn and externalities

Subsidy burn is the direct cost of incentives (e.g., discounts, credits). Surge and ETA externalities are indirect costs: increased surge pricing or longer wait times for non-incentivized users, which can be monetized using elasticity models or observed changes in rider/driver behavior.

4. Model long-term habit formation and LTV lift

Use cohort analysis and survival models to estimate how the incentive changes retention and order frequency over time. LTV lift = (New LTV - Baseline LTV) * Number of users. Amortize acquisition cost by spreading it over the expected customer lifetime (e.g., using discounted cash flow).

5. Construct confidence intervals

Propagate uncertainty from each component (e.g., via Monte Carlo simulation or bootstrapping) to generate a distribution of ROI. Report the mean and 90% or 95% confidence interval, and conduct sensitivity analysis on key assumptions.

Key Points to Mention

  • Incremental analysis: use control groups to isolate true lift and avoid overestimating ROI.
  • Cannibalization: measure negative impact on other products or segments, e.g., Uber Eats vs. Rides.
  • Subsidy burn: track direct incentive costs and their impact on unit economics.
  • Externalities: account for surge pricing and ETA changes that affect user experience and future demand.
  • LTV lift: estimate long-term value changes using cohort retention curves and habit formation models.
  • Confidence intervals: use bootstrapping or Bayesian methods to quantify uncertainty and support decision-making.

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q4

Design a two-stage randomization scheme at the city-week level and then the rider level to separately identify direct effects versus spillover effects. Define your estimands clearly.

A/B Testing & ExperimentationSystem Design
Author's notes

This one I actually liked.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by clearly defining the direct and spillover estimands in the context of a two-stage randomization: first randomize cities to treatment/control, then within treated cities randomize riders to treatment/control. Explain how the second stage identifies direct effects (treated riders in treated cities vs. control riders in treated cities) and how the first stage identifies spillover effects (control riders in treated cities vs. control riders in control cities). Emphasize the need for careful design to avoid contamination and ensure sufficient power at both levels.

Pro tip: Highlight the importance of pre-registering your analysis plan and using cluster-robust standard errors to account for correlation within cities and weeks. Also, discuss how to handle potential interference between riders in the same city by measuring spillover through network or geographic proximity.

1. Define the randomization units and treatment arms

Specify that cities are randomly assigned to either treatment or control at the first stage, and then within each treated city, riders are randomly assigned to treatment or control at the second stage. This creates four groups: treated riders in treated cities, control riders in treated cities, treated riders in control cities (if any), and control riders in control cities.

2. Clearly define the estimands

Direct effect: difference in outcomes between treated and control riders within treated cities. Spillover effect: difference in outcomes between control riders in treated cities and control riders in control cities. Total effect: difference between treated riders in treated cities and control riders in control cities.

3. Discuss identification assumptions and potential biases

State assumptions such as no interference between cities (SUTVA at city level) and no spillover from control cities. Address potential biases like selection bias if rider randomization is not truly random, and confounding if city-level characteristics differ.

4. Outline the analysis plan

Use regression models with fixed effects for city and week, and cluster-robust standard errors at the city level. For direct effect, compare treated vs. control riders within treated cities; for spillover, compare control riders in treated vs. control cities. Consider interaction terms to test heterogeneity.

5. Address power and sample size considerations

Calculate power for both stages: number of cities needed for spillover detection and number of riders per city for direct effect. Discuss trade-offs between number of cities and riders per city given budget constraints.

Key Points to Mention

  • Definition of direct, spillover, and total effects with clear notation (e.g., Y_i(city treatment, rider treatment)).
  • Two-stage randomization design: city-level randomization first, then rider-level within treated cities.
  • Identification assumptions: no interference between cities, no spillover from control cities, and random assignment within cities.
  • Statistical methods: cluster-robust standard errors, fixed effects, and potential use of instrumental variables if compliance is imperfect.
  • Power analysis: ensuring sufficient number of cities and riders per city to detect effects.
  • Practical challenges: contamination, network effects, and implementation logistics at Uber scale.

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q5

With one primary metric, three key secondary metrics, and around twenty diagnostic checks, what multiple testing correction strategy would you use and why?

A/B Testing & ExperimentationTechnical Trade-offs
Author's notes

Went with hierarchical testing and explained why flat Bonferroni is too conservative when you have 20 diagnostics that are mostly sanity checks.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by clarifying the testing hierarchy: one primary metric for the decision, three secondary metrics for directional support, and twenty diagnostics for health checks. Then propose a tiered correction strategy: strict control for the primary (e.g., alpha=0.05), moderate correction for secondary (e.g., Holm-Bonferroni or Benjamini-Hochberg), and minimal or no correction for diagnostics, focusing instead on effect sizes and practical significance.

Pro tip: Emphasize that diagnostics are not for hypothesis testing but for detecting guardrail violations; use false discovery rate (FDR) control for secondary metrics to balance power and false positives, and always pre-register the correction plan to avoid p-hacking concerns.

1. Clarify the metric hierarchy and decision rules

Confirm that the primary metric drives the launch decision, secondary metrics provide supporting evidence, and diagnostics are for sanity checks. Establish that the primary metric is tested at a strict alpha (e.g., 0.05) without correction.

2. Apply family-wise error rate (FWER) control for secondary metrics

Use a method like Holm-Bonferroni or Bonferroni to control the probability of any false positive among the three secondary metrics, since they are a small family and false positives are costly.

3. Use false discovery rate (FDR) or no correction for diagnostics

For the twenty diagnostic checks, apply Benjamini-Hochberg (FDR) if you must flag anomalies, or skip formal correction and rely on effect sizes and domain thresholds to identify issues.

4. Consider sequential testing and peeking

If the experiment is monitored continuously, use group sequential methods or alpha spending to control error rates over time, especially for the primary metric.

5. Pre-register and document the correction plan

State the correction strategy before analysis to maintain statistical rigor and avoid criticism of p-hacking. Include sensitivity analyses to show robustness.

Key Points to Mention

  • Family-wise error rate (FWER) vs. false discovery rate (FDR) and when to use each
  • Holm-Bonferroni or Bonferroni for small families of secondary metrics
  • Benjamini-Hochberg procedure for larger sets of diagnostics
  • Sequential testing adjustments (e.g., O'Brien-Fleming, alpha spending) if peeking
  • Pre-registration of the analysis plan to ensure validity
  • Practical significance and effect sizes for diagnostics, not just p-values

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q6

How would you pre-register and execute a heterogeneity analysis to detect effect moderation by city tier and weather conditions, while controlling for sign and magnitude errors?

A/B Testing & ExperimentationProduct Analytics & Metrics
Author's notes

Causal forests came up naturally here.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by outlining a pre-registration plan that specifies hypotheses, subgroups (city tier, weather conditions), primary and secondary metrics, and correction methods for multiple comparisons. Then describe the execution: pre-register the analysis plan, collect data, run heterogeneity tests (e.g., interaction effects), and apply appropriate statistical controls to mitigate sign and magnitude errors. Emphasize the importance of pre-registration in preventing p-hacking and ensuring valid inference.

Pro tip: Pre-register not just the subgroups but also the minimum detectable effect (MDE) for each subgroup to avoid post-hoc power issues. Use a hierarchical model or Bayesian approach to borrow strength across subgroups, reducing the risk of false positives from multiple comparisons.

1. Define hypotheses and subgroups

Clearly state the primary hypothesis and the expected moderation effects by city tier (e.g., Tier 1 vs. Tier 2) and weather conditions (e.g., rain vs. clear). Specify the direction and magnitude of expected effects.

2. Pre-register analysis plan

Document the analysis plan including metrics, subgroup definitions, statistical tests (e.g., interaction terms in regression), correction methods (e.g., Bonferroni, FDR), and decision rules. Register on a public platform like AsPredicted or OSF.

3. Execute experiment and collect data

Run the experiment ensuring proper randomization and data collection. Blind analysts to treatment assignment where possible to reduce bias.

4. Analyze heterogeneity with controls

Use regression models with interaction terms to test moderation. Apply multiple comparison corrections and check for sign (direction) and magnitude errors by comparing observed effects to pre-registered expectations.

5. Interpret and report

Report subgroup effects with confidence intervals, noting any deviations from pre-registered hypotheses. Discuss limitations and potential for false positives/negatives.

Key Points to Mention

  • Pre-registration prevents p-hacking and ensures confirmatory analysis.
  • Multiple comparison corrections (e.g., Bonferroni, Benjamini-Hochberg) control family-wise error rate.
  • Interaction terms in regression models test effect moderation.
  • Sign errors: effect direction opposite to hypothesis; magnitude errors: effect size different from expected.
  • Power analysis for subgroups to ensure adequate sample size.
  • Bayesian hierarchical models can improve estimates for small subgroups.

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q7

What diagnostics would you run to validate the experiment: pre-trend checks, placebo tests, negative control outcomes? And what are your explicit criteria for declaring success versus triggering a rollback?

A/B Testing & ExperimentationRoot Cause Analysis
Author's notes

Pre-trends I covered well.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by outlining a layered diagnostic plan that validates the experiment's assumptions before interpreting results. Then, define clear, pre-registered success and rollback criteria that balance statistical significance with practical impact and guardrail metrics. Emphasize that these criteria should be set before the experiment starts to avoid post-hoc rationalization.

Pro tip: At Uber, where experiments run at massive scale, even tiny effects can be statistically significant but practically meaningless. Always pair p-values with confidence intervals and effect sizes, and tie rollback decisions to guardrail metrics like cancellation rates or driver acceptance rates, not just the primary metric.

1. Pre-trend and Sample Ratio Mismatch (SRM) Checks

Verify that the treatment and control groups were balanced before the experiment and that the randomization didn't introduce bias. Run an SRM check to ensure the observed split matches the intended ratio.

2. Placebo and Negative Control Tests

Use placebo tests (e.g., A/A tests or pre-experiment periods) to confirm no spurious effects. Apply negative control outcomes (metrics that shouldn't be affected) to detect systematic bias or confounding.

3. Guardrail and Heterogeneity Diagnostics

Monitor guardrail metrics (e.g., latency, crash rates, customer satisfaction) to ensure no harm. Check for heterogeneous treatment effects across key segments (e.g., cities, user types) to avoid masking negative impacts.

4. Define Success Criteria

Pre-specify the primary metric, minimum detectable effect (MDE), statistical significance threshold (e.g., p < 0.05), and practical significance (e.g., lift > 1%). Success requires both statistical and practical significance, with no guardrail violations.

5. Define Rollback Criteria

Set explicit thresholds for guardrail metrics (e.g., >2% increase in cancellations) and for primary metric underperformance (e.g., negative lift with p < 0.05). Rollback if any guardrail is breached or if the primary metric shows significant harm.

Key Points to Mention

  • Sample Ratio Mismatch (SRM) check to validate randomization
  • Pre-trend analysis using pre-experiment data to ensure parallel trends
  • Placebo tests (A/A tests) and negative control outcomes to detect bias
  • Guardrail metrics and their thresholds for rollback decisions
  • Statistical significance (p-value, confidence intervals) and practical significance (effect size, business impact)
  • Pre-registration of success and rollback criteria to avoid p-hacking and post-hoc rationalization

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.