← Uber Interview Insights

Uber·Data Scientist·Technical Phone Screen·Senior

Senior
Jul 2026

Summary

Uber data science interview with a single brutally deep question on experiment design for a two-sided marketplace. The whole thing was essentially one extended case study that kept branching into sub-problems. Dense.

Questions Asked (1)

Q1

Uber is launching a new networked product in a two-sided marketplace with riders and drivers. Design a causal experiment that accounts for interference and network effects. Cover your unit of randomization and why, how you'd limit and measure spillovers, your primary and guardrail metrics, power calculations under clustering, bias-reduction techniques, SUTVA violation diagnostics, and how the design changes if driver supply is effectively unlimited. Walk through a concrete rollout and analysis plan.

A/B Testing & ExperimentationProduct Analytics & MetricsTechnical Trade-offs
Author's notes

This question is basically seven questions duct-taped together.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by framing the two-sided marketplace and the interference problem, then propose a cluster-randomized design (e.g., by city or driver cohort) to contain spillovers. Walk through metrics, power, bias reduction, and SUTVA diagnostics, and finally discuss how the design simplifies if driver supply is unlimited.

Pro tip: Emphasize that in two-sided markets, the unit of randomization should align with the unit of interference; often randomizing by driver or geographic cluster is more practical than by rider. Also, pre-register your analysis plan to avoid p-hacking.

1. Define the experiment and unit of randomization

Clarify the product change and choose a randomization unit that minimizes interference, such as driver cohorts or geographic clusters. Explain why individual rider randomization is problematic due to shared driver supply.

2. Design to limit and measure spillovers

Use techniques like cluster randomization, saturation design, or switchback to contain spillovers. Plan to measure spillovers via differences in outcomes between treated and control within clusters or across boundaries.

3. Select metrics and power calculations

Define primary metrics (e.g., completed rides, rider wait time) and guardrail metrics (e.g., driver utilization, cancellation rate). Adjust power calculations for clustering using intra-cluster correlation (ICC) and design effect.

4. Address bias and SUTVA violations

Use bias-reduction techniques like CUPED or stratification. Diagnose SUTVA violations by checking for interference through network effects, e.g., comparing outcomes of control units near treated units.

5. Adapt design for unlimited driver supply and rollout plan

If driver supply is unlimited, interference reduces, allowing individual-level randomization. Outline a phased rollout: pilot in one city, then expand, with continuous monitoring and analysis using cluster-robust standard errors.

Key Points to Mention

  • Cluster randomization by city or driver cohort to handle interference
  • Use of saturation design or switchback experiments to measure spillovers
  • Primary metrics: rider wait time, completed rides; guardrail metrics: driver earnings, cancellation rate
  • Power calculations adjusted for intra-cluster correlation (ICC) and design effect
  • Bias reduction: CUPED, stratification, or regression adjustment
  • SUTVA diagnostics: compare control units near treated vs. far, or use network exposure models

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.