← Uber Interview Insights

Uber·Research Scientist·Onsite - Multi Round·Senior

SeniorPrefer not to say
Jul 2026

Summary

Onsite scientist round at Uber, no coding, just a big open-ended experiment design prompt. The whole session was basically one long reasoning exercise about marketplace experimentation, and the bar for metric definition speed was higher than I expected.

Questions Asked (5)

Q1

Pick a marketplace metric, define it precisely, and design a full experiment to measure its causal effect, including your choice of experimentation framework, treatment structure, rollout phases, guardrails, and how you'd interpret the confidence intervals.

A/B Testing & ExperimentationProduct Analytics & Metrics
Author's notes

This is the whole round basically.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Choose a metric like 'completed trips per active rider' and define it with precise numerator and denominator, time window, and exclusions. Then design a randomized controlled experiment using switchback or cluster randomization to handle interference, with phased rollout, guardrail metrics, and clear interpretation of confidence intervals for practical significance.

Pro tip: Acknowledge that marketplace metrics often suffer from interference and network effects, so standard A/B tests may be biased; propose a switchback or cluster-randomized design and discuss how you'd validate assumptions.

1. Define the metric precisely

Specify the metric's numerator, denominator, time window, and any exclusions (e.g., completed trips per active rider per day, excluding canceled trips and new users).

2. Choose experimentation framework

Select a design that accounts for interference, such as switchback (time-based randomization) or cluster randomization (by city or driver), and justify why it's appropriate for the marketplace metric.

3. Design treatment structure and rollout phases

Define treatment and control conditions, then plan a phased rollout (e.g., pilot in one city, then expand) to monitor for early signals and operational risks.

4. Set guardrail metrics and monitoring

Identify guardrail metrics (e.g., driver earnings, rider wait time, cancellation rate) and establish thresholds for acceptable degradation to ensure the experiment doesn't harm the ecosystem.

5. Interpret confidence intervals and make decisions

Analyze the confidence intervals for the treatment effect, considering both statistical and practical significance, and decide whether to launch, iterate, or stop based on the interval's width and position relative to the minimum detectable effect.

Key Points to Mention

  • Precise metric definition with numerator, denominator, and time window
  • Interference and network effects in marketplaces, requiring switchback or cluster randomization
  • Phased rollout to mitigate risk and allow for early stopping
  • Guardrail metrics to protect user experience and ecosystem health
  • Interpretation of confidence intervals: statistical significance vs. practical significance, and width indicating precision
  • Power analysis and minimum detectable effect to ensure adequate sample size

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q2

Why not just run a standard A/B test here instead of a switchback design?

A/B Testing & ExperimentationTechnical Trade-offs
Author's notes

Follow-up after I committed to switchback.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Acknowledge that standard A/B tests are the default but explain that switchback designs are necessary when there are interference or spillover effects between units. Focus on the specific context (e.g., Uber's marketplace) where such effects are present, and discuss the trade-offs between the two designs.

Pro tip: Mention that switchback designs can also help with novelty effects and reduce the required sample size by using within-unit comparisons over time, but be careful about time-varying confounders.

1. Define the standard A/B test

Briefly describe what a standard A/B test entails: randomizing units (e.g., users, drivers) into control and treatment groups, assuming no interference between units.

2. Identify interference/spillover

Explain that in Uber's marketplace, units interact (e.g., drivers and riders), so treating one unit can affect others, violating the no-interference assumption (SUTVA).

3. Introduce switchback design

Describe switchback: randomizing treatment over time periods for the entire system (or geographic area), so all units receive both treatments at different times, mitigating interference.

4. Compare trade-offs

Discuss pros and cons: switchback reduces interference bias but may introduce time trends, carryover effects, and requires careful analysis; standard A/B is simpler but biased under interference.

5. Conclude with when to use each

Summarize that switchback is preferred when interference is strong (e.g., pricing, dispatch algorithms), while standard A/B is fine when interference is negligible (e.g., UI changes).

Key Points to Mention

  • SUTVA (Stable Unit Treatment Value Assumption) and its violation in marketplaces
  • Interference/spillover effects between drivers and riders
  • Time-based randomization in switchback designs
  • Potential biases in switchback: time trends, carryover effects, seasonality
  • Trade-off between bias and variance: switchback may have higher variance but lower bias
  • Examples from Uber: surge pricing, dispatch, matching algorithms

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q3

You see an effect at the 90% confidence level but not at 95%. How do you think about that result?

A/B Testing & ExperimentationProduct Analytics & Metrics
Author's notes

Tricky one.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Acknowledge that the result is suggestive but not conclusive, and emphasize the importance of considering practical significance, statistical power, and business context. Discuss how you would communicate the uncertainty and decide on next steps, such as running a longer experiment or analyzing segments.

Pro tip: Avoid fixating on the p-value threshold; instead, focus on the effect size and confidence interval to assess practical impact. Frame the result as an opportunity to gather more data or refine the experiment rather than a binary win/lose.

1. Interpret the result

Explain that a 90% confidence level means there is a 10% chance of a false positive, so the effect is not statistically significant at the conventional 95% level. Clarify that this does not mean the effect is absent, but that the evidence is weaker.

2. Assess practical significance

Look at the effect size and confidence interval to determine if the observed effect, if real, would be meaningful for the business. Consider the cost of a false positive versus the potential gain.

3. Consider power and sample size

Evaluate whether the experiment was adequately powered to detect the effect at 95% confidence. If not, the result may be due to insufficient sample size, and extending the experiment could provide clarity.

4. Check for validity threats

Rule out issues like multiple testing, peeking, or segment-specific effects that could explain the result. Ensure the experiment was designed and analyzed correctly.

5. Decide on next steps

Recommend actions such as running the experiment longer, conducting a follow-up study, or making a decision based on risk tolerance and business priorities. Communicate the uncertainty clearly to stakeholders.

Key Points to Mention

  • Difference between statistical significance and practical significance
  • Confidence intervals and effect size estimation
  • Statistical power and sample size considerations
  • Multiple testing and peeking pitfalls
  • Business context and risk tolerance in decision-making
  • Communication of uncertainty to stakeholders

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q4

What guardrail metric would you choose for this experiment, and why?

A/B Testing & ExperimentationProduct Analytics & Metrics
Author's notes

Pre-picked driver-side income as my guardrail.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by clarifying the experiment's primary goal and the potential negative side effects it could cause. Then propose a guardrail metric that directly measures a critical downside risk, explaining why it's the most important to monitor and how you would set thresholds for action.

Pro tip: Choose a guardrail metric that is sensitive to the change and tied to long-term user value, not just a vanity metric. Also, mention that you would monitor multiple guardrails but prioritize one as the primary guardrail based on the biggest risk.

1. Clarify the experiment's goal and potential risks

Ask or infer what the experiment is trying to improve and what negative side effects it might inadvertently cause. This sets the context for choosing a relevant guardrail.

2. Identify candidate guardrail metrics

Brainstorm metrics that capture potential harms, such as user retention, satisfaction, latency, or revenue. Consider both short-term and long-term impacts.

3. Select the most critical guardrail

Choose one metric that best represents the biggest risk to the business or user experience. Explain why it's more important than others in this context.

4. Define thresholds and monitoring plan

Specify what level of degradation would be unacceptable and how you would monitor it (e.g., statistical significance, practical significance). Mention what action you would take if the guardrail is breached.

5. Validate and iterate

Discuss how you would validate the guardrail's sensitivity and consider if additional guardrails are needed as the experiment evolves.

Key Points to Mention

  • Alignment with business objectives and user experience
  • Sensitivity to the experimental change
  • Leading vs. lagging indicators
  • Statistical power and minimum detectable effect
  • Trade-offs between primary and guardrail metrics
  • Examples of guardrail metrics (e.g., retention, churn, NPS, latency, revenue)

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q5

How would you choose the window length for a switchback experiment, and what are the tradeoffs?

A/B Testing & ExperimentationTechnical Trade-offs
Author's notes

Knew this one cold.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by defining the tradeoff between statistical power and operational constraints: shorter windows increase sample size but may introduce interference and carryover effects, while longer windows reduce interference but increase variance and reduce sensitivity. Then propose a data-driven approach that balances these factors, such as using pilot data to estimate autocorrelation and interference decay, and selecting the window length that minimizes mean squared error of the treatment effect estimate.

Pro tip: Mention that in practice, you often run a pilot with varying window lengths to empirically measure carryover and interference, and then choose the shortest window that keeps bias below a pre-specified threshold. This shows you value both rigor and pragmatism.

1. Clarify the goal and constraints

Identify the primary metric, the desired minimum detectable effect, and operational constraints like switching costs and user experience. This sets the stage for balancing statistical and practical considerations.

2. Assess interference and carryover

Determine how long treatment effects persist and how much interference occurs between units. Use domain knowledge or pilot data to estimate the decay rate of these effects.

3. Quantify bias-variance tradeoff

Model how window length affects bias (from interference/carryover) and variance (from fewer independent units). Aim to minimize mean squared error of the treatment effect estimate.

4. Validate with pilot or simulation

Run a pilot experiment with multiple window lengths or simulate data to empirically choose the optimal window. Check that the chosen length keeps bias within acceptable limits.

5. Monitor and adjust

After deployment, monitor for unexpected interference or changes in user behavior, and be prepared to adjust the window length if needed.

Key Points to Mention

  • Interference between units (e.g., network effects, shared resources) and how it biases treatment effect estimates.
  • Carryover effects: treatment effects that persist into subsequent periods, requiring washout periods.
  • Statistical power and variance: shorter windows yield more independent observations but may increase variance due to noise.
  • Bias-variance tradeoff: choosing window length to minimize mean squared error of the estimated treatment effect.
  • Operational constraints: switching costs, user experience, and engineering feasibility.
  • Pilot experiments or simulations to empirically determine the optimal window length.

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.