← Uber Interview Insights

Uber·Data Scientist·Technical Phone Screen·Senior

SeniorPrefer not to say
Jun 2026

Summary

Deep dive into marketplace experimentation for Uber's pricing team. The whole interview was basically one long case study about A/B testing a new pricing model for a ride-hailing platform, with follow-ups stacked on top of each other. Solid question set but it gets pretty granular fast.

Questions Asked (9)

Q1

Before designing the pricing experiment, how would you think through a driver's mental state when they receive a trip request prompt?

Product Sense & IdeationA/B Testing & Experimentation
Author's notes

This was a weird opener and I didn't expect it.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by framing the driver's mental state as a key input to the experiment design, not just a qualitative aside. Walk through the driver's context, motivations, and decision process when receiving a trip request, then connect those insights to specific experimental design choices like randomization unit, metrics, and guardrails.

Pro tip: Show that you understand the difference between a driver's short-term reaction (accept/decline) and long-term engagement, and propose metrics that capture both without confounding the experiment.

1. Contextualize the moment

Describe the driver's immediate situation: where they are, what they were doing, time of day, and how the request interrupts them. Consider factors like current earnings goal, fatigue, and location.

2. Identify motivations and constraints

List what drives the driver's decision: earnings, trip destination, estimated time, surge, acceptance rate, and personal safety. Also consider constraints like gas level, shift end time, and passenger rating.

3. Map the decision process

Outline the cognitive steps: notice prompt, evaluate options, weigh trade-offs, decide to accept or decline, and post-decision feelings. Highlight where friction or uncertainty occurs.

4. Translate to experiment design

Connect mental state insights to design choices: what to randomize (driver vs. trip), what metrics to track (acceptance rate, cancellation, driver satisfaction), and what guardrails to set (e.g., avoid overburdening drivers).

5. Anticipate behavioral biases

Consider biases like loss aversion, present bias, or reactance that could affect responses to pricing changes. Plan how to measure or control for them in the experiment.

Key Points to Mention

  • Driver's primary goal is often maximizing earnings per hour while minimizing downtime and uncertainty.
  • The mental state includes stress from time pressure, cognitive load from evaluating multiple factors, and emotional response to the prompt.
  • Acceptance rate and cancellation rate are key behavioral metrics that reflect mental state and should be primary or guardrail metrics.
  • Experiment design should account for driver heterogeneity (e.g., full-time vs. part-time) and avoid confounding with trip characteristics.
  • Long-term effects on driver retention and satisfaction must be monitored, not just short-term acceptance.
  • Ethical considerations: avoid experiments that exploit driver biases or lead to unsafe driving behavior.

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q2

What primary metrics and guardrail metrics would you define for riders, drivers, and the overall marketplace when testing a new pricing model?

Product Analytics & MetricsA/B Testing & ExperimentationPricing & Monetization
Author's notes

Pretty standard metrics question but the three-sided framing (riders, drivers, marketplace) made it trickier than it sounds.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by framing the pricing model's goal (e.g., increase revenue, improve marketplace efficiency) and then define metrics for each side: riders, drivers, and the marketplace. For each group, specify primary metrics that directly measure success and guardrail metrics that ensure no harm to user experience or long-term health. Emphasize the need to balance trade-offs and monitor both short-term and long-term effects.

Pro tip: Highlight that guardrail metrics should include both quantitative and qualitative measures, and consider the potential for cannibalization or unintended consequences on other parts of the business. Also, mention the importance of segmenting metrics by rider/driver cohorts to detect heterogeneous effects.

1. Clarify the goal of the pricing model

Ask or state the primary objective of the new pricing model (e.g., increase revenue, improve utilization, or balance supply-demand). This will guide metric selection.

2. Define rider metrics

Identify primary metrics like conversion rate, ride frequency, or average spend per rider. Guardrail metrics could include rider cancellation rate, wait time, or satisfaction (CSAT).

3. Define driver metrics

Primary metrics: driver earnings per hour, acceptance rate, or online hours. Guardrails: driver cancellation rate, satisfaction, or churn rate.

4. Define marketplace metrics

Primary: gross bookings, take rate, or match rate. Guardrails: ETA, unfulfilled requests, or overall marketplace efficiency (e.g., utilization).

5. Consider interactions and long-term effects

Discuss how metrics interact (e.g., higher prices may reduce rider demand but increase driver earnings) and include long-term guardrails like retention or lifetime value.

Key Points to Mention

  • Primary vs. guardrail metrics: primary measure success, guardrails prevent negative side effects.
  • Rider metrics: conversion, retention, cancellation, wait time, CSAT.
  • Driver metrics: earnings, acceptance rate, cancellation, churn, satisfaction.
  • Marketplace metrics: gross bookings, match rate, ETA, unfulfilled requests, take rate.
  • Segment by cohorts (e.g., new vs. existing users, high vs. low frequency) to detect heterogeneous effects.
  • Long-term considerations: retention, lifetime value, and potential cannibalization of other products.

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q3

If you run the experiment in a single city with standard treatment and control groups, what interference problems could arise from marketplace network effects?

A/B Testing & ExperimentationTechnical Trade-offs
Author's notes

Classic SUTVA violation setup.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by defining the interference problem in terms of marketplace network effects, then explain how treatment and control groups in the same city can affect each other through rider-driver interactions. Finally, discuss the implications for experiment validity and potential solutions like cluster randomization or switchback testing.

Pro tip: Emphasize that interference violates SUTVA (Stable Unit Treatment Value Assumption) and can bias estimates, so it's crucial to design experiments that account for spillovers, such as using geographic or temporal separation.

1. Define the interference problem

Explain that in a two-sided marketplace like Uber, treatment and control groups are not independent because riders and drivers interact across groups, leading to spillover effects.

2. Identify specific network effects

Describe how changes in treatment group behavior (e.g., pricing, incentives) can affect driver supply and rider demand in the control group, altering their outcomes.

3. Consequences for experiment validity

Discuss how interference can bias treatment effect estimates, increase variance, and lead to false conclusions about the effectiveness of the treatment.

4. Mitigation strategies

Propose solutions such as cluster randomization (e.g., by city or neighborhood), switchback experiments, or using instrumental variables to isolate direct effects.

Key Points to Mention

  • SUTVA violation and its implications
  • Marketplace network effects: rider-driver matching, supply-demand equilibrium
  • Spillover effects: treatment group affecting control group outcomes
  • Bias in treatment effect estimation (e.g., dilution or amplification)
  • Alternative experimental designs: cluster randomization, switchback, geo-based experiments
  • Importance of considering interference in experiment design and analysis

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q4

How would you design a switchback experiment for a single city to avoid those interference issues?

A/B Testing & ExperimentationTechnical Trade-offs
Author's notes

Switchback design: alternate treatment and control across time windows rather than splitting users.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by defining the interference problem in the single-city context, then propose a switchback design where treatment and control alternate over time within the same city. Explain how the design mitigates interference, and discuss key considerations like time block randomization, washout periods, and analysis methods.

Pro tip: Emphasize that switchback experiments require careful handling of carryover effects and time-varying confounders; using a washout period and including time fixed effects in the analysis can strengthen the causal inference.

1. Define the interference problem

Explain why traditional A/B tests fail in a single city due to spillover effects between users or regions, such as network effects or shared resources.

2. Propose switchback design

Describe how to randomly assign the entire city to treatment or control in alternating time blocks (e.g., hours, days) to ensure all users experience both conditions.

3. Address carryover effects

Include washout periods between switches to allow the system to return to baseline, and consider using buffer periods or excluding transition data.

4. Randomization and analysis

Randomize the order of treatment and control blocks, and analyze using methods like difference-in-differences or fixed effects models to account for time trends.

5. Validate and monitor

Check for balance in covariates across time blocks, monitor for unexpected interference, and run power analysis to determine block length and number of switches.

Key Points to Mention

  • Interference in single-city experiments (e.g., network effects, shared supply/demand)
  • Time-based randomization instead of user-based
  • Washout periods to mitigate carryover effects
  • Analysis techniques: fixed effects, difference-in-differences, time-series models
  • Power considerations: block length, number of switches, and autocorrelation
  • Practical challenges: operational constraints, user experience, and implementation complexity

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q5

How would you handle limited sample size in a switchback test?

A/B Testing & Experimentation
Author's notes

Short answer: fewer windows means less power.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by acknowledging that switchback tests are used when individual randomization isn't feasible, but limited sample size reduces power. Then discuss strategies to increase effective sample size, such as using within-unit comparisons, leveraging pre-period data, and applying variance reduction techniques like CUPED. Finally, emphasize the importance of sensitivity analysis and clear communication of limitations.

Pro tip: At Uber, where switchback tests are common for marketplace experiments, mention that you can increase power by using shorter switchback intervals if carryover effects are minimal, and by modeling the time series structure to account for autocorrelation.

1. Clarify the constraints and goals

Understand why sample size is limited (e.g., few units, short duration) and what decision the test needs to inform. This shapes acceptable trade-offs between power and bias.

2. Maximize effective sample size

Use within-unit comparisons (each unit serves as its own control), increase switch frequency if carryover is negligible, and extend duration if possible. Consider pooling data across similar units or time periods.

3. Apply variance reduction techniques

Use CUPED, stratification, or regression adjustment with pre-experiment covariates to reduce noise. Model time trends and autocorrelation to improve precision.

4. Assess and communicate uncertainty

Conduct power analysis to set realistic expectations, use confidence intervals and Bayesian methods to quantify uncertainty, and perform sensitivity checks for carryover and time effects.

5. Consider alternative designs or decision rules

If power remains low, propose sequential testing, switch to a different design (e.g., cluster randomization), or define a decision rule that accounts for limited evidence (e.g., require larger effect sizes).

Key Points to Mention

  • Within-unit comparisons and paired analysis to control for unit-level confounding
  • CUPED or other variance reduction using pre-experiment data
  • Modeling time trends and autocorrelation in switchback designs
  • Trade-offs between switch frequency and carryover effects
  • Power analysis and minimum detectable effect (MDE) calculations
  • Bayesian methods or sequential testing for small samples

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q6

What are the tradeoffs of using full-day switchback windows versus shorter windows?

A/B Testing & ExperimentationTechnical Trade-offs
Author's notes

Longer windows reduce carryover noise between periods but also mean fewer switches, so less statistical power for a given calendar duration.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by defining what a switchback window is and why it's used in ride-sharing experiments. Then compare full-day vs. shorter windows across dimensions like bias, variance, operational feasibility, and sensitivity to time-of-day effects. Conclude with a recommendation based on the specific context and tradeoffs.

Pro tip: Mention that the choice depends on the expected effect size and the autocorrelation structure of the metric; shorter windows reduce bias from time trends but increase variance, so you need to balance statistical power with validity.

1. Define switchback windows and their purpose

Explain that switchback experiments alternate treatment and control over time windows to handle interference and time-varying confounders in marketplaces like Uber.

2. Discuss full-day windows

Highlight advantages: captures full daily cycles, reduces variance by aggregating more data, and simplifies operations. Disadvantages: fewer switch points, potential confounding with day-of-week effects, and lower power if the experiment is short.

3. Discuss shorter windows

Highlight advantages: more switch points increase effective sample size, better control for time trends, and ability to detect transient effects. Disadvantages: may not capture full daily patterns, higher risk of carryover effects, and operational complexity.

4. Compare on key dimensions

Evaluate tradeoffs in terms of bias, variance, statistical power, operational feasibility, and sensitivity to time-of-day effects. Mention that shorter windows can reduce bias but increase variance, while full-day windows do the opposite.

5. Recommend based on context

Conclude that the optimal window depends on the metric, expected effect size, and duration of the experiment. Suggest using simulations or pilot data to determine the best window length.

Key Points to Mention

  • Interference and carryover effects in marketplace experiments
  • Bias-variance tradeoff: shorter windows reduce time-trend bias but increase variance
  • Statistical power and sample size considerations
  • Time-of-day and day-of-week effects (e.g., rush hours, weekends)
  • Operational complexity of frequent switches
  • Recommendation to use simulations or historical data to choose window length

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q7

What are the tradeoffs between using a larger versus smaller treatment group in this kind of experiment?

A/B Testing & ExperimentationTechnical Trade-offs
Author's notes

Bigger treatment group gives more power but increases exposure risk if the new pricing model is bad.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by framing the tradeoff as a balance between statistical power and practical constraints, then discuss how larger groups improve sensitivity to detect small effects but increase cost, risk, and potential user experience impact. Conclude by emphasizing that the optimal size depends on the minimum detectable effect, business context, and cost-benefit analysis.

Pro tip: Mention that at companies like Uber, even small effect sizes can translate to millions in revenue, so the tradeoff often hinges on the cost of false negatives versus false positives. Also, highlight that larger treatment groups can introduce novelty effects or dilute the user experience if the treatment is risky.

1. Define the goal and constraints

Clarify the experiment's objective, the minimum detectable effect (MDE) that matters for business impact, and any practical constraints like budget, time, and user availability.

2. Explain statistical power and sensitivity

Discuss how larger treatment groups increase statistical power, reduce variance, and enable detection of smaller effects, while smaller groups may miss meaningful differences.

3. Discuss cost and risk considerations

Cover how larger groups increase costs (e.g., engineering, opportunity cost) and risks (e.g., negative user experience, revenue loss if treatment is bad), while smaller groups limit exposure but may prolong experimentation.

4. Address practical tradeoffs and business context

Explain that the optimal size depends on the expected effect size, the cost of errors (false positive vs. false negative), and the stage of product development (e.g., early exploration vs. scaling).

5. Conclude with a balanced recommendation

Summarize that there's no one-size-fits-all answer; recommend using power analysis to determine the minimum sample size needed, and consider sequential testing or adaptive designs to optimize.

Key Points to Mention

  • Statistical power and minimum detectable effect (MDE)
  • Cost implications: engineering, opportunity cost, and user experience
  • Risk of negative impact on users and revenue
  • Duration of experiment and time to decision
  • Novelty effects and long-term vs. short-term tradeoffs
  • Business context: Uber's scale and sensitivity to small effect sizes

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q8

What are the tradeoffs between longer and shorter treatment windows in a switchback design?

A/B Testing & ExperimentationTechnical Trade-offs
Author's notes

Shorter windows = more switches = more power, but also more carryover contamination.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by defining switchback designs and why they are used when unit-level randomization is infeasible. Then systematically compare longer and shorter treatment windows across statistical, operational, and practical dimensions, highlighting the bias-variance tradeoff and context-dependent factors. Conclude with a recommendation framework based on the specific goals and constraints of the experiment.

Pro tip: Emphasize that the optimal window length depends on the carryover effect's decay rate and the outcome's time sensitivity—mentioning this shows you understand the underlying dynamics rather than just listing pros and cons.

1. Define switchback design and its purpose

Explain that switchback designs alternate treatments over time for the same units, often used when individual randomization is impossible (e.g., pricing, driver incentives). This sets the context for discussing window length tradeoffs.

2. Discuss statistical power and variance

Longer windows increase the number of observations per treatment period, reducing variance and increasing power, but may introduce time-varying confounders. Shorter windows allow more treatment switches, increasing effective sample size but may suffer from carryover effects and higher variance due to fewer observations per period.

3. Address carryover and interference effects

Shorter windows risk contamination from previous treatments (carryover effects) if the effect persists, biasing estimates. Longer windows mitigate carryover but may miss short-term effects and reduce the number of switches, potentially confounding with temporal trends.

4. Consider operational and practical constraints

Longer windows require stable environments and may be impractical if treatments need frequent changes or if user behavior changes rapidly. Shorter windows demand more frequent switching, which can be operationally complex and may cause user confusion or system instability.

5. Synthesize tradeoffs and recommend based on context

Summarize that the choice depends on the magnitude and decay of carryover effects, the desired power, the stability of the environment, and operational feasibility. Recommend a data-driven approach, such as piloting different window lengths or using washout periods.

Key Points to Mention

  • Bias-variance tradeoff: longer windows reduce variance but increase bias from time trends; shorter windows reduce bias but increase variance.
  • Carryover effects: shorter windows are more susceptible if treatment effects persist; washout periods can help.
  • Statistical power: longer windows provide more data per period, increasing power, but fewer independent periods may reduce effective sample size.
  • Temporal confounders: longer windows may confound treatment with time-varying factors (e.g., seasonality, trends).
  • Operational complexity: shorter windows require more frequent switches, which can be costly and disruptive.
  • Context-dependence: optimal window length depends on the specific experiment, outcome, and environment; no one-size-fits-all answer.

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q9

If the experiment is later expanded to 12 cities, how would you redesign the experiment to increase sample size and improve validity?

A/B Testing & ExperimentationProduct Analytics & MetricsTechnical Trade-offs
Author's notes

This is where it got interesting.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by acknowledging the goal: increasing sample size and improving validity when expanding from a smaller number of cities to 12. Then outline a redesign that leverages the additional cities for greater statistical power, while addressing potential heterogeneity and interference. Emphasize the importance of maintaining randomization, controlling for city-level differences, and considering practical constraints like cost and operational feasibility.

Pro tip: Mention that with more cities, you can use a stratified or matched-pair design to balance key city characteristics, and consider running the experiment longer to capture more data points. Also, highlight the trade-off between sample size and validity: more cities can increase external validity but may introduce more noise, so careful design is crucial.

1. Clarify objectives and constraints

Restate the goal of increasing sample size and improving validity, and identify any constraints such as budget, time, or operational limitations. This ensures the redesign is practical and aligned with business needs.

2. Choose an appropriate experimental design

Decide between a completely randomized design across cities, a stratified design, or a matched-pair design based on city characteristics. Consider using switchback or geo-based randomization if interference is a concern.

3. Determine sample size and power

Calculate the required sample size per city or overall to detect the desired effect size with sufficient power. Account for intra-city correlation and adjust the design accordingly, possibly increasing the number of users or the duration.

4. Address validity threats

Identify and mitigate threats to internal and external validity, such as selection bias, confounding variables, and spillover effects. Use techniques like stratification, matching, or covariate adjustment to control for city-level differences.

5. Plan for analysis and monitoring

Pre-register the analysis plan, including how to handle multiple comparisons and heterogeneity. Set up monitoring to detect issues early and ensure data quality across all cities.

Key Points to Mention

  • Stratified randomization by city characteristics (e.g., market size, demographics) to improve balance and validity.
  • Matched-pair design: pair similar cities and randomize treatment within pairs to reduce confounding.
  • Sample size calculation: account for intra-class correlation (ICC) and design effect when clustering by city.
  • Consider running the experiment longer or increasing user allocation per city to boost sample size.
  • Use of switchback or geo-based randomization to mitigate interference and spillover effects.
  • Pre-registration of analysis plan and correction for multiple comparisons to maintain validity.

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.