I started rattling off conversion rate and GMV like a reflex, which was the wrong move.
Start by framing the program's goal: to increase rider engagement and marketplace liquidity without eroding profitability. Then define metrics across rider behavior, marketplace efficiency, and unit economics, providing clear formulas for each. Finally, emphasize the importance of A/B testing to measure incremental impact and guard against cannibalization.
Pro tip: Always distinguish between correlation and causation by highlighting the need for a control group; propensity models can create selection bias, so measure incremental lift, not just raw differences.
Clarify the primary goal (e.g., increase ride frequency, improve retention) and secondary goals (e.g., maintain driver utilization, protect margins). This ensures metrics align with business strategy.
Choose metrics that capture changes in rider behavior, such as redemption rate, incremental rides per redeemed offer, and retention lift. Provide formulas like Redemption Rate = (# Offers Redeemed) / (# Offers Sent).
Include metrics that reflect overall marketplace efficiency, such as driver utilization, ETA, and match rate. For example, Driver Utilization = (Total Driver Hours on Trip) / (Total Driver Hours Online).
Define profitability metrics like incremental gross bookings, cost per incremental ride, and ROI. For instance, ROI = (Incremental Gross Profit - Program Cost) / Program Cost.
Explain how to measure incrementality via A/B testing, including holdout groups and statistical significance. Mention metrics like lift and p-value to validate impact.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.
This was the part I found most interesting and also most stressful.
Structure your answer by first clarifying the causal estimand (e.g., ATE or LATE) and the assumptions required for each identification strategy. Then, for each scenario, explain how you would implement the method, check assumptions, and interpret the estimate in the context of Uber's incentive program. Emphasize the trade-offs between internal validity and generalizability across the three designs.
Pro tip: Show awareness that geo holdouts often suffer from interference and spillover effects, so you'd complement them with switchback or cluster randomization if possible. Also, mention that IV estimates a LATE, which may not generalize to all drivers, and that RD requires a sharp threshold and no manipulation of the running variable.
Clarify what causal effect you want to estimate (e.g., effect of incentive on driver hours or completions) and for which population (e.g., all drivers or those near the threshold). Specify whether you're targeting ATE, ATT, or LATE.
Explain that randomizing at the geo level allows estimating the ATE if geos are independent and no spillover. Discuss power, balance checks, and potential interference (e.g., drivers crossing geos). Suggest using cluster-robust standard errors and possibly a difference-in-differences if pre-period data available.
Describe using offer score thresholds to compare drivers just above and below the cutoff. Check for manipulation of the running variable (e.g., McCrary test), continuity of covariates, and bandwidth selection. Estimate LATE at the threshold, and discuss external validity.
Propose operational friction (e.g., app glitches, payment delays) as an instrument for treatment take-up. Argue relevance and exclusion restriction: friction affects participation but not outcomes except through the program. Estimate LATE for compliers, and discuss potential violations (e.g., friction correlated with driver quality).
Discuss how the three estimates might differ due to different complier populations and assumptions. Recommend triangulation and sensitivity analyses (e.g., placebo tests, alternative specifications) to strengthen causal claims.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.
Honestly the externalities piece is where I added the most value.
Start by defining the ROI formula as incremental profit divided by total investment, explicitly subtracting cannibalization, subsidy burn, and externality costs from gross incremental revenue. Then explain how to estimate LTV lift using holdout groups and cohort analysis, amortize acquisition cost over expected lifetime, and quantify uncertainty via bootstrapping or Bayesian methods to produce confidence intervals.
Pro tip: Emphasize that cannibalization and externalities are often the largest hidden costs; propose using a geo-based holdout or switchback experiment to isolate true incrementality, and always report ROI as a distribution (e.g., 90% credible interval) rather than a point estimate to reflect decision risk.
Write ROI = (Incremental Gross Profit - Cannibalization Cost - Subsidy Burn - Surge/ETA Externality Cost + Long-term Habit Value) / Total Investment. Clearly label each component and explain how it is measured.
Use a randomized controlled experiment (e.g., switchback or geo holdout) to measure the treatment effect on orders and revenue. Cannibalization is the negative spillover to non-incentivized products or segments, estimated via difference-in-differences or causal inference methods.
Subsidy burn is the direct cost of incentives (e.g., discounts, credits). Surge and ETA externalities are indirect costs: increased surge pricing or longer wait times for non-incentivized users, which can be monetized using elasticity models or observed changes in rider/driver behavior.
Use cohort analysis and survival models to estimate how the incentive changes retention and order frequency over time. LTV lift = (New LTV - Baseline LTV) * Number of users. Amortize acquisition cost by spreading it over the expected customer lifetime (e.g., using discounted cash flow).
Propagate uncertainty from each component (e.g., via Monte Carlo simulation or bootstrapping) to generate a distribution of ROI. Report the mean and 90% or 95% confidence interval, and conduct sensitivity analysis on key assumptions.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.
Start by clearly defining the direct and spillover estimands in the context of a two-stage randomization: first randomize cities to treatment/control, then within treated cities randomize riders to treatment/control. Explain how the second stage identifies direct effects (treated riders in treated cities vs. control riders in treated cities) and how the first stage identifies spillover effects (control riders in treated cities vs. control riders in control cities). Emphasize the need for careful design to avoid contamination and ensure sufficient power at both levels.
Pro tip: Highlight the importance of pre-registering your analysis plan and using cluster-robust standard errors to account for correlation within cities and weeks. Also, discuss how to handle potential interference between riders in the same city by measuring spillover through network or geographic proximity.
Specify that cities are randomly assigned to either treatment or control at the first stage, and then within each treated city, riders are randomly assigned to treatment or control at the second stage. This creates four groups: treated riders in treated cities, control riders in treated cities, treated riders in control cities (if any), and control riders in control cities.
Direct effect: difference in outcomes between treated and control riders within treated cities. Spillover effect: difference in outcomes between control riders in treated cities and control riders in control cities. Total effect: difference between treated riders in treated cities and control riders in control cities.
State assumptions such as no interference between cities (SUTVA at city level) and no spillover from control cities. Address potential biases like selection bias if rider randomization is not truly random, and confounding if city-level characteristics differ.
Use regression models with fixed effects for city and week, and cluster-robust standard errors at the city level. For direct effect, compare treated vs. control riders within treated cities; for spillover, compare control riders in treated vs. control cities. Consider interaction terms to test heterogeneity.
Calculate power for both stages: number of cities needed for spillover detection and number of riders per city for direct effect. Discuss trade-offs between number of cities and riders per city given budget constraints.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.
Went with hierarchical testing and explained why flat Bonferroni is too conservative when you have 20 diagnostics that are mostly sanity checks.
Start by clarifying the testing hierarchy: one primary metric for the decision, three secondary metrics for directional support, and twenty diagnostics for health checks. Then propose a tiered correction strategy: strict control for the primary (e.g., alpha=0.05), moderate correction for secondary (e.g., Holm-Bonferroni or Benjamini-Hochberg), and minimal or no correction for diagnostics, focusing instead on effect sizes and practical significance.
Pro tip: Emphasize that diagnostics are not for hypothesis testing but for detecting guardrail violations; use false discovery rate (FDR) control for secondary metrics to balance power and false positives, and always pre-register the correction plan to avoid p-hacking concerns.
Confirm that the primary metric drives the launch decision, secondary metrics provide supporting evidence, and diagnostics are for sanity checks. Establish that the primary metric is tested at a strict alpha (e.g., 0.05) without correction.
Use a method like Holm-Bonferroni or Bonferroni to control the probability of any false positive among the three secondary metrics, since they are a small family and false positives are costly.
For the twenty diagnostic checks, apply Benjamini-Hochberg (FDR) if you must flag anomalies, or skip formal correction and rely on effect sizes and domain thresholds to identify issues.
If the experiment is monitored continuously, use group sequential methods or alpha spending to control error rates over time, especially for the primary metric.
State the correction strategy before analysis to maintain statistical rigor and avoid criticism of p-hacking. Include sensitivity analyses to show robustness.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.
Start by outlining a pre-registration plan that specifies hypotheses, subgroups (city tier, weather conditions), primary and secondary metrics, and correction methods for multiple comparisons. Then describe the execution: pre-register the analysis plan, collect data, run heterogeneity tests (e.g., interaction effects), and apply appropriate statistical controls to mitigate sign and magnitude errors. Emphasize the importance of pre-registration in preventing p-hacking and ensuring valid inference.
Pro tip: Pre-register not just the subgroups but also the minimum detectable effect (MDE) for each subgroup to avoid post-hoc power issues. Use a hierarchical model or Bayesian approach to borrow strength across subgroups, reducing the risk of false positives from multiple comparisons.
Clearly state the primary hypothesis and the expected moderation effects by city tier (e.g., Tier 1 vs. Tier 2) and weather conditions (e.g., rain vs. clear). Specify the direction and magnitude of expected effects.
Document the analysis plan including metrics, subgroup definitions, statistical tests (e.g., interaction terms in regression), correction methods (e.g., Bonferroni, FDR), and decision rules. Register on a public platform like AsPredicted or OSF.
Run the experiment ensuring proper randomization and data collection. Blind analysts to treatment assignment where possible to reduce bias.
Use regression models with interaction terms to test moderation. Apply multiple comparison corrections and check for sign (direction) and magnitude errors by comparing observed effects to pre-registered expectations.
Report subgroup effects with confidence intervals, noting any deviations from pre-registered hypotheses. Discuss limitations and potential for false positives/negatives.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.
Start by outlining a layered diagnostic plan that validates the experiment's assumptions before interpreting results. Then, define clear, pre-registered success and rollback criteria that balance statistical significance with practical impact and guardrail metrics. Emphasize that these criteria should be set before the experiment starts to avoid post-hoc rationalization.
Pro tip: At Uber, where experiments run at massive scale, even tiny effects can be statistically significant but practically meaningless. Always pair p-values with confidence intervals and effect sizes, and tie rollback decisions to guardrail metrics like cancellation rates or driver acceptance rates, not just the primary metric.
Verify that the treatment and control groups were balanced before the experiment and that the randomization didn't introduce bias. Run an SRM check to ensure the observed split matches the intended ratio.
Use placebo tests (e.g., A/A tests or pre-experiment periods) to confirm no spurious effects. Apply negative control outcomes (metrics that shouldn't be affected) to detect systematic bias or confounding.
Monitor guardrail metrics (e.g., latency, crash rates, customer satisfaction) to ensure no harm. Check for heterogeneous treatment effects across key segments (e.g., cities, user types) to avoid masking negative impacts.
Pre-specify the primary metric, minimum detectable effect (MDE), statistical significance threshold (e.g., p < 0.05), and practical significance (e.g., lift > 1%). Success requires both statistical and practical significance, with no guardrail violations.
Set explicit thresholds for guardrail metrics (e.g., >2% increase in cancellations) and for primary metric underperformance (e.g., negative lift with p < 0.05). Rollback if any guardrail is breached or if the primary metric shows significant harm.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.