← Upstart Interview Insights

Upstart·Data Scientist·Technical Phone Screen·Senior

Senior
May 2026

Summary

Data science interview at Upstart centered on a pretty meaty A/B testing scenario involving two treatments and conflicting metric signals. The whole thing felt like one long case question with follow-ups that kept getting harder.

Questions Asked (4)

Q1

Treatment t1 shows no significant change in either Gross Booking or Variable Consideration, while t2 shows a significant GB increase but a significant VC decrease. How do you explain these results to a PM and what do you recommend doing next?

A/B Testing & ExperimentationStakeholder ManagementProduct Analytics & Metrics
Author's notes

I spent too long on the mechanics of the stats and not enough on the story for the PM.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

First, acknowledge the mixed results and emphasize the need to understand the trade-off between GB and VC. Then, propose a structured approach to diagnose the underlying reasons and recommend next steps based on business impact and statistical validity.

Pro tip: Frame the recommendation around the company's north star metric and unit economics, showing you understand the business beyond just statistical significance.

1. Clarify Metrics and Business Context

Define Gross Booking (GB) and Variable Consideration (VC) and their relationship to key business outcomes. Confirm with the PM whether the goal is to maximize GB, VC, or a combination (e.g., total revenue or profit).

2. Analyze the Trade-off

Quantify the trade-off: calculate the net effect on a composite metric (e.g., GB * VC margin) and assess if the GB increase outweighs the VC decrease. Consider segment-level analysis to see if effects vary by user groups.

3. Investigate Root Causes

Explore potential reasons for the VC decrease: e.g., t2 may attract lower-quality borrowers, increase adverse selection, or change pricing dynamics. Check for interactions with other variables and ensure the experiment was properly randomized.

4. Evaluate Statistical and Practical Significance

Confirm that the observed changes are statistically significant and not due to chance. Also assess practical significance: is the magnitude of change meaningful for the business?

5. Recommend Next Steps

Based on the analysis, recommend actions: e.g., iterate on t2 to mitigate VC decrease, run further experiments, or abandon t2 if the trade-off is unfavorable. Suggest monitoring long-term effects and considering guardrail metrics.

Key Points to Mention

  • Define and align on success metrics with the PM, considering both GB and VC.
  • Quantify the trade-off and compute net impact on a composite metric like revenue or profit.
  • Check for statistical significance and practical significance of the results.
  • Investigate potential root causes for the VC decrease, such as adverse selection or pricing changes.
  • Consider segment-level analysis to identify heterogeneous treatment effects.
  • Recommend iterative testing or further analysis before making a final decision.

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q2

Given that t2 has a GB confidence interval of [+0.1%, +2.3%] with a +$0.48 lift and a VC confidence interval of [-2.5%, -1.5%] with a -$0.20 loss, would you launch it? Walk through your reasoning.

A/B Testing & ExperimentationTechnical Trade-offsPricing & Monetization
Author's notes

Net lift is positive so on pure math you'd say launch, but I got tripped up on whether VC is a margin metric or a revenue metric and how to weight the two.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

First, clarify what t2 and VC represent (e.g., treatment variants or metrics) and the business context. Then, evaluate the confidence intervals and lifts to assess statistical significance and practical impact, considering trade-offs between GB and VC. Finally, weigh the risks and potential upside to make a recommendation, possibly suggesting further analysis or a phased launch.

Pro tip: Demonstrate that you consider both statistical significance and business impact, and that you're comfortable making decisions under uncertainty. Mention the importance of aligning with business goals and risk tolerance.

1. Clarify the context and metrics

Identify what t2 and VC represent (e.g., treatment groups, metrics like Gross Bookings and Variable Contribution) and the business objectives. Understand the units and the significance of the confidence intervals.

2. Assess statistical significance and effect sizes

For t2, the GB confidence interval is entirely positive, indicating a statistically significant positive effect on GB. For VC, the confidence interval is entirely negative, indicating a statistically significant negative effect on VC. Quantify the expected lifts: +$0.48 for GB and -$0.20 for VC.

3. Evaluate business impact and trade-offs

Consider the net impact: Does the increase in GB outweigh the decrease in VC? Calculate the net monetary impact if possible, and consider strategic importance of GB vs. VC. Also, consider the confidence intervals: the ranges suggest uncertainty, but both are significant.

4. Consider risks and alternatives

Think about potential risks: Is the VC loss acceptable? Could it be mitigated? Are there long-term effects? Consider if a phased launch or further testing could provide more information.

5. Make a recommendation

Based on the analysis, decide whether to launch, not launch, or launch with modifications. Justify your decision with data and business reasoning.

Key Points to Mention

  • Statistical significance: both intervals exclude zero, so effects are significant.
  • Effect sizes: +$0.48 GB lift vs. -$0.20 VC loss; net impact depends on weighting.
  • Business context: importance of GB vs. VC for Upstart's model.
  • Risk tolerance: willingness to accept VC loss for GB gain.
  • Potential for further testing or phased rollout to mitigate risk.
  • Alignment with overall company strategy and goals.

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q3

How would you segment users or orders to find cohorts where t2 drives positive GB impact without the negative VC impact?

A/B Testing & ExperimentationProduct Analytics & MetricsRoot Cause Analysis
Author's notes

Cohort analysis question, fairly standard framing.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by clarifying the business context: define t2, GB (likely gross bookings or gross balance), and VC (likely variable cost or variable contribution). Then propose a segmentation strategy that first identifies cohorts where t2 has a positive GB impact, and within those, further segments to isolate where the VC impact is neutral or positive. Use a combination of exploratory analysis, causal inference (e.g., A/B test or observational methods), and statistical testing to validate the segments.

Pro tip: Focus on actionable segments: avoid over-segmenting to the point of noise. Use a holdout or validation set to confirm that the positive GB impact without negative VC impact holds out-of-sample, and consider the cost of targeting each segment.

1. Clarify definitions and objectives

Confirm what t2, GB, and VC stand for, and the exact business goal (e.g., maximize GB while keeping VC non-negative). Ensure alignment on the time horizon and any constraints.

2. Identify relevant dimensions for segmentation

List potential segmentation variables such as user demographics, order characteristics, t2 exposure levels, and temporal factors. Prioritize dimensions that are likely to interact with t2's effect on GB and VC.

3. Analyze t2 impact on GB and VC by segment

For each segment, estimate the causal effect of t2 on GB and VC using A/B test data or quasi-experimental methods. Look for segments where the GB effect is positive and the VC effect is non-negative.

4. Validate and refine segments

Use statistical tests and cross-validation to ensure the effects are robust and not due to chance. Refine segments by merging small ones or dropping unstable ones.

5. Recommend actionable segments

Summarize the segments that meet the criteria, along with expected impact and confidence intervals. Suggest next steps for targeting or further experimentation.

Key Points to Mention

  • Define t2, GB, and VC clearly and align with business stakeholders.
  • Use causal inference methods (e.g., A/B testing, propensity score matching) to estimate segment-level effects.
  • Consider interaction effects between t2 and user/order characteristics.
  • Apply multiple testing corrections when analyzing many segments.
  • Validate findings on a holdout set to avoid overfitting.
  • Prioritize segments by potential impact and ease of implementation.

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q4

If 20 experiments are running at the same time, how do you set launch criteria that control for false discoveries and still let you make reliable decisions?

A/B Testing & ExperimentationTechnical Trade-offsAdaptability & Ambiguity
Author's notes

This is where I actually felt okay.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by acknowledging the multiple testing problem and the need to balance false discovery control with statistical power. Then outline a structured approach: choose an appropriate error control method (e.g., FDR), consider business impact and power, and implement monitoring and decision rules. Emphasize that launch criteria should be pre-registered and aligned with business goals.

Pro tip: Don't just default to Bonferroni; discuss the trade-off between false positives and false negatives in the context of business risk, and suggest using a false discovery rate (FDR) approach like Benjamini-Hochberg when many tests are run, as it often provides a better balance for product experimentation.

1. Define the decision framework

Clarify what decisions will be made based on the experiments (e.g., launch, iterate, kill) and the cost of false positives vs. false negatives. This guides the choice of error control method.

2. Choose an error control method

Select a method to control false discoveries, such as Bonferroni, Holm-Bonferroni, or Benjamini-Hochberg (FDR). Consider the number of tests, correlation, and whether you need strong FWER control or can tolerate some false positives.

3. Adjust for power and sample size

Ensure each experiment has adequate power to detect meaningful effects after multiplicity adjustment. This may require larger sample sizes or prioritizing fewer, higher-impact tests.

4. Pre-register and monitor

Pre-register the analysis plan, including the error control method and decision thresholds. Monitor experiments for early stopping or futility, using sequential testing or alpha spending if needed.

5. Make decisions with business context

Interpret results in light of business impact, not just statistical significance. Use confidence intervals and effect sizes to assess practical significance and make reliable launch decisions.

Key Points to Mention

  • Multiple testing problem and inflated Type I error rate
  • False Discovery Rate (FDR) vs. Familywise Error Rate (FWER) control
  • Benjamini-Hochberg procedure and its assumptions
  • Statistical power and sample size implications
  • Pre-registration and avoiding p-hacking
  • Business impact and cost-benefit analysis of false positives/negatives

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.