← PayPal Interview Insights

PayPal·Data Scientist·Technical Phone Screen·Senior

Senior
May 2026

Summary

A deep product analytics case for a Data Scientist role, structured around reducing cancellation rates in airport ride pickups. The question was layered and genuinely hard, covering metrics, causal inference, experiment design, and stakeholder communication all in one go.

Questions Asked (4)

Q1

How would you define supply, demand, and marketplace health for airport pickups, and what primary, diagnostic, and guardrail metrics would you use to measure a reduction in cancellation rates?

Product Analytics & MetricsA/B Testing & Experimentation
Author's notes

I spent way too long on definitions and not enough on the guardrails, which is honestly where the interesting tradeoffs live.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by defining supply, demand, and marketplace health in the context of airport pickups, emphasizing the balance between driver availability and rider demand. Then, outline a metric framework with primary, diagnostic, and guardrail metrics to measure cancellation rate reduction, linking each to business outcomes. Use a structured, hypothesis-driven approach to show how you'd validate the impact of any intervention.

Pro tip: Tie your metrics to PayPal's two-sided marketplace dynamics and emphasize how reducing cancellations improves both rider and driver experiences, ultimately driving loyalty and transaction volume. Mention the importance of segmenting by airport, time of day, and user type to uncover actionable insights.

1. Define supply, demand, and marketplace health

Explain supply as the number of available drivers, demand as the number of ride requests, and marketplace health as the efficient matching of the two with minimal cancellations and wait times. Highlight that health is a balance: too much supply leads to idle drivers, too little leads to unmet demand.

2. Identify primary metric for cancellation reduction

Choose a primary metric that directly reflects the goal, such as overall cancellation rate (cancellations per completed ride) or the complement, match rate. Ensure it's sensitive to changes and aligned with business objectives.

3. Select diagnostic metrics

Pick metrics that explain why cancellations occur, such as driver wait time, rider wait time, cancellation reason codes, and supply-demand ratio by time and location. These help diagnose root causes and guide interventions.

4. Define guardrail metrics

Identify metrics that ensure the reduction in cancellations doesn't harm other aspects, like driver earnings, rider satisfaction (CSAT), completed rides, and overall marketplace liquidity. Monitor these to avoid unintended consequences.

5. Design measurement and experimentation plan

Propose an A/B test or quasi-experimental design to measure the impact of the intervention on the primary metric, while tracking diagnostic and guardrail metrics. Include segmentation and statistical power considerations.

Key Points to Mention

  • Two-sided marketplace dynamics: supply (drivers) and demand (riders) must be balanced for health.
  • Primary metric: cancellation rate (or its inverse, match rate) as the key success measure.
  • Diagnostic metrics: wait times, cancellation reasons, supply-demand ratio, and geographic/time segmentation.
  • Guardrail metrics: driver earnings, rider satisfaction, completed rides, and marketplace liquidity.
  • Experimentation: A/B testing with proper randomization and control for confounders.
  • Business impact: link cancellation reduction to improved retention, loyalty, and revenue.

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q2

What are your hypotheses for why cancellations happen at airports, separately for drivers and riders, and what signals in the data would you look for to test each one?

Product Analytics & MetricsRoot Cause AnalysisProduct Sense & Ideation
Author's notes

This part I actually liked.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by clarifying the marketplace context and defining cancellation events, then structure your answer by separating driver-side and rider-side hypotheses. For each hypothesis, specify the data signals and metrics you would analyze to test it, and prioritize hypotheses based on potential impact and ease of validation.

Pro tip: Demonstrate product sense by linking cancellations to marketplace health metrics like liquidity and utilization, and mention how you'd design experiments or use causal inference to move beyond correlations.

1. Clarify and Define

Ask clarifying questions to understand the product (e.g., ride-hailing), what constitutes a cancellation, and the time window. Define key terms and scope the analysis.

2. Driver-Side Hypotheses

Brainstorm reasons drivers cancel: long pickup distance, low expected fare, safety concerns, or better opportunities elsewhere. For each, identify data signals like pickup ETA, fare estimate, driver location, and historical acceptance rates.

3. Rider-Side Hypotheses

Brainstorm reasons riders cancel: long wait times, price surges, finding alternative transport, or accidental bookings. For each, identify data signals like wait time, price changes, rider history, and app usage patterns.

4. Data Signals and Testing

For each hypothesis, specify the metrics and data cuts (e.g., by time of day, location, user segment) to test it. Suggest analytical methods like cohort analysis, regression, or A/B tests to validate.

5. Prioritize and Recommend

Prioritize hypotheses based on impact and feasibility, and propose next steps such as experiments or product changes to reduce cancellations.

Key Points to Mention

  • Distinguish between driver-initiated and rider-initiated cancellations, as drivers and riders have different incentives.
  • Consider marketplace dynamics: supply-demand imbalance, surge pricing, and driver utilization.
  • Use metrics like cancellation rate, time-to-cancel, and cancellation reason codes (if available).
  • Segment analysis by geography, time, user tenure, and trip characteristics.
  • Propose experiments (e.g., incentives, better matching) to test hypotheses causally.
  • Mention potential data limitations and how to mitigate them (e.g., missing reason codes, selection bias).

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q3

Given queue interference, time-varying confounding, and a small number of airports, how would you design an experiment or quasi-experiment to estimate the causal impact of an intervention on cancellation rates?

A/B Testing & ExperimentationTechnical Trade-offsAdaptability & Ambiguity
Author's notes

This is where I struggled.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by acknowledging the complexities: queue interference (SUTVA violations), time-varying confounding, and limited airports. Propose a quasi-experimental design like a difference-in-differences with staggered adoption or synthetic control, leveraging variation in intervention timing across airports. Emphasize robustness checks and sensitivity analyses to address confounding and interference.

Pro tip: Consider using a causal inference method that explicitly models interference, such as network-based or spatial models, and pre-register your analysis plan to enhance credibility. Also, discuss how you would validate the design using placebo tests or negative controls.

1. Define the causal estimand and assumptions

Clarify the target effect (e.g., ATE on cancellation rates) and state assumptions like no interference (SUTVA) and no unmeasured confounding. Discuss how queue interference violates SUTVA and how time-varying confounding complicates estimation.

2. Choose a quasi-experimental design

Given few airports, consider designs like difference-in-differences with staggered intervention rollout, synthetic control, or interrupted time series. If randomization is possible, use cluster randomization with airports as clusters, but account for interference via design (e.g., saturation design).

3. Address confounding and interference

Use methods like propensity score weighting, fixed effects, or instrumental variables to handle time-varying confounding. For interference, model spillovers using spatial or network models, or use designs that isolate interference (e.g., partial interference).

4. Plan analysis and robustness checks

Pre-specify the analysis plan, including sensitivity analyses for unmeasured confounding (e.g., E-value), placebo tests, and negative controls. Use bootstrapping or randomization inference for valid inference with few clusters.

5. Interpret and communicate results

Discuss limitations, especially regarding generalizability and potential biases. Provide confidence intervals and effect sizes, and relate findings to business impact (e.g., cancellation rate reduction).

Key Points to Mention

  • SUTVA violation due to queue interference and potential spillover effects
  • Time-varying confounding and methods to adjust for it (e.g., marginal structural models, g-computation)
  • Quasi-experimental designs suitable for small number of units (synthetic control, difference-in-differences with staggered adoption)
  • Cluster randomization and interference-aware designs (e.g., saturation design, partial interference)
  • Sensitivity analysis for unmeasured confounding (E-value, Rosenbaum bounds)
  • Inference with few clusters (randomization inference, wild bootstrap)

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q4

How would you build proxy metrics for driver and rider satisfaction specific to the airport context, what data would you need, and how would you communicate the tradeoffs of a proposed intervention to stakeholders?

Stakeholder ManagementProduct Analytics & MetricsCross-functional Alignment
Author's notes

Proxy metrics were fine, I talked about re-acceptance rate after a cancellation for drivers and app re-open behavior for riders.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by defining the ideal satisfaction metric (e.g., post-ride CSAT) and then propose proxy metrics that are measurable and correlated, such as wait time, cancellation rate, and driver acceptance rate. Explain the data needed to validate these proxies and how you would communicate tradeoffs to stakeholders using a framework that balances rider and driver needs.

Pro tip: Acknowledge that airport dynamics are unique (e.g., queue systems, regulations) and propose segmenting metrics by terminal, time of day, and driver type to avoid one-size-fits-all solutions. This shows you understand the complexity and can tailor your approach.

1. Define satisfaction and identify proxies

Clarify what 'satisfaction' means for riders and drivers in the airport context (e.g., rider: timely pickup, driver: fair earnings). Propose proxy metrics like wait time, cancellation rate, driver acceptance rate, and trip completion rate.

2. Determine data requirements

List the data needed: GPS logs, timestamps, driver/rider app interactions, cancellation reasons, surge pricing, airport queue data, and external factors like flight schedules. Mention the need for historical data to establish baselines and correlations.

3. Validate proxies with statistical methods

Describe how you would validate proxies: correlation analysis with direct satisfaction surveys, regression models, and A/B tests. Emphasize the importance of statistical significance and avoiding spurious correlations.

4. Communicate tradeoffs to stakeholders

Use a structured approach: quantify impact on both rider and driver metrics, present scenarios (e.g., reducing wait time may increase driver idle time), and align with business goals. Visualize tradeoffs with charts or dashboards.

5. Propose and iterate on interventions

Suggest a pilot intervention (e.g., dynamic pricing, dedicated pickup zones) and define success metrics. Emphasize iterative testing and stakeholder feedback to refine proxies and interventions.

Key Points to Mention

  • Correlation vs. causation when using proxies
  • Segmenting metrics by airport, terminal, time, and driver/rider demographics
  • Balancing rider and driver satisfaction (e.g., wait time vs. driver earnings)
  • Data privacy and ethical considerations when using location data
  • Stakeholder alignment through clear communication of tradeoffs and business impact
  • Use of A/B testing and control groups to measure intervention effectiveness

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.