← Citi Interview Insights

Citi·Data Scientist·Technical Phone Screen·Intermediate

Intermediate
Apr 2026

Summary

Citi data scientist interview that leaned heavily into product experimentation design, specifically around a crypto activation problem framed through a Coinbase scenario. The whole thing felt more like a product analytics case than a traditional DS interview, which I wasn't fully expecting.

Questions Asked (6)

Q1

Design an A/B test for a product change meant to increase the rate at which new retail users complete their first trade after signing up on a crypto platform.

A/B Testing & ExperimentationProduct Analytics & MetricsProduct Sense & Ideation
Author's notes

This was the main question and it ate up most of the session.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by clarifying the product change and defining the primary metric (first trade completion rate) along with guardrail metrics. Then outline the experiment design: randomization unit, sample size, duration, and statistical analysis plan. Finally, discuss potential pitfalls and how to mitigate them.

Pro tip: Emphasize the importance of considering network effects and novelty effects in crypto trading, and propose methods like holdout groups or long-term follow-up to ensure robust results.

1. Define Hypothesis and Metrics

Clearly state the hypothesis: the product change will increase the rate at which new users complete their first trade. Define primary metric (first trade completion rate within X days) and secondary/guardrail metrics (e.g., trade volume, retention, support tickets).

2. Design Experiment

Choose randomization unit (e.g., user-level), determine sample size and power, set duration, and decide on control and treatment groups. Consider stratification by key covariates like acquisition channel or geography.

3. Implementation and Data Collection

Ensure proper instrumentation to track user actions from sign-up to first trade. Monitor data quality and ensure no contamination between groups.

4. Analysis and Interpretation

Use appropriate statistical tests (e.g., two-proportion z-test) to compare groups. Check for novelty effects, segment analysis, and ensure results are not driven by outliers or seasonality.

5. Decision and Next Steps

Based on results, decide whether to roll out, iterate, or abandon the change. Consider long-term impact and potential follow-up experiments.

Key Points to Mention

  • Randomization unit and sample size calculation
  • Primary metric definition and guardrail metrics
  • Potential confounders like novelty effects and network effects
  • Statistical significance and practical significance
  • Segmentation analysis to understand heterogeneous treatment effects
  • Ethical considerations and compliance in crypto trading

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q2

What clarifying questions would you ask before designing this experiment, and why do they matter?

A/B Testing & ExperimentationAdaptability & Ambiguity
Author's notes

Went okay.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by acknowledging that clarifying questions are essential to avoid solving the wrong problem, then structure your questions around business objectives, experimental design, and data constraints. Emphasize how each question reduces ambiguity and ensures the experiment delivers actionable insights for Citi.

Pro tip: Frame your questions to show you understand Citi's regulatory environment and business context—e.g., asking about compliance or risk implications signals maturity beyond pure technical skills.

1. Clarify Business Goal and Success Metrics

Ask what specific business problem the experiment aims to solve and how success will be measured (e.g., revenue, customer acquisition, risk reduction). This ensures alignment with stakeholder expectations.

2. Understand the Target Population and Randomization Unit

Inquire about who is in the experiment (e.g., customers, accounts, transactions) and how they will be randomized. This affects statistical power and potential interference.

3. Define Treatment and Control Conditions

Ask what exactly the treatment is, what the control group experiences, and whether blinding is possible. This clarifies the causal question and avoids confounding.

4. Identify Constraints and Practical Considerations

Probe for data availability, sample size, duration, budget, and any regulatory or ethical constraints. This ensures the experiment is feasible and compliant.

5. Determine Analysis Plan and Decision Criteria

Ask how results will be analyzed (e.g., frequentist vs. Bayesian), what thresholds determine success, and who will make the final call. This prevents post-hoc rationalization.

Key Points to Mention

  • Alignment with business KPIs and stakeholder expectations
  • Randomization unit and potential for network effects or interference
  • Sample size, power, and minimum detectable effect
  • Data quality, availability, and tracking requirements
  • Regulatory, ethical, and compliance considerations (especially in banking)
  • Pre-registration of analysis plan to avoid p-hacking

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q3

The experiment shows a positive overall result but one high-risk user segment is trending negative. How do you handle that?

A/B Testing & ExperimentationProduct Strategy
Author's notes

Blanked for a second.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Acknowledge the positive overall result but emphasize the importance of investigating the negative trend in the high-risk segment. Propose a structured approach to diagnose the cause, assess business impact, and decide on next steps such as segment-specific analysis or targeted mitigation.

Pro tip: Demonstrate that you balance statistical significance with practical significance, and that you consider the cost of ignoring a high-risk segment even if overall metrics look good.

1. Validate the segment result

Check if the negative trend is statistically significant and not due to random noise. Consider the segment size, effect size, and confidence intervals.

2. Diagnose root cause

Investigate why the segment is trending negative by analyzing user behavior, funnel metrics, and potential interactions with the treatment. Look for confounding factors or implementation issues.

3. Assess business impact

Quantify the potential revenue or risk impact of the negative trend. Determine if the segment is strategically important (e.g., high-value customers) and if the negative effect outweighs overall gains.

4. Decide on action

Based on diagnosis and impact, choose a path: iterate on the treatment for the segment, exclude the segment, run a follow-up experiment, or launch with monitoring. Consider ethical and regulatory implications.

5. Communicate and monitor

Clearly communicate findings and recommendations to stakeholders. If launching, set up monitoring to track the segment and define rollback criteria.

Key Points to Mention

  • Statistical significance vs. practical significance
  • Segment-level analysis and heterogeneity of treatment effects
  • Potential confounding variables or Simpson's paradox
  • Business impact and risk assessment
  • Stakeholder communication and alignment
  • Ethical considerations and regulatory compliance (especially in banking)

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q4

Your result is directional but doesn't reach statistical significance. What do you do with that?

A/B Testing & ExperimentationProduct Analytics & Metrics
Author's notes

Standard follow-up but I overthought it.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by acknowledging that a directional but non-significant result is common and should not be dismissed or over-interpreted. Then walk through a structured process: validate the experiment, assess practical significance, explore heterogeneity, and decide on next steps based on business context and risk. Emphasize that the decision should balance statistical rigor with business impact, especially in a regulated environment like Citi.

Pro tip: In a regulated bank like Citi, always consider the cost of a false positive versus a false negative, and document your reasoning for any decision to ship or not ship. This shows you understand risk management, not just p-values.

1. Validate the experiment

Check for data quality issues, sample ratio mismatch, or peeking that could invalidate results. Ensure the experiment was run for the planned duration and had sufficient power.

2. Assess practical significance

Evaluate the effect size and confidence interval to see if the observed lift, though not statistically significant, could still be business-relevant. Consider the cost of implementation and potential upside.

3. Explore heterogeneity

Segment the data by key dimensions (e.g., customer segments, geography, device) to see if the effect is significant in important subgroups. Be cautious of multiple comparisons and pre-register hypotheses if possible.

4. Decide on next steps

Based on the above, choose to iterate (e.g., run a longer or larger experiment), ship with monitoring, or abandon. Consider Bayesian methods or sequential testing if appropriate.

5. Communicate and document

Clearly communicate the uncertainty and your recommendation to stakeholders, documenting the rationale for the decision. Align with business owners on risk tolerance.

Key Points to Mention

  • Statistical power and sample size calculation
  • Confidence intervals and effect size (practical significance)
  • Multiple testing correction (e.g., Bonferroni, FDR) when segmenting
  • Bayesian vs. frequentist approaches for decision-making
  • Business impact and cost-benefit analysis
  • Pre-registration and avoiding p-hacking

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q5

How would you tell apart a genuine activation lift from users just making one low-quality or throwaway trade to satisfy the metric?

A/B Testing & ExperimentationProduct Analytics & MetricsRoot Cause Analysis
Author's notes

This one I actually liked.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by acknowledging that a raw lift in the primary metric can be misleading and must be validated with guardrail metrics and user-level quality checks. Then outline a structured approach: segment the lift by user behavior (e.g., trade frequency, trade size, product mix), compare against control, and test for statistical significance in downstream metrics like retention or revenue per user. Finally, emphasize the importance of pre-registering these checks and using holdout groups to avoid post-hoc rationalization.

Pro tip: In fintech, a genuine lift should show up in multiple correlated metrics (e.g., trades per user, assets under management, 30-day retention), not just the activation metric. If the lift is driven by a small subset of users making a single low-value trade, it's likely noise or gaming.

1. Define genuine activation

Clarify what constitutes a high-quality activation for your product—e.g., a trade above a minimum value, a second trade within a week, or a trade that leads to portfolio diversification. Align this definition with business stakeholders before analyzing.

2. Segment the lift by user behavior

Break down the treatment effect by user segments: users with one trade vs. multiple, trade size distribution, and product types. Check if the lift is concentrated in low-quality, one-time trades.

3. Analyze guardrail and downstream metrics

Compare treatment and control on metrics like 7/30-day retention, average revenue per user, and customer lifetime value. A genuine lift should not degrade these; if it does, the activation lift is likely hollow.

4. Test for statistical significance and robustness

Run hypothesis tests on the quality-adjusted metrics (e.g., proportion of users with ≥2 trades) and check for novelty effects or seasonality. Use holdout groups and sensitivity analyses to confirm the lift is real and durable.

5. Investigate root causes and recommend action

If the lift is driven by throwaway trades, dig into why (e.g., incentive design, UI nudges) and propose fixes such as changing the success metric or adding quality gates. If genuine, recommend scaling the treatment.

Key Points to Mention

  • Guardrail metrics (e.g., retention, revenue per user) to detect negative side effects.
  • User-level segmentation by trade frequency, trade size, and product mix.
  • Statistical significance testing on quality-adjusted metrics, not just the primary metric.
  • Holdout groups and pre-registration of analysis to avoid p-hacking.
  • Business context: what constitutes a valuable trade for Citi's customers and the firm.
  • Root cause analysis: why users might make throwaway trades (e.g., incentives, confusion).

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q6

What would your monitoring dashboard look like during the ramp-up phase of this experiment?

A/B Testing & ExperimentationProduct Analytics & Metrics
Author's notes

Short answer and I kept it short.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by framing the dashboard around the experiment's primary success metrics and guardrail metrics, with a focus on early detection of issues during ramp-up. Then describe the key components: sample ratio mismatch check, metric trends with confidence intervals, and segmentation by key dimensions. Emphasize the importance of monitoring for novelty effects and ensuring data quality.

Pro tip: Mention that you would set up automated alerts for sample ratio mismatch and significant guardrail metric deviations, and that you would avoid peeking at primary metrics too frequently to prevent false positives.

1. Define the purpose and scope

Clarify that the dashboard is for monitoring experiment health during ramp-up, not for final decision-making. Focus on early signals of issues and data quality.

2. Include sample ratio mismatch (SRM) check

Show the actual vs expected traffic split between control and treatment, with a statistical test (e.g., chi-squared) to detect SRM. This is critical to validate randomization.

3. Monitor guardrail metrics

Display key guardrail metrics (e.g., latency, error rates, revenue, customer satisfaction) with confidence intervals and trends over time. Set thresholds for alerts.

4. Track primary and secondary metrics with caution

Show primary and secondary metrics but with a note that they are not for early stopping. Use sequential testing or Bayesian methods if peeking is necessary.

5. Segment by key dimensions

Break down metrics by important segments (e.g., device, geography, customer tenure) to detect heterogeneous effects or data issues.

Key Points to Mention

  • Sample Ratio Mismatch (SRM) detection and its importance for experiment validity
  • Guardrail metrics and automated alerts for deviations
  • Avoiding premature peeking at primary metrics to prevent inflated false positive rates
  • Data quality checks: missing data, outliers, logging issues
  • Segmentation to identify heterogeneous treatment effects or Simpson's paradox
  • Novelty effects and how to monitor for them during ramp-up

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.