This was the main question and it ate up most of the session.
Start by clarifying the product change and defining the primary metric (first trade completion rate) along with guardrail metrics. Then outline the experiment design: randomization unit, sample size, duration, and statistical analysis plan. Finally, discuss potential pitfalls and how to mitigate them.
Pro tip: Emphasize the importance of considering network effects and novelty effects in crypto trading, and propose methods like holdout groups or long-term follow-up to ensure robust results.
Clearly state the hypothesis: the product change will increase the rate at which new users complete their first trade. Define primary metric (first trade completion rate within X days) and secondary/guardrail metrics (e.g., trade volume, retention, support tickets).
Choose randomization unit (e.g., user-level), determine sample size and power, set duration, and decide on control and treatment groups. Consider stratification by key covariates like acquisition channel or geography.
Ensure proper instrumentation to track user actions from sign-up to first trade. Monitor data quality and ensure no contamination between groups.
Use appropriate statistical tests (e.g., two-proportion z-test) to compare groups. Check for novelty effects, segment analysis, and ensure results are not driven by outliers or seasonality.
Based on results, decide whether to roll out, iterate, or abandon the change. Consider long-term impact and potential follow-up experiments.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.
Start by acknowledging that clarifying questions are essential to avoid solving the wrong problem, then structure your questions around business objectives, experimental design, and data constraints. Emphasize how each question reduces ambiguity and ensures the experiment delivers actionable insights for Citi.
Pro tip: Frame your questions to show you understand Citi's regulatory environment and business context—e.g., asking about compliance or risk implications signals maturity beyond pure technical skills.
Ask what specific business problem the experiment aims to solve and how success will be measured (e.g., revenue, customer acquisition, risk reduction). This ensures alignment with stakeholder expectations.
Inquire about who is in the experiment (e.g., customers, accounts, transactions) and how they will be randomized. This affects statistical power and potential interference.
Ask what exactly the treatment is, what the control group experiences, and whether blinding is possible. This clarifies the causal question and avoids confounding.
Probe for data availability, sample size, duration, budget, and any regulatory or ethical constraints. This ensures the experiment is feasible and compliant.
Ask how results will be analyzed (e.g., frequentist vs. Bayesian), what thresholds determine success, and who will make the final call. This prevents post-hoc rationalization.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.
Acknowledge the positive overall result but emphasize the importance of investigating the negative trend in the high-risk segment. Propose a structured approach to diagnose the cause, assess business impact, and decide on next steps such as segment-specific analysis or targeted mitigation.
Pro tip: Demonstrate that you balance statistical significance with practical significance, and that you consider the cost of ignoring a high-risk segment even if overall metrics look good.
Check if the negative trend is statistically significant and not due to random noise. Consider the segment size, effect size, and confidence intervals.
Investigate why the segment is trending negative by analyzing user behavior, funnel metrics, and potential interactions with the treatment. Look for confounding factors or implementation issues.
Quantify the potential revenue or risk impact of the negative trend. Determine if the segment is strategically important (e.g., high-value customers) and if the negative effect outweighs overall gains.
Based on diagnosis and impact, choose a path: iterate on the treatment for the segment, exclude the segment, run a follow-up experiment, or launch with monitoring. Consider ethical and regulatory implications.
Clearly communicate findings and recommendations to stakeholders. If launching, set up monitoring to track the segment and define rollback criteria.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.
Start by acknowledging that a directional but non-significant result is common and should not be dismissed or over-interpreted. Then walk through a structured process: validate the experiment, assess practical significance, explore heterogeneity, and decide on next steps based on business context and risk. Emphasize that the decision should balance statistical rigor with business impact, especially in a regulated environment like Citi.
Pro tip: In a regulated bank like Citi, always consider the cost of a false positive versus a false negative, and document your reasoning for any decision to ship or not ship. This shows you understand risk management, not just p-values.
Check for data quality issues, sample ratio mismatch, or peeking that could invalidate results. Ensure the experiment was run for the planned duration and had sufficient power.
Evaluate the effect size and confidence interval to see if the observed lift, though not statistically significant, could still be business-relevant. Consider the cost of implementation and potential upside.
Segment the data by key dimensions (e.g., customer segments, geography, device) to see if the effect is significant in important subgroups. Be cautious of multiple comparisons and pre-register hypotheses if possible.
Based on the above, choose to iterate (e.g., run a longer or larger experiment), ship with monitoring, or abandon. Consider Bayesian methods or sequential testing if appropriate.
Clearly communicate the uncertainty and your recommendation to stakeholders, documenting the rationale for the decision. Align with business owners on risk tolerance.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.
Start by acknowledging that a raw lift in the primary metric can be misleading and must be validated with guardrail metrics and user-level quality checks. Then outline a structured approach: segment the lift by user behavior (e.g., trade frequency, trade size, product mix), compare against control, and test for statistical significance in downstream metrics like retention or revenue per user. Finally, emphasize the importance of pre-registering these checks and using holdout groups to avoid post-hoc rationalization.
Pro tip: In fintech, a genuine lift should show up in multiple correlated metrics (e.g., trades per user, assets under management, 30-day retention), not just the activation metric. If the lift is driven by a small subset of users making a single low-value trade, it's likely noise or gaming.
Clarify what constitutes a high-quality activation for your product—e.g., a trade above a minimum value, a second trade within a week, or a trade that leads to portfolio diversification. Align this definition with business stakeholders before analyzing.
Break down the treatment effect by user segments: users with one trade vs. multiple, trade size distribution, and product types. Check if the lift is concentrated in low-quality, one-time trades.
Compare treatment and control on metrics like 7/30-day retention, average revenue per user, and customer lifetime value. A genuine lift should not degrade these; if it does, the activation lift is likely hollow.
Run hypothesis tests on the quality-adjusted metrics (e.g., proportion of users with ≥2 trades) and check for novelty effects or seasonality. Use holdout groups and sensitivity analyses to confirm the lift is real and durable.
If the lift is driven by throwaway trades, dig into why (e.g., incentive design, UI nudges) and propose fixes such as changing the success metric or adding quality gates. If genuine, recommend scaling the treatment.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.
Start by framing the dashboard around the experiment's primary success metrics and guardrail metrics, with a focus on early detection of issues during ramp-up. Then describe the key components: sample ratio mismatch check, metric trends with confidence intervals, and segmentation by key dimensions. Emphasize the importance of monitoring for novelty effects and ensuring data quality.
Pro tip: Mention that you would set up automated alerts for sample ratio mismatch and significant guardrail metric deviations, and that you would avoid peeking at primary metrics too frequently to prevent false positives.
Clarify that the dashboard is for monitoring experiment health during ramp-up, not for final decision-making. Focus on early signals of issues and data quality.
Show the actual vs expected traffic split between control and treatment, with a statistical test (e.g., chi-squared) to detect SRM. This is critical to validate randomization.
Display key guardrail metrics (e.g., latency, error rates, revenue, customer satisfaction) with confidence intervals and trends over time. Set thresholds for alerts.
Show primary and secondary metrics but with a note that they are not for early stopping. Use sequential testing or Bayesian methods if peeking is necessary.
Break down metrics by important segments (e.g., device, geography, customer tenure) to detect heterogeneous effects or data issues.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.