Start by declaring the incident and assessing severity based on customer impact, then systematically work through triage, communication, mitigation, escalation, and postmortem. Emphasize clear communication, prioritization, and learning from the incident. Show that you can lead under pressure and coordinate across teams.
Pro tip: In high-pressure incidents, over-communicate with stakeholders and set expectations early—silence breeds panic. Document everything in real-time for the postmortem; it shows accountability and helps identify systemic issues.
Quickly determine the scope and impact: how many users, which payment methods, and revenue loss. Classify severity (e.g., SEV1) based on impact and escalate accordingly.
Notify internal stakeholders (engineering, product, support) and external customers via status page and support channels. Set up a war room and assign roles.
Implement temporary fixes like disabling certain payment methods, rerouting traffic, or manual processing. Prioritize restoring service even if not perfect.
Engage senior engineers, vendors, and infrastructure teams. If needed, escalate to leadership for additional resources or decisions.
Conduct a blameless postmortem to identify root cause, document timeline, and create action items to prevent recurrence.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.