Make sure an urgent page always reaches a human
The most expensive failure isn’t a slow response, it’s a page nobody received. Before anything else, verify:- No gaps in your schedules. Coverage across time zones, holidays accounted for, overrides in place for planned time off.
- More than one level in every escalation policy. A single-level policy is a single point of failure. Level two should fire a short time after level one is paged and doesn’t respond.
- Critical alerts page a human target. Routing an alert with a High urgency to a Slack channel or a group with no on-call attached is the most common gap we find.
- Test the path. Trigger a test alert on each critical service quarterly and confirm who it woke up.
Define a Severity matrix to quickly communicate incident impact
Severities are a simple way for your teammates to easily understand the impact of an incident. Define three or four levels. Define each one by what the customer experiences, not by which team is involved or how stressed anyone is. Write the response expectation into the definition: who gets paged, how fast, who gets told. A useful test: two engineers reading the same alert at 3am should pick the same severity without discussing it.Ask the Rootly AI Agent to help you choose the right severity:
@Rootly what severity should this incident be set to? Update it for me.The Rootly AI Agent uses the context of the current incident, past incidents, and the Severity matrix you’ve configured in Rootly to propose an appropriate severity.Assign individuals to roles as soon as possible
Coordination failures cost more time than technical ones. At minimum:- Incident commander: owns the response, doesn’t debug. The single most valuable role to make non-optional.
- Communications lead: owns updates to stakeholders and customers so the commander doesn’t context-switch.
- Scribe: captures the timeline as it happens, not from memory afterward.
If you aren’t sure who to assign to a role, ask the Rootly AI Agent:
@Rootly assign the incident roles for me.Automate the first ten minutes with a playbook of steps
The teams that respond fastest aren’t improvising faster; they’ve removed the need to improvise completely by building a playbook. A good playbook is a short checklist of the first actions for a given failure mode, attached to the incident automatically, with each step already assigned to a role rather than a named person. Assigning to roles matters: named people go on vacation, roles don’t. This ensures a repeatable process every time an incident is declared, so your responding teams can focus without missing any critical steps.Structure your communications strategy
Responders, internal stakeholders, and customers need different things at different intervals. Most teams collapse them into one channel and then either spam responders or leave executives to interrupt the response team asking for updates. Responders need the working channel, and don’t mind the noise. Internal stakeholders need a scheduled update on a fixed cadence tied to severity. Every 30 minutes for SEV1, even if the update is “no change.” Customers need a status page and honest, low-detail updates. Template the updates. Nobody writes well under pressure, and hand-written updates are the first thing to slip when things get hard.Use the Rootly AI Agent to help draft an update:
@Rootly update my status page with a summary of the current work the team is doing to resolve the issue.Automate the operational work, and use an agent for the rest
There are two kinds of work in an incident: solving the problem, and everything around solving the problem. The second kind is what makes a response feel heavy, yet it’s necessary to keep teams moving in unison. Rather than relying on a teammate to complete the operational tasks, which will distract them from their mitigation work, and increases the chances that things fall through the cracks, use an AI Agent embedded right in your incident channel to handle this work. For example, useful things to ask the Rootly Agent in Slack:- “What key events have happened so far?” - keeps teammates up-to-date, without distracting the rest of the team to draft an update.
- “Is SEV2 right for this?” - a second opinion on severity with the incident’s context.
- “Has anyone else on the team mitigated a similar incident in the past week?” - make sure all the right stakeholders are working together.
- “Did anything get missed? Create the action items for me and assign them to the relevant team.” - nothing falls through the cracks, preventing a repeat incident from occurring in the future.
- “Summarize where we are for someone joining” - status, current theory, open action items, who’s involved.
- “Draft a customer update” - and publish it to the status page after you confirm.
- “Update the impacted services” - resolved from real service discovery, not guessed names.
Reserve time to learn and reflect for the future
Define the trigger criteria in advance so nobody has to argue about whether an incident “deserves” a retrospective. For example, trigger criteria can include a severity threshold, customer impact, or incident duration. Then: Run it within a week, while memory is fresh Use one template so retros are comparable to each other Focus on contributing factors and systems, not individual decisions Leave with action items that have owners and due dates, or don’t botherChecklist to get started
Before you start adopting these practices, make sure you have the following set up in Rootly:- Rootly’s AI features and AI Agent. These will be used to help you focus on the incident, without losing the outcomes you benefit from in your incident response process.
- Connectors: these power your AI Agent in Rootly to gather context from your external systems, like your CI/CD platform or cloud provider.
- Google Meet or Zoom integration. Whatever your video conferencing tool is, Rootly will automatically spin up a room for each incident. Bonus: enable to AI Scribe to take notes and record your call!
- Review your Retrospective template. You’ll already have a default template built by Rootly based on industry best practices: add any additional section(s) you’d like to capture in your retrospective for your team to reflect on.