Global
Applies to every AI feature.
Proactive features
When Rootly speaks up on its own in the incident channel.
Rootly Agent in Slack
Questions asked in an incident channel.
Incident summarization
Summaries across the incident lifecycle.
AI in Retrospectives
Every AI block in every retrospective.
Rootly Agent in Web
Questions asked from the web app.
Rootly Agent in Mobile
Questions asked from the mobile app.
Global
Applies to every AI feature.Why we recommend this
Why we recommend this
These defaults come from resilience engineering. It is the field that studies how complex systems and the people running them cope with failure, drawn from decades of research in industries where the stakes made getting this right unavoidable.Its starting point is that no system this complex fits in one person’s head. Modern software runs beyond the mental model of any individual or team, so incidents are one of the few times you get to see what it actually does.Resilience engineering treats failure as a permanent condition rather than a defect to eliminate. A good record is less about stopping one thing from happening again and more about improving how well your people handle the next thing, which will be different.It also holds that the people running a system are the source of its resilience. They adapt around the gaps between documented process and what is actually in front of them, and on most days that adaptation is why nothing broke.Incidents rarely have one cause. Several conditions have to line up, most of them individually harmless. Settling on a single cause usually just marks where the investigation stopped.Hindsight makes decisions look obvious that never were. Describing what people knew at the moment they acted is what gives you something you can act on.
Proactive Features
Set on the Global page, next to the opt-in toggle. Applies when Rootly speaks up on its own in the incident channel, without anyone asking, during the incident and for the hour after it resolves.Why we recommend this
Why we recommend this
A proactive agent is only worth having if it speaks when one message would shorten the incident and stays silent otherwise. An unanswered question that blocks the next step is the clearest case: a short reply in its thread that names who can answer costs the channel nothing and often saves minutes. One nudge per question keeps the agent from becoming noise.Every claim the agent makes has to trace back to something it read. A confident guess in an incident channel is worse than silence, because people act on it.The minutes after resolution are when follow-ups get lost: people leave the channel and whatever was promised but not written down goes with them. One look at what is missing, said once, is cheap; nagging a team that has just finished is not.
Slack
Rootly Agent in Slack
Controls the agent in your incident channels.Why we recommend this
Why we recommend this
Resilience engineering treats an incident as a period of active sensemaking. People are building an understanding of a system nobody fully holds, and how that understanding forms determines how the response goes.Understanding narrows faster than anyone decides to narrow it. A theory proposed early gets repeated, and within the hour it is being treated as established. Naming a theory as a theory and saying who is testing it keeps it open.Investigations usually run several theories at once. Reporting only the loudest one closes down the search before anyone chose to.Two things breaking at the same time is weak evidence of a connection. Treating it as strong reliably costs an hour.Discarded hypotheses are the hardest thing to recover by scrolling. They are also what someone joining needs most.Silence means nobody wrote it down. It says nothing about whether something happened.
Incident Summarization
Covers/rootly summary and /rootly catch up.
Why we recommend this
Why we recommend this
Complex systems fail in ways that are partial and uneven. Scope is genuinely unclear early on, so a summary that sounds more certain than the record does real damage.Saying who is unaffected is as useful as saying who is affected. It is how a reader works out whether this touches them.Readers arrive with a theory already formed. Naming the obvious suspected cause and giving its status is often the most useful line in an update.Put the time of the last update at the top. The reader needs to know how current this is before they read the rest of it.Detection time only becomes measurable once it is written down. Recording the gap between first signal and declaration turns it into something you can improve.Catch-up serves a different job from a summary. Someone joining needs current state and open questions ahead of a chronology.
Retros
AI in Retrospectives
Applies to every AI block in every retrospective. It runs longest because a retrospective has the most to get right.Why we recommend this
Why we recommend this
Resilience engineering holds that a review exists to build understanding in the people who run the system. That is the whole return on the time spent, and a finding written for them is a finding they did not reach themselves.It starts from the assumption that people act sensibly given what they can see at the time. Reconstructing what they were seeing is what produces conditions you can change.Question wording steers the answer. “Why did you” pulls toward motives and people. “What were you seeing” pulls toward the system.Hindsight makes a written sequence read as inevitable. Naming what was genuinely uncertain at the time stops a reader concluding someone should have seen it coming.Note what surprised people. If a responder says “that shouldn’t be possible”, that is worth capturing. It usually means something about the system is not what the team thinks it is.The things that limited the damage are capabilities. Resilience engineering treats them as findings in their own right, because they are what you rely on next time.
Web
Rootly Agent in Web
Covers questions asked from the web app.Why we recommend this
Why we recommend this
Resilience engineering has a name for the problem with looking across incidents: what you look for determines what you find. The filter chooses the conclusion before the analysis starts.State the filter, the window and the count. It turns an assertion into something a reader can check and argue with.Ambiguous questions produce confident wrong answers. Restating the interpretation first catches that before anyone anchors on a number.If one factor were sufficient, the problem would already be fixed. Recurring problems recur because several conditions keep lining up.Near misses are the fastest route to what actually varies. Comparing them against the ones that escalated shows what made the difference.An incomplete record is worth naming. Saying what would need to be captured turns a dead end into an improvement.
Mobile
Rootly Agent in Mobile
Covers questions asked from the mobile app.Why we recommend this
Why we recommend this
The reader is deciding one thing: whether to get involved right now.When an answer gets shorter, words like “unconfirmed” are the first to go. What is left sounds certain even when it is a guess, so those words stay in and other detail gets cut instead.Early causal guesses are often wrong, and a short summary gets read quickly, remembered, and repeated. So the summary carries no cause.Someone woken at 3am does timezone arithmetic badly. Relative time comes first, absolute in brackets.Silence during an incident is information. Hiding it behind a confident summary makes a stalled incident look handled.
Frequently Asked Questions
Do I have to use these?
Do I have to use these?
They’re active by default, and most teams leave them that way. Switch a box
to Custom when you want something different.
Will my own text be overwritten when Rootly updates a set?
Will my own text be overwritten when Rootly updates a set?
No. Updates only reach a box still set to Best practice. Custom text is
yours.
Can I shorten a set?
Can I shorten a set?
Switch to Custom and cut. Keep the rules that describe something you’d
otherwise correct by hand.
Should I keep the global set and a feature set?
Should I keep the global set and a feature set?
Yes. They’re sent together and neither replaces the other. Where a rule
appears in both, keep it in one place so the two don’t drift apart. On a
direct conflict, the feature rule wins.