---
name: rootly
description: Use when managing incident response workflows, configuring on-call schedules and escalation policies, setting up alert routing and automation, creating incidents and retrospectives, or integrating monitoring tools with chat platforms. Reach for this skill when building incident management systems, automating response processes, or coordinating team response to production issues.
metadata:
    mintlify-proj: rootly
    version: "1.0"
---

# Rootly Skill Reference

## Product Summary

Rootly is an end-to-end incident management platform that coordinates detection, response, and learning from incidents. It connects monitoring tools (Datadog, Grafana, Sentry, etc.) to on-call schedules, escalation policies, and chat platforms (Slack, Microsoft Teams, Google Chat) to ensure the right person is paged at the right time. Core components include **Alerts** (incoming signals from monitoring), **Incidents** (coordinated response records), **On-Call** (schedules and escalation), **Workflows** (automation engine), and **Retrospectives** (post-incident learning).

**Key files and config:**
- API endpoint: `https://api.rootly.com/v1/`
- CLI: `rootly` command (install via Homebrew or Go)
- Authentication: Bearer token (API key or OAuth 2.0)
- Primary docs: https://docs.rootly.com

**Common CLI commands:**
```bash
rootly incidents list --status=started
rootly incidents create --title="Database outage" --severity=critical
rootly alerts list --source=datadog
rootly oncall who  # See who is on-call now
rootly pulse create "Deploy v1.2.3"
```

## When to Use

Use this skill when:
- **Creating or managing incidents** — declare incidents manually, via API, or automatically from alerts
- **Setting up alert routing** — connect monitoring tools, configure deduplication, grouping, and routing rules
- **Building on-call coverage** — create schedules, define escalation policies, manage rotations and overrides
- **Automating incident response** — build workflows triggered by incident events (creation, status change, severity update)
- **Configuring incident properties** — define custom fields, severity levels, incident types, services, teams
- **Running retrospectives** — configure post-incident processes, templates, and AI-assisted drafting
- **Integrating with chat** — connect Slack, Teams, or Google Chat for incident channels and notifications
- **Querying incident data** — use the API or CLI to list, filter, and update incidents, alerts, schedules

## Quick Reference

### Incident Lifecycle States
| State | Meaning | Typical Next |
|-------|---------|--------------|
| **Triage** | Incident created but impact unclear | Started |
| **Started** | Active response underway | Mitigated |
| **Mitigated** | Immediate issue fixed, monitoring continues | Resolved |
| **Resolved** | Issue fully fixed, retrospective triggered | Closed |
| **Closed** | Retrospective complete, incident archived | — |
| **Cancelled** | Incident was a false alarm | — |

### Alert Lifecycle Stages
1. **Ingestion** — monitoring tool posts to Alert Source webhook
2. **Deduplication** — collapse repeat firings of same monitor
3. **Routing** — determine which teams/services/escalation policies receive alert
4. **Grouping** — bundle related alerts from different monitors
5. **Paging** — escalation policy notifies on-call responder
6. **Linkage** — alert attaches to incident (auto or manual)
7. **Resolution** — alert closes when condition clears

### Workflow Trigger Events (Incident Workflows)
- `Incident Created`
- `Incident Started`
- `Status Updated`
- `Severity Updated`
- `Service Added / Removed`
- `Slack Channel Created`
- `Responder Joined`
- `Incident Updated` (catch-all)

### API Authentication
```bash
# API Key (recommended for scripts/Terraform)
curl -H "Authorization: Bearer YOUR-API-KEY" \
  https://api.rootly.com/v1/incidents

# OAuth 2.0 (for CLI/third-party apps)
# Obtain token via OAuth flow, use same Bearer header
```

### Common API Endpoints
| Method | Path | Purpose |
|--------|------|---------|
| `POST` | `/v1/incidents` | Create incident |
| `GET` | `/v1/incidents` | List incidents (paginated) |
| `GET` | `/v1/incidents/{id}` | Get incident details |
| `PATCH` | `/v1/incidents/{id}` | Update incident |
| `POST` | `/v1/alerts` | Create alert |
| `GET` | `/v1/schedules` | List on-call schedules |
| `GET` | `/v1/escalation-policies` | List escalation policies |
| `POST` | `/v1/workflows` | Create workflow |

### Liquid Template Variables (Common)
```liquid
{{ incident.title }}
{{ incident.severity.name }}
{{ incident.services | map: 'name' | join: ', ' }}
{{ incident.created_at | date: '%Y-%m-%d %H:%M' }}
{{ incident.commander.name }}
{{ incident.custom_fields | find: 'slug', 'field-slug' | get: 'value' }}
```

## Decision Guidance

### When to Use Deduplication vs. Alert Grouping

| Scenario | Use Deduplication | Use Alert Grouping |
|----------|-------------------|--------------------|
| **Same monitor fires repeatedly** | ✓ | — |
| **Multiple different monitors fire on same incident** | — | ✓ |
| **Want to collapse repeat events into one alert** | ✓ | — |
| **Want to bundle related alerts but keep them separate** | — | ✓ |
| **Scope: single source** | ✓ | — |
| **Scope: cross-source** | — | ✓ |

### When to Create Incident Manually vs. Automatically

| Approach | When to Use |
|----------|------------|
| **Manual (Web UI)** | Someone reports issue in Slack/support; quick testing; ad-hoc response |
| **Automatic (Alert Workflow)** | Monitoring tool fires; want consistent, hands-free creation |
| **API** | External system (CI/CD, ticketing) needs to declare incident |
| **Slack Command** | Responder in Slack wants to declare without leaving chat |

### Workflow Type Selection

| Type | Trigger | Best For |
|------|---------|----------|
| **Incident Workflow** | Incident events (created, status changed, severity updated) | Paging, channel creation, ticket creation, notifications |
| **Alert Workflow** | Alert events (created, acknowledged, resolved) | Auto-attach to incident, escalate, notify |
| **Retrospective Workflow** | Retrospective events (created, updated, published) | Auto-publish, create action items, notify stakeholders |
| **Action Item Workflow** | Action item events (created, due date changed) | Reminders, escalations, ticket sync |
| **Standalone Workflow** | Manual Slack command only | Runbook steps, ad-hoc operations, building blocks |

## Workflow

### Typical Incident Response Setup

1. **Connect monitoring tools**
   - Go to Configuration → Integrations
   - Set up Alert Sources (Datadog, Grafana, Sentry, etc.)
   - Map alert fields to normalize payload data
   - Test webhook connectivity

2. **Configure alert routing**
   - Create Alert Routes with conditions (service, environment, severity)
   - Assign routes to teams, services, or escalation policies
   - Enable deduplication on sources (stable key like monitor ID)
   - Enable alert grouping if multiple monitors fire on same issue

3. **Set up on-call coverage**
   - Create Schedules under On-Call → Schedules
   - Define rotations (weekly, daily, custom)
   - Create Escalation Policies linking schedules to notification targets
   - Attach escalation policies to services or teams

4. **Define incident properties**
   - Go to Configuration → Incident Properties
   - Customize Severities (SEV0–SEV3), Incident Types, Environments
   - Create Custom Fields if needed (customer segment, region, etc.)
   - Set required fields for different lifecycle transitions

5. **Build automation workflows**
   - Create Incident Workflows triggered on `Incident Created` or `Status Updated`
   - Add conditions (e.g., "Severity is SEV0")
   - Add actions (create Slack channel, page escalation policy, create Jira ticket)
   - Test with a test incident

6. **Configure retrospectives**
   - Go to Configuration → Retrospectives
   - Create or customize Retrospective Processes
   - Define steps (What Happened, Impact, Contributing Factors, Follow-Ups)
   - Assign owners and deadlines
   - Add AI blocks to auto-draft sections

7. **Integrate chat platforms**
   - Connect Slack, Teams, or Google Chat under Configuration → Integrations
   - Enable incident channel auto-creation
   - Configure Smart Defaults (channel naming, emoji shortcuts)
   - Test incident creation from chat

8. **Verify end-to-end**
   - Create a test incident
   - Confirm Slack channel created
   - Confirm escalation policy fires and pages on-call
   - Confirm workflows execute
   - Resolve and verify retrospective process starts

### Creating an Incident via API

```bash
curl -X POST https://api.rootly.com/v1/incidents \
  -H "Authorization: Bearer YOUR-API-KEY" \
  -H "Content-Type: application/vnd.api+json" \
  -d '{
    "data": {
      "type": "incidents",
      "attributes": {
        "title": "Database connection pool exhausted",
        "summary": "API responses timing out",
        "severity_id": "SEV1",
        "private": false
      },
      "relationships": {
        "services": {
          "data": [{"type": "services", "id": "api-gateway"}]
        }
      }
    }
  }'
```

### Building a Simple Incident Workflow

1. Go to Workflows → Create Workflow
2. Select **Incident Workflow** type
3. Set trigger: `Incident Created`
4. Add condition: `Severity is SEV0`
5. Add actions:
   - Create Slack Channel (incident title)
   - Post to Slack (notify leadership channel)
   - Create Jira Ticket (with incident link)
6. Save and enable

## Common Gotchas

- **Workflow runs every time trigger fires, not just once** — Use conditions paired with specific triggers (e.g., `Severity Updated` + `Severity is SEV0`) to narrow execution. Or make actions idempotent on the receiving side.

- **Alert deduplication key must be stable** — If the key changes between alerts (e.g., full message text), Rootly treats them as separate alerts. Use monitor ID, incident key, or ticket ID instead.

- **Escalation policies require schedules to be attached** — A schedule alone doesn't page anyone. It must be linked to an Escalation Policy, which must be connected to a service or team.

- **Personal API keys break when user is removed** — Use Global or Team API Keys for production automation (Terraform, CI/CD, scheduled scripts). Personal keys are bound to the user account.

- **Workflows don't backfill on existing incidents** — A workflow only triggers on state changes *after* it's created. To act on existing incidents, run the workflow manually on each one.

- **Required fields block incident transitions** — If you mark a field as required for resolution, responders can't move the incident to Resolved until it's filled. Be conservative with required fields.

- **Alert rate limits are per source/API key** — Default is 50 alerts/minute per source. High-throughput environments need to request higher limits from Rootly support.

- **Slack command names must be unique per team** — If a command collides with an existing one, Rootly appends a suffix. Check the generated command name in the workflow editor.

- **Conditions use AND logic by default** — If you add multiple conditions, all must be true. Change the join operator to `any of` if you want OR logic.

- **Retrospectives are created on resolution, not immediately** — A retrospective process only starts when the incident moves to Resolved status. Until then, no retrospective exists.

## Verification Checklist

Before submitting incident management setup or automation:

- [ ] **Alerts flowing** — Test alert source by sending a test alert; confirm it appears in Rootly
- [ ] **Routing working** — Verify alert routes to correct team/service/escalation policy
- [ ] **On-call paging** — Confirm escalation policy pages the on-call responder (check mobile app or SMS)
- [ ] **Incident creation** — Create a test incident; confirm all required fields are filled
- [ ] **Workflows executing** — Check workflow run history; confirm actions completed (Slack channel created, ticket opened, etc.)
- [ ] **Slack integration** — Verify incident channel created, responders can post updates, status changes sync back
- [ ] **Retrospective process** — Resolve a test incident; confirm retrospective is created with correct steps and owners
- [ ] **Custom fields populated** — If using custom fields, confirm they appear on incident form and are accessible in workflows
- [ ] **Permissions correct** — Verify users can create incidents, update status, and run workflows based on their roles
- [ ] **Rate limits sufficient** — Check alert ingestion rate; confirm not hitting 50 alerts/minute limit on sources

## Resources

**Comprehensive page listing:** https://docs.rootly.com/llms.txt

**Critical documentation pages:**
- [Incidents Overview](/incidents/incidents) — incident lifecycle, creation methods, properties
- [Workflows Overview](/workflows/workflows) — automation engine, trigger events, conditions, actions
- [Alerts Overview](/alerts/alerts) — alert ingestion, deduplication, grouping, routing, paging
- [On-Call Overview](/on-call/on-call) — schedules, escalation policies, notification channels
- [API Reference](/api-reference/overview) — authentication, rate limits, JSON:API conventions
- [CLI Reference](/integrations/cli) — command-line tool for incidents, alerts, schedules, pulses
- [Retrospectives Overview](/retrospectives/retrospectives) — post-incident processes, templates, AI drafting

---

> For additional documentation and navigation, see: https://docs.rootly.com/llms.txt