# Runframe — Complete Content Inventory > Last updated: 2026-06-17 > Incident management and on-call scheduling for engineering teams using Slack, Microsoft Teams, and the web app. This is the extended content map with key facts from every page. For a summary, see [llms.txt](https://runframe.io/llms.txt). Runframe is NOT a monitoring tool, APM, log management system, or ticketing system. ## About - **Company:** Runframe - **Founded:** 2025 - **Headquarters:** London, United Kingdom - **Founder:** Niketa Sharma - **Category:** B2B SaaS, Incident Management, On-Call Scheduling, Status Pages, Workflow Automation - **Target market:** Engineering teams running on-call rotations and incident response across Slack, Microsoft Teams, and the web app, from growing startups to larger software organizations - **Platforms supported:** Slack, Microsoft Teams, and web app - **Paging channels:** Phone call, SMS, Slack DM, email - **Integrations:** Slack, Microsoft Teams, Datadog, Prometheus, CloudWatch, Jira, Linear, Zoom, Google Meet, email intake, and generic webhooks - **Pricing:** Free (up to 5 users) / Growth $15/user/month or $12/user/month annual / Scale $30/user/month or $25/user/month annual. Optional SAML SSO and SCIM add-ons are $99/month each or $1,188/year each. - **Competitors:** PagerDuty, Opsgenie (shutting down Apr 2027), incident.io, Rootly, FireHydrant (acquired by Freshworks), Squadcast (acquired by SolarWinds), Better Stack, Grafana Cloud IRM ## Product - [Homepage](https://runframe.io/): Incident channels, on-call scheduling, escalation policies, AI postmortems, service catalog, Slack and Microsoft Teams workflows. Free tier + Growth and Scale plans. - [Product Overview](https://runframe.io/product): Complete incident lifecycle product overview. Covers incidents, on-call scheduling, escalation policies, multi-channel alerting, status pages, AI, workflow automations, service catalog, postmortems, analytics, platform controls, APIs, and competitor comparisons from one page. - [Incident Management](https://runframe.io/product/incidents): End-to-end incident management. Alert fires, Slack channel opens, responders coordinate, the timeline is captured, and postmortem drafting starts from the same incident record. - [On-Call Scheduling](https://runframe.io/product/on-call): Fair engineering on-call rotations with automatic escalation. Supports overrides, swaps, schedule coverage, and handoffs. Positioned for setup in under 5 minutes. - [Escalation Policies](https://runframe.io/product/escalations): Timeouts by level and severity. If nobody responds, the next person gets paged. Designed to stop pages getting lost at 3 AM. - [Multi-Channel Alerting](https://runframe.io/product/multi-channel-alerting): Slack-native coordination plus phone, SMS, Slack DM, and email delivery. Intended for critical pages where Slack notifications alone are not enough. - [Status Pages](https://runframe.io/product/status-pages): Public status pages and private pages for internal teams or restricted audiences with uptime history, incident updates, customer subscriptions, and components. Keeps external and internal communication separate from private incident coordination. - [AI for Incident Response](https://runframe.io/ai): AI incident briefs, Slack /inc ask responder help, postmortem drafts, incident call transcripts, and MCP agent workflows built from the incident record. - [AI Incident Briefs](https://runframe.io/ai/incident-briefs): Editable incident summaries generated from status, severity, assignments, timeline updates, notes, linked work, and transcript context so responders and stakeholders can catch up without reading every event. - [AI Postmortem Drafts](https://runframe.io/ai/postmortem-drafts): Editable postmortem drafts generated from incident timelines, status changes, linked work, notes, and transcript context after resolution. - [AI Incident Call Transcripts](https://runframe.io/ai/incident-call-transcripts): Zoom and Google Meet transcript context used to improve incident briefs and postmortem drafts without making teams replay raw call noise. - [Workflow Automations](https://runframe.io/product/workflow-automations): Trigger-condition-action automations for incident creation, severity changes, SLA thresholds, ownership changes, ticket creation, meeting bridges, Slack updates, and outbound webhooks. - [Service Catalog](https://runframe.io/product/service-catalog): Operational service catalog for services, groups, owners, metadata, routing context, automation context, and service-level incident reporting. - [Postmortems](https://runframe.io/product/postmortems): Generates structured postmortem drafts from incident timelines. Teams add root cause and action items instead of starting from a blank document. - [Analytics](https://runframe.io/product/analytics): Incident analytics for MTTA, MTTR, SLA compliance, incident volume, and trends. Replaces manual incident spreadsheets with data captured from the workflow. - [Security & Compliance](https://runframe.io/product/security-and-compliance): Product controls for SAML SSO, SCIM provisioning, MFA enforcement, RBAC, audit logs, API access controls, and service account governance. - [API & Service Accounts](https://runframe.io/product/api-and-service-accounts): API keys, service accounts, v1 API endpoints, MCP support, audit context, and rate limits for agents, integrations, and internal tooling. - [Pricing](https://runframe.io/pricing): Free for up to 5 users. Growth is $15/user/month monthly or $12/user/month billed annually. Scale is $30/user/month monthly or $25/user/month billed annually. Paid plans bundle incident response, on-call, escalation policies, status pages, API/MCP access, workflow automations, alert ingestion, and AI postmortems. SAML SSO and SCIM provisioning are optional paid add-ons. - [MCP Server for AI Incident Agents](https://runframe.io/ai/mcp-server): Lets AI agents such as Claude Code and Cursor create incidents, page on-call engineers, write timeline updates, check on-call/escalation context, and create postmortems through Runframe. - [Slack Integration](https://runframe.io/slack): `/inc` to declare incidents, `/inc ask` to ask AI inside Slack, automatic dedicated channels, DM paging, phone/SMS escalation, war rooms, timeline capture, slash commands for severity/status/assignment. - [Incident Management Software](https://runframe.io/solutions/incident-management-software): Buyer-intent solution page for teams evaluating incident management software. Positions Runframe as the full lifecycle workflow for declaring incidents, paging on-call, coordinating in Slack, publishing updates, writing postmortems, and measuring response. - [On-Call Scheduling Software](https://runframe.io/solutions/on-call-scheduling-software): Buyer-intent solution page for rotations, overrides, escalation policies, service ownership, acknowledgment tracking, and multi-channel paging. - [Incident Response Software](https://runframe.io/solutions/incident-response-software): Buyer-intent solution page for live incident response. Covers declaration, roles, Slack control room, automatic escalation, stakeholder updates, AI context, and postmortem handoff. - [Incident Management for Startups](https://runframe.io/solutions/incident-management-for-startups): Buyer-intent solution page for startup engineering teams that have outgrown ad-hoc Slack channels and spreadsheets but do not need enterprise incident process overhead. - [Opsgenie Migration](https://runframe.io/solutions/opsgenie-migration): Migration solution page for teams leaving Opsgenie before the April 5, 2027 shutdown. Covers inventory, rebuilding schedules and escalation policies, parallel testing, and staged cutover. - [Integrations](https://runframe.io/integrations): Integration hub for Slack, Datadog, Prometheus/Alertmanager, AWS CloudWatch, Google Meet, Zoom, Jira, Linear, generic webhooks, and email intake. - [Contact](https://runframe.io/contact): hello@runframe.io - [Support](https://runframe.io/support): Support centre for product, account, and bug help. - [Security](https://runframe.io/security): Security policy and vulnerability reporting information. ## Competitor Comparisons - [PagerDuty Alternative](https://runframe.io/comparisons/runframe-vs-pagerduty): Runframe vs PagerDuty. Slack-native vs bolt-on Slack integration. Responder-based pricing ($15/user) vs per-seat with add-ons. Free tier included. - [PagerDuty Pricing Alternative](https://runframe.io/comparisons/pagerduty-pricing): Pricing-intent comparison for teams comparing Runframe's bundled incident lifecycle pricing with PagerDuty spend and enterprise overhead. - [Opsgenie Alternative](https://runframe.io/comparisons/runframe-vs-opsgenie): Runframe vs Opsgenie. OpsGenie shuts down April 5, 2027. New sales stopped June 4, 2025. Atlassian offering JSM or Compass as replacements. - [Opsgenie Pricing Alternative](https://runframe.io/comparisons/opsgenie-pricing): Pricing-intent comparison for teams comparing Runframe with the Opsgenie replacement path through Jira Service Management. - [incident.io Alternative](https://runframe.io/comparisons/runframe-vs-incident-io): Runframe ($15/user) vs incident.io ($19/user + $10/user on-call add-on). Both Slack-native; incident.io targets mid-market to enterprise. - [Incident.io Pricing Alternative](https://runframe.io/comparisons/incident-io-pricing): Pricing-intent comparison for teams comparing Runframe's bundled on-call and incident lifecycle pricing with incident.io's separate module pricing. - [FireHydrant Alternative](https://runframe.io/comparisons/runframe-vs-firehydrant): Runframe ($15/user/month) vs FireHydrant ($9,600/year for up to 20 responders). FireHydrant acquired by Freshworks. - [Rootly Alternative](https://runframe.io/comparisons/runframe-vs-rootly): Runframe (fixed per-user pricing) vs Rootly (usage-based transparent pricing). Both focus on incident coordination workflows. - [Grafana OnCall Alternative](https://runframe.io/comparisons/runframe-vs-grafana-oncall): Runframe vs Grafana OnCall/Grafana Cloud IRM. Useful for teams comparing managed Slack-native incident response against Grafana-centric alerting and IRM workflows. - [Squadcast Alternative](https://runframe.io/comparisons/runframe-vs-squadcast): Runframe vs Squadcast. Compares Slack-native incident lifecycle, on-call, escalation, and post-incident workflows against Squadcast after the SolarWinds acquisition. - [Incident Management Tools with On-Call](https://runframe.io/comparisons/best-incident-management-tools-with-on-call): Refreshed May 30, 2026 as an 8-tool commercial category page for incident management software with on-call scheduling. Compares Runframe, incident.io, PagerDuty, Rootly, FireHydrant, Squadcast, Grafana Cloud IRM, and Better Stack by alert routing, bundled vs add-on on-call, Slack/Teams workflow, status pages, 15-person monthly cost, startup fit, enterprise fit, Slack-native fit, and Grafana-heavy fit. - [Incident Management Tools for Startups](https://runframe.io/comparisons/best-incident-management-tools-for-startups): Refreshed May 30, 2026 for early and growth-stage engineering teams with fast setup needs, low configuration overhead, on-call scheduling, status pages, and startup pricing. - [Grafana OnCall Alternatives](https://runframe.io/comparisons/grafana-oncall-alternatives): Grafana OnCall OSS was archived March 24, 2026. Compares Grafana Cloud IRM and 5 standalone alternatives with archive/status language, pricing, migration checklist, and common questions. ## Integrations - [Slack](https://runframe.io/integrations/slack): Native Slack incident workflow. `/inc create`, `/inc ask`, `/inc page`, severity changes, status moves, assignment, updates, resolve, and close from Slack. Auto-created incident channels, AI answers from incident context, and timeline capture. - [Datadog](https://runframe.io/integrations/datadog): Paste one webhook URL in Datadog. Monitor alerts create structured incidents with severity, service, and alert context attached. Setup positioned as 2 minutes. - [Prometheus](https://runframe.io/integrations/prometheus): Add a webhook receiver to Alertmanager. Alerts create incidents, labels map to severity/service, and Runframe pages the on-call engineer. - [CloudWatch](https://runframe.io/integrations/cloudwatch): CloudWatch alarms flow through SNS to a Runframe webhook. Incidents are created with alarm context; no AWS credentials shared. - [Generic Webhooks](https://runframe.io/integrations/webhooks): Any tool, script, or pipeline that can POST JSON can create an incident. Supports metadata, secret URL, and deduplication. - [Email Intake](https://runframe.io/integrations/email): Authorized senders can create incidents by email. Subject becomes title, body becomes context, and replies add comments. - [Google Meet](https://runframe.io/integrations/google-meet): One-click incident war rooms. Meet links are attached to the incident record and expire after 24 hours. - [Zoom](https://runframe.io/integrations/zoom): One-click Zoom bridges from incidents. Active bridges can be reused and responders join from the incident record. - [Jira](https://runframe.io/integrations/jira): Create Jira tickets from incidents. Jira status changes flow back to Runframe via webhooks. Supports incident-linked post-incident action items. - [Linear](https://runframe.io/integrations/linear): Create Linear issues from incidents. Linear status changes sync back and post-incident work remains visible from the incident. ## Documentation - [Docs Home](https://runframe.io/docs): Documentation hub for setup, guides, integrations, and API reference. - [Quickstart](https://runframe.io/docs/getting-started/quickstart): Getting started with Runframe. - [Incident Guide](https://runframe.io/docs/guides/incidents): How to create and manage incidents. - [On-Call Guide](https://runframe.io/docs/guides/on-call): How to configure on-call. - [Escalations Guide](https://runframe.io/docs/guides/escalations): How to configure escalation policies. - [Postmortems Guide](https://runframe.io/docs/guides/postmortems): How postmortems work. - [Slash Commands Guide](https://runframe.io/docs/guides/slash-commands): Slack slash command usage. - [API Reference](https://runframe.io/docs/api-reference): API overview. Includes authentication, incidents, on-call, escalation policies, postmortems, services, teams, users, and webhooks. - [Integration Docs](https://runframe.io/docs/integrations): Setup docs for Slack, Datadog, Prometheus, CloudWatch, custom webhooks, email, Google Meet, Jira, Linear, and Zoom. ## Free Tools - [MTTR Calculator](https://runframe.io/tools/mttr-calculator): Enter incident start and end timestamps to calculate mean time to resolution. Formula: MTTR = total resolution time / number of incidents. DORA benchmarks: Elite <1 hour, High 1-24 hours, Medium 1-7 days, Low >6 months. - [Incident Severity Matrix Generator](https://runframe.io/tools/incident-severity-matrix-generator): Interactive builder for SEV0-SEV4 severity matrices. Generates impact criteria, response time expectations, escalation rules, and notification requirements. Customizable for team size and industry. - [On-Call Schedule Generator](https://runframe.io/tools/oncall-builder): Free on-call schedule generator and visual rotation builder for weekly, biweekly, or custom calendars. Supports timezone distribution, primary/secondary responders, holiday coverage, follow-the-sun patterns, and handoff templates. ## Blog ### Build vs Buy - [Build, Open Source, or Buy Incident Management in 2026](https://runframe.io/blog/incident-management-build-or-buy): 3-year TCO for a 20-person team: build from scratch $233K-$395K, open source self-host $99K-$360K, buy commercial $11K-$83K. Building costs 3-8x more than buying. The main cost is a 0.25 FTE senior engineer ($250K-$400K fully-loaded, per Levels.fyi 2025) at $62K-$100K/year in maintenance. AI cut initial build from $19K-$31K to $8K-$15K (1-2 weeks) but initial build is only 3-6% of 3-year cost. Netflix archived Dispatch Sep 2025. Grafana closed-sourced OnCall Mar 2025. Remaining OSS: Incidental (MIT, v0.1.0), incident-bot (MIT, Python/PostgreSQL), IncidentFox (Apache 2.0 core, BSL 1.1 production security). Buy triggers: 8+ on-call, 4+ incidents/month, 3+ teams involved, customer-facing SLAs, compliance requirements. - [Incident Management for Early-Stage Teams](https://runframe.io/blog/incident-management-for-early-stage-teams): Incident management defaults for early-stage engineering teams. Covers severity levels, on-call, escalation, and postmortems in the right order from 15 to 100 engineers. ### AI & Agents - [Your Agent Can Manage Incidents Now](https://runframe.io/blog/your-agent-can-manage-incidents-now): Runframe shipped an MCP server for Claude Code and Cursor. AI agents can manage incidents, check on-call, escalate, page responders, and create postmortems through Runframe. - [Your AI Agent Already Knows Your System Better Than Ours Ever Will](https://runframe.io/blog/your-ai-already-knows-your-system-better-than-ours): Runframe's position on AI incident management: customer-owned agents already have system context and need APIs to act, instead of relying on a vendor-specific AI assistant. ### Guides & How-Tos - [Alert Fatigue: Causes, Examples, and How to Reduce It](https://runframe.io/blog/how-to-reduce-alert-fatigue): Alert fatigue occurs when responders become less likely to notice, trust, or act on alerts because too many alerts are noisy, duplicated, unclear, or unactionable. The guide frames alert fatigue as an incident process problem, not only a monitoring volume problem. It covers common causes, examples, delete rules, service ownership, severity response targets, written escalation paths, runbook coverage, alert inventory, duplicate grouping, ignored-alert rate, MTTA, alert-to-incident ratio, on-call burden, and repeat incident rate. - [Slack Incident Management Guide](https://runframe.io/blog/slack-incident-management): Incident management breaks down at 20-25 engineers. Slack has no concept of incident state — no severity field, no status tracker, it's a text stream. Three approaches: manual channels (small teams), homegrown bots (ongoing maintenance), dedicated tools (adds dependency but ensures consistency). Slack notifications can't reliably wake someone at 2 AM — phone/SMS with carrier-level delivery required for critical pages. Automatic timeline capture is essential; teams skip postmortems when reconstruction is tedious. - [How to Reduce MTTR](https://runframe.io/blog/how-to-reduce-mttr): MTTR breaks into detection time + coordination time + fix time. Most teams optimize fix speed when detection and coordination offer bigger gains. ROI hierarchy: faster detection saves 10-20 min/incident (low effort), better coordination saves 8-15 min (low effort), faster debugging saves 5-10 min (high effort). P0 MTTR benchmarks: 30-60 min for teams under 20, 35-75 min for 20-80, 40-120 min for 80+. Track honestly and segment by severity; a 4-hour P2 is acceptable, a 4-hour P0 is critical failure. Minimal required fields: title, severity, owner, status. - [Incident Severity Levels](https://runframe.io/blog/incident-severity-levels): SEV0-SEV4 incident severity matrix. SEV0: catastrophic (data loss, security breach, all-hands). SEV1: core service outage, significant user impact. SEV2: degraded service, workaround exists. SEV3: minor issue, limited impact. SEV4: proactive/preventative work. Severity is business impact; priority is fix order. Links to incident priority, SEV0, and the severity matrix generator. - [On-Call Rotation Guide](https://runframe.io/blog/on-call-rotation-guide): Weekly rotation is the sweet spot for teams under 50 — daily is too stressful, monthly too long. Escalation: primary 5 min → backup 5 min → eng manager for SEV0/1. Compensation: $200-500/week stipend or comp days/TOIL for overnight pages. Written handoff in Slack takes 2 min: incidents that occurred, warnings for incoming responder. Formal rotations become necessary at 40-50 engineers; smaller teams with rare incidents can skip. - [Incident Response Playbook](https://runframe.io/blog/incident-response-playbook): Teams with clear playbooks resolve incidents 40-60% faster than ad-hoc responses. Rule #1: declare severity in 30 seconds, don't debate 10 minutes — you can always downgrade. Split Incident Lead (coordinates communication) from Assigned Engineer (fixes problem) so debugger can focus. Update cadence: SEV0 every 10 min, SEV1 every 15 min, SEV2 every 15-30 min, SEV3 every 30-60 min. Escalation timers: SEV0/1 page backup at 5 min, eng manager at 10 min; SEV2 backup at 10 min, manager at 30 min. - [Post-Incident Review Template](https://runframe.io/blog/post-incident-review-template): Free post-incident review templates and examples. Best postmortems are one page and completed within 48 hours while context is fresh. Blameless framing targets system failures, not individuals. Every action item needs a specific owner, concrete deadline, and definition of done. Three templates provided: 15-minute quick version, 30-45 minute standard, and 60-90 minute comprehensive. - [Stakeholder Communication Templates](https://runframe.io/blog/incident-stakeholder-communication-templates): Core principle: one owner, one source of truth, consistent cadence. The Incident Commander handles all outbound updates; engineers focus on fixes. Update frequency: SEV0 every 15 min, SEV1 every 30-60 min, SEV2 every 60-120 min. Always include next update timestamp. Describe customer symptoms ("checkout failing with payment errors") not internals ("database replication lag on shard 3"). "Unknown at this time" with committed next-update timing beats fabricated ETAs. Eight template categories: status page, customer email, executive summary, support script, sales/CSM note, internal engineering update, social media response, post-incident closure. ### Concepts & Comparisons - [SLA vs SLO vs SLI](https://runframe.io/blog/sla-vs-slo-vs-sli): SLI = the actual metric you measure (error rate, latency — things customers notice). SLO = your internal reliability target (e.g., 99.7%). SLA = your external contractual promise with consequences (e.g., 99.5% with credits). Buffer strategy: maintain SLO above SLA (99.7% internal / 99.5% external = 0.2% buffer for unexpected issues). Error budget formula: 100% - SLO target. For 99.5% SLO: 0.5% budget = 216 minutes/month = 3.6 hours allowed downtime. Set SLO slightly below current performance, not aspirationally. Cost of nines increases exponentially: 99.9% (8.77 hrs/year downtime) to 99.99% (52 min/year) requires order-of-magnitude infrastructure investment. - [Runbook vs Playbook](https://runframe.io/blog/runbook-vs-playbook): Runbooks document technical execution — step-by-step commands and procedures. Answer: "how do I fix this?" Playbooks document roles, escalation, and communication. Answer: "who handles this?" Runbooks consulted during investigation/remediation phases. Playbooks guide entire incident lifecycle from declaration through resolution. Runbooks update when infrastructure changes. Playbooks update when team structure/processes change. Build playbooks first — they address coordination overhead. Runbooks follow later for recurring failure scenarios. - [Incident Management vs Incident Response](https://runframe.io/blog/incident-management-vs-incident-response): Incident response = tactical work during active incidents (declare, coordinate, fix, communicate). Incident management = ongoing strategic lifecycle (postmortems, runbooks, on-call, trend analysis). Response measures MTTR. Management measures repeat-incident rate, action-item closure rate, and MTTD. MTTR can improve while reliability worsens if recurrence stays high — a company can have 42-minute MTTR yet experience the same database outage every quarter. Strong response + weak management = reactive cycles. Strong management + weak response = great plans failing during crises. Need both. - [Best PagerDuty Alternatives 2026](https://runframe.io/blog/best-pagerduty-alternatives): Refreshed May 30, 2026 around current search demand. Best alternatives are Runframe, incident.io, Rootly, Grafana Cloud IRM, Better Stack, and FireHydrant. Runframe $15/user/month ($12 annual, free tier) - Slack-native incident management with on-call included, no add-on fees, and a focused rollout path for teams that want the core incident lifecycle without enterprise overhead. Includes sections for Slack-integrated alternatives, startup vs enterprise fit, on-call inclusion, pricing, and PagerDuty vs Rootly / FireHydrant / incident.io / OpsGenie comparisons. - [Best OpsGenie Alternatives 2026](https://runframe.io/blog/best-opsgenie-alternatives): Alternatives teams actually switch to before the April 5, 2027 OpsGenie shutdown. Compares pricing, features, and migration options. ### Strategy & Trends - [Scaling Incident Management](https://runframe.io/blog/scaling-incident-management): Four predictable stages: Stage 1 single Slack channel (5-15 people), Stage 2 Python scripts (15-40), Stage 3 "should buy a tool" limbo (40-100 — most stuck 6-12 months until a crisis forces Stage 4), Stage 4 formal tool (100+). The 40-50 person inflection point: informal "whoever's around" coordination fails, formal on-call with primary/backup becomes necessary. Setup complexity (not cost) is the real barrier — enterprise tools require decisions teams haven't made yet. Teams want reasonable defaults (3 severity levels, 5-min escalation, auto-channel creation) that work with zero configuration. - [Engineering Productivity & Incident Management](https://runframe.io/blog/engineering-productivity-incident-management): Tool-switching during incidents breaks focus and compounds MTTR. Dedicated incident threads work best for 20-100 person teams with ~10 min overhead/incident. Centralize all status, decisions, and handoffs in one location (typically Slack). Incident coordination is the harder problem — the technical fix is usually straightforward, getting everyone aligned is harder. - [OpsGenie Migration Guide](https://runframe.io/blog/opsgenie-migration-guide): OpsGenie fully shuts down April 5, 2027. New sales stopped June 4, 2025. Migration timeline: 4-8 weeks basic, 8-16 weeks for 20+ integrations — initial estimates typically underestimate by 2-3x. CSV schedule exports don't import cleanly; most teams rebuild manually (1-2 weeks). Atlassian's path: JSM (ITSM-heavy) or Compass (engineering-focused). Success pattern: run parallel systems 4-8 weeks, train before cutover, audit integrations early, budget 2-3x initial cost estimates. - [OpsGenie End of Support Guide](https://runframe.io/blog/opsgenie-shutdown-guide): OpsGenie support ends April 5, 2027. Explains the end-of-life timeline, new-sales cutoff, Atlassian JSM/Compass migration paths, third-party alternatives, and what teams should do this week. Unique retention detail: OpsGenie Enterprise had effectively unlimited alert data retention, while JSM after migration retains alert data for 1 month on Free, 1 year on Standard, and 3 years on Premium. - [State of Incident Management 2025](https://runframe.io/blog/state-of-incident-management-2025): Despite 51% AI deployment (86% expected by 2027), operational toil rose to 30% from 25% — first increase in five years. 78% of developers spend 30%+ of time on manual toil (~$9.4M annual lost productivity per 250-engineer team). 73% of organizations experienced outages from ignored alerts; ~67% of daily alerts disregarded. Market consolidation: OpsGenie shutting down Apr 2027, Freshworks acquired FireHydrant, SolarWinds acquired Squadcast. High-impact IT outages cost ~$2M/hour; median $76M annually from unplanned downtime. ## Learn — Reference Encyclopedia ### Core Metrics - [MTTR — Mean Time to Resolution](https://runframe.io/learn/mttr): The average time to fully resolve an incident from detection to service restoration. Formula: MTTR = total resolution time / number of incidents. Benchmarks: Excellent <30 min (top 5%), Good <1 hour, Average 1-24 hours. Elite DORA performers achieve <1 hour. Anything >24 hours = Low performer. - [MTTA — Mean Time to Acknowledge](https://runframe.io/learn/mtta): Average time from alert firing to human acknowledgment. Formula: MTTA = total acknowledgment times / number of incidents. Benchmarks: Excellent <1 min (top 5%), Good <5 min. Anything >15 min implies broken paging or on-call processes. - [MTTD — Mean Time to Detect](https://runframe.io/learn/mttd): Average time from issue occurring to alert firing. Formula: MTTD = total detection times / number of incidents. Benchmark: <1 min is gold standard. Practically always some lag (e.g., 30-second polling interval). - [MTBF — Mean Time Between Failures](https://runframe.io/learn/mtbf): Average time between system failures. Formula: MTBF = total uptime / number of failures. Benchmark: Excellent >720 hours (30 days) at top 5%. Scheduled maintenance does not count — MTBF measures unexpected outages only. - [SLO — Service Level Objective](https://runframe.io/learn/slo): Internal reliability target, typically a percentage over a time period. Formula: SLO = (successful requests / total requests) × 100%. Product Manager (representing user) and Engineering Lead (representing reality) must agree on the target. Set slightly below current performance, not aspirationally. - [SLI — Service Level Indicator](https://runframe.io/learn/sli): A measurable metric indicating service performance, used to track SLOs. Formula: SLI = (good events / total events) × 100%. Keep simple: 1-3 SLIs per user journey (login, checkout, search). - [SLA — Service Level Agreement](https://runframe.io/learn/sla): Contractual commitment to customers on service performance/availability. Formula: SLA = uptime target (e.g., 99.9%). Usually only for paid enterprise plans. Free users rarely get SLAs. Always set SLO higher than SLA for buffer. - [Error Budget](https://runframe.io/learn/error-budget): The allowed unreliability before violating the SLO. Formula: Error budget = 100% - SLO. For 99.5% SLO: 0.5% budget = 216 minutes/month. Resets monthly or quarterly (rolling window). When exhausted, freeze feature releases and prioritize reliability. ### Incident Severity - [Incident Severity Matrix](https://runframe.io/learn/incident-severity-matrix): Framework for classifying incidents by impact and urgency. Use SEV0-SEV4 or P0-P4 — doesn't matter which, just pick one standard. SEV0 means "absolute critical emergency." Google uses P0-P4; many DevOps teams use SEV1-SEV5. - [Severity 0 — Critical](https://runframe.io/learn/severity-0): System completely unusable or data integrity at risk. All-hands response. Anyone can declare a SEV0 — better to false alarm than delay. - [Severity 1 — High](https://runframe.io/learn/severity-1): Major functionality broken or degraded, significant user impact, workaround may exist. Performance counts as SEV1 if >3s latency for 50%+ of users. - [Severity 2 — Medium](https://runframe.io/learn/severity-2): Degraded performance, minor functionality broken for some users, workarounds available. Page if daytime; email/ticket if night (team-dependent). - [Severity 3 — Low](https://runframe.io/learn/severity-3): Minor bug or non-critical issue. No immediate user impact. Do not page anyone. - [Severity 4 — Trivial](https://runframe.io/learn/severity-4): Cosmetic issues, typos, internal-only problems. Zero user impact. Track to quantify "polish" work and prioritize quality weeks. - [Incident Priority](https://runframe.io/learn/incident-priority): Priority determines fix order (changes with context); severity determines business impact (stays fixed). P0 and SEV0 are the same concept — different notation. Both mean "drop everything." ### Incident Roles - [Incident Commander](https://runframe.io/learn/incident-commander): Leads coordination and decision-making during an incident. It's a role, not a rank — a junior engineer can be IC for a VP. IC decides strategy; SMEs handle tactics. - [Communication Lead](https://runframe.io/learn/communication-lead): Manages all internal/external communication during incidents. Not needed for every incident — IC handles SEV3/4 alone. Dedicated Comms Lead essential for SEV0/1. - [Incident Scribe](https://runframe.io/learn/incident-scribe): Documents timeline, actions, and decisions in real-time. Cannot also fix things — if the Scribe starts coding, they stop writing and documentation gaps appear immediately. - [Subject Matter Expert](https://runframe.io/learn/subject-matter-expert): Technical specialist for the specific service or component causing the incident. IC can override SME on strategy ("prioritize data safety over uptime") but defers on tactics ("how to query the DB"). ### Incident Process - [Incident Triage](https://runframe.io/learn/incident-triage): Initial phase where severity, impact, and required expertise are determined. Severity is the classification (SEV1, SEV2); triage is the process of determining that classification. - [Incident Lifecycle](https://runframe.io/learn/incident-lifecycle): End-to-end journey from occurrence through post-incident review. Detection is the most critical stage — you can't fix what you don't know is broken. Improving MTTD often has the highest ROI for reducing overall MTTR. - [Incident Resolution](https://runframe.io/learn/incident-resolution): Point where service is restored to full functionality. Resolution ≠ mitigation: resolution restores normal operation, mitigation reduces impact but may not fully restore. A workaround is temporary mitigation. - [Incident Automation](https://runframe.io/learn/incident-automation): Using technology to reduce manual intervention. Auto-remediation (e.g., auto-restart) can create fail loops if not careful. Always have a kill switch. - [Post-Incident Review](https://runframe.io/learn/post-incident-review): Meeting to analyze what happened and identify improvements. Required for SEV0/1, recommended for SEV2, skip for SEV3. - [Blameless Postmortem](https://runframe.io/learn/blameless-postmortem): Post-incident analysis focused on system/process failures, not individual blame. Blameless ≠ no consequences — malice or negligence is an HR issue. But 99% of incidents are honest mistakes in bad systems. - [Root Cause Analysis](https://runframe.io/learn/root-cause-analysis): Systematic method for identifying underlying causes. Techniques: 5 Whys, fishbone diagram, fault tree analysis. Rarely one root cause in complex systems (Swiss Cheese Model) — usually a combination of factors. ### Documentation & Planning - [Runbook](https://runframe.io/learn/runbook): Step-by-step guide for specific operational tasks. Keep to 1 page — no one reads longer during an outage. Update when infrastructure changes. - [Playbook](https://runframe.io/learn/playbook): Comprehensive strategies and procedures for handling scenarios. For high-stakes, complex scenarios where bad decisions are costly. Not needed for everything. - [Escalation Policy](https://runframe.io/learn/escalation-policy): Predefined rules for when/how to escalate. Rule of thumb: if stuck 15-30 minutes, or if severity increases (SEV2 → SEV1), escalate. ### On-Call Management - [On-Call](https://runframe.io/learn/on-call): Being available to respond to urgent issues outside standard hours. On-call ≠ on-duty: on-call means available to be contacted, on-duty means actively working. On-call engineers aren't working unless an incident occurs. - [On-Call Rotation](https://runframe.io/learn/on-call-rotation): Schedule determining alert responsibility per time period. Ideal frequency: no more than 1 week every 6-8 weeks. More frequent leads to fatigue. - [On-Call Schedule](https://runframe.io/learn/on-call-schedule): Concrete timetable of who is on-call when. No "best" schedule — depends on team size, alert volume, timezone distribution. Start weekly and adjust. - [On-Call Responder](https://runframe.io/learn/on-call-responder): Engineer currently designated to receive alerts. Responsibilities: acknowledge within SLA (typically 5-15 min), triage severity, mitigate or escalate, document actions, participate in postmortems. Goal is rapid response, not solo heroics. - [On-Call Handoff](https://runframe.io/learn/on-call-handoff): Structured transfer of context between on-call engineers. Include: active incidents with status, muted alerts with reasons, ongoing investigations, pending items, contact info for key stakeholders, recent deployments/changes. - [Follow-the-Sun](https://runframe.io/learn/follow-the-sun): Global on-call model using timezone distribution to avoid night shifts. Minimum: 3 teams in different timezones (8-hour separation), each with 3-5 engineers. Total minimum: 9-15 engineers. - [On-Call Burnout](https://runframe.io/learn/on-call-burnout): Exhaustion from frequent sleep interruption, excessive alerts, and availability stress. Early signs: dreading on-call week, cynicism about alerts, slower response, irritability, insomnia. - [Sustainable On-Call](https://runframe.io/learn/sustainable-on-call): On-call practices prioritizing human health alongside reliability. Requirements: predictable schedules, fair distribution, low alert noise, compensation, clear escalation paths, team input into process. - [Fair Rotations](https://runframe.io/learn/fair-rotations): Equitable distribution of on-call burden. Track: total on-call days, holidays covered, incidents handled, nights awoken, total pages. Use weighted fairness score. Review quarterly. - [Load Distribution](https://runframe.io/learn/load-distribution): Equalizing on-call effort across team. Multiple metrics needed: total incidents, severity score, nights awoken, hours, holidays. Combine into single "burden score." ### Observability & Reliability - [Monitoring](https://runframe.io/learn/monitoring): Collecting, analyzing, and alerting on system health data. Monitoring is for known problems ("Is CPU high?"). Observability is for unknown problems ("Why is checkout slow only for iOS users in Germany?"). - [Observability](https://runframe.io/learn/observability): How well internal system states can be inferred from external outputs (logs, metrics, traces). It's a practice, not a tool. Start with structured logs. Tools like Honeycomb, Datadog, Jaeger help but aren't required. - [Availability](https://runframe.io/learn/availability): Proportion of time a system is operational. Formula: Availability = uptime / (uptime + downtime). Five nines (99.999%) = <5 min/year downtime. 99.9% = 8.77 hours/year. For most SaaS, 99.9% is adequate — users tolerate ~45 min maintenance/month. - [Reliability](https://runframe.io/learn/reliability): Probability a system functions correctly under stated conditions for a specified period. Improve by: measuring with SLOs, then testing failure modes with chaos engineering. - [Toil](https://runframe.io/learn/toil): Manual, repetitive, automatable operational work. Not all ops work is toil — responding to a novel 3 AM outage is operational work but not toil (it's not repetitive/predictable). Toil should be systematically reduced through automation. - [Four Golden Signals](https://runframe.io/learn/four-golden-signals): Google SRE's core monitoring signals: Latency, Traffic, Errors, Saturation. Most important: Errors — if users see errors, nothing else matters, fix that first. - [Status Page](https://runframe.io/learn/status-page): Public/private dashboard communicating service health to users. Host on a third party (Atlassian Statuspage, BetterStack) so it stays up when you go down. - [Alert Fatigue](https://runframe.io/learn/alert-fatigue): Desensitization from excessive/low-quality alerts causing missed critical alerts. Threshold: >1-2 pages per 12-hour shift creates fatigue. On-call should be silent unless things are actually broken. Deep dive: [Alert Fatigue: Causes, Examples, and How to Reduce It](https://runframe.io/blog/how-to-reduce-alert-fatigue). ### Operations & Testing - [War Room](https://runframe.io/learn/war-room): Dedicated space (physical or virtual) for coordinating major incident response. Stakeholders can listen (muted) but must not interrupt — IC can remove them if they do. - [Downtime](https://runframe.io/learn/downtime): Period when system is unavailable. Formula: Downtime = time restored - time detected. Benchmark: Excellent <5 min/month (top 5%). Cost calculation: (revenue/hour × hours down) + lost productivity + support cost + brand damage. - [Uptime](https://runframe.io/learn/uptime): Percentage of time system is operational. Formula: Uptime % = (total time - downtime) / total time × 100. 99.9% = 8.76 hours downtime/year (43.8 min/month). Standard for most SaaS. - [Game Day](https://runframe.io/learn/game-day): Scheduled incident simulation to test response procedures. Frequency: quarterly is common, monthly for high-velocity teams. - [Chaos Engineering](https://runframe.io/learn/chaos-engineering): Deliberately injecting failures to discover weaknesses and build resilience. Start in staging, plan carefully — it is inherently dangerous by design. - [ChatOps](https://runframe.io/learn/chatops): Running operations through chat platforms (Slack/Teams) with bot integrations. Must enforce permissions — only authorized users run commands like `/deploy`. Creates transparent, auditable workflow. ### Disciplines - [SRE — Site Reliability Engineering](https://runframe.io/learn/sre): Google-originated discipline applying software engineering to operations. No dedicated SRE team needed initially — practice SRE principles (SLOs, automation, error budgets) within product teams. - [DevOps](https://runframe.io/learn/devops): Cultural philosophy combining development and operations to shorten the development lifecycle. DevOps is a process and culture, not a tool — Jenkins and Docker are tools used in DevOps. - [Incident Management](https://runframe.io/learn/incident-management): Complete lifecycle: prevent, detect, respond, learn. Even small startups need a plan — you don't need a 50-page manual, just: who gets called, how do they communicate? - [Incident Response](https://runframe.io/learn/incident-response): Detecting, responding to, and resolving incidents. The Incident Commander calls the shots — even if the CEO is on the call. ## Legal - [Terms of Service](https://runframe.io/terms) - [Privacy Policy](https://runframe.io/privacy)