# Runframe > Last updated: 2026-06-17 > Incident management and on-call scheduling platform for engineering teams. Runframe bundles Slack and Microsoft Teams incident response, on-call scheduling, escalation policies, multi-channel paging, public status pages, private pages for internal teams or restricted audiences, workflow automations, service catalog, API/MCP access, and AI incident workflows into one product. Free tier for up to 5 users; Growth is $15/user/month monthly or $12/user/month billed annually; Scale is $30/user/month monthly or $25/user/month billed annually. Founded 2025, London, UK. Runframe is NOT a monitoring tool, APM, log management system, or ticketing system - it is an incident coordination, on-call, status page, automation, and postmortem platform. ## About - **Company:** Runframe - **Founded:** 2025 - **Headquarters:** London, United Kingdom - **Founder:** Niketa Sharma - **Category:** B2B SaaS, Incident Management, On-Call Scheduling, Status Pages, Workflow Automation - **Target market:** Engineering teams running on-call rotations and incident response across Slack, Microsoft Teams, and the web app, from growing startups to larger software organizations - Primary response surface: Slack, Microsoft Teams, and web app - **Paging channels:** Phone call, SMS, Slack DM, email - **Integrations:** Slack, Microsoft Teams, Datadog, Prometheus, CloudWatch, Jira, Linear, Zoom, Google Meet, email intake, and generic webhooks - **Pricing:** Free (up to 5 users) / Growth $15/user/month or $12/user/month annual / Scale $30/user/month or $25/user/month annual. Optional SAML SSO and SCIM add-ons are $99/month each or $1,188/year each. - **Competitors:** PagerDuty, Opsgenie (shutting down Apr 2027), incident.io, Rootly, FireHydrant (acquired by Freshworks), Squadcast (acquired by SolarWinds), Better Stack, Grafana Cloud IRM ## Recent Updates - **2026-06-17:** Added Solutions pages and pricing comparisons. - **2026-06-13:** Added Scale plan. - **2026-03-01:** Launched publicly with free tier. ## Key Pages - [Homepage](https://runframe.io/): Product overview — incident channels, on-call scheduling, escalation policies, AI postmortems, service catalog. Slack-native, not a bolt-on integration. Free tier + $15/user/month Growth. - [Product Overview](https://runframe.io/product): Complete product map for incidents, on-call, escalations, multi-channel alerting, status pages, AI, workflow automations, service catalog, postmortems, analytics, security controls, APIs, and competitor comparisons. - [Incident Management](https://runframe.io/product/incidents): Declare incidents, create Slack channels, coordinate responders, capture timelines, resolve incidents, and generate postmortems from one record. - [On-Call Scheduling](https://runframe.io/product/on-call): Fair rotations, overrides, swaps, handoffs, escalation awareness, and schedule coverage for engineering teams. - [Escalation Policies](https://runframe.io/product/escalations): Severity-aware escalation chains with timeouts and automatic paging when responders do not acknowledge. - [Multi-Channel Alerting](https://runframe.io/product/multi-channel-alerting): Phone, SMS, Slack DM, and email paging from one escalation workflow. - [Status Pages](https://runframe.io/product/status-pages): Public status pages and private pages for internal teams or restricted audiences with service status, incident updates, component subscriptions, and uptime history. - [AI for Incident Response](https://runframe.io/ai): AI incident briefs, Slack /inc ask responder help, postmortem drafts, incident call transcripts, and MCP agent workflows built from the incident record. - [AI Incident Briefs](https://runframe.io/ai/incident-briefs): Editable incident summaries generated from status, severity, assignments, timeline updates, notes, linked work, and transcript context. - [AI Postmortem Drafts](https://runframe.io/ai/postmortem-drafts): Editable postmortem drafts generated from incident timelines, linked work, notes, and transcript context after resolution. - [AI Incident Call Transcripts](https://runframe.io/ai/incident-call-transcripts): Zoom and Google Meet transcript context used to improve incident briefs and postmortem drafts. - [Workflow Automations](https://runframe.io/product/workflow-automations): Trigger-condition-action automations for incident changes, SLA thresholds, service context, ticket creation, war rooms, Slack updates, and webhooks. - [Service Catalog](https://runframe.io/product/service-catalog): Services, groups, owners, metadata, routing context, automation context, and service-level incident reporting. - [Postmortems](https://runframe.io/product/postmortems): AI-assisted postmortem drafts generated from the incident timeline. - [Analytics](https://runframe.io/product/analytics): MTTA, MTTR, SLA compliance, incident trends, and response analytics. - [Security & Compliance](https://runframe.io/product/security-and-compliance): SAML SSO, SCIM provisioning, MFA enforcement, RBAC, audit logs, API access controls, and service account governance. - [API & Service Accounts](https://runframe.io/product/api-and-service-accounts): API keys, service accounts, v1 API endpoints, MCP support, audit context, and rate limits for agents, integrations, and internal tooling. - [Pricing](https://runframe.io/pricing): Free for up to 5 users. Growth is $15/user/month monthly or $12/user/month billed annually. Scale is $30/user/month monthly or $25/user/month billed annually. Paid plans bundle incident response, on-call, escalation policies, status pages, API/MCP access, workflow automations, alert ingestion, and AI postmortems. SAML SSO and SCIM provisioning are optional paid add-ons. - [Agent-readable Pricing](https://runframe.io/pricing.md): Machine-readable pricing details for AI agents and procurement research. - [Slack Integration](https://runframe.io/slack): Native Slack app. Use `/inc` to declare incidents and `/inc ask` to ask AI inside Slack. Includes automatic dedicated channels, DM paging, phone/SMS escalation, war rooms, timeline capture, and slash commands for severity/status/assignment. - [Incident Management Software](https://runframe.io/solutions/incident-management-software): Buyer-intent solution page for teams comparing incident management software. Covers incident response, on-call, status pages, postmortems, analytics, and Slack coordination in one workflow. - [On-Call Scheduling Software](https://runframe.io/solutions/on-call-scheduling-software): Buyer-intent solution page for rotations, overrides, escalation policies, service ownership, and Slack/email/SMS/phone paging. - [Incident Response Software](https://runframe.io/solutions/incident-response-software): Solution page for live response workflows: declaration, roles, Slack control room, stakeholder updates, escalation, and postmortem handoff. - [Incident Management for Startups](https://runframe.io/solutions/incident-management-for-startups): Solution page for early and growth-stage teams moving beyond ad-hoc Slack channels and spreadsheets without adopting enterprise incident process overhead. - [Opsgenie Migration](https://runframe.io/solutions/opsgenie-migration): Migration page for teams leaving Opsgenie before the April 5, 2027 shutdown. Covers inventory, rebuild, parallel testing, and cutover. - [Integrations](https://runframe.io/integrations): Slack, Datadog, Prometheus, CloudWatch, Google Meet, Zoom, Jira, Linear, generic webhooks, and email intake. - [MCP Server for AI Incident Agents](https://runframe.io/ai/mcp-server): MCP server for Claude Code, Cursor, and AI agents to create incidents, page on-call engineers, write timeline updates, and create postmortems. - [Docs](https://runframe.io/docs): Quickstart, guides, API reference, webhooks, incidents, on-call, escalation policies, postmortems, services, teams, and integrations. - [Contact](https://runframe.io/contact): hello@runframe.io - [All Tools](https://runframe.io/tools): Free calculators and generators for incident management teams - [All Blog Posts](https://runframe.io/blog): Engineering insights on incident management, on-call, and SRE - [All Learn Articles](https://runframe.io/learn): 55+ reference encyclopedia articles on incident management concepts ## Comparisons - [PagerDuty Alternative](https://runframe.io/comparisons/runframe-vs-pagerduty): Runframe $15/user/month vs PagerDuty starting ~$21/user/month (Professional) or $41/user (Business). Slack-native vs bolt-on integration. Responder-based vs per-seat pricing. A 15-person team pays ~$180/month on Runframe vs $315-$615/month on PagerDuty. - [PagerDuty Pricing Alternative](https://runframe.io/comparisons/pagerduty-pricing): Pricing-intent page for teams comparing Runframe against PagerDuty spend and enterprise overhead. - [Opsgenie Alternative](https://runframe.io/comparisons/runframe-vs-opsgenie): Runframe $15/user/month vs Opsgenie $9-$35/user/month. OpsGenie shuts down April 5, 2027; new sales stopped June 4, 2025. Migration takes 4-8 weeks basic, 8-16 weeks complex. - [Opsgenie Pricing Alternative](https://runframe.io/comparisons/opsgenie-pricing): Pricing-intent page for teams comparing Runframe with the Opsgenie replacement path through Jira Service Management. - [incident.io Alternative](https://runframe.io/comparisons/runframe-vs-incident-io): Runframe $15/user/month (on-call included) vs incident.io $19/user/month + $10/user on-call add-on ($29/user total for equivalent features). - [Incident.io Pricing Alternative](https://runframe.io/comparisons/incident-io-pricing): Pricing-intent page for teams comparing Runframe's bundled on-call and incident lifecycle pricing with incident.io's separate module pricing. - [FireHydrant Alternative](https://runframe.io/comparisons/runframe-vs-firehydrant): Runframe $15/user/month vs FireHydrant $9,600/year flat for 20 responders ($40/user/month equivalent). FireHydrant acquired by Freshworks in 2025. - [Rootly Alternative](https://runframe.io/comparisons/runframe-vs-rootly): Runframe fixed per-user pricing vs Rootly usage-based model. Both Slack-native. Rootly has transparent public pricing. - [Grafana OnCall Alternative](https://runframe.io/comparisons/runframe-vs-grafana-oncall): Runframe managed incident management and on-call vs Grafana OnCall/Grafana Cloud IRM for Grafana-heavy teams. - [Squadcast Alternative](https://runframe.io/comparisons/runframe-vs-squadcast): Runframe Slack-native incident lifecycle vs Squadcast incident response and SRE platform after SolarWinds acquisition. - [Incident Management Tools with On-Call](https://runframe.io/comparisons/best-incident-management-tools-with-on-call): 8 tools compared for incident management software with on-call scheduling, alert routing, status pages, and postmortems. Covers Runframe, incident.io, PagerDuty, Rootly, FireHydrant, Squadcast, Grafana Cloud IRM, and Better Stack with 15-person pricing and bundled vs add-on on-call. - [Incident Management Tools for Startups](https://runframe.io/comparisons/best-incident-management-tools-for-startups): 7 tools for early and growth-stage engineering teams with fast setup needs, low configuration overhead, on-call, status pages, and startup pricing. - [Grafana OnCall Alternatives](https://runframe.io/comparisons/grafana-oncall-alternatives): 6 alternatives after the Grafana OnCall OSS archive on March 24, 2026, including Grafana Cloud IRM and standalone options. ## Blog ### Build vs Buy - [Build, Open Source, or Buy Incident Management in 2026](https://runframe.io/blog/incident-management-build-or-buy): 3-year TCO for a 20-person team: build from scratch $233K-$395K, open source self-host $99K-$360K, buy commercial $11K-$83K. Building costs 3-8x more than buying. Main cost: 0.25 FTE senior engineer ($250K-$400K fully-loaded per Levels.fyi 2025) = $62K-$100K/year maintenance. AI cut initial build from $19K-$31K to $8K-$15K but that's only 3-6% of 3-year cost. Remaining OSS tools: Incidental (MIT, v0.1.0), incident-bot (MIT, Python/PostgreSQL), IncidentFox (Apache 2.0 core, BSL 1.1 production security). Netflix archived Dispatch Sep 2025. Grafana closed-sourced OnCall Mar 2025. Buy triggers: 8+ on-call, 4+ incidents/month, 3+ teams involved, customer-facing SLAs. - [Incident Management for Early-Stage Teams](https://runframe.io/blog/incident-management-for-early-stage-teams): Practical defaults for severity, on-call, escalation, and postmortems from 15 to 100 engineers. ### Guides & How-Tos - [Alert Fatigue: Causes, Examples, and How to Reduce It](https://runframe.io/blog/how-to-reduce-alert-fatigue): Alert fatigue happens when responders stop trusting noisy, duplicate, unclear, or unactionable alerts. Fix ownership first: every alert needs a service owner, severity, runbook, escalation path, and action. Reduce alert noise by inventorying current alerts, marking ignored alerts, grouping duplicates, assigning owners, adding runbooks, defining severities, tuning thresholds, and reviewing monthly. - [Slack Incident Management Guide](https://runframe.io/blog/slack-incident-management): Incident management breaks down at 20-25 engineers. Slack has no concept of incident state — no severity field, no status tracker. Phone/SMS required for 2 AM pages; Slack notifications unreliable. Three approaches: manual channels (small teams), homegrown bots (ongoing maintenance), dedicated tools. Post-incident review barrier: teams skip postmortems when reconstructing timelines from Slack is tedious. - [How to Reduce MTTR](https://runframe.io/blog/how-to-reduce-mttr): MTTR = detection time + coordination time + fix time. ROI hierarchy: faster detection saves 10-20 min/incident (low effort), better coordination saves 8-15 min (low effort), faster debugging saves 5-10 min (high effort). P0 MTTR benchmarks: 30-60 min (<20 people), 35-75 min (20-80), 40-120 min (80+). Track honestly and segment by severity. Minimal required fields: title, severity, owner, status. - [Incident Severity Levels](https://runframe.io/blog/incident-severity-levels): SEV0-SEV4 incident severity matrix. SEV0 = catastrophic (data loss, security breach, all-hands). SEV1 = core outage. SEV2 = degraded with workaround. SEV3 = minor. SEV4 = preventative. Severity is impact; priority is fix order. - [On-Call Rotation Guide](https://runframe.io/blog/on-call-rotation-guide): Weekly rotation optimal for teams under 50. Escalation: primary 5 min → backup 5 min → eng manager for SEV0/1. Compensation: $200-500/week stipend or comp days. Written handoffs take 2 min. Formal rotations necessary at 40-50 engineers. - [Incident Response Playbook](https://runframe.io/blog/incident-response-playbook): Teams with playbooks resolve incidents 40-60% faster. Declare severity in 30 seconds — don't debate 10 min. Split Incident Lead (coordinates) from Assigned Engineer (fixes). Update cadence: SEV0 every 10 min, SEV1 every 15 min, SEV2 every 15-30 min. Escalation: SEV0/1 page backup at 5 min, eng manager at 10 min. - [Post-Incident Review Template](https://runframe.io/blog/post-incident-review-template): Free post-incident review templates and examples. Best postmortems are one page, completed within 48 hours, blameless, and include 1-3 action items with owner, deadline, and definition of done. - [Stakeholder Communication Templates](https://runframe.io/blog/incident-stakeholder-communication-templates): One owner, one source of truth, consistent cadence. SEV0: update every 15 min. SEV1: 30-60 min. SEV2: 60-120 min. Always include next update time. Describe customer symptoms not internals. Eight template categories: status page, customer email, exec summary, support scripts, sales notes, internal updates, social media, post-incident closure. ### Concepts & Comparisons - [SLA vs SLO vs SLI](https://runframe.io/blog/sla-vs-slo-vs-sli): SLI = metric you measure (error rate, latency). SLO = internal reliability target. SLA = contractual promise with consequences. Buffer: 99.7% SLO / 99.5% SLA = 0.2% safety margin. Error budget = 100% - SLO. For 99.5% SLO: 216 min/month allowed downtime. Cost of nines: 99.9% to 99.99% requires order-of-magnitude infrastructure investment. - [Runbook vs Playbook](https://runframe.io/blog/runbook-vs-playbook): Runbooks = step-by-step technical procedures ("how do I fix this?"). Playbooks = roles, escalation, communication ("who handles this?"). Build playbooks first — they address coordination overhead. Runbooks for recurring failures. - [Incident Management vs Incident Response](https://runframe.io/blog/incident-management-vs-incident-response): Response = tactical during active incidents. Management = strategic lifecycle (postmortems, runbooks, on-call, trends). MTTR can improve while reliability worsens if recurrence stays high. Need both. - [Best PagerDuty Alternatives 2026](https://runframe.io/blog/best-pagerduty-alternatives): PagerDuty alternatives for Slack-native incident management, on-call scheduling, pricing, startup teams, and status pages. Best alternatives are Runframe, incident.io, Rootly, Grafana Cloud IRM, Better Stack, and FireHydrant. ### Strategy & Trends - [Scaling Incident Management](https://runframe.io/blog/scaling-incident-management): Four stages: Slack channel (5-15) → scripts (15-40) → "should buy a tool" limbo (40-100, stuck 6-12 months) → formal tool (100+). 40-50 person inflection point. Setup complexity is the real barrier, not cost. Teams want zero-config defaults. - [Engineering Productivity & Incident Management](https://runframe.io/blog/engineering-productivity-incident-management): Coordination is harder than the technical fix. Tool-switching during incidents breaks focus and compounds MTTR. Dedicated threads work for 20-100 person teams with ~10 min overhead. Centralize everything in one place. - [OpsGenie Migration Guide](https://runframe.io/blog/opsgenie-migration-guide): OpsGenie shuts down April 5, 2027. Migration: 4-8 weeks basic, 8-16 weeks for 20+ integrations — teams underestimate 2-3x. CSV exports don't import cleanly. Run parallel systems 4-8 weeks. - [OpsGenie End of Support Guide](https://runframe.io/blog/opsgenie-shutdown-guide): OpsGenie support ends April 5, 2027. Explains the end-of-life timeline, Atlassian JSM/Compass migration paths, third-party alternatives, and what teams should do next. - [Best OpsGenie Alternatives 2026](https://runframe.io/blog/best-opsgenie-alternatives): Alternatives teams actually switch to before the April 2027 OpsGenie shutdown. - [State of Incident Management 2025](https://runframe.io/blog/state-of-incident-management-2025): Toil rose to 30% despite 51% AI deployment. 78% of devs spend 30%+ time on manual toil (~$9.4M lost/year per 250-eng team). 73% of orgs had outages from ignored alerts. Market consolidation: OpsGenie shutting down, Freshworks acquired FireHydrant, SolarWinds acquired Squadcast. Outages cost ~$2M/hour. ### AI & Agents - [Your Agent Can Manage Incidents Now](https://runframe.io/blog/your-agent-can-manage-incidents-now): Runframe MCP server for managing incidents from Claude Code and Cursor: on-call, escalation, paging, and postmortems. - [Your AI Agent Already Knows Your System Better Than Ours Ever Will](https://runframe.io/blog/your-ai-already-knows-your-system-better-than-ours): Argument for giving customer-owned AI agents incident APIs instead of forcing vendor-specific incident AI. ## Free Tools - [MTTR Calculator](https://runframe.io/tools/mttr-calculator): MTTR = total resolution time / number of incidents. Enter timestamps, get your MTTR. DORA benchmarks: Elite <1 hour, High 1-24 hours, Medium 1-7 days, Low >6 months. - [Severity Matrix Generator](https://runframe.io/tools/incident-severity-matrix-generator): Build a custom SEV0-SEV4 matrix with impact criteria, response expectations, and escalation rules for your team size. - [On-Call Schedule Generator](https://runframe.io/tools/oncall-builder): Free on-call schedule generator and rotation builder for weekly, biweekly, and custom calendars with timezone support, primary/secondary responders, and handoff templates. ## Learn — Key Reference Articles Definitions, formulas, benchmarks, and FAQs for core incident management concepts. Full inventory of all 55+ articles with detailed descriptions: [llms-full.txt](https://runframe.io/llms-full.txt) ### Core Metrics - [MTTR](https://runframe.io/learn/mttr): Mean Time to Resolution. Formula: MTTR = total resolution time / number of incidents. Benchmarks: Excellent <30 min (top 5%), Good <1 hour, Average 1-24 hours, Low >24 hours (DORA). - [MTTA](https://runframe.io/learn/mtta): Mean Time to Acknowledge. Formula: MTTA = total ack times / count. Elite <1 min, Good <5 min. Over 15 min = broken paging process. - [MTTD](https://runframe.io/learn/mttd): Mean Time to Detect. Time from failure to alert. Gold standard: <1 min. Practical minimum ~30 seconds (polling interval). - [MTBF](https://runframe.io/learn/mtbf): Mean Time Between Failures. Formula: total uptime / number of failures. Excellent: >720 hours (30 days). Scheduled maintenance doesn't count. - [Error Budget](https://runframe.io/learn/error-budget): Allowed unreliability = 100% - SLO. For 99.5% SLO: 0.5% = 216 min/month. When exhausted, freeze features and fix reliability. Resets monthly or quarterly. ### Service Levels - [SLO](https://runframe.io/learn/slo): Service Level Objective. Internal reliability target. Formula: (successful requests / total) × 100%. Set by PM + eng lead agreement. Target slightly below current performance. - [SLI](https://runframe.io/learn/sli): Service Level Indicator. The metric you actually measure. Formula: (good events / total) × 100%. Keep 1-3 per user journey (login, checkout, search). - [SLA](https://runframe.io/learn/sla): Service Level Agreement. Contractual commitment with penalties. Usually only for paid enterprise plans. Always set SLO higher than SLA for buffer. ### Severity & Response - [Incident Severity Matrix](https://runframe.io/learn/incident-severity-matrix): SEV0 (critical) through SEV4 (trivial). Use SEV0-SEV4 or P0-P4 — doesn't matter which, just pick one and stick with it. - [Incident Commander](https://runframe.io/learn/incident-commander): Leads coordination during incidents. A role, not a rank — a junior can be IC for a VP. - [Escalation Policy](https://runframe.io/learn/escalation-policy): When and how to escalate. Rule: if stuck 15-30 min or severity increases, escalate. - [Blameless Postmortem](https://runframe.io/learn/blameless-postmortem): Focus on system/process failures, not individual blame. Blameless ≠ no consequences — 99% of incidents are honest mistakes in bad systems. ### On-Call - [On-Call Rotation](https://runframe.io/learn/on-call-rotation): Schedule determining who handles alerts. Ideal: no more than 1 week every 6-8 weeks. More frequent → fatigue. - [Follow-the-Sun](https://runframe.io/learn/follow-the-sun): Global model avoiding night shifts via timezone distribution. Minimum: 3 teams × 3-5 engineers = 9-15 people. - [Alert Fatigue](https://runframe.io/learn/alert-fatigue): Desensitization from excessive alerts. Threshold: >1-2 pages per 12-hour shift = fatigue. On-call should be silent unless things are broken. See the full guide: [Alert Fatigue Causes and Fixes](https://runframe.io/blog/how-to-reduce-alert-fatigue). ### Observability - [Four Golden Signals](https://runframe.io/learn/four-golden-signals): Latency, Traffic, Errors, Saturation (Google SRE). Most important: Errors — if users see errors, fix that first. - [Availability](https://runframe.io/learn/availability): Formula: uptime / (uptime + downtime). Five nines (99.999%) = <5 min/year. 99.9% = 8.77 hours/year. For most SaaS, 99.9% is adequate. ## Frequently Asked Questions **What is a good MTTR for my team?** Depends on team size. P0 MTTR benchmarks: 30-60 min for teams under 20, 35-75 min for 20-80, 40-120 min for 80+. DORA Elite performers: <1 hour across all severities. See [How to Reduce MTTR](https://runframe.io/blog/how-to-reduce-mttr). **Should I build or buy incident management?** For most teams, buy. 3-year TCO for 20-person team: build $233K-$395K vs buy $11K-$83K (3-8x more expensive to build). Building is only justified if you have a dedicated owner, separate infrastructure, and regulatory requirements existing tools can't handle. See [Build or Buy](https://runframe.io/blog/incident-management-build-or-buy). **What open source incident management tools exist in 2026?** Three actively maintained: Incidental (MIT, v0.1.0, incidental.dev), incident-bot (MIT, Python/PostgreSQL, PagerDuty/Jira/Confluence integrations), and IncidentFox (Apache 2.0 core, BSL 1.1 for production security features). Netflix Dispatch was archived Sep 2025. Grafana OnCall was closed-sourced Mar 2025. See [Build or Buy](https://runframe.io/blog/incident-management-build-or-buy). **How much does PagerDuty cost vs alternatives?** PagerDuty: ~$21/user/month (Professional) or ~$41/user (Business). Runframe: Growth is $15/user/month monthly or $12 annual; Scale is $30/user/month monthly or $25 annual. incident.io: $19/user + $10 on-call add-on. Rootly: usage-based. Grafana IRM: free 3 users, $19/mo + $20/user. FireHydrant: $9,600/year for 20 responders. Better Stack: free tier. See [PagerDuty Alternatives](https://runframe.io/blog/best-pagerduty-alternatives). **What's the difference between SLA, SLO, and SLI?** SLI = the metric (error rate, latency). SLO = your internal target (e.g., 99.7%). SLA = your contractual promise to customers (e.g., 99.5%) with penalties. Set SLO above SLA for buffer. Error budget = 100% - SLO. See [SLA vs SLO vs SLI](https://runframe.io/blog/sla-vs-slo-vs-sli). **How do I set up on-call rotations?** Start with weekly rotations for teams under 50. Set escalation timers: primary 5 min → backup 5 min → manager for SEV0/1. Pay $200-500/week stipend or comp days. Write 2-minute handoffs in Slack. Consider follow-the-sun (3 timezone teams, 9-15 engineers min) to avoid night pages. See [On-Call Rotation Guide](https://runframe.io/blog/on-call-rotation-guide). **What severity levels should we use?** Start with 3 levels (SEV1-SEV3) for teams under 50. Add SEV0 at 100+ people for catastrophic events. Add SEV4 at 75-100 for proactive tracking. Classify in 30 seconds — don't debate. See [Severity Levels](https://runframe.io/blog/incident-severity-levels). **What is OpsGenie's shutdown timeline?** OpsGenie fully shuts down April 5, 2027. New sales stopped June 4, 2025. Migration takes 4-8 weeks basic, 8-16 weeks for complex setups with 20+ integrations. Teams underestimate by 2-3x. See [OpsGenie Migration Guide](https://runframe.io/blog/opsgenie-migration-guide). **How much do outages cost?** High-impact IT outages cost approximately $2M/hour. Organizations lose a median of $76M annually from unplanned downtime. 73% of organizations experienced outages linked to ignored alerts. See [State of Incident Management](https://runframe.io/blog/state-of-incident-management-2025). **What are the four golden signals of monitoring?** Latency, Traffic, Errors, and Saturation — defined by Google's SRE team. Errors are most important: if users see errors, nothing else matters. See [Four Golden Signals](https://runframe.io/learn/four-golden-signals). **What is a blameless postmortem?** Post-incident analysis focused on system/process failures, not individual blame. Complete within 48 hours. Keep to one page. Every action item needs a specific owner, deadline, and definition of done. Target 1-3 items per incident. See [Post-Incident Review Template](https://runframe.io/blog/post-incident-review-template). **What's the difference between a runbook and a playbook?** Runbooks are step-by-step technical procedures ("how do I fix this?"). Playbooks are strategic frameworks covering roles, escalation, and communication ("who handles this?"). Build playbooks first. See [Runbook vs Playbook](https://runframe.io/blog/runbook-vs-playbook).