Skip to content

Glossary

The reliability vocabulary, in plain language.

Clear definitions for the incident, on-call, monitoring, and status terms teams use every day.

Alert Fatigue

Alert fatigue is the desensitization that sets in when responders receive too many alerts — especially noisy or low-value ones — making them slower to react or more likely to miss the alerts that matter.

Read definition

Error Budget

An error budget is the allowed amount of unreliability over a period — the gap between 100% and an SLO target. It gives teams a shared, measurable limit for acceptable failures.

Read definition

Escalation Policy

An escalation policy is a defined set of rules for who gets notified about an alert and in what order, escalating to additional responders if it isn't acknowledged or resolved within a set time.

Read definition

Heartbeat Monitoring

Heartbeat monitoring watches for a regular "check-in" signal from a job or service and raises an alert when the expected signal doesn't arrive in time — catching silent failures of things that don't serve traffic.

Read definition

Incident Management

Incident management is the practice of detecting, coordinating a response to, and resolving unplanned events that degrade a service — then capturing what happened so the team improves over time.

Read definition

MTTA (Mean Time to Acknowledge)

MTTA is the average time between when an alert fires and when a responder acknowledges it. It measures how quickly a team notices and takes ownership of a problem.

Read definition

MTTR (Mean Time to Resolve)

MTTR is the average time it takes to fully resolve an incident, from detection to restored service. It's a common measure of how effectively a team responds to and recovers from disruptions.

Read definition

On-Call

On-call is an arrangement where designated responders are available to react to alerts and incidents during scheduled shifts, ensuring someone is always ready to address problems with a service.

Read definition

Runbook

A runbook is a documented set of steps for handling a specific operational task or failure scenario, giving responders a reliable, repeatable procedure to follow instead of improvising under pressure.

Read definition

Service Catalog

A service catalog is a central inventory of the services a team runs, recording each service's owner, dependencies, monitors, escalation policy, and reliability targets so responsibility and context are always clear.

Read definition

SLA (Service Level Agreement)

An SLA is a formal commitment between a provider and its customers that defines the expected level of service — often an availability target — along with how it's measured and what happens if it isn't met.

Read definition

SLO (Service Level Objective)

An SLO is an internal reliability target for a service — such as a percentage of successful requests or availability over a window — that a team commits to and measures against to guide day-to-day decisions.

Read definition

Status Page

A status page is a public or private web page that shows the current health of a service and its components, and communicates incidents and maintenance to users in one trusted place.

Read definition

Uptime

Uptime is the proportion of time a service is available and functioning as expected, usually expressed as a percentage over a period. It's a core measure of reliability.

Read definition

From definitions to a working incident workflow.

Free plan available · no credit card required