Glossary
The reliability vocabulary, in plain language.
Clear definitions for the incident, on-call, monitoring, and status terms teams use every day.
Alert Fatigue
Alert fatigue is the desensitization that sets in when responders receive too many alerts — especially noisy or low-value ones — making them slower to react or more likely to miss the alerts that matter.
Read definitionError Budget
An error budget is the allowed amount of unreliability over a period — the gap between 100% and an SLO target. It gives teams a shared, measurable limit for acceptable failures.
Read definitionEscalation Policy
An escalation policy is a defined set of rules for who gets notified about an alert and in what order, escalating to additional responders if it isn't acknowledged or resolved within a set time.
Read definitionHeartbeat Monitoring
Heartbeat monitoring watches for a regular "check-in" signal from a job or service and raises an alert when the expected signal doesn't arrive in time — catching silent failures of things that don't serve traffic.
Read definitionIncident Management
Incident management is the practice of detecting, coordinating a response to, and resolving unplanned events that degrade a service — then capturing what happened so the team improves over time.
Read definitionMTTA (Mean Time to Acknowledge)
MTTA is the average time between when an alert fires and when a responder acknowledges it. It measures how quickly a team notices and takes ownership of a problem.
Read definitionMTTR (Mean Time to Resolve)
MTTR is the average time it takes to fully resolve an incident, from detection to restored service. It's a common measure of how effectively a team responds to and recovers from disruptions.
Read definitionOn-Call
On-call is an arrangement where designated responders are available to react to alerts and incidents during scheduled shifts, ensuring someone is always ready to address problems with a service.
Read definitionRunbook
A runbook is a documented set of steps for handling a specific operational task or failure scenario, giving responders a reliable, repeatable procedure to follow instead of improvising under pressure.
Read definitionService Catalog
A service catalog is a central inventory of the services a team runs, recording each service's owner, dependencies, monitors, escalation policy, and reliability targets so responsibility and context are always clear.
Read definitionSLA (Service Level Agreement)
An SLA is a formal commitment between a provider and its customers that defines the expected level of service — often an availability target — along with how it's measured and what happens if it isn't met.
Read definitionSLO (Service Level Objective)
An SLO is an internal reliability target for a service — such as a percentage of successful requests or availability over a window — that a team commits to and measures against to guide day-to-day decisions.
Read definitionStatus Page
A status page is a public or private web page that shows the current health of a service and its components, and communicates incidents and maintenance to users in one trusted place.
Read definitionUptime
Uptime is the proportion of time a service is available and functioning as expected, usually expressed as a percentage over a period. It's a core measure of reliability.
Read definitionFrom definitions to a working incident workflow.
Free plan available · no credit card required