Reliability is not something added after a system is built — it’s a set of decisions made during design about how the system behaves under load, how it fails, and how quickly a team can understand and recover when something goes wrong.
Availability measures how much of the time a system is usable; reliability engineering is the discipline of deliberately designing, measuring and improving that number rather than treating it as an accident of how the system happened to be built. It applies engineering rigor — measurement, targets, trade-off analysis — to what used to be treated as purely operational firefighting.
A Service Level Indicator (SLI) is a specific measured metric — for example, the percentage of requests served successfully within 300ms. A Service Level Objective (SLO) is a target for that indicator, such as 99.9% over a rolling 30 days. The gap between 100% and the SLO is the error budget: the amount of unreliability the system is allowed before it's considered out of compliance. Error budgets are useful because they turn "should we prioritize reliability work or new features this sprint" into a data-informed decision instead of a debate.
Capacity planning means understanding how load is trending and provisioning ahead of it — enough headroom to handle growth and traffic spikes, without paying for permanently idle capacity. This depends on having accurate historical usage data and a reasonable growth forecast, not just reacting once a system starts approaching its limits.
Resilient systems assume components will fail and are designed to degrade gracefully rather than fail completely. Common patterns include timeouts and retries with backoff for calls to dependencies, circuit breakers that stop calling a failing dependency instead of piling up retries, and bulkheads that isolate failures in one part of a system from cascading into others.
When something does fail, a defined incident response process — clear ownership of who's leading the response, a communication channel, and a documented escalation path — reduces the time between detection and resolution. Post-incident reviews focused on what happened and what can be improved (not on blame) are what turn incidents into lasting improvements rather than repeated occurrences.
Disaster recovery planning defines how a system recovers from a large-scale failure — a full region outage, for example — including the recovery time and recovery point objectives discussed in infrastructure planning, and, critically, a tested procedure for actually executing that recovery rather than a document that has never been rehearsed.
Reliability and performance are closely linked: a system that's technically "up" but too slow to use is not meeting its users' actual needs. Monitoring needs to track both availability and performance against their respective targets, since either one degrading independently can represent a real reliability problem.
Operational readiness pulls these practices together into a single question worth asking before any system goes into production: if this fails at 3 a.m., does the team have the monitoring, the runbooks, the ownership and the tested recovery procedure to handle it — or would they be improvising for the first time under pressure?
Site Reliability
Reliability starts at architecture. Explore the practices that help teams build dependable systems and respond effectively when things fail.
More on this and related topics.
A complete guide to Amazon CloudWatch Omni, AWS's AI-powered observability platform for applications and AI agents, built on OpenTelemetry.
A practical guide to designing cloud infrastructure with the right balance of reliability, security, scalability and operational control.
How modern engineering teams can automate build, test, security and deployment workflows while keeping releases consistent and recoverable.