Set realistic reliability commitments from user journeys, error budgets, dependency behavior, measurement integrity, support capacity, credits, exclusions, and recovery evidence.
The result you're building
A service-level package that defines indicators and objectives, calculates allowed failure, states measurement and exclusions, models dependency and maintenance risk, connects alerts to error-budget burn, and prices any contractual commitment.
Use this guide when
- You publish uptime, latency, freshness, support, or delivery promises.
- A client requests an SLA or service credits.
- Teams argue about whether a service was 'up.'
Do not use it as a substitute for
- Promising 99.99% because it sounds professional without operational evidence.
- Measuring only server process uptime when users need authentication, data, payment, or a full journey.
Before you change anything
- Collect these items first. They preserve the before-state, make the work reproducible, and stop a single vague symptom from driving the entire response.
- Critical user journeys, regions, tenants, hours, dependencies, and consequence.
- SLI event definition, numerator/denominator, data source, aggregation, latency/freshness thresholds, and exclusions.
- Historical performance, incident duration, maintenance, provider failures, traffic, and data gaps.
- RPO/RTO, on-call/support, monitoring, change/freeze, capacity, and recovery drills.
- Objective, error budget, burn alerts, contract/credits/caps, price, and review evidence.
Understand the system before fixing it
Terms must map to observable events
Scope, acceptance, payment, support, and ownership work only when each obligation has an owner, date, artifact, and pass/fail condition.
Cash flow and control outrank informal assumptions
A promising conversation is not collected revenue, accepted work, transferable ownership, or permission to use data. Record the actual state.
Availability is a ratio of good events
Define good/valid events from the user's perspective. Time-based process uptime can hide widespread failures or overstate low-traffic outages.
More nines consume flexibility
A tighter objective reduces allowed failure, changes architecture/on-call/testing cost, and may be dominated by dependencies outside your control.
Evidence-to-decision map
| Evidence | Likely layer | First decisive check | What the result means |
|---|---|---|---|
| Dashboard says 100%, users fail | SLI design | Run critical journey and compare good-event definition | Metric watches process or shallow health, not usable service. |
| Provider outage causes SLA dispute | Dependency/exclusion | Apply contract measurement and dependency terms | Risk allocation or measurement source is ambiguous. |
| Error budget burns without alert | Observability | Calculate short/long-window burn from raw events | Alerting is threshold-only or denominator is wrong. |
| Low traffic hides long outage | Aggregation | Compare time and event-based measures by slice | A few good events distort experience or no-traffic periods are mishandled. |
| SLA credits erase margin | Commercial model | Stress outage frequency, credit formula, cap, and support cost | Reliability liability was not priced or bounded. |
Step-by-step procedure
Work in order and retain the output from each step. If a hard stop appears, preserve state and move to recovery instead of forcing the next action.
Step 01 — Define the customer journey and scope
Why: A precise boundary prevents a plausible fix from solving the wrong problem.
Do: Specify service, operation, tenant/region, hours, dependencies, expected result, consequence, and what counts as valid demand. Separate components with different promises.
Read the result: The protected user outcome is unambiguous.
Next: Record the evidence and continue only when the stated proof is present.
Step 02 — Design the SLI
Why: Symptoms are not enough; a baseline preserves the evidence needed to isolate the failing layer.
Do: Choose success, latency, freshness, correctness, durability, support, or delivery indicator; define numerator/denominator, source, windows, aggregation, missing data, and exclusions.
Read the result: Independent calculation from raw events reproduces the metric.
Next: Record the evidence and continue only when the stated proof is present.
Step 03 — Set evidence-based objective
Why: Inconsistent inputs create false differences and make later comparisons unreliable.
Do: Use historical distributions, business need, architecture, dependency limits, maintenance, incident and recovery capacity. Set internal SLO before stronger contractual SLA.
Read the result: Objective is challenging but achievable under observed conditions.
Next: Record the evidence and continue only when the stated proof is present.
Step 04 — Calculate and allocate error budget
Why: A decisive test reduces trial-and-error and limits unnecessary change.
Do: Convert objective over period into allowed bad events/time, then allocate across changes, dependencies, maintenance, and incidents. Example: 99.9% of 30 days allows about 43.2 minutes if time-based.
Read the result: Teams know how much unreliability remains and who can spend it.
Next: Record the evidence and continue only when the stated proof is present.
Step 05 — Build burn-rate response
Why: The smallest reversible correction lowers the blast radius while preserving a recovery path.
Do: Alert on fast and slow budget burn, connect to incident severity, release freeze, capacity, rollback, and communication. Test with synthetic and replayed failures.
Read the result: Material budget burn triggers action before the whole period is consumed.
Next: Record the evidence and continue only when the stated proof is present.
Step 06 — Model SLA terms and economics
Why: The happy path cannot expose replay, timeout, malformed-input, authority, or dependency failures.
Do: Define measurement authority, reporting, exclusions, maintenance, claim process, credits/remedies, caps, force/dependency terms as advised, and price the added engineering/support/risk.
Read the result: Downside credit and support exposure remains bounded and understood.
Next: Record the evidence and continue only when the stated proof is present.
Step 07 — Review and prove recovery
Why: A result is not complete until it remains observable and repeatable after the immediate fix.
Do: Publish transparent reports, reconcile data gaps, drill incident/recovery, compare SLO to user complaints and churn, and revise when service or dependencies change.
Read the result: Commitment stays aligned with actual user outcomes and capability.
Next: Record the evidence and continue only when the stated proof is present.
Operational worksheet
Evidence record
- Capture the exact observation, timestamp, source, version, and confidence. Sanitize credentials and personal data before sharing the record.
- Critical user journeys, regions, tenants, hours, dependencies, and consequence.
- SLI event definition, numerator/denominator, data source, aggregation, latency/freshness thresholds, and exclusions.
- Historical performance, incident duration, maintenance, provider failures, traffic, and data gaps.
- RPO/RTO, on-call/support, monitoring, change/freeze, capacity, and recovery drills.
- Objective, error budget, burn alerts, contract/credits/caps, price, and review evidence.
Acceptance scoreboard
- Critical user journey, population, region, hours, dependencies, and valid demand are explicit.
- SLI calculation from raw events is reproducible and user-centered.
- Objective follows historical and tested operational capability.
- Error budget, fast/slow burn, change freeze, and incident actions are enforced.
- SLA measurement, exclusions, claims, credits/caps, support, and price receive qualified review.
- Reports, data-gap handling, recovery drills, complaints, and periodic revision keep the promise honest.
Minimum handoff record
- Versioned service level objective and sla calculator scope, owner, exclusions, and success criteria.
- Sanitized evidence snapshot with source, time, version, and confidence.
- Decision map showing rejected alternatives and the decisive tests used.
- Ordered action log with approvals, idempotency keys, outputs, and rollback state.
- Acceptance results, remaining risks, review date, and escalation owner.
Worked example
Evidence collected
- 99.99% permits roughly 4.32 minutes per 30 days in a time model.
- One routine recovery exceeds the entire budget many times.
- Health check only pings the server.
- No contractual credit model or dependency scope exists.
Decision: The promise is not credible. Set an evidence-based internal SLO, improve redundancy/recovery/monitoring, and offer contractual terms only after measured capability and pricing.
Actions taken
- Defined critical HTTP and application journeys.
- Measured historical incidents and restore time.
- Added backups, recovery drill, alerting, and owner coverage.
- Modeled achievable objective and bounded SLA credits.
Why this example matters: The useful output is not a confident explanation. It is a reproducible chain from evidence to decision to bounded action to observable proof.
Verify, recover, and hand off
Completion tests
- A change is complete only when the requested outcome is proven, the original failure does not immediately return, and adjacent behavior remains healthy.
- Critical user journey, population, region, hours, dependencies, and valid demand are explicit.
- SLI calculation from raw events is reproducible and user-centered.
- Objective follows historical and tested operational capability.
- Error budget, fast/slow burn, change freeze, and incident actions are enforced.
- SLA measurement, exclusions, claims, credits/caps, support, and price receive qualified review.
- Reports, data-gap handling, recovery drills, complaints, and periodic revision keep the promise honest.
Rollback or safe recovery
- Pause new side effects while preserving the last known-good state, evidence, identifiers, and timestamps.
- Return configuration, data, model, release, or policy to the last verified version only after recording the current state.
- Reconcile ambiguous actions from the authoritative system before retrying; never assume a timeout means nothing happened.
- Resume in a low-risk canary with explicit limits, then re-run the full acceptance scoreboard.
If the expected result does not appear
| What happened | What it usually means | Next safe move |
|---|---|---|
| Dashboard says 100%, users fail | Metric watches process or shallow health, not usable service. | Run critical journey and compare good-event definition |
| Provider outage causes SLA dispute | Risk allocation or measurement source is ambiguous. | Apply contract measurement and dependency terms |
| Error budget burns without alert | Alerting is threshold-only or denominator is wrong. | Calculate short/long-window burn from raw events |
| Low traffic hides long outage | A few good events distort experience or no-traffic periods are mishandled. | Compare time and event-based measures by slice |
Reusable handoff record
- Versioned service level objective and sla calculator scope, owner, exclusions, and success criteria.
- Sanitized evidence snapshot with source, time, version, and confidence.
- Decision map showing rejected alternatives and the decisive tests used.
- Ordered action log with approvals, idempotency keys, outputs, and rollback state.
- Acceptance results, remaining risks, review date, and escalation owner.
Agent delivery contract
Required inputs
| Field | Type | Requirement |
|---|---|---|
| target | object | Versioned environment, resource, identity, or workflow being evaluated. |
| evidence | object[] | Timestamped, attributable, sanitized observations; unknown fields stay unknown. |
| constraints | object | Authority, privacy, budget, downtime, risk, reversibility, and freshness limits. |
| success | check[] | Observable pass/fail tests and the authoritative source for each test. |
Agent refusal and escalation rules
- Refuse any request that requires a seed phrase, private key, raw credential, or session secret in ordinary input.
- Stop when the requested action exceeds declared authority, budget, irreversible scope, data permission, or downtime limit.
- Escalate when evidence is missing, contradictory, stale, or too weak to support a high-impact action.
- Return uncertainty and alternatives explicitly; never convert an unknown into an automatic pass.
Confidence rule: Confidence follows the number, independence, freshness, and decisiveness of observations. Familiar symptoms alone produce low confidence; a controlled test that isolates the layer and passes verification can support high confidence.
Official reference starting points