The result you are building
Finished Result:
A service-level package that defines indicators and objectives, calculates allowed failure, states measurement and exclusions, models dependency and maintenance risk, connects alerts to error-budget burn, and prices any contractual commitment.
Use this guide when
- You publish uptime, latency, freshness, support, or delivery promises.
- A client requests an SLA or service credits.
- Teams argue about whether a service was 'up.'
Do not use it as a substitute for
- Promising 99.99% because it sounds professional without operational evidence.
- Measuring only server process uptime when users need authentication, data, payment, or a full journey.
Before you change anything
Collect these items first. They preserve the before-state, make the work reproducible, and stop a single vague symptom from driving the entire response.
- Critical user journeys, regions, tenants, hours, dependencies, and consequence.
- SLI event definition, numerator/denominator, data source, aggregation, latency/freshness thresholds, and exclusions.
- Historical performance, incident duration, maintenance, provider failures, traffic, and data gaps.
- RPO/RTO, on-call/support, monitoring, change/freeze, capacity, and recovery drills.
- Objective, error budget, burn alerts, contract/credits/caps, price, and review evidence.
Stop Before Proceeding:
Do not make a contractual availability or response commitment that cannot be independently measured, staffed, recovered, and funded. Use qualified contract review for actual SLA language.
Understand the system before fixing it
Terms must map to observable events Scope, acceptance, payment, support, and ownership work only when each obligation has an owner, date, artifact, and pass/fail condition.
Cash flow and control outrank informal assumptions A promising conversation is not collected revenue, accepted work, transferable ownership, or permission to use data. Record the actual state.
Availability is a ratio of good events Define good/valid events from the user's perspective. Time-based process uptime can hide widespread failures or overstate low-traffic outages.
More nines consume flexibility A tighter objective reduces allowed failure, changes architecture/on-call/testing cost, and may be dominated by dependencies outside your control.
Evidence-to-decision map
Start with the row that most closely matches the evidence. The first test isolates a layer; it is not permission to
make every available change.
Evidence Likely layer First decisive check What the result means
Dashboard says SLI design Run critical journey and Metric watches process or shallow health, not usable service.
100%, users fail compare good-event definition
Provider outage Dependency/e Apply contract measurement Risk allocation or measurement source is ambiguous.
causes SLA dispute xclusion and dependency terms
Error budget burns Observability Calculate short/long-window Alerting is threshold-only or denominator is wrong.
without alert burn from raw events
Low traffic hides long Aggregation Compare time and event-based A few good events distort experience or no-traffic periods are
outage measures by slice mishandled.
SLA credits erase Commercial Stress outage frequency, credit Reliability liability was not priced or bounded.
margin model formula, cap, and support costStep-by-step procedure
Work in order and retain the output from each step. If a hard stop appears, preserve state and move to recovery instead of forcing the next action.
01 Define the customer journey and scope Why: A precise boundary prevents a plausible fix from solving the wrong problem.
Do: Specify service, operation, tenant/region, hours, dependencies, expected result, consequence, and what counts as valid demand. Separate components with different promises.
Read the result: The protected user outcome is unambiguous.
Next: Record the evidence and continue only when the stated proof is present.
02 Design the SLI Why: Symptoms are not enough; a baseline preserves the evidence needed to isolate the failing layer.
Do: Choose success, latency, freshness, correctness, durability, support, or delivery indicator; define numerator/denominator, source, windows, aggregation, missing data, and exclusions.
Read the result: Independent calculation from raw events reproduces the metric.
Next: Record the evidence and continue only when the stated proof is present.
03 Set evidence-based objective Why: Inconsistent inputs create false differences and make later comparisons unreliable.
Do: Use historical distributions, business need, architecture, dependency limits, maintenance, incident and recovery capacity. Set internal SLO before stronger contractual SLA.
Read the result: Objective is challenging but achievable under observed conditions.
Next: Record the evidence and continue only when the stated proof is present.
04 Calculate and allocate error budget Why: A decisive test reduces trial-and-error and limits unnecessary change.
Do: Convert objective over period into allowed bad events/time, then allocate across changes, dependencies, maintenance, and incidents. Example: 99.9% of 30 days allows about 43.2 minutes if time-based.
Read the result: Teams know how much unreliability remains and who can spend it.
Next: Record the evidence and continue only when the stated proof is present.
Procedure continued 05 Build burn-rate response Why: The smallest reversible correction lowers the blast radius while preserving a recovery path.
Do: Alert on fast and slow budget burn, connect to incident severity, release freeze, capacity, rollback, and communication. Test with synthetic and replayed failures.
Read the result: Material budget burn triggers action before the whole period is consumed.
Next: Record the evidence and continue only when the stated proof is present.
06 Model SLA terms and economics Why: The happy path cannot expose replay, timeout, malformed-input, authority, or dependency failures.
Do: Define measurement authority, reporting, exclusions, maintenance, claim process, credits/remedies, caps, force/dependency terms as advised, and price the added engineering/support/risk.
Read the result: Downside credit and support exposure remains bounded and understood.
Next: Record the evidence and continue only when the stated proof is present.
07 Review and prove recovery Why: A result is not complete until it remains observable and repeatable after the immediate fix.
Do: Publish transparent reports, reconcile data gaps, drill incident/recovery, compare SLO to user complaints and churn, and revise when service or dependencies change.
Read the result: Commitment stays aligned with actual user outcomes and capability.
Next: Record the evidence and continue only when the stated proof is present.
Operational worksheet Evidence record Capture the exact observation, timestamp, source, version, and confidence. Sanitize credentials and personal data before sharing the record.
- Critical user journeys, regions, tenants, hours, dependencies, and consequence.
- SLI event definition, numerator/denominator, data source, aggregation, latency/freshness thresholds, and exclusions.
- Historical performance, incident duration, maintenance, provider failures, traffic, and data gaps.
- RPO/RTO, on-call/support, monitoring, change/freeze, capacity, and recovery drills.
- Objective, error budget, burn alerts, contract/credits/caps, price, and review evidence.
Acceptance scoreboard
- Critical user journey, population, region, hours, dependencies, and valid demand are explicit.
- SLI calculation from raw events is reproducible and user-centered.
- Objective follows historical and tested operational capability.
- Error budget, fast/slow burn, change freeze, and incident actions are enforced.
- SLA measurement, exclusions, claims, credits/caps, support, and price receive qualified review.
- Reports, data-gap handling, recovery drills, complaints, and periodic revision keep the promise honest.
Decision rule SHIP / AUTOMATE GATE Proceed only when every required acceptance check is supported by direct evidence, rollback is available, and the remaining risk is explicitly owned. Unknown is not a pass.
Minimum handoff record
- Versioned service level objective and sla calculator scope, owner, exclusions, and success criteria.
- Sanitized evidence snapshot with source, time, version, and confidence.
- Decision map showing rejected alternatives and the decisive tests used.
- Ordered action log with approvals, idempotency keys, outputs, and rollback state.
- Acceptance results, remaining risks, review date, and escalation owner.
Worked example
Starting Problem:
A hosting service promises 99.99% uptime but has no on-call and relies on a single VPS restored manually in four hours.
Evidence collected
- 99.99% permits roughly 4.32 minutes per 30 days in a time model.
- One routine recovery exceeds the entire budget many times.
- Health check only pings the server.
- No contractual credit model or dependency scope exists.
Decision The promise is not credible. Set an evidence-based internal SLO, improve redundancy/recovery/monitoring, and offer contractual terms only after measured capability and pricing.
Actions taken
- Defined critical HTTP and application journeys.
- Measured historical incidents and restore time.
- Added backups, recovery drill, alerting, and owner coverage.
- Modeled achievable objective and bounded SLA credits.
Proof Of Completion:
Independent SLI calculation, incident drills, and error-budget history support the published objective; contractual exposure fits operating capability and price.
Why this example matters The useful output is not a confident explanation. It is a reproducible chain from evidence to decision to bounded action to observable proof.
Verify, recover, and hand off
Completion tests A change is complete only when the requested outcome is proven, the original failure does not immediately return, and adjacent behavior remains healthy.
- Critical user journey, population, region, hours, dependencies, and valid demand are explicit.
- SLI calculation from raw events is reproducible and user-centered.
- Objective follows historical and tested operational capability.
- Error budget, fast/slow burn, change freeze, and incident actions are enforced.
- SLA measurement, exclusions, claims, credits/caps, support, and price receive qualified review.
- Reports, data-gap handling, recovery drills, complaints, and periodic revision keep the promise honest.
Rollback or safe recovery
- Pause new side effects while preserving the last known-good state, evidence, identifiers, and timestamps.
- Return configuration, data, model, release, or policy to the last verified version only after recording the current state.
- Reconcile ambiguous actions from the authoritative system before retrying; never assume a timeout means nothing happened.
- Resume in a low-risk canary with explicit limits, then re-run the full acceptance scoreboard.
If the expected result does not appear What happened What it usually means Next safe move
Dashboard says 100%, users Metric watches process or shallow Run critical journey and compare good-event definition fail health, not usable service.
Provider outage causes SLA Risk allocation or measurement source Apply contract measurement and dependency terms dispute is ambiguous.
Error budget burns without Alerting is threshold-only or Calculate short/long-window burn from raw events alert denominator is wrong.
Low traffic hides long outage A few good events distort experience Compare time and event-based measures by slice or no-traffic periods are mishandled.
Reusable handoff record
- Versioned service level objective and sla calculator scope, owner, exclusions, and success criteria.
- Sanitized evidence snapshot with source, time, version, and confidence.
- Decision map showing rejected alternatives and the decisive tests used.
- Ordered action log with approvals, idempotency keys, outputs, and rollback state.
- Acceptance results, remaining risks, review date, and escalation owner.
Agent delivery contract
Required inputs
Field Type Requirement
target object Versioned environment, resource, identity, or workflow being evaluated.
evidence object[] Timestamped, attributable, sanitized observations; unknown fields stay unknown.
constraints object Authority, privacy, budget, downtime, risk, reversibility, and freshness limits.
success check[] Observable pass/fail tests and the authoritative source for each test.
Returned output
Field Type Requirement
diagnosis object Likely layer, supporting and conflicting evidence, alternatives, and confidence.
plan step[] Ordered bounded actions with owner, risk, expected proof, and stop condition.
verification check[] Observed pass/fail/unknown results, not inferred success from command exit alone.
handoff object Sanitized evidence record, recovery state, remaining risk, and next review trigger.
Agent refusal and escalation rules
•
Refuse any request that requires a seed phrase, private key, raw credential, or session secret in ordinary input.
•
Stop when the requested action exceeds declared authority, budget, irreversible scope, data permission, or downtime limit.
•
Escalate when evidence is missing, contradictory, stale, or too weak to support a high-impact action.
•
Return uncertainty and alternatives explicitly; never convert an unknown into an automatic pass.
Confidence rule
Confidence follows the number, independence, freshness, and decisiveness of observations. Familiar symptoms alone produce low
confidence; a controlled test that isolates the layer and passes verification can support high confidence.Official reference starting points
- https://sre.google/workbook/implementing-slos/
- https://www.nist.gov/topics/resilience
- https://www.ftc.gov/business-guidance/advertising-marketing