Memory hook: Measure reliability from the user’s side.
Must remember
An SLI measures service behavior, an SLO sets a target and an SLA defines an external commitment/remedy. An error budget is the tolerated unreliability implied by an SLO over its window. Use it to make release/reliability tradeoffs; do not invent a universal acceptable percentage for every business.
Alert on symptoms that require action, with ownership and runbooks. Burn-rate-style alerts compare how quickly the budget is being consumed over suitable windows. Metrics, logs, traces and profiles answer complementary questions. Avoid paging for every transient utilization spike when users are unaffected.
CI validates changes; delivery/deployment moves approved artifacts through environments. Canary releases limit initial exposure; blue/green provides separate environments; rolling updates replace incrementally. Automated rollback needs trustworthy health signals and compatible data/schema. Feature flags separate activation from deployment but require lifecycle cleanup and secure access.
Use load testing for capacity, penetration/security testing for authorized attack resistance and chaos experiments for controlled failure hypotheses. Define blast radius, stop conditions and recovery before experiments. A test in staging may miss production-scale bottlenecks; justify what conclusions it supports.
Incident response should restore service, communicate clearly and preserve enough evidence for a blameless root-cause review. Fix system conditions, not only the person who made the last change. Reduce toil with bounded automation and improve runbooks through actual exercises.
Review capacity, quotas, dependencies, support plans, cost allocation and sustainability continuously. Gemini Cloud Assist and other tools can help investigate or propose changes, but verify recommendations against evidence and the intended environment. Operational excellence is sustained ownership, not a one-time dashboard installation.
Review details
For a request-based SLO, budget is allowed bad-event fraction × valid requests. A 99.9% target across 1,000,000 requests permits 1,000 bad requests. For a time-based 99.9% target over 30 days, the equivalent allowed bad time is 43.2 minutes. Do not mix request and time denominators.
Burn rate is the observed bad-event fraction divided by the budget's allowed bad-event fraction. With a 99.9% objective, a 1% error fraction burns at 10×. Combine suitable short and long windows to distinguish urgent sustained budget consumption from transient noise; the exact alert policy follows the service and response needs.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Rapid releases consume reliability tolerance | Use an explicit error-budget policy and stabilize the service. |
| Risky new version | Canary/staged deployment with meaningful user-health signals. |
| Repeated manual incidents | Automate understood recovery and remove the root cause. |
Traps
- An SLA and an internal SLO can have different targets and consequences.
- Rollback cannot automatically undo an incompatible data migration.
Active recall
1. SLI versus SLO?
An SLI is a measurement; an SLO is its target over a defined window.
2. What does an error budget represent?
The unreliability tolerated by the service objective.
3. Why can CPU-only paging be noisy?
Resource use can vary without user impact; the signal needs operational context.
4. What limits a chaos experiment?
An explicit hypothesis, scope, stop conditions and recovery plan.
5. What makes a postmortem useful?
Evidence-based contributing causes and owned improvements that are verified afterward.