Operations notes for thinly staffed systems
Keep the service alive. Leave the drama out.
Clear incident runbooks, handoff checklists, and recovery habits for small teams that do not have a full SRE bench waiting offstage.
Markdown download. No signup wall. No tracking pixel.
NormalPrimary user journey works; watch the trend.
DegradedUsers can continue, but risk or delay is rising.
DownCritical journey fails; incident lead takes control.
The runbook shelf
Start with the artifact you need next
Each guide produces something concrete: a log, a handoff packet, a go/no-go record, a postmortem, or an alert policy.
Incident Starter Kit
Take control of the first 15 minutes and establish a clean recovery gate.
RB–02Service Handoff Checklist
Transfer a live system without transferring a pile of undocumented risk.
RB–03Deployment Checklist
Define the change, verification, stop conditions, and rollback before release.
RB–04Postmortem Guide
Turn a service failure into owned, verifiable system changes—not blame.
RB–05Monitoring Basics
Cover the user journey first, then alert only when a human can act.
The twenty-minute handoff test
Can the next operator find the truth without calling the last one?
A service is not handed off until a new operator can answer five questions from written evidence.
- OwnerWho decides, who responds, and who pays the vendor?
- DeployHow does a change reach production, and how is it rolled back?
- RestoreWhere are backups, and when was restore last proven?
- SignalWhich check represents the real user journey?
- SupplierWhich credentials, renewals, and contracts can stop the service?
When the runbook is not enough
Inherited app? Unknown deploy? Fragile vendor chain?
Hoyack helps small teams stabilize, document, and adopt services that arrived without a safe operating path.
Book a service rescue review
Bring the service, the symptoms, and what you know. The first review is for scoping the operating risk and the next useful step—not promising a magic fix.