Alerts
🔴 checkout · p95 latency above 800ms for 5m · 3.1% of requests affected · runbook: /runbooks/checkout-latency
How I work
Your team stops waiting on infrastructure. Things are done right the first time, written down as they happen, and left cleaner than I found them.
01
Your infrastructure, in code you own
You get
The first pull request
02
A merged pull request goes to production
You get
A pipeline run
Merged 09:14, in production 09:20. Rollback is reverting the merge; it takes as long as this did.
03
Alerts that page on customers, not CPU
You get
An alert, and the morning after
#alerts
Alerts
🔴 checkout · p95 latency above 800ms for 5m · 3.1% of requests affected · runbook: /runbooks/checkout-latency
Alerts
🟢 resolved · on-call followed the runbook: reconciliation job paused, checkout recovered in 4m
ShlomoDelivOps
Saw the night. The job now has its own node pool so it cannot evict checkout again — PR up, no downtime to apply.
04
The baseline an auditor expects, before they ask
You get
The findings report
The other 37 are logged with a reason each. The four above are pull requests this week.
05
A bill somebody owns
You get
The monthly cost message
#infra
ShlomoDelivOps
March: $18,420, down 9% from February. By service: EKS $7,900 · RDS $4,100 · data transfer $2,300 · S3 $1,600 · other $2,520. Next: the cross-AZ transfer on the ingest workers, worth about $900 a month. Commitments still on hold until that settles.
Founder
first time I have understood our bill. go ahead.
Illustrative, and typical. Real pull requests, alerts and reports stay in clients' workspaces.
What I run day to day. If yours isn't here, ask — it is usually a short conversation.
AWS is the deepest, since 2018. The others as clients bring them.
Thirty minutes. Tell me what it looks like now and I'll tell you what I'd do in week one.
A few companies at a time
Shlomo commented
Zero changes to production. It is now code in your repo, and the next PR can change it safely.