How I work

What changes when
someone owns it.

Your team stops waiting on infrastructure. Things are done right the first time, written down as they happen, and left cleaner than I found them.

What you can count on

  1. I own it, end to endDesign, build, run, and the answer when it breaks. You don't manage the work; you see it land.
  2. Everything is in codeTerraform in your repo, changed through pull requests. Nothing is clicked, nothing lives only in someone's head.
  3. You talk to the person doing the workThe engineer who reads your message is the one who fixes it.
  4. Measure before you cutPer-service cost breakdown first. Incident history first.
  5. Boring beats cleverThe architecture your team can run without me.
  6. Built so you can leaveRunbooks written as the work happens. A structured handover when you hire in-house.

01

Terraform

Your infrastructure, in code you own

You get

  • An infrastructure repo in your GitHub, owned by you
  • A written account review before anything changes
  • CI that plans on every PR

The first pull request

acme / infrastructure · pull request #12
open

import: bring the production VPC and RDS under Terraform

  • terraform
  • import
  • no-op
module.network.aws_vpc.prod# imported, not recreated
module.network.aws_subnet.private[0..2]
module.db.aws_db_instance.main# imported
module.db.aws_db_parameter_group.main
Plan: 0 to add, 0 to change, 0 to destroy. 6 resources imported.
  • terraform plan41s
  • tflint8s
  • checkov / policy19s

Shlomo commented

Zero changes to production. It is now code in your repo, and the next PR can change it safely.

02

Pipelines

A merged pull request goes to production

You get

  • One pipeline template, reused by every service
  • Each environment in its own Terraform, sharing one set of modules
  • Deploy notifications in your Slack

A pipeline run

acme / checkout-api · actions · main @ 4f2c1e9
passed
  1. build1m 12s
  2. test2m 04s
  3. scan38s
  4. deploy staging54s
  5. smoke21s
  6. deploy production58s

Merged 09:14, in production 09:20. Rollback is reverting the merge; it takes as long as this did.

03

Monitoring

Alerts that page on customers, not CPU

You get

  • Alerts on what customers feel, not CPU
  • A runbook linked from every alert
  • One dashboard per service, same shape

An alert, and the morning after

#alerts

Alerts

🔴 checkout · p95 latency above 800ms for 5m · 3.1% of requests affected · runbook: /runbooks/checkout-latency

Alerts

🟢 resolved · on-call followed the runbook: reconciliation job paused, checkout recovered in 4m

ShlomoDelivOps

Saw the night. The job now has its own node pool so it cannot evict checkout again — PR up, no downtime to apply.

  • 🙏 2

04

Security

The baseline an auditor expects, before they ask

You get

  • IAM, encryption, logging and WAF in Terraform
  • Policy checks in CI that stop regressions
  • The evidence trail SOC 2 auditors ask for

The findings report

acme · account review · security
early
41findings
4that matter
37noise, noted
  1. Root account has no MFAFix today
  2. 3 IAM users with admin keysReplace with roles
  3. S3 bucket public: exports-prodClose, then audit access
  4. RDS not encrypted at restSnapshot, re-create encrypted

The other 37 are logged with a reason each. The four above are pull requests this week.

05

Cost

A bill somebody owns

You get

  • A tagged account with a per-service breakdown
  • A monthly report written for founders
  • Budgets and anomaly alerts before the invoice

The monthly cost message

#infra

ShlomoDelivOps

March: $18,420, down 9% from February. By service: EKS $7,900 · RDS $4,100 · data transfer $2,300 · S3 $1,600 · other $2,520. Next: the cross-AZ transfer on the ingest workers, worth about $900 a month. Commitments still on hold until that settles.

Founder

first time I have understood our bill. go ahead.

Illustrative, and typical. Real pull requests, alerts and reports stay in clients' workspaces.

Platforms and tools

What I run day to day. If yours isn't here, ask — it is usually a short conversation.

Clouds
  • AWS
  • GCP
  • Azure

AWS is the deepest, since 2018. The others as clients bring them.

Infrastructure as code
  • Terraform
  • Pulumi
Containers
  • Kubernetes
  • EKS
  • GKE
  • ECS
CI/CD
  • GitHub Actions
  • GitLab CI
  • Jenkins
  • ArgoCD
Observability
  • Grafana
  • Prometheus
  • OpenTelemetry
  • eBPF
  • Datadog
  • CloudWatch
AI and ML
  • Bedrock
  • SageMaker
  • GPU nodes on EKS
Infrastructure I have built or run, 2018–present

Not sure which one you need first?

Thirty minutes. Tell me what it looks like now and I'll tell you what I'd do in week one.

A few companies at a time