Blog
Why Uptime and Away
What this site is for, who it is for, and the four threads I plan to pull on: AI in operations, SRE, platform engineering and service management.
I have spent most of my career on the side of technology that nobody notices until it stops working. Reliability, platforms, the processes that get a change from someone’s laptop into production without waking anyone up at 3 a.m. It is good work, and it produces a steady stream of lessons that are too long for a chat message and too specific for a conference talk. This site is where those go.
The four threads
AI in operations. Not the demo. The part where you try to put a model in front of an incident timeline, a change calendar or a backlog of tickets, and discover what it is actually good at. I will write about what worked, what quietly made things worse, and how to tell the difference before the post-mortem does.
Site reliability engineering. SLOs that people actually use, on-call that does not burn through a team in a year, and the unglamorous tooling that makes both possible.
Platform engineering. Building the paved road, and the harder problem of getting people to drive on it. Internal developer platforms, golden paths, and the organisational design that decides whether any of it sticks.
Service management. The unfashionable one. Change, incident and problem management still decide how fast a company can move. Done badly, they are theatre. Done well, they are how you ship on a Friday afternoon and sleep.
And the “away” part
When I am not on a laptop I am usually on a road or a trail somewhere with a camera. Those go in Travel and Photos. They have their own sections and their own feeds, so if you are here purely for the SRE content you will not be ambushed by mountains.
How to follow along
There is an RSS feed for everything, or one per section. Posts will be irregular but, I hope, worth the wait. Thanks for reading.