Right now, no matter where you go, you can’t avoid the discussion on AI technology. There are two sides of the debate- one says it’s the end of knowledge work as we know it- adopt it everywhere, or get left behind. The other says it’s a bubble built on stolen data, unsustainable power draw, and unreasonable expectations.

Both camps are loud. Both have valid points to make. And a lot of the loudest voices aren’t impartial- whether they’re selling tokens, graphics cards, data centers, or selling fear, clicks, and ad space.

Meanwhile, the actual failure modes are piling up: AI slop flooding every feed, embarrassing production incidents, CEOs making irresponsible calls about what to automate and who to cut, and software organizations tokenmaxxing- blowing their budget on AI consumption- with little to show for it. A widely-cited MIT study found 95% of enterprise AI deployments fail to deliver measurable value, and the study’s own conclusion points to organizational failure, not the technology itself.

Let’s bring it back to the topic of site reliability engineering. Just like virtual machines, cloud, and containerization, this technology is here to stay. How do we use it effectively to solve problems, the same way we practice our craft everywhere else?

My guest today is Stephen Roylance, a systems and information infrastructure engineer with 30 years across health care, e-commerce, and big tech- including time we spent working together in Meta Production Engineering. These days, he and his wife Susan run a small applied AI consultancy out of a solar-powered home in the hill country of Western Massachusetts.

Stephen’s spent three decades watching hype cycles come and go. So I wanted his perspective- not the hyped version, not the hot-take version, but the experienced practitioner’s version- on how we put this technology to work without repeating the mistakes we’re already seeing.

Links: