Almost four years ago, right after ChatGPT launched, I started writing this newsletter, The Hitchhiker’s Guide to AI, to learn about the then-nascent AI space and share what I learned with other curious builders and operators in my network. You subscribed then because you were as curious about AI as I was. At the time, most of us were trying to figure out generative AI, large language models, and AI agents for the first time. What could these models do? Where was this going? How much of the hype was real? I kept writing for a while, and that eventually led me to the conviction to found an AI startup. Writing took a back seat to building to building. For the last three years, I’ve been deep in deploying AI agents into production for some of the world's leading fintechs, like Airwallex, Coinbase, IG, and Worldpay. We started with AI for compliance, where the stakes are high, and mistakes have real consequences. Today, with our latest product, Grep.ai, we’re applying what we learned to a much broader category: high-stakes, repetitive knowledge work. A few weeks ago, a customer asked me how other companies were approaching AI. What strategies were working? How were they rolling agents out? Where were they getting stuck? I realized I answer versions of these questions all the time. I talk to teams building their own agents, executives deciding where to deploy them, and operators responsible for making them work once the demo is over. I’ve just never written much of it down. So I’m starting now. I’m also renaming this publication The Operator’s Guide to AI. When I started writing, we were all still trying to understand AI. Today, many of you deploy it. You have agents in production, teams trying to scale them, a board asking about ROI, and your name on the line when something goes wrong. How do you decide which work is ready for agents? How do you evaluate them? Where do humans stay in the loop? What does reliability look like when an agent is doing thousands of tasks a day? How do you know if any of this is creating real value? That’s what I want to write about here. I’ll share what we’re learning from running agents in production at Grep, what I’m seeing from companies doing their own rollouts, and the practices that seem to hold up once AI meets real work. After three years of building, I have a lot to write down! The first edition, which will land later this week, will focus on why AI agents fail in production, the most common pitfalls, and why it’s often the harness, not the model, that is the culprit. If you’re deploying AI into production, automating mission-critical work with agents, or leading AI transformation at your company, I hope you find this newsletter helpful! About the author: |