Durable execution
The agent resumes, retries, and finishes. A dropped model call does not lose the work.
AI made software easy to build and hard to run. I productionalise agents that fail silently, burn money, and break under real traffic. Durable, observable, cost-capped.
Vibecoding got you a working prototype in hours. Now your agent returns wrong answers and nothing flags it. The bill climbs and nobody knows why. The gap between demo and production is where projects die.
The agent resumes, retries, and finishes. A dropped model call does not lose the work.
A test set catches wrong answers before your customers do.
Hard limits on spend and response time. The agent cannot burn the budget.
I work on live agents with real traffic. If you need a demo or a chatbot, I'll refer you to someone else.
Need a demo? I'll refer you to someone else.
We start with the logs.
We look at what broke in production and what the eval missed.
Durable execution, real evals, and hard caps on spend and latency.
You get a sanitised postmortem you can share with your team.
I take a public AI failure, rebuild the assumptions that caused it, and show exactly where production broke them. Each one covers what broke, what the eval missed, and what would have caught it sooner.
The first reconstruction is in progress.
View all writingTell me what broke. I'll tell you if I can fix it.
Start a conversation