Notes from the field.
Engineering, strategy, and operational lessons from the work we ship. Opinionated and specific. No "AI is transforming everything" intros.
Articles
Why your AI agent works in the demo and breaks in production.
Most agent demos run for 30 seconds in a sandbox. Production agents have to work for years. Four patterns that hold up.
RAG is not a product. It's an architecture.
Most teams treat RAG as plumbing. It isn't. The failure modes live in places that never show up in a demo.
The eval is the product.
Without evals you can't deploy with confidence, can't improve with discipline, can't catch the regression that's about to embarrass you.
Build, buy, or wait: a framework for enterprise AI investment.
The interesting AI initiatives are never trivially "buy" or "build." They're the ones in between.
The four ways enterprise AI engagements fail.
Wrong problem, wrong data, wrong adoption, wrong governance. Spotting them early is the difference between deployed and shelved.
Frontier models vs. open-source: when each one wins.
Most production systems we ship use both. The question isn't "which one." It's "where does each one earn its place."
What to ask when interviewing an AI consultancy.
Most consultancies sell strategy slides. Three questions separate builders from deck-makers.
The hidden cost of generic AI SaaS.
You don't pay in license fees. You pay in workflow drift — and that bill compounds.
How to run an AI proof-of-concept that doesn't waste a quarter.
A POC that can't ship is an expensive demo. Three rules to make yours actually inform the decision.
Designing AI for people who hate it.
Frontline workers, clinicians, analysts — the people who actually use enterprise AI often start out skeptical.
Notes from a 200,000-employee frontline AI deployment.
Frontline AI at retail scale isn't a feature. It's a platform decision. Three lessons.