A wood-framed building left standing after a storm, with a large heap of blown-down loose lumber piled in front of it.
Going up, or blown down? From the pile alone you can't tell. The frame is what tells you.

All the work Startuplandia does for people is done under a contract that indemnifies us from liability for the software we produce. In every case that’s the correct move, and I wouldn’t change course. But if, contractually speaking, we lack that liability — how then do we know our work is trustable?

The question turned concrete for me in mid-August, when I was confronted with the Harvest problem. Harvest — the time-tracking and invoicing service we’d run on for over a decade — sent notice of a steep pricing increase (monthly $79 → $939, +1,089%).

Harvest support email: Basic plan at $79.00/monthly with 8 seats, scheduled to renew on the Enterprise plan with Unlimited billing at $939.00/monthly.

That put me squarely in front of the classic build-versus-buy decision. And what pulled on me as I weighed it was something I’d been sitting with for a while: Harvest really hasn’t improved much in a decade.

So I decided to build it myself

Faced with the increase, I really saw only two roads: find another vendor, or build my own. Much like my decision to move to Austin, I knew what had to be done.

Building appealed to me for reasons that had been forming long before Harvest forced the question. I’ve become fixated on owning and custom-tailoring every operational routine Startuplandia runs on — and, more specifically, on instrumenting and productizing the line between human time and agent compute. Time tracking and invoicing sit directly on that line.

I also had a handful of things about Harvest I’d disliked for years, small annoyances I’d rebuild given the chance. And the deeper I looked, the more the case made itself: in over a decade and forty thousand logged hours, the surface area of what we actually used had never meaningfully changed. A tool that stable, doing work that well understood, is prime agentic-rebuild territory. Invoicing and time tracking had started to feel complacent.

Underneath all of it is a belief I’m increasingly certain of — that owning your core workflows will be a defining source of differentiation from here forward. The old build-versus-buy conversation would have called it madness to build something I’d then have to support indefinitely. I decided to build it anyway.

And I had almost no time. The notice landed August 20th, and I wasn’t about to track hours in two systems, so the new one had to be live by September 1st — under two weeks.

What happened

I began the build on Friday, August 28th. By Monday the 31st I had a fully built, unit-tested time-tracking system in place — a stable system of record, and honestly a huge win. I archived every team member out of Harvest and invited them into the new one.

The new Startuplandia time-tracking system, live and in production.

The whole thing took under twenty hours of cooperative work with AI.

Along the way I fixed the handful of Harvest annoyances I’d been carrying, and a few of them mattered more than “annoyance” lets on: long-form markdown in time descriptions, image uploads to document a log, an easy way for teammates to see all their monthly logged time since we self-manage, retainer-based billing running alongside hourly, and automatic, intelligent charting of gross margin in all its texture and complexity. By September 1st the whole team was onboarded and tracking time in a system plainly better than the one it replaced.

The reception surprised me — “great layout,” “markdown is sweet.” Small things, but morale is not a small thing.

Team chat feedback on the new system: 'Great layout' and 'markdown is sweet.'

And I realized the complacency I’d tolerated around invoicing had been holding back how I thought about — and managed — our gross margin. Those twenty hours corrected a perspective I’d let harden for a decade — and sooner or later you’ll have to do the same, and take a hard look at the things you’ve simply accepted as being the way they are.

It all runs on one loop: specify, run, evaluate

This is the part that matters most — and it reaches well past this one build. What follows is drawn from fifteen years as a product engineer and, more sharply, from the last year of working almost entirely this way: since Claude Code arrived late in 2025, across my own systems and the three founders’ vibe-coded projects I help support.

The work runs on a loop I think of as specify, run, evaluate. I specify what I want, the agent runs, and I evaluate what comes back before anything moves forward. The mastery lives almost entirely in the specification — and I don’t only mean code and architecture. It’s specifying the feature and UX requirements, and the general principles the system has to hold to, with the same care. Before I wrote a line of this system I wrote a detailed pseudo-code scope: the domain models, and the relationships between them, spelled out.

The loop has to run on the smallest unit possible. The more you ask an agent to produce in one pass, the less of it you can actually judge.

And it has to run on the data model first — alone — before any UI, UX, or authentication and roles are built on top of it. Trying to do all of that at once is the fatal flaw. Get the data model right and whole categories of bugs never get written; more than once on this build, a single question about the data model dissolved a bug I’d otherwise have spent a day patching.

I don’t really review the code

A word on what evaluate means here, because I think the usual framing gets it wrong. I don’t really code-review in the traditional sense. Claude knows the syntax far better than I do, and I’ve made peace with that. My work sits upstream of it: I specify the architecture, the needs, the features, the principles, and the constraints — and then my verification lives in more nuanced territory. I’m observing the outputs, evaluating the UX, and challenging how Claude articulates what we’re building and why. It’s closer to pair programming, where I’m the product designer and Claude is the engineer.

“Code review,” as most people mean it, misses this distinction. I’m not trying to interpret a language I no longer speak. I’m using and evaluating its outputs — which is a different kind of scrutiny, and I think a truer one for how this work actually happens.

Most agentic work aims at the wrong part of the loop

Which is why I think much of what’s happening right now — vibe coding, agentic coding, whatever you want to call it — is aimed at the wrong part of that loop. Almost all the energy goes into the run: get the agent producing, more code, more UI, more output, faster. But every unit of output is a unit someone eventually has to evaluate. Past a certain point you’ve generated far more signal than any human can actually judge — and unevaluated output isn’t progress, it only looks like it.

This is software’s Kaizen moment

I’ve come to think software engineering is having its real Kaizen moment. Not Agile — that was a paper-mâché dragon, fierce-looking and hollow. This is the actual thing: software starting to operate like a factory line.

The most important principle in that system — the one the West took the better part of thirty years to absorb after it took hold in postwar Japan — is that you build quality in at every gate rather than inspecting for it at the end. And the reason it won wasn’t only better cars; it was cost. Catching a defect at its station is cheap. Letting it ride to the end of the line — where it takes dedicated inspectors and rework to pull back out — is the most expensive quality there is, which is why Toyota came to treat end-of-line inspection as waste. Specify, run, evaluate is the same idea, gate for gate.

By the time the rest of the industry grasped what was happening, the gap was undeniable: through the 1980s, American cars were shipping defects at nearly twice the rate of Japanese ones.1 Toyota wasn’t a little ahead — it was operating on a different level.

It’s also why I think the current fascination with reviewing pull requests en masse won’t last. The Toyota way tells you why: over any real horizon, a human can’t meaningfully process, parse, or correct the sheer volume of code that piles up at the end of an agentic process. Quality has to be caught inside the line — not inspected at the door on the way out.

The focus is the point — not the automation

This is where I part hardest from the mainstream. The popular posture treats the agent’s focus as the valuable thing and the human’s as optional — set the agent going and step away, do something else while it works. I’ve come to believe almost the exact opposite. When I’m working on something new and sensitive — mission-critical logic like how much money I owe people and how much the business makes — the agent isn’t there to free up my attention. It’s the mechanism that concentrates it.

It was put better than I can, decades before any of this:

“My whole life has been spent trying to teach people that intense concentration for hour after hour can bring out in people resources they didn’t know they had.”

— Edwin Land, founder of Polaroid and an inspiration to Steve Jobs

That is exactly what this feels like. In a working session the agent drops me into a pocket of focus so deep that almost everything outside the immediate problem falls away — a level of concentration I genuinely could not reach before AI. Inside that pocket, Claude and I are working through first principles, constraints, features, and trade-offs; I’m reading the outputs, weighing the UX, pushing on how the work gets explained back to me. The agent isn’t doing the thinking for me. It’s holding open the space that lets me think harder than I otherwise could.

I’ve come to consider that kind of sustained focus an essential skill — maybe the essential skill — in an age engineered to dissociate our attention.

What twenty-two cents proved

So: this process let me rebuild one of the most sensitive systems my business runs on in under twenty hours. My honest estimate of what that same work would have cost me before AI is something like seventy hours a month for three months — call it two hundred and ten hours. Roughly ten-to-one leverage. An adversarial code review of the result clocked it at about sixteen hours of human time against two hundred and three tests, live in production, and called the quality “disproportionately high for the effort.” I’ll take that verdict — but it isn’t the one that matters to me.

Late in the build I had the system produce a report whose entire purpose was to be auditable by hand. It looked right. But I sat down and multiplied one line out myself — 1.28 times 66.5 — and landed twenty-two cents off from what the system showed. Twenty-two cents, on a report I’d just told myself a person could check by hand. The agent’s first instinct was that the gap was rounding noise, essentially close enough. I didn’t accept that. What I said back was, “by the letter of the law you’re correct, but you don’t see the spirit of the law.” A report built for a hand-audit has to reconcile when a human actually does the arithmetic — not merely in some column no one will reach for.

And then the part I’m proudest of: the first fix was still wrong. It corrected the figure I’d caught but quietly mis-reconciled another row — a double-rounding artifact I only found because I kept pulling the thread. That second catch is what produced a rule the system now lives by: minutes are the only exact basis for money, and every number a person can see has to reconcile back to them. That is evaluation aimed at the spirit of the thing, not the syntax — and it is the whole reason I trust what came out.

The one that matters is this: I trust the system enough that it now governs the primary financial and profitability engine of my business — what I owe people, what we make. I can’t think of a stronger way to answer the question I opened with. That is what “do I trust our work” looks like when it’s real.

And it puts the real question back to you as an entrepreneur — greed or stability, speed or correctness. Because if ten-to-one somehow isn’t enough for you, I’d honestly ask how you know a hundred-to-one will be. The market is racing toward that number — more leverage, agents overseeing agents — but a ten-to-one build already earned enough trust to run my money. So what is the hundred-to-one race actually producing? A great deal of it, when you look closely, is thin UX and broken information hierarchy. And here is the part worth sitting with: what happens when a hundred-to-one is also the rate at which errors compound — and, in the end, the rate at which customers leave?

The contract that releases us from liability hasn’t changed, and I still wouldn’t change it. But trust was never going to live in a clause. It lives in whether I’d put my own most important work — my payroll, what I owe people, what the business makes — on the same way of working I’m asking you to trust. I did. Which is why I’d rather not ask you to trust me at all. I’d rather show you the work I already trust.

  1. A 1980s survey by Rachel Dardis found roughly 89 of 100 American cars carried defects against about 47 of 100 Japanese cars — nearly double. The landmark study of the broader quality gap is James Womack, Daniel Jones & Daniel Roos, The Machine That Changed the World (MIT International Motor Vehicle Program, 1990). ↩

John Davison

Founder & CTO, StartupLandia

← All articles