Prototypes Need a Protocol
A vibe-coded demo looks like proof the idea has value and proof it can be built. Usually it’s neither. Drug trials solved a version of this problem long ago: decide in advance what evidence each stage has to produce.
I’ve made this mistake myself. I vibe-coded a working demo, showed it in a meeting, and the room liked it. I took the reaction as a green light. But I hadn’t decided what the demo was supposed to prove, so nobody, including me, could say what should happen next: fund it, hand it to engineering, or let it quietly become a tool a few teams depend on and nobody owns.
The debate that usually follows is about who should be allowed to build. The more useful question is what the demo actually proved.
In an earlier post I wrote about what goes missing when a prototype reaches engineering. This post is about what should happen before that handoff.
None of this structure is new. Robert Cooper’s Stage-Gate process has put go/kill decisions between phases of product development for more than thirty years. NASA’s Technology Readiness Levels describe nine steps from basic research to “flight proven.” Eric Ries’s Lean Startup framed the MVP as a way to start learning as quickly as possible. Marty Cagan’s four big risks (value, usability, feasibility, and viability) separated whether people want something from whether it can be built. All of them agree that evidence should increase as investment increases.
Those frameworks were built when starting was expensive, so the gates protected budgets. Now starting costs almost nothing. What’s still expensive is finding out whether it’s worth building: real users, real data, and enough time to see a result. The ladder still applies, but what it protects has changed.
Clinical trials add one thing worth borrowing: each phase answers a different question and is judged only on that question. Phase I asks whether a drug is safe, Phase II whether it works, Phase III how it compares with existing treatment in a larger group, and Phase IV watches it after release. Researchers have already applied these phases to evaluating AI in medicine. The same logic works for the build process in any org.
That’s what went wrong with my demo. I built it to show the idea had value. What it produced was a room of colleagues who liked it, which is evidence the idea is persuasive, not that users need it. And because it ran, people heard “it works” as “it’s feasible,” a claim I never meant to make. The demo proved neither value nor feasibility, but it looked like both. It was one stage asked to answer questions it wasn’t designed for.
Five practices follow.
Decide what counts as success before the demo. Clinical trials register their endpoints in advance. Medical journals began requiring this in 2005 to stop researchers from reporting only the results that looked good. A demo invites the same thing: you remember what impressed the room. Before anyone builds, write down the question, the result that counts as a yes, and the result or date that counts as a no. A good endpoint names a cost. Ask people if they want something and they’ll say yes. Ask if they want it at the cost of something else and you get an honest answer. “The room liked it” isn’t an endpoint. “The team would give up next quarter’s slot for it” is. Annie Duke calls the no side kill criteria in Quit. Astro Teller at X adds an order: “tackle the monkey first,” the hardest, most uncertain part before the easy, impressive one.
Match fidelity to the question. Requiring production-quality code for a proof of concept is like running a 3,000-person trial to check whether a compound binds to its target. The opposite mistake is worse: treating polish as evidence. Extreme Programming calls code written only to answer a question a spike, and you expect to throw it away. State the fidelity level up front and review the work against it.
Check safety before efficacy. Teams often find the data access, security, or compliance problem after a successful pilot, when the most time has been spent. Trials check safety first because a drug that isn’t safe doesn’t need an efficacy study. A short review of what the prototype touches would stop a lot of pilots that were never going to ship.
Treat the prototype as evidence, not the product. In my experience, the reluctance to throw a prototype away isn’t about what it cost to build. It’s the feeling that the work is half done. It usually isn’t. The prototype built the half that demos well. The security, scale, and maintenance half hasn’t started, and the prototype’s shortcuts can make it harder. Fred Brooks framed the choice in 1975, in “Plan to Throw One Away”: the only question is “whether to plan in advance to build a throwaway, or to promise to deliver the throwaway to customers.” Twenty years later he called that advice “too simplistic” and endorsed building in small increments instead, but the choice he named still applies to a prototype. The chemist who found the compound doesn’t run Phase III, and the person who built the prototype usually shouldn’t own it in production. Name the next owner early, and assume a rewrite.
Count kills as a result. About 8% of drugs that enter Phase I are approved, and nobody in pharma calls the rest waste. Orgs that treat a stopped prototype as a failure get fewer stopped prototypes and more unowned tools running in production. Track how many ideas started, how many were stopped, and at which stage.
The comparison has two limits. The first is reversibility. A drug can’t be taken back once a patient takes it, but most software can be rolled back. Jeff Bezos’s one-way and two-way doors apply here: the bar for evidence should depend on whether a mistake can be undone and what it would cost, not on a fixed sequence. An internal tool that touches nothing sensitive can go straight to a pilot. Something touching customer data or money can’t. It’s the same point I made about the cost of finding errors: how much checking you need depends on what a miss costs.
The second is time. Take the logic from clinical trials, not the timeline. Each stage should take days or weeks. Trials have faster variants too, such as small exploratory studies before Phase I and adaptive designs that make planned changes based on interim results.
This also isn’t a case for gating exploration. Nobody pre-registers basic lab research, and nobody should need a protocol to tinker. The protocol gates commitment, not creation. It starts when a prototype asks for something scarce: users’ time, real data, or engineering capacity.
Here’s the protocol, short enough to use:
Question: Which risk does this stage retire (value, feasibility, or safety), and what does it need to answer?
Endpoint: What result counts as a yes?
Kill criteria: What result, or what date, counts as a no?
Blast radius: What does this touch if it goes wrong?
Next owner: Who takes it if it passes?
Go back to the demo in the meeting. With a protocol, the reaction in the room matters less. Everyone knows what the demo was meant to prove, whether it did, and who takes the next step. The question was never who gets to build. It’s what the building is supposed to prove.
← Back to writing