Engineering
The quiet cost of the prototype that never ships
Aman Mundra · July 7, 2026 · 9 min read

Contents
Prototypes are seductive. A demo gets applause, a slide deck gets a budget, and both feel like progress. The numbers on what happens next are brutal enough to be famous: MIT's NANDA study found about 95 percent of GenAI pilots deliver no measurable P&L impact against $30-40 billion of enterprise spend; IDC and Lenovo count 4 of every 33 POCs reaching production; and S&P Global watched the share of enterprises abandoning most of their AI initiatives jump from 17 to 42 percent in a single year.
Most commentary on these numbers is about model quality. The studies themselves say otherwise: the failures are organizational and architectural - pilots designed in ways that structurally cannot become products. This post prices the costs that never make the postmortem (they compound quietly), then lays out how the shipping minority sequences the work differently. The one-line version: a prototype that never ships is not a cheap experiment; it is an expensive way to learn that a demo works in a demo.
Table of contents
1. The numbers, and what they actually say
2. Why demos die: the four structural gaps
3. The quiet costs: what purgatory actually spends
4. The 5% pattern: production-readiness as the entry gate
5. A pilot contract worth signing
6. When a throwaway prototype is the right call
7. Takeaways
1. The numbers, and what they actually say
Stack the 2025-2026 evidence and a consistent picture emerges: McKinsey finds nearly two-thirds of organizations stuck in pilot mode; BCG finds 74 percent yet to show tangible value; a March 2026 survey has 78 percent of enterprises running agent pilots and 14 percent scaling any of them; Gartner projects over 40 percent of agentic AI projects cancelled by 2027.
Two honest footnotes before building on this. First, the viral MIT number has methodological critics - "no P&L impact yet" is not the same claim as "failed," and 150 interviews is not a census. Second, and more damning for the excuse-making: MIT's own finding was that the core issue was not model quality but the learning gap in enterprise integration. Whatever the true failure rate is - 60, 88, or 95 percent - no serious source locates the problem in the models. The pattern held before GenAI (the same purgatory statistics haunted classic ML and IoT) and it will hold after, because the mechanism is structural.
2. Why demos die: the four structural gaps
The recurring mechanics across the failure analyses, in the order teams usually discover them:
- The demo-to-app gap. A production system needs auth, tenancy, access controls, a database, monitoring, rate limits, audit logs, and a deployment pipeline - all of which the POC skipped, because skipping them is what made it fast. The pilot's speed was borrowed from production's budget.
- The curated-data gap. Pilots run on the clean extract someone hand-prepared; production must reason over the real corpus scattered across dozens of systems with permissions, staleness, and contradictions. Half the "AI failed" stories are this gap wearing a trench coat (we wrote about the remedy in Don't wait for clean data - the answer is not a readiness gate, it is piloting on real data from day one).
- The ownership gap. The pilot belongs to innovation; production belongs to nobody. No named owner for the model's decisions, no budget line for its operation, no on-call. Systems without owners do not get adopted; they get demoed.
- The trust gap. Users who weren't in the loop during the pilot meet the system at rollout, with no error-handling workflow, no recourse path, and every incentive to route around it. Adoption is a change-management deliverable, and it was not in the pilot's scope.
Note what all four have in common: none is discovered by making the demo more impressive. That is why iterating on the prototype - the instinctive response to stalled pilots - deepens purgatory instead of escaping it. There is even a name for the aggregate behavior: innovation theater.
3. The quiet costs: what purgatory actually spends
The direct budget is the visible cost, and the smallest. The compounding ones:
- Credibility interest. Every stalled pilot raises the evidence bar for the next proposal. Sponsors who funded three demos that went nowhere do not fund the fourth - including the one that would have worked. The organization's cost of AI capital goes up with each unshipped prototype.
- Talent drain. Strong engineers leave work that never ships. The people most capable of crossing the production gap are precisely the ones most demoralized by never being asked to - the failure analyses call this out explicitly.
- Pilot fatigue. Deloitte's term for the terminal stage: teams that live through repeated purgatory cycles become progressively worse at running pilots - institutional knowledge decays, cultural appetite dies, and the org loses the ability to ship even when it finally wants to.
- The option cost of the counterfactual. The measured-value engagement the team didn't run while polishing the demo. This one never appears in any ledger, and it is usually the largest number in the room.
Pricing these changes the meeting. "The pilot cost ₹40 lakh" invites a shrug; "the pilot cost ₹40 lakh, our next three proposals' credibility, and the quarter we didn't spend on the invoice-automation case with a measurable payback" invites a process change.
4. The 5% pattern: production-readiness as the entry gate
What the studies find on the other side of the divide is remarkably consistent - and none of it is a better model:
- Governance and evaluation infrastructure first. Organizations with unified AI governance put an order of magnitude more projects into production, and systematic evaluation frameworks correlate with nearly six times higher production success. Evals are not QA garnish; they are the difference between "the demo felt good" and "we can prove it should ship."
- Production-readiness as the starting requirement. The escapees only greenlight pilots with a credible path to a deployed, governed application - real data source, real auth model, real owner - named on day one, not discovered in month six.
- The three pillars. People and workflow change, an AI-ready data foundation, and an automated MLOps path to deploy - simultaneously, not sequentially. Two of three still equals purgatory.
- Capability framing over headcount framing. The purgatory cohort asks "how many people can this eliminate"; the scaling cohort asks what their people can now do that they couldn't - which is also the framing that survives contact with the workforce whose adoption you need.
5. A pilot contract worth signing
The instrument we use to keep engagements out of the graveyard - one page, agreed before any code:
PILOT CONTRACT (before week 1)
Decision it serves .... named business decision + owner
Production data ....... real source, real permissions, week 1
(no curated extracts)
Success metric ........ business number + threshold that
triggers the production build
Kill criteria ......... conditions under which we stop and
write it up (a documented kill is a
SUCCESS - it cost 6 weeks, not 6 months)
Path to prod .......... auth model, owner, on-call, budget line
sketched NOW, not after the demo
Expiry date ........... pilots don't get extensions; they get
verdicts: ship, kill, or re-scope
The two lines that do the most work are the kill criteria and the expiry date. Purgatory is not a failed pilot - it is a pilot that was never allowed to fail, kept alive by sunk cost and demo applause. Making "kill with a write-up" an explicitly honorable outcome is what converts prototypes from theater into experiments: an experiment produces a decision either way. On a typical portfolio, we would rather kill four pilots in six weeks each and ship the fifth than run five demos for a year - the arithmetic of section 3 says that trade wins even before the shipped one pays back.
6. When a throwaway prototype is the right call
The honest boundary: production-first is not always right, and pretending otherwise is its own theater.
- Genuine feasibility unknowns. "Can any model read these engineering drawings at all?" deserves a week of throwaway spike work before anyone drafts a contract. The keyword is throwaway - time-boxed, and deleted on schedule regardless of how good the demo felt.
- Vendor bake-offs. Comparative evaluation on your workload is prototype-shaped by nature. The deliverable is the eval harness and the decision memo - both of which survive into production - not the demos.
- Stakeholder alignment props. Sometimes a demo exists to make an abstract capability concrete for a budget conversation. Fine - but name it as a prop in the plan, budget it in days, and never let it be mistaken for progress toward a system.
The discipline in all three cases is the same: declare the artifact's fate in advance. Prototypes are only cheap when they are allowed to die.
7. Takeaways
- The purgatory numbers (95 percent no-impact pilots, 4-of-33 POC conversion, 42 percent abandonment) describe an organizational failure, not a model failure - the studies say so themselves.
- Demos die at four structural gaps: production engineering, real data, ownership, and user trust. None is closed by making the demo better.
- The real costs compound quietly: credibility interest, talent drain, pilot fatigue, and the counterfactual project you didn't run. Price them in the proposal.
- The shipping minority gates entry on production-readiness: evals, governance, a named owner, and real data from week one - the six-times-higher success rate belongs to the eval-infrastructure crowd.
- Run pilots on a contract: success threshold, kill criteria, expiry date. A documented kill is a cheap success; an immortal prototype is the expensive failure.
- Throwaway spikes are legitimate - when they are time-boxed and actually thrown away.
References
- MIT report: 95 percent of generative AI pilots are failing (Fortune)
- That viral MIT study - don't believe the hype (Marketing AI Institute)
- Scaling AI from pilot purgatory (Astrafy)
- Why 78 percent of AI agent pilots never reach production (Zen van Riel)
- Why 88 to 95 percent of enterprise AI pilots never reach production (SoftwareSeni)
- Pilot purgatory: why AI initiatives fail to reach production (Velosio)
- Why your AI pilots never reach production (MindStudio)
- MIT finds 95 percent of GenAI pilots fail because companies avoid friction (Forbes)
Hero image: Hal Gatewood on Unsplash.
We build to production from the start, so the value is real and measured. Explore our other insights or get in touch if you would like to talk it through.










