Skip to content
Welzin
← All posts

Engineering

The quiet cost of the prototype that never ships

Aman Mundra · July 7, 2026 · 9 min read

The quiet cost of the prototype that never ships
Image source
Contents

Summarize using AI

Prototypes are seductive. A demo gets applause, a slide deck gets a budget, and both feel like progress. The numbers on what happens next are brutal enough to be famous: MIT's NANDA study found about 95 percent of GenAI pilots deliver no measurable P&L impact against $30-40 billion of enterprise spend; IDC and Lenovo count 4 of every 33 POCs reaching production; and S&P Global watched the share of enterprises abandoning most of their AI initiatives jump from 17 to 42 percent in a single year.

Most commentary on these numbers is about model quality. The studies themselves say otherwise: the failures are organizational and architectural - pilots designed in ways that structurally cannot become products. This post prices the costs that never make the postmortem (they compound quietly), then lays out how the shipping minority sequences the work differently. The one-line version: a prototype that never ships is not a cheap experiment; it is an expensive way to learn that a demo works in a demo.

Table of contents
1. The numbers, and what they actually say
2. Why demos die: the four structural gaps
3. The quiet costs: what purgatory actually spends
4. The 5% pattern: production-readiness as the entry gate
5. A pilot contract worth signing
6. When a throwaway prototype is the right call
7. Takeaways

1. The numbers, and what they actually say

Stack the 2025-2026 evidence and a consistent picture emerges: McKinsey finds nearly two-thirds of organizations stuck in pilot mode; BCG finds 74 percent yet to show tangible value; a March 2026 survey has 78 percent of enterprises running agent pilots and 14 percent scaling any of them; Gartner projects over 40 percent of agentic AI projects cancelled by 2027.

Two honest footnotes before building on this. First, the viral MIT number has methodological critics - "no P&L impact yet" is not the same claim as "failed," and 150 interviews is not a census. Second, and more damning for the excuse-making: MIT's own finding was that the core issue was not model quality but the learning gap in enterprise integration. Whatever the true failure rate is - 60, 88, or 95 percent - no serious source locates the problem in the models. The pattern held before GenAI (the same purgatory statistics haunted classic ML and IoT) and it will hold after, because the mechanism is structural.

2. Why demos die: the four structural gaps

The recurring mechanics across the failure analyses, in the order teams usually discover them:

Note what all four have in common: none is discovered by making the demo more impressive. That is why iterating on the prototype - the instinctive response to stalled pilots - deepens purgatory instead of escaping it. There is even a name for the aggregate behavior: innovation theater.

3. The quiet costs: what purgatory actually spends

The direct budget is the visible cost, and the smallest. The compounding ones:

  • Credibility interest. Every stalled pilot raises the evidence bar for the next proposal. Sponsors who funded three demos that went nowhere do not fund the fourth - including the one that would have worked. The organization's cost of AI capital goes up with each unshipped prototype.
  • Talent drain. Strong engineers leave work that never ships. The people most capable of crossing the production gap are precisely the ones most demoralized by never being asked to - the failure analyses call this out explicitly.
  • Pilot fatigue. Deloitte's term for the terminal stage: teams that live through repeated purgatory cycles become progressively worse at running pilots - institutional knowledge decays, cultural appetite dies, and the org loses the ability to ship even when it finally wants to.
  • The option cost of the counterfactual. The measured-value engagement the team didn't run while polishing the demo. This one never appears in any ledger, and it is usually the largest number in the room.

Pricing these changes the meeting. "The pilot cost ₹40 lakh" invites a shrug; "the pilot cost ₹40 lakh, our next three proposals' credibility, and the quarter we didn't spend on the invoice-automation case with a measurable payback" invites a process change.

4. The 5% pattern: production-readiness as the entry gate

What the studies find on the other side of the divide is remarkably consistent - and none of it is a better model:

5. A pilot contract worth signing

The instrument we use to keep engagements out of the graveyard - one page, agreed before any code:

PILOT CONTRACT (before week 1)
  Decision it serves .... named business decision + owner
  Production data ....... real source, real permissions, week 1
                          (no curated extracts)
  Success metric ........ business number + threshold that
                          triggers the production build
  Kill criteria ......... conditions under which we stop and
                          write it up (a documented kill is a
                          SUCCESS - it cost 6 weeks, not 6 months)
  Path to prod .......... auth model, owner, on-call, budget line
                          sketched NOW, not after the demo
  Expiry date ........... pilots don't get extensions; they get
                          verdicts: ship, kill, or re-scope

The two lines that do the most work are the kill criteria and the expiry date. Purgatory is not a failed pilot - it is a pilot that was never allowed to fail, kept alive by sunk cost and demo applause. Making "kill with a write-up" an explicitly honorable outcome is what converts prototypes from theater into experiments: an experiment produces a decision either way. On a typical portfolio, we would rather kill four pilots in six weeks each and ship the fifth than run five demos for a year - the arithmetic of section 3 says that trade wins even before the shipped one pays back.

6. When a throwaway prototype is the right call

The honest boundary: production-first is not always right, and pretending otherwise is its own theater.

  • Genuine feasibility unknowns. "Can any model read these engineering drawings at all?" deserves a week of throwaway spike work before anyone drafts a contract. The keyword is throwaway - time-boxed, and deleted on schedule regardless of how good the demo felt.
  • Vendor bake-offs. Comparative evaluation on your workload is prototype-shaped by nature. The deliverable is the eval harness and the decision memo - both of which survive into production - not the demos.
  • Stakeholder alignment props. Sometimes a demo exists to make an abstract capability concrete for a budget conversation. Fine - but name it as a prop in the plan, budget it in days, and never let it be mistaken for progress toward a system.

The discipline in all three cases is the same: declare the artifact's fate in advance. Prototypes are only cheap when they are allowed to die.

7. Takeaways

  • The purgatory numbers (95 percent no-impact pilots, 4-of-33 POC conversion, 42 percent abandonment) describe an organizational failure, not a model failure - the studies say so themselves.
  • Demos die at four structural gaps: production engineering, real data, ownership, and user trust. None is closed by making the demo better.
  • The real costs compound quietly: credibility interest, talent drain, pilot fatigue, and the counterfactual project you didn't run. Price them in the proposal.
  • The shipping minority gates entry on production-readiness: evals, governance, a named owner, and real data from week one - the six-times-higher success rate belongs to the eval-infrastructure crowd.
  • Run pilots on a contract: success threshold, kill criteria, expiry date. A documented kill is a cheap success; an immortal prototype is the expensive failure.
  • Throwaway spikes are legitimate - when they are time-boxed and actually thrown away.

References

  1. MIT report: 95 percent of generative AI pilots are failing (Fortune)
  2. That viral MIT study - don't believe the hype (Marketing AI Institute)
  3. Scaling AI from pilot purgatory (Astrafy)
  4. Why 78 percent of AI agent pilots never reach production (Zen van Riel)
  5. Why 88 to 95 percent of enterprise AI pilots never reach production (SoftwareSeni)
  6. Pilot purgatory: why AI initiatives fail to reach production (Velosio)
  7. Why your AI pilots never reach production (MindStudio)
  8. MIT finds 95 percent of GenAI pilots fail because companies avoid friction (Forbes)

Hero image: Hal Gatewood on Unsplash.

We build to production from the start, so the value is real and measured. Explore our other insights or get in touch if you would like to talk it through.

Frequently asked questions

Why do most GenAI pilots never reach production?

Most GenAI pilots stall for organizational and architectural reasons, not because the models are weak. MIT's own study located the core issue in the enterprise integration learning gap, and the failures cluster at four structural gaps the demo skipped: production engineering (auth, tenancy, monitoring), real messy data instead of a curated extract, a named owner with a budget line, and user trust. Making the demo more impressive closes none of these, which is why iterating on the prototype deepens the stall rather than escaping it.

What does pilot purgatory actually cost beyond the direct budget?

The visible budget is the smallest cost; the expensive ones compound quietly. They include credibility interest (every stalled pilot raises the evidence bar for the next proposal), talent drain (strong engineers leave work that never ships), pilot fatigue (teams get worse at running pilots after repeated cycles), and the option cost of the measured-value project you did not run while polishing the demo. Pricing these in the proposal, rather than just the rupees spent, is what turns a shrug into a process change.

What should a pilot contract include to avoid an endless prototype?

A one-page pilot contract, agreed before any code, should name the business decision and owner it serves, commit to real production data with real permissions from week one, set a success metric and threshold that triggers the production build, define kill criteria, sketch the path to production (auth model, owner, on-call, budget), and set a hard expiry date. The two lines doing the most work are the kill criteria and the expiry date: a documented kill that cost six weeks is an honorable success, while a prototype kept alive by sunk cost and demo applause is the expensive failure.

When is a throwaway prototype actually the right call?

A throwaway prototype is legitimate for genuine feasibility unknowns (can any model read these engineering drawings at all), vendor bake-offs where the real deliverable is the eval harness and decision memo, and stakeholder-alignment props that make an abstract capability concrete for a budget conversation. The discipline in every case is to declare the artifact's fate in advance: time-box it, budget it in days, and actually delete it on schedule regardless of how good the demo felt. Prototypes are only cheap when they are allowed to die.

  • Engineering

    Supercharge Your Frontend with Lovable AI

    In this article, we’ll guide you step by step to build your very first React project using Lovable AI, connect it to Supabase, and publish it online. By the...

    9 min read
  • Engineering

    How To Optimize the Performance of a Web App

    A fast and smooth app keeps users happy, lowers bounce rates, and encourages people to come back. But if your app is slow, users quickly lose patience,...

    7 min read
  • Engineering

    A Beginner’s Introduction to React.js

    React.js, commonly known as React, is an open-source JavaScript library developed by Facebook for building user interfaces, specifically for single-page...

    7 min read
  • AI Models

    Claude Fable 5: what it is, how good it is, and when to use it

    Anthropic's most capable model is the first from its Mythos tier, a rung above Opus. The positioning, the specs that matter, the benchmarks, and the practitioner's rule for when its 2x price is actually worth paying.

    16 min read
  • AI Tooling

    Build a Karpathy-style LLM wiki in Obsidian: the implementation guide

    Everyone publishes the concept. Almost nobody publishes the fields, the ingest contract, or the dispatcher. This is a working LLM wiki taken apart piece by piece - the permission model, the three frontmatter fields that carry the system, the five-step ingest, the invariant that makes automation safe, and the two things that actually break.

    23 min read
  • GenAI & Agents

    Evaluate Your Large Language Model

    Large Language Models (LLMs) like GPT-4, LLaMA, and Claude now power critical real-world applications - from customer support bots that might accidentally...

    17 min read
  • GenAI & Agents

    LangChain vs LangGraph vs LangSmith

    Large Language Models(LLMs) are changing how we create AI-powered applications. However, the actual impact depends on the frameworks we use to manage and...

    7 min read