What an AI-Built MVP Still Needs to Reach Production

Execution has gotten cheaper, but the judgment that keeps a system standing under real users hasn't, and that's where an IT services partner still earns its value.

thought leadership8 min read

A serious software release (the kind with real payment flows, permissions, and systems built to recover when something breaks) once came with a price tag near half a million dollars before anyone found out whether the idea actually worked. 

AI has cut that cost dramatically: a team can now get a working prototype in front of real users for a fraction of what it took even a few years ago.

However, AI is not at the point where you can avoid spending money forever. Instead, it has just moved the cost from getting an MVP built to the judgment required to keep it working reliably for the real users who depend on it. The falling cost of that first part is its own story; this one is about what doesn't get cheaper.

mvp decisions

I find that making this judgment correctly requires domain expertise, or more precisely, a memory of past consequences: the record of every time a sensible assumption met a real system and lost. AI only has that memory if we have made it explicit, structured, and available in context, and that is exactly where the risk behind a serious release now lives.

The Rule Nobody Wrote Down

We see this same risk across many of our core industries. Ticketing, for example, carries its own version of the same gap. Producing the code on the ticket is trivial; the harder questions are what the scanner should do after a transfer or refund, or what happens when two entrances are offline and the same ticket appears at both.

Writing the code is usually not the difficult part of a serious release. The problem is the rule nobody thought to write down, the one that only shows up once real customers are moving through the system.

We saw the same pattern in an internal R&D proof of concept for AI-assisted business analysis. AI cut the first-draft time for a product requirements document from roughly 24 hours to about 3.5 hours, but our reviewers rated the result at only around 70% accuracy, and parts had to be rebuilt because of hallucinations, factual inconsistencies, and missing project context. Softjourn's R&D proof of concept produced a useful draft much faster, but it also showed why an experienced reviewer was still necessary: speeding up the first draft did not speed up the reviewer's ability to catch a confident mistake.

Expertise Has to Change Too

It would be comforting to think domain knowledge stays valuable simply because AI struggles with some of it today, but I do not think that this will hold true forever: models, context tools, test generation, dependency discovery, and architectural analysis will all keep improving from here.

METR ran into this problem while trying to repeat its developer-productivity research. By early 2026, it believed newer tools were probably creating real gains, but the size of the effect had become difficult to measure. Developers increasingly avoided tasks they might have to complete without AI, and some changed which work they selected in the first place. As METR put it in a 2026 update on measuring AI uplift, the method had to change because the way people worked had changed.

developers working

Experience has a weakness of its own: too often, it lives in the head of the person who was there last time. That person becomes indispensable and eventually becomes the bottleneck. We need to turn what they know into usable rules, examples, tests, and warnings: material that both people and AI can work with.

A better model and better context still leave a business decision for somebody else: even if the model spots a failure first, it cannot accept the loss on behalf of the company or make the uncomfortable call to the client.

The NIST AI Risk Management Framework reflects this distinction, putting roles, accountability, monitoring, and executive responsibility at the center of AI risk management: tools can take on more work, but they do not take on accountability.

A Prototype is Not a Production System

The cleanest way to separate the two is to ask different questions. A prototype asks, "Can this idea work under the conditions we have defined?" Production asks, "Can we depend on this system when the conditions are no longer controlled?"

Testing is predictable in one important way: we know when and what we are testing. Production is not like that; problems show up when nobody is waiting for them. For example, a permission is wrong, data doesn’t arrive, or some external service is down. First we have to find out, then figure out how to recover. Next, if rollback is the right answer, who makes that decision? And if the team has to figure this out during the incident, we went live too early.

We learned this while building an internal AI-searchable knowledge platform. The team's estimate was that prompts, model calls, and vector search (the part most people would recognize as "the AI") took about 10% of the effort. The less visible work was preparing data, building the infrastructure, connecting the pieces, handling errors, deploying, and testing.

middle graphic rnd case study

Ten percent is our internal estimate from that R&D project, not a general ratio for the industry. But it does explain why an AI feature can be cheap while a production-oriented AI system is not. The expensive part begins where the demo loses control over the conditions.

What This Means for IT Services Business

Our industry grew up selling scarce engineering capacity. Clients needed teams because producing software took a great deal of time. That assumption is weakening: not all at the same pace, but enough to change buying behavior.

As an industry, we have a choice to make: we can keep charging for the hours a project takes –the way software services firms always have– or we can get more precise about pricing the actual result a client is buying. I do not think clinging to hourly billing is a winning position for much longer.

Developers are not disappearing, but their work is moving. Less value will come from manually producing a predictable block of code, while more will come from framing the problem, making context usable, reviewing what AI produces, testing the system against reality, and preparing it for production.

Code will keep getting easier to obtain, and I don't expect clients to pay a premium merely because a team can produce it. However, I would suggest investing in IT teams who have already debugged a production outage nobody could reproduce in staging, and who have already worked through the industry problems you're facing.

How We Handle the Age of AI

At Softjourn, one practical response has been to separate a four-to-eight-week proof of concept from the larger production commitment. Before we begin, we agree on the hypothesis, the evidence we hope to see, and what would lead to a go, hold, or no-go decision. Four to eight weeks is not the important part; making the investment decision before funding the whole program is.

If a client is still asking whether the product idea deserves serious investment, we’d spend the next few weeks proving or killing the idea, which an AI-assisted proof of concept can likely do, depending on what the project actually requires.

If the client already has a credible AI-generated MVP, I would pause before adding another sprint. Before going full speed ahead, let someone outside the build take a look at the architecture, security, dependencies, tests, and operational ownership. That production-readiness and technical due-diligence conversation may save more money than another burst of code, before the system meets real users and real money.

code audit team for mvp

The New Bargain

The disappearance of the $500K MVP does not mean a client can avoid investing time and money in development. It means the client no longer has to place the whole bet before learning whether the problem is worth solving. 

This changes the bargain between a client and an IT services company; clients will pay less for routine execution and expect more judgment before, during, and after it to ensure reliability once users are there. The hard part is still there; what is changing is what clients are paying us to be good at.

If you're weighing whether your team's AI-built prototype is ready for real users, or still deciding whether an idea is worth the investment, that's the conversation to have with us before you commit the budget. Contact Softjourn to talk through where your project actually stands.

What Our Clients Say

  • Your team has provided us with outstanding service and outcomes. We couldn't be happier with your work or our progress. All of the members of your team have each shown themselves experts in their respective areas and have been a pleasure to work with.

    Ben Melton

    Product Owner at CapStorm

    Read case study →
  • The partnership, commitment, and skill of the Softjourn team enabled us to navigate this product transformation effectively.
    Eric Rauch

    Eric Rauch

    Co-Founder of Pivot, Pivot

    Read case study →
  • The Softjourn team was very quick to response to issues as well. I'm happy with the result.

    Mike Kenefsky

    Operations Director at PM Vitals, PM Vitals

  • Softjourn's pragmatic approach spotted potential blockers early on, ensuring we stayed on track.
    Sam Mogil

    Sam Mogil

    CEO & Co-Founder, SquadUP

    Read case study →
  • Softjourn's pragmatic approach spotted potential blockers early on, ensuring we stayed on track.
    Richard Bates

    Richard Bates

    Director of Product at Spektrix, Spektrix

    Read case study →
  • Wonderful work on our platform – everything looks great, and you did such a great job!

    Myers-Briggs

    Team Leaders, Myers-Briggs

    Read case study →

Partnership & Recognition

Want to Know More?

Fill out your contact information so we can call you