A colleague walks into a planning meeting with a working application they built over the weekend. It addresses a genuine operational problem. The interface looks polished. The demonstration takes five minutes and the value appears obvious.
The natural question is: how quickly can we launch it?
Generative AI has changed who can create software and how quickly an idea can become tangible. Teams can explore and test concepts before committing significant budget, while developers can produce and revise code faster.
That is a valuable shift. It can also create a dangerous illusion.
The prototype has compressed the time needed to produce evidence. It has not removed the work required to understand users, protect data, integrate systems, test failure modes or operate the service.
The 2025 DORA research into AI-assisted software development describes AI as an amplifier of an organisation’s existing strengths and weaknesses. It found that 90% of technology professionals surveyed used AI at work and more than 80% believed it had increased productivity. Yet 30% reported little or no trust in AI-generated code.
For business leaders, the priority is not to slow experimentation. It is to create a deliberate route from prompt to production.
A prototype is evidence, not a finished product
A useful prototype can answer important questions. Does the proposed experience make sense to users? Can available data support the interaction? Is the potential benefit large enough to justify further investment? Where do existing processes create friction?
It should be judged against those questions.
Problems begin when a prototype is treated as nearly finished because the visible experience appears complete. Most missing work sits behind the interface: permissions, integrations, data quality, security, performance, support and lifecycle cost.
AI-generated software does not need a completely separate delivery discipline. It needs established product and engineering disciplines applied with a clear understanding of new risks.
NIST’s secure software guidance for generative AI makes an important point: organisations should not distinguish between human-written and AI-generated source code when assessing vulnerabilities. All code needs appropriate evaluation before use. The speed or method of generation does not change the standard the finished service must meet.
Six tests of production readiness
Before a promising build becomes a live service, leaders should expect evidence across six connected areas.
1. The service solves a defined problem
A polished demonstration can attract support before the underlying problem has been properly examined. Start by defining who the service is for, what they are trying to achieve and which part of the current experience needs to change.
This matters because automating the wrong step can make a weak process faster without improving the overall outcome. A new interface may also move work elsewhere, create extra checks or introduce a new failure point for employees or customers.
The product owner should be able to explain the intended outcome, affected users, redesigned process and measures of success.
2. The architecture can support the real environment
A prototype often works with sample data, one user and a limited set of actions. Production introduces identity, permissions, live data, existing systems, concurrent use and contractual or regulatory constraints.
The technical team needs to decide what should be retained, rebuilt or replaced. It should map data flows, define access controls and identify dependencies on models, libraries, vendors and infrastructure.
This is also where legacy systems become a strategic issue. A prototype may bypass them during a demonstration, but the live service still needs reliable access to the organisation’s digital core.
3. Security and privacy are designed in
Security cannot be added as a final review. The UK National Cyber Security Centre’s guidelines for secure AI system development organise the work across secure design, development, deployment, and operation and maintenance. That lifecycle view is essential.
Teams need to consider conventional software risks and AI-specific threats such as sensitive information disclosure, prompt injection or excessive agent permissions. The OWASP Top 10 for LLM Applications 2025 provides a practical reference.
The aim is not to apply every possible control. It is to understand the solution’s likely impact and build proportionate protection into its design.
4. Quality is demonstrated through evidence
“It worked in the demo” is not a release criterion.
The team needs test scenarios that reflect normal use, edge cases, misuse and foreseeable failure. For an AI-enabled service, this may include the accuracy and relevance of outputs, inconsistent behaviour, harmful responses, unauthorised data access, performance under load and the effectiveness of human review.
Testing should also cover integrations, permissions, accessibility, recovery, monitoring and rollback. The production decision should rest on agreed evidence, not excitement about the prototype.
This is an area where human expertise remains central. In the 2025 Stack Overflow Developer Survey, 75.8% of respondents said they did not plan to use AI for deployment and monitoring, while 58.7% said the same about committing and reviewing code. AI can support these activities, but accountability for quality cannot be delegated to the tool that generated the work.
5. Someone owns the live service
Moving into production changes the ownership question. The person who created the prototype may not be the right person to operate the service.
Every live solution needs a business owner accountable for the outcome and named responsibility for day-to-day performance, incidents, user feedback, supplier changes and retirement decisions.
For a mid-sized organisation, these responsibilities can sit within existing product, operations or technology roles. The structure does not need to be elaborate, but it must be explicit. This is the practical delivery layer beneath the operating model described in.
6. The organisation is ready to adopt the change
A technically sound application can still fail if it does not fit the way people work.
Leaders should identify which roles, decisions and hand-offs will change. Users need to know when to rely on the service, apply judgement or report a problem. Managers need measures of the intended outcome, not simply activity.
The UK Government AI Playbook offers a useful principle for any organisation: manage the full AI lifecycle and maintain meaningful human control at the right stages. Production readiness includes the people and operating environment around the technology.
Use two speeds, with a clear gate between them
Organisations do not need to apply full production controls to every early experiment. That would remove much of the value of rapid prototyping.
Instead, use two connected delivery modes.
Exploration mode should make it simple to test a problem, user experience or technical assumption in a controlled environment. Use limited data, restrict access, record the owner and make it clear that the output is not a live service.
Production mode should begin only after an explicit investment decision. At that point, the organisation agrees the business outcome, scope, ownership, architecture, risk level, evidence required for release and ongoing operating model.
A practical route has four stages:
- Frame: define the user, problem, intended outcome, constraints and owner.
- Prototype: test the highest-risk assumptions with the smallest useful build.
- Decide: review the evidence and choose whether to stop, revise, buy, integrate or build for production.
- Harden and release: complete the engineering, assurance, operating and change work required for the agreed level of risk.
Teams then know when they are learning and when they are creating something the organisation will depend on.
Calls9 helps organisations make that assessment and complete the work needed for production. We can review the business case, user experience and existing code, identify what can be retained, and create a clear route to release. We can then develop the production version, connect it to existing data and systems, test it, put the right controls in place and support its continued improvement.
The questions leaders should ask before approving production
Senior leaders do not need to review code. They do need to test whether the production decision is complete.
Ask:
- What business or service outcome will this improve, and how will we measure it?
- What did the prototype prove, and what remains untested?
- Which data, systems, suppliers and permissions will the live service depend on?
- What could go wrong for users or the organisation, and how will the service fail safely?
- What evidence must exist before release?
- Who owns the outcome, the live product, technical operation and significant risk decisions?
- What will adoption require from users, managers and support teams?
- What will it cost to run, monitor and improve the service after launch?
If these questions cannot be answered, the organisation does not yet have a production plan. It has a promising prototype.
Calls9 can work through these questions with leadership, product and technical teams. We assess the prototype, identify what is missing and recommend whether to strengthen it, rebuild parts of it, integrate it with existing systems or take a different route.
If you have created an AI-assisted prototype and want to understand what it will take to make it production-ready, contact Calls9.
* This articles' cover image is generated by AI




