Software engineering · 4 min read

Taking an AI Prototype to Production

Follow an invoice through a prototype to check data access, extraction, review, and saving. Find the work needed before the first release.

By Permadyn AIPublished September 9, 2026Updated September 20, 2026

An invoice-review prototype can work on every sample used to build it and still save the wrong total when someone uploads two invoices in one scan. A production review needs to find cases like that before staff rely on the application.

A prototype built with an AI coding tool needs the same code review as other software. If the application also calls a model, it needs a way to evaluate those responses. Code review checks the implementation; model evaluation checks how it responds to inputs the developer may not have anticipated.

Follow an invoice all the way through

For a document review application, try uploading a permitted file, extracting fields, correcting a mistake, and saving the approved record. Check what was actually written to the database. Repeat with a missing invoice number, a duplicate upload, and a user who lacks permission to approve it.

Make a short list of what the prototype is simulating. A hardcoded account, sample response, or button that updates only the screen can be useful during design, but each needs a decision before release.

The first production version might accept one document type and require a person to review every extraction. You can test that narrower release against the documents staff actually receive.

Review the code in parts

Some of the prototype may be worth keeping. Inspect the interface, data model, business rules, authentication, integrations, and deployment separately.

DecisionReasonFollow-up
KeepThe behavior meets the release requirementsRecord the tests and assumptions
RepairThe design works, but a required check is missingName the change and how to test it
ReplaceThe structure cannot support a requirementExplain the failure and plan the migration

For example, a useful review screen may sit on top of a database query that ignores account permissions. Keeping the screen does not require keeping the query. Nor does a poor query justify rebuilding the whole application.

Check permissions before the model sees the data

Extraction should produce candidate fields. Application code validates them, the reviewer approves the change, and the server checks permission and current record state before saving it. Generated text must not become an unrestricted database command or an approval flag.

The same rule applies to an internal search assistant: filter source material by the user's access before passing it to the model. A prompt asking the model to withhold restricted documents cannot do that job. AWS describes this in its RAG authorization guidance.

Keep examples of mistakes

Build a test collection from the documents and questions the application will handle. Include incomplete inputs, conflicting values, and cases that need a person's help. Record the expected result and acceptable variation for each.

Count different failures separately. A misplaced comma, an incorrect invoice total, and an unauthorized update have different consequences. Check the saved record as well as the model response.

Reserve some cases for evaluation and keep them out of prompt tuning. Rerun them when the model, prompt, retrieval logic, or application changes. Report how many examples you tested and which document types they cover; a small passing sample leaves plenty untested.

Interrupt a save

Close the browser during submission. Let a model request time out. Disconnect an integration after it accepts a request but before it returns confirmation.

The application needs to determine whether anything changed before retrying. Otherwise a temporary network failure can create a second invoice or leave the user unsure which version was saved. If the underlying record changes during review, require a fresh approval.

A manual path can keep work moving during an outage. In this example, staff could still enter fields and review the source document when extraction is unavailable.

Estimate the running cost

Measure the time and cost of retrieval, model calls, retries, and staff review together. Set limits on input size, simultaneous requests, and repeated calls. Decide what users will see when a provider or spending limit is reached.

A fixed sequence with one model step may be enough. Anthropic's workflow and agent guidance discusses the tradeoff between prescribed steps and agents that choose their own actions.

Before launch, assign responsibility for hosting, model configuration, tests, and incidents. Decide what gets logged and how long it is kept. Test a restoration and a rollback in the release environment, then introduce the application to a small group of users. Their corrections and support requests become the next set of cases to investigate.

If you need a review of an existing prototype, our product and system assessment identifies the work required for a first release.

The invoice example is fictional. Technical references were checked September 8, 2026.

Working on something similar?