Software engineering · 4 min read
Taking an AI Prototype to Production
Follow an invoice through a prototype to check data access, extraction, review, and saving. Find the work needed before the first release.
An invoice-review prototype can work on every sample used to build it and still save the wrong total when someone uploads two invoices in one scan. A production review needs to find cases like that before staff rely on the application.
A prototype built with an AI coding tool needs the same code review as other software. If the application also calls a model, it needs a way to evaluate those responses. Code review checks the implementation; model evaluation checks how it responds to inputs the developer may not have anticipated.
Follow an invoice all the way through
For a document review application, try uploading a permitted file, extracting fields, correcting a mistake, and saving the approved record. Check what was actually written to the database. Repeat with a missing invoice number, a duplicate upload, and a user who lacks permission to approve it.
Make a short list of what the prototype is simulating. A hardcoded account, sample response, or button that updates only the screen can be useful during design, but each needs a decision before release.
The first production version might accept one document type and require a person to review every extraction. You can test that narrower release against the documents staff actually receive.
Review the code in parts
Some of the prototype may be worth keeping. Inspect the interface, data model, business rules, authentication, integrations, and deployment separately.
| Decision | Reason | Follow-up |
|---|---|---|
| Keep | The behavior meets the release requirements | Record the tests and assumptions |
| Repair | The design works, but a required check is missing | Name the change and how to test it |
| Replace | The structure cannot support a requirement | Explain the failure and plan the migration |
For example, a useful review screen may sit on top of a database query that ignores account permissions. Keeping the screen does not require keeping the query. Nor does a poor query justify rebuilding the whole application.
Check permissions before the model sees the data
Extraction should produce candidate fields. Application code validates them, the reviewer approves the change, and the server checks permission and current record state before saving it. Generated text must not become an unrestricted database command or an approval flag.
The same rule applies to an internal search assistant: filter source material by the user's access before passing it to the model. A prompt asking the model to withhold restricted documents cannot do that job. AWS describes this in its RAG authorization guidance.
Keep examples of mistakes
Build a test collection from the documents and questions the application will handle. Include incomplete inputs, conflicting values, and cases that need a person's help. Record the expected result and acceptable variation for each.
Count different failures separately. A misplaced comma, an incorrect invoice total, and an unauthorized update have different consequences. Check the saved record as well as the model response.
Reserve some cases for evaluation and keep them out of prompt tuning. Rerun them when the model, prompt, retrieval logic, or application changes. Report how many examples you tested and which document types they cover; a small passing sample leaves plenty untested.
Interrupt a save
Close the browser during submission. Let a model request time out. Disconnect an integration after it accepts a request but before it returns confirmation.
The application needs to determine whether anything changed before retrying. Otherwise a temporary network failure can create a second invoice or leave the user unsure which version was saved. If the underlying record changes during review, require a fresh approval.
A manual path can keep work moving during an outage. In this example, staff could still enter fields and review the source document when extraction is unavailable.
Estimate the running cost
Measure the time and cost of retrieval, model calls, retries, and staff review together. Set limits on input size, simultaneous requests, and repeated calls. Decide what users will see when a provider or spending limit is reached.
A fixed sequence with one model step may be enough. Anthropic's workflow and agent guidance discusses the tradeoff between prescribed steps and agents that choose their own actions.
Before launch, assign responsibility for hosting, model configuration, tests, and incidents. Decide what gets logged and how long it is kept. Test a restoration and a rollback in the release environment, then introduce the application to a small group of users. Their corrections and support requests become the next set of cases to investigate.
If you need a review of an existing prototype, our product and system assessment identifies the work required for a first release.
The invoice example is fictional. Technical references were checked September 8, 2026.