AI Engineering
Acceptance Tests for AI-Generated Code That Catch Real Failures
Review AI-generated code with acceptance tests, boundary cases, security checks, and independent fixtures instead of trusting a convincing explanation.
In this article
AI-generated code often arrives with a plausible explanation and a few passing tests. That can feel reassuring until you notice that the tests repeat the implementation's assumptions. A function may work for the example in the prompt while failing at an empty input, a permission boundary, or a second concurrent request.
Acceptance tests for AI-generated code should describe what the product promises, independently of how the assistant chose to implement it. You don't need a special testing framework for this. You need clear behaviour, representative fixtures, and someone willing to challenge the happy path.
Write the contract before reviewing the implementation
State the inputs, outputs, side effects, and failure behaviour in plain language. For a price-calculation feature, define currency, rounding rules, eligible discounts, tax treatment, and invalid inputs. For an export feature, define who can export, which records are included, and what happens when the dataset is empty.
Avoid vague requirements such as “handle errors gracefully.” Specify whether a rejected request changes data, which error category the caller receives, and what information must not appear in the response. A concrete contract makes it easier to notice omissions that attractive code might hide.
The OWASP Secure Coding with AI Cheat Sheet recommends treating generated code as material requiring verification. Use assistance to draft possibilities, but keep the acceptance criteria anchored in your application rather than the generated implementation.
Start with an example table
For a fictional discount function, list cases before writing assertions:
| Case | Input | Expected behaviour |
|---|---|---|
| Ordinary order | 100.00, valid 10% discount | 90.00 before tax |
| Empty basket | 0.00 | No negative total |
| Invalid discount | 150% | Reject input |
| Repeated request | Same order identifier | No second side effect |
The figures illustrate requirements, not an accounting recommendation. Real financial calculations need a documented rounding policy and appropriate numeric representation. Ask the product owner to approve the examples before an assistant turns them into test code.
Use Text Diff to compare a sanitised requirement revision. A changed sentence about rounding can matter more than dozens of changed implementation lines. Review requirements and code together so the test suite doesn't silently retain the old promise.
Cover boundaries, not just ordinary values
For every input, test absence, emptiness, minimum and maximum valid values, and values just outside the permitted range. Include Unicode where text is supported, unexpected whitespace where parsing matters, and malformed structures where external callers provide data.
For collections, test zero, one, and several items. For time-dependent features, use an injectable clock so tests don't depend on the day they run. For paginated APIs, test the last page and a cursor that no longer refers to an existing record. These cases are often inexpensive to write and expensive to discover in production.
Don't assume an exception proves safe failure. Check whether a database write happened before the exception, whether a file was left behind, and whether a queued job still runs. An acceptance test should inspect the observable side effects, not just the returned status.
Add adversarial tests for access boundaries
Create two test users with different records. Confirm that one can't retrieve, modify, export, or infer the other's data. Use synthetic identifiers rather than real customer data. Also test unauthenticated requests and a user whose access was removed after a session began.
Where code accepts a file path, URL, query filter, or object identifier, inspect how it constrains the operation. A parser accepting an input doesn't mean the operation is authorised. Don't ask an assistant to generate attack traffic against a live third-party system; run defensive tests inside an environment you control.
Review logging too. A failure path that prints a full request may reveal credentials even when the main function rejects it correctly. The log redaction guide covers a separate acceptance surface that ordinary unit tests can miss.
Test the property, then test the example
Example-based tests show specific behaviour. Property-based tests check broader rules across generated inputs. A sorting function should preserve item count and order correctly. A reversible encoding should recover supported input. A price calculation should never create a negative charge when the contract forbids one.
Keep properties independent of the implementation. Calling the same helper to calculate both the actual result and the expected result can make a bug pass twice. For important calculations, use hand-reviewed fixtures or a separately constructed reference implementation.
When inspecting structured fixtures, JSON Formatter can make synthetic records readable. It isn't a substitute for a schema check or an assertion. Run those checks in your test environment, with fixtures committed alongside the code they protect.
Review failure evidence before merging
Run the tests yourself in a clean environment. Inspect failures instead of letting the assistant repeatedly weaken assertions until the suite turns green. If a test changes, ask whether the product contract changed or the implementation was wrong. Keep that answer in the review record.
Check whether the proposed code adds dependencies, migrations, network calls, or permissions. Those changes may need separate approvals even when behaviour tests pass. For database changes, see safe AI-assisted migrations.
Conclusion
A useful acceptance suite makes the product's promise visible. Write independent expectations, cover boundaries and permissions, inspect side effects, and run the tests in a controlled environment. Generated code earns trust through evidence, not through the confidence of the explanation beside it.