Engineering practice with a working example
Give AI-written code a test it can fail
Imagine a service that creates a task when a request arrives. The first request works. Its response gets lost, the caller retries, and now two people are assigned the same work. The code compiles. The first test passes. The person who receives the duplicate task still has a problem.
An AI assistant can write either version of that service quickly. It can also explain either version with confidence. Before asking for the patch, give it a condition the implementation has to survive: receiving the same request again must leave one task.
Same request.
Second arrival.
Keep the second request in the room
In the example below, one implementation appends every arrival and the other checks whether the request key is already present. One implementation appends every arrival. The other checks whether that key is already present. Run each against the same inputs. The expected answers are fixed before either function runs.
The repeated request exposes the difference immediately. The other cases matter too. Two distinct requests should still produce two tasks, and retrying the first after another arrival should leave the count unchanged. A fix that discards every request would fail these checks even though it prevents duplicates.
A real service also needs a test where concurrency and restarts actually happen, such as at the storage layer. This in-memory version cannot show those failures.
Interactive example · runs in your browser
Run the retry example
Two small functions receive the same fixed request sequences. The test compares the task count each one produces with the count you expect.
function receive(tasks, key) {
return [...tasks, key];
}A task is represented by its request key. Each check starts with an empty list.
Replay / same request, twice
Run the checks to inspect the result.
No checks run yet.
A → AA → BA → B → AIllustrative functions, real local checks. No AI call or backend is involved. These checks cover sequential arrivals in memory; they do not test concurrency or persistence.
The patch owes you an explanation of the state
In a real service, somebody has to decide what makes two requests the same. A new attempt might carry the same request key and the same task, or reuse that key with different content. That is a product contract the implementation must follow. Asking an assistant to “handle retries” leaves too much room for it to choose the rule for you.
Then follow the state. Where is the key stored? When is the task created? What happens if the process stops between those operations? An array is enough for this demonstration. A service that must survive concurrent workers and restarts needs to establish the same behavior in its persistent storage.
Reading the callers helps keep the patch small. If several entry points create tasks, fixing only one leaves the other paths exposed. Ask the assistant to trace those paths, identify the shared operation and explain where the repeated request will be recognized. The review is easier when the answer names actual code.
Leave room for a reviewer to disagree
Give the reviewer the requirement, the changed code and the failing example. Let them form their own view before reading the author’s explanation. They may notice that the key expires too early, that an older caller never sends it, or that a failure between writes still creates duplicate work.
An AI reviewer can help find these possibilities, but each finding needs a concrete path through the code. Reproduce the ones that matter. A second confident paragraph from another model adds little by itself. A request sequence that breaks the proposed fix is useful.
GitHub’s review decisions give the team a place to record approval or requested changes. Keep unresolved findings attached to the change. The next person should be able to see which concern was addressed and which limitation the team deliberately accepted.
Try the released path
Once the patch is deployed, repeat the relevant workflow in the intended environment with safe test data. Confirm that the caller sends the key, the service keeps the expected state and the downstream worker receives the intended task. A local test cannot observe a configuration error in that path.
Keep the handover specific. Record the version checked, the request sequence, the observed result and any part that still needs validation. For this example, the useful evidence is a repeated request with one resulting task. That gives the next engineer something they can run again when the implementation changes.