Lesson 1 of 1 in Reviewing Code You Did Not Write
Passing Tests Is the Weakest Signal
Generated code arrives already working. That is the problem, not the reassurance.
3 min read
Not yet reviewed
Code you wrote yourself carries something you cannot get from code you accepted: you remember the decisions. You know which branch you were unsure about, which value you meant to validate and did not, and where you told yourself you would come back.
Generated code has none of that, and it arrives passing its tests, which is where the trouble starts. Tests check the cases somebody thought of. A model produces code for the cases in the prompt, and a test written by the same model checks the same cases.
Note
A green test suite on generated code tells you the code and the tests agree. They were produced from one description of the problem, so agreement is close to guaranteed and says very little about whether the description was complete.
What to check, in order
1. Every value that came from outside, and where it stops being a string
2. Every decision - who is allowed to do this, and what decided
3. Every dependency it introduced
4. Every case the prompt did not mention
- Line 1The same first step as reviewing your own code, and it matters more here: you did not choose where those values go, so you have no memory of the path to check against.
- Line 2Models produce plausible authorization, which is the dangerous kind. A check that reads correctly and asks the wrong question passes review because it is present.
- Line 3The one with no equivalent in ordinary review. Generated code imports things, and the next chapter is about what some of them are.
- Line 4Empty input, a value at the boundary, a second request arriving while the first is still running. The prompt said what should happen; it rarely says what should not.
The question that finds most of it
For each piece of generated code, ask: what did I not say? The output addresses the prompt. Everything the prompt left implicit — the empty list, the duplicate submission, the user who is not the owner — was filled in by a model guessing at your intent from the average of everything it has read.
But it usually guesses well. I have shipped a lot of this and almost none of it has been wrong.
It does guess well, and that is precisely the risk. A tool that is wrong half the time gets checked every time. A tool that is right most of the time trains you to stop looking, and then the wrong one goes through with everything else.
Generated code comes with generated tests, and they all pass. What have you learned?
Tip
Write one test yourself before reading the generated ones — for the case you think is most likely to be wrong. If it passes you have learned something real, and if it fails you have learned more.