Research
Why AI-generated code fails at authorisation more than at anything else
Assistants are trained to produce code that runs. Authorisation is the one property that cannot be inferred from the surrounding code, because it lives in business rules the model was never shown.
In short
- Most vulnerability classes have a local, syntactic fix a model can learn. Authorisation does not.
- Whether this user may read this row is a business rule, and it is not present in the prompt or the file.
- So review AI-written code by asking who is allowed, not by looking for unsafe functions.
The claim
Code assistants have got measurably good at the vulnerability classes with a local fix.
Ask for a database query and you will usually get a parameterised one. Ask for a password
field and you will usually get a hashing library rather than md5. These patterns are
everywhere in the training data, the fix sits inside the same expression as the bug, and
the correct form is the common form.
Authorisation is different in kind, and the difference is why it survives.
Every other class can be judged from the code. Authorisation can only be judged against a rule that is not in the code.
Whether a signed-in user may read order 8231 is not a property of the route handler. It
is a fact about the product: orders belong to the customer who placed them, unless the
account is a shared team account, unless the caller is a support agent, unless the order
is in a state that makes it visible to the fulfilment partner. None of that is in the
file. Usually none of it is written down anywhere.
So a model asked to “fetch the order by id” produces a function that fetches the order by id. That is a correct implementation of the request. It is also an IDOR, and the model had no way to know, because the missing information was never in the prompt.
Why this is structural, not a temporary gap
Three things keep it in place even as models improve.
The request under-specifies the requirement. A developer writing the same function has context the prompt does not carry — they know what a tenant is in this product. The assistant infers intent from the visible file, and ownership rules are rarely visible in it.
The generated code passes its tests. Happy-path tests are written from the same understanding that produced the code. An IDOR does not fail a test that signs in as the owner, and that is the test that gets written.
Review reads for correctness, not for authority. A reviewer scanning a diff sees a query, a null check, a response. Nothing looks wrong, because nothing is wrong except an absence — and absences do not appear in diffs.
That last one is the reason this class hides in human-written code too. Assistants have not created a new bug; they have industrialised an old one, by making it cheap to produce plausible handlers faster than anyone reviews them.
What I have not measured
I want to be exact about the status of this piece, because the alternative is the thing this whole site exists to avoid.
This is an argument, not a finding. I have not yet published a sample of AI-generated repositories with counts by vulnerability class, and there is no N here. When I have one, it goes in this section with the method stated — how repositories were selected, what counted as a finding, and what I could not check.
Until then, treat this as a hypothesis worth testing rather than a measured result. If it turns out to be wrong when the sample exists, this page will say so.
The checklist
What I would actually check, in order, on a codebase written with heavy assistance. None of this requires tooling.
Enumerate the endpoints that take an identifier. Anything with :id, ?id=, or an
id in a request body. This is the candidate list.
For each one, ask who is allowed to call it — not “is the caller authenticated”. Write the sentence down. If nobody can state the rule in one sentence, that endpoint is where to start, because the code cannot enforce a rule the team cannot articulate.
Check the write paths as hard as the read paths. PATCH and DELETE built from the
same scaffold as a leaky GET are usually leaky in the same way, and they are worse.
Look for ownership in the query, not next to it. A where clause scoped by the caller
cannot return the wrong row. A fetch followed by an if can, as soon as someone moves the
if.
Check middleware’s matcher against the routes it is assumed to cover. Path-based protection plus a route added later outside the pattern is a common, quiet gap.
Confirm the failure mode is 404, not 403. Answering 403 on an object the caller
may not see confirms that it exists.
Grep the built client bundle for credential shapes, not just the repository. Public environment prefixes ship secrets to every visitor by design.
What this does not cover
This is a code-reading method. It says nothing about infrastructure, network exposure, dependency vulnerabilities, or anything that only shows up at runtime under real traffic. It also assumes you can read the source — it is not a black-box technique.
