“Is our code too old for coding agents?”
I have heard that question in almost every first conversation since spring, and it always comes with the same undertone: as if the answer were known and only the extent still open.
It is the wrong question.
I currently work in both worlds. In a grown system of PHP and Go that is older than the careers of some of its developers, and in repositories that have existed for a few weeks and in which agents have written every line from the first one. When I count where an agent got something wrong past me this year, old code does not lose. What loses is the code where no command could say no.
That is the quantity that matters. Not the age, not the language, not the style. This article replaces the one question with four that can be answered within a week, and says in which order I tackle them in an old system.
What an agent can do in legacy code and where it fails is covered in A coding agent in a legacy codebase. This one is about what has to stand around it.
The green tick is the sum of the questions somebody asked
A while ago an agent delivered Go code to me. Compiled, review unremarkable, tests green. In a loop it opened a file on every pass and closed it at the end of the function instead of right after use. In a test with ten files nobody notices. In production the open handles pile up until the operating system refuses to hand out any more and the service stops.
The code was not written wrong. It was thought wrong, and none of the checks had asked about it.
Since then I read a green tick differently. It is not a statement about the code. It is the sum of the questions somebody previously poured into tests, linters and type checks. If nobody ever asked about open handles, there is no answer about open handles. The right linter catches this case, by the way, as soon as it is switched on. I just did not know it was the one I needed until after the outage.
That is the core of the whole subject. An agent works as safely as the checks around it ask. It does not ask questions nobody has asked it. Verifiability is the quantity, not age.
Why old code is often at an advantage here
And this is where the original question gets interesting, because twelve years of operation are twelve years of questions asked.
Every incident that once cost a night has, with a bit of luck, left behind a test, a monitor, a comment with a ticket number. The special case for that one key account sits in the code with a date on it. The database has constraints that came out of real data errors. A system that has been running for a long time has seen its edge cases. A repository that has existed for three months has not seen a single one.
Add what old code does not have: it no longer grows by a third every week. With agents that is not a drawback. A codebase that changes slowly can be described, and the description is still true the next day.
What old code usually lacks is the command that runs all of that together. The tests exist, but three directories of them have not run since 2019. The static analysis exists, but only on the machine of the colleague who set it up. That is exactly where the work begins, and it is smaller than a modernization.
How a safety net gets into code that never had tests is in Testing legacy code when there are no tests.
Question 1: Is there a command that can say no?
Not “are there tests”. Is there one single command that runs tests, static analysis and linters and exits with a non-zero code when something is wrong?
That command is the interface between the agent and the truth. An agent considers itself done when it is done. The command considers it done when it stays silent. The second yardstick is the one that counts, and in my setup it is the only one: acceptance happens through exit codes, not through an agent deciding it has finished.
check:
vendor/bin/phpstan analyse --no-progress
vendor/bin/php-cs-fixer check
vendor/bin/phpunit --testsuite=unit
go vet ./... && staticcheck ./...The command runs locally and in CI, and it is the same command in both places. I have seen projects where vendor/bin/phpunit called directly was green and make test was red, because only the Makefile knew the right configuration. An agent takes the path that is green.
What the command checks grows out of incidents. On my own website, agents wrote fifty articles in September. Two of them went live with a meta description of 21 and of 5 characters: a straight quotation mark in the text had ended the attribute early. Until then the build only checked that a description existed. Since then it aborts below fifty characters, and the case has never come back. Not because the agents got better. Because the question is now being asked.
How static analysis gets into code that never knew it, one level at a time, is in Introducing PHPStan into legacy code step by step.
Question 2: Does the check fail loudly?
The most dangerous state is not the missing check. It is the check that has stopped checking and keeps reporting green.
Three cases from a single month. In my knowledge repository a rule required proving before each commit that a purely additive contribution deleted nothing, using a grep for lines that start with one minus sign and not with two. But a deleted bullet line appears in the diff exactly like that: one minus from the diff, one from the list. The pattern excluded precisely those. In a repository made of bullet lists, the check was blind to the most likely deletions, and had been from the day it was written down. It surfaced as four swallowed lines in a commit that had seven shown to it.
On one of my websites, speculation rules sat as an inline script in the head, and the content security policy allowed no inline scripts. A browser discards such a block without a word: page normal, code in the source, effect zero. Build green, tests green, and my own check across all 32 pages green too, because it had measured presence in the HTML and not effect.
And a script that checks links and metadata before every commit had skipped half a file for four days and reported a reliable “0 findings” all that time.
That third case, and the canary that has stood against it since, is described in An agent is only as good as its context.
Three ways to fail, one pattern. None of the checks was broken in the sense of an error. One had a pattern that was too narrow, the second measured the wrong layer, the third had quietly stopped. All three looked like success while they did.
Two habits came out of that. Every automated check gets a canary: a deliberately planted error that every run must find, or it aborts instead of reporting green. And every search pattern is tried against a real hit before it becomes a rule. Untested, a search pattern is a claim, not a check.
The question I ask when building any check: what does its failure look like? If the answer is “like success”, it is not finished.
Question 3: Is what the agent reads true?
Every tool reads a file from the repository at start-up, AGENTS.md or CLAUDE.md depending on the vendor. Whatever is in it, the agent believes without asking. That is the purpose of the file, and it is its danger.
In September I checked these files in two grown repositories. Both contained load-bearing false statements: the wrong PHP version, the wrong Go version, an autoloader convention that never applied like that, references to a GitLab that had long been switched off, and commands that were not installed on any machine. Nobody had lied. The files had once been written from one colleague's knowledge and never held against reality again. A human would have stopped trusting them at the third wrong command. An agent did not.
Three rules I have applied since:
- A human writes the first version, from the last three surprises, not from the architecture. An agent can describe the code. What is not in the code it cannot know, and that is exactly what belongs in the file.
- Every sentence must be executable. “Tests run with
make test” can be verified. “We take care to keep the layers cleanly separated” cannot. The second sentence sits in almost every one of these files and has never changed a decision. - Keep it short, link the rest. The file keeps what applies to every change. Domain knowledge moves to its own directory that the agent reads when needed. A file of four hundred lines is not read, it is skimmed, by people and by models alike.
And the file grows for the same reason the check command does: whenever an agent misunderstood something a human would have known, the missing sentence gets added. After a few months it holds what used to live in three heads.
What belongs in the knowledge layer beneath and how it differs from rules is in An agent is only as good as its context.
Question 4: Is the rule a sentence or a mechanism?
The most uncomfortable lesson of this year: rules get broken. By sessions that know them.
My knowledge repository has a fixed list of permitted metadata values. When four invented ones turned up, I reverted them and wrote the rule into every instruction file. In the two days that followed, three new ones appeared. Then the check was added, and the fourth was reported immediately. More emphasis in the wording would have been the wrong lever. Whoever does not apply a convention does not apply a more sharply worded one either. Wherever a sentence can be phrased as a check, the check is the rule and the sentence is only the reasoning behind it.
That reaches down to the tooling level. In one of the two repositories above, the directory holding the agent rules was excluded from version control entirely. So the rules could only ever be personal; every developer had their own, or none. Now it lives in the repository, and instead of the sentence “PHP only runs in the container” there is a hook that rejects every call to php and composer on the host before it runs. The sentence was ignored. The hook cannot be.
And it applies to isolation. In early September two of my sessions ran in the same directory. Twice on the same day a git add -A swept up the other one's half-finished work and filed it under a foreign commit message. A whole section of evaluated documents would have been lost; the diff showed 57 deletions where the edit was purely additive, and only that made it visible. After the first incident I extended the rule. The second session broke it anyway, because it had started before the addition and its instructions came from the time before. Since then every session works in its own Git worktree. That is not a rule. That is a file system.
Whoever builds does not sign off
The four questions describe what stands around the agent. That leaves the question of who says in the end that the change is right.
Not the agent that wrote it. Whoever writes code checks it with the same assumptions with which they built in the mistake. That holds for people, and it holds for models, even within one model family: similar training, similar blind spots.
That is why I split it up like a building site. A strong model plans and breaks the task down. Several cheap instances implement, each in its own worktree, and none of them may commit. Merging happens centrally, task by task, and acceptance goes through the command from question 1. The review is then done by a model from a different family, and for everything that has no rule, a human at the end.
Two things keep that economical. The deterministic checks remain mandatory and run around every task; the AI review is the supplement, not the replacement. And not every change needs the big review. I review at milestones, with the smallest model that still reliably finds the relevant classes of error. Otherwise the control costs more than the work.
How the task beforehand becomes a spec against which implementation and acceptance can be measured is in Agentic coding with spec-driven development.
Where I start in an old system
None of this requires a modernization. The order that has proven itself fits into a week, and the agent itself is already useful along the way.
- Day 1: the command. Everything that already exists in tests, analysis and linters, behind one target. Whatever is red is not fixed but excluded and written down. By the evening there is a command that can say no.
- Day 2: the same command in CI, on every pull request, without exception. Only now does it count.
- Day 3: the canary. A deliberately broken test case, a deliberately wrong type. If CI does not report it, it is not checking what you think it checks.
- Day 4: the instructions, by hand, from the last three surprises. Every sentence is executed once before it is allowed to stay.
- Day 5: the first task for the agent, one whose result the command can judge: characterization tests around the spot that will be touched next. A test that does not pass against the unchanged code is wrong. That is a referee you cannot talk round.
From then on the net grows out of incidents. Every time a human finds something no command found, it becomes a check. That is also the number by which I measure progress: how often in a month a human found something the checks did not see. It should fall. It never reaches zero.
Frequently asked questions
Do we have to catch up on test coverage first? No. A coverage of 80 percent says nothing about whether the tests can say no. A net around the spot being touched is enough, and it is the first task for the agent.
Is a linter enough? To begin with, yes, if it runs in CI and turns red. It catches less than it promises and more than nothing. Which checks are missing, the next incident will show.
Our system is too big for the context window. The agent does not need the system. It needs the spot it is changing and a command that says, across the whole system, whether the change broke something. Size is a problem of scoping, not of age; why a larger context window does not solve it is in the glossary.
And if the command says no although the change is right? Then the check is wrong, and that is a finding that surfaces before the commit instead of in production. The agent may not adjust the check. It reports it.
What I would do this week
Open the oldest repository that still earns you money and run the one command that checks everything. If it does not exist, that is your week. If it exists and is green: deliberately introduce a mistake and see whether it turns red.
Only after that is it worth asking which agent to use.
Old code is no obstacle for an agent. Unchecked code is.
How a setup like this comes about in a team is described under Introducing Claude Code.
This article belongs to a series about systems that already exist. The retrospective orders every article in it by situation.

