Several developers use agents.
Prompts and output should rely on repository rules rather than individual experience.
Introducing Claude Code
Claude Code is already there in many teams, usually without a decision. Output rises immediately, and with it something that only shows up later: no single commit is obviously wrong, but after three weeks a domain object imports from the infrastructure layer and a module boundary your team spent half a year on has a silent hole. Better prompts do not fix that. Structure does.
Fixed price entry on your repository. No licence sales, no tool training: this is about the structure the tool works in.
Remote from Germany. Straight with me, no agency in between.
The starting point
Most teams optimize the wrong thing when they bring in agents. They want faster generation, so they write longer prompts, add more context and hope the output lands closer to what is needed. If it does not, they fix it by hand and try again. The pattern is familiar from twenty years of development and holds for people as much as for agents: an undisciplined process produces unpredictable output.
The failure nobody talks about is not a crash. The agent produces working code, the tests pass, the linter is happy, it gets merged. Three weeks later a domain object reaches straight into the infrastructure layer. A payment handler makes database calls it should know nothing about. No single step was obviously wrong, but the sum has hollowed out the architecture.
That is also why satisfaction and output can move apart: teams ship more and trust the result less. The structure that helps does not sit in the prompt but a layer above it, in what the agent is allowed to do at all and what its result is measured against.
Does any of this sound familiar?Who this is for
For CTOs, engineering managers and tech leads whose developers already use Claude Code or will adopt it as a team. The constraint is not model access but a shared quality system.
Prompts and output should rely on repository rules rather than individual experience.
Architecture, security and tests cannot become personal interpretations as output accelerates.
The team needs specs, checks and workflows that continue on real tickets after the engagement.
What I do
How does your team work with the agent today, where does most of the rework come from, and which architecture boundaries have already become porous. That is the baseline to compare against later.
Between intent and generation comes a layer that records what is to be built, within which boundaries, and how the result is measured. Three types for three sizes, as templates in your repository, not as theory.
Module boundaries, permitted dependency directions and layering rules belong in the spec, not in the prompt. A prompt is a request, a rule in the workflow is a condition. Only the second holds when twenty runs happen in a day.
The agent writes the test before the implementation, and the test has to fail first. That is not a matter of preference but the one condition that prevents a test describing what the code happens to do.
Instead of training, a real run on your code: from spec through implementation to the merged branch, together with the people who will work with it afterwards. Whatever comes up goes into the templates.
Who reviews what, which rules apply to merged branches, and how to spot drift before it is three weeks old. In writing, short, and somewhere it actually gets read.
Is your team writing faster than it can review?
Send me the key details. You get an assessment of what the structure would mean for you, before you commission anything.
How it runs
Four steps, and after the second something sits in your repository that keeps working without me.
How many people work with the agent, on which code, and what has gone wrong so far. After that I tell you whether the effort pays off for you. If it does not, I say so.
Access to one repository is enough. The templates are built on your code, with your module boundaries and your conventions, not on a sample.
A real task from spec to merged branch, with the people who will work with it afterwards. Recorded on request.
What applies goes into the repository in writing. From then on it runs without me, and I only look in again if you want me to.
The other side
The total is not printed below, because I do not know it. I do know the items, and with agents they accrue faster than with people, because more code appears in less time.
The entry package below costs about as much as two weeks of rework on generated code. It replaces none of these figures, but it stops them from growing.
Entry offer
Fixed price. Further repositories or teams are quoted by scope before they start.
Not a talk about agents but a structure in your code: after a week the spec layer sits in your repository, the architecture rules hold, and one run has gone through together.
What you get
What you do not get
The outcome
Dependency directions and module boundaries are a rule in the workflow, not a request in the prompt. A violation shows up at merge time, not three weeks later.
Whether someone has worked with agents for a year or a week matters less to the result than before. The structure carries it, not the practice in phrasing prompts.
A test written before the implementation, that failed first, checks the requirement. One written afterwards checks the code. The difference shows up at the first real bug.
How AI is used in the development process, who reviews and against what, is written down. That answers the question before it is asked.
Technologies I use
Further reading
The failure mode nobody talks about, the spec layer with three templates, why architecture rules belong in the spec rather than the prompt, and where the limits of the approach are. The method behind this page, written out in detail.
Read the articleAlso from real projects
The layer beneath the spec layer: how I filed my own knowledge so that agents can work with it. With the two lessons this page rests on: a rule no mechanism checks does not hold, and a check that reports green in silence is worse than none.
Read the article
Who you are talking to
I am Tim Rutte. More than 20 years in software development, and I work with agents myself every day, on systems that run in production. You talk to the person who touches your code, from the first call to the handover.
Common questions
That is the normal case. The tools are usually there already, often without a decision, and the question only comes up afterwards: output is up, review cannot keep pace, and the architecture shows places nobody decided on. That is where this starts, not at the installation.
Because a prompt is a request. It works for one run, not for twenty a day, and it works for the person who wrote it, not for the whole team. What holds are conditions in the workflow: a spec with acceptance criteria, architecture rules that get checked, and a test that has to fail before the implementation.
The approach does not, the templates partly do. Spec layer, architecture rules and the test condition are independent of the tool; how they are wired in differs. If you have several tools in use, say so beforehand and it gets cut accordingly.
The entry package costs 2,400 euros net at a fixed price. Within one week, the spec layer is in place in one repository, including a run through together. Further repositories or larger teams are quoted by scope before they start. For comparison: that is about two weeks of rework on generated code, and those come back every month.
Generation does not get faster, and that is not the goal. What gets faster is the path from generation to merged branch, because there is less rework and review finds less. Measuring only the time to the first result measures the wrong thing.
On three things recorded beforehand: how much time goes into rework, how many comments a review produces on average, and how often an architecture rule is violated. The third is the most interesting, because it usually is not measured at all.
The work happens in your environment, on your repository, with your access. If certain areas are off limits for external tools, that belongs in the rules that come out of this. What leaves your systems and what does not is part of the baseline, not an afterthought.
Yes, but as a project of its own. In a modernization or a handover I use the same structure; there it is part of the work rather than a separate offer. This page is for teams that want to carry on themselves.
Other services
The infrastructure behind AI systems that actually ship: MCP servers, controlled tool access, LLM integration with real permissions and cost control.
Learn moreMCP servers that hold up in production: the identity of the user instead of a shared technical account, tools with limits, a complete audit trail and a cost cap per use case.
Learn moreFinOps for inference: make cost per operation and per user visible, choose the model by task, keep context and repetition in check. Measure first, then act.
Learn more