All case studies

Case Study · SaaS · Agentic Engineering

An operating system for trade businesses.
Built with agentic engineering from day one.

A multi-tenant SaaS platform brings customers, jobs, quotes, invoices and site documentation together. Built with coding agents, steered by specifications, architecture rules and quality criteria that can be checked.

Book a call
Agentic engineering flow diagram: requirement, specification, architecture context, scoped task, implementation by the agent, automated checks, review by Codex and CodeRabbit, integration
5 monthsto a working platform
2,000+pull requests
3developers with coding agents

The starting point

One business, five tools, no overview.

A trade business with a handful of people rarely runs on one system. The quote is written in Excel, the job arrives on WhatsApp, the invoice goes out by email, and the photos from the site live on somebody’s phone. In between: paper, and separate programs that know nothing about each other.

The consequences are unspectacular and expensive. The same data gets entered several times. Information gets lost between the tools. Which invoice is still open and what is due today is known, if at all, by whoever keeps it all in their head. And the admin eats the time that is missing on site.

The task

One platform instead of five islands.

The goal is a central SaaS platform for the commercial and operational work of a trade business with one to fifteen employees: customer management, jobs, quotes and invoices, open items, site documentation and reporting.

The product is multi-tenant by design. Every business works on the same platform, and one business’s data never shows up at another. That separation is not a later expansion stage. It is the foundation everything else stands on.

At the same time, the architecture has to carry what is still to come: voice control, a phone assistant and, later, agentic features inside the product itself. Language models are already integrated, Gemini and Mistral. For the rest, what gets built today is the foundation, not the feature.

My responsibility

An architect who implements.

The product is being built by a founding team of four, three of whom develop. My responsibility is the technical architecture and the hands-on implementation. I did not build it alone, and this page does not claim I did.

Architecture. The technical foundation of the platform, including tenant separation, and the rules every change follows.

Structuring requirements. Business wishes become specifications that an agent can implement and a review can measure against.

Building core parts of the product. Myself, with coding agents.

The agentic engineering workflow. How a requirement becomes a checked change in the system, and where a change fails before it lands.

Quality control and delivery. The rules everything is checked against, automated checks, a multi-stage review by agents and delivery close to production.

My approach

The model first.
Then the process.
Then the speed.

01

Domain foundation and tenant separation

Before an agent writes a single line, it has to be clear what a job is, when a quote becomes an invoice and when an item is open. A domain spread across spreadsheets, chat threads and people’s heads becomes a product model with fixed terms. Tenant separation is part of it from the start, because retrofitting it later is expensive.

02

An agentic development process

The code is written with coding agents. What they write is determined by specifications and documented architecture rules, and what gets merged is determined by automated checks and a multi-stage review by Codex and CodeRabbit. The flow behind it is laid out step by step below.

03

A platform close to production, built to extend

Next to production sits a complete staging environment with its own data stores. CI runs on dedicated runners, delivery happens in containers. Refactoring and operations run as workstreams of their own, so speed does not turn into technical debt and the platform stays open for voice control and AI features.

The actual core

Eight steps per task.
The agent is one of them.

01

Requirement

What a business wants to get done, in its own words. Whether and when it gets built is decided by a person. Prioritization is product work, not a job for a model.

02

Specification

Before implementation, it is settled what gets built and how to tell it is done. For user interfaces that includes every state: list, empty state, detail, edit, success and error. Whatever is not specified, the agent invents, and that is exactly what it should not do. How such a spec layer for agentic coding is structured is covered in the article.

03

Architecture context

Architecture decisions are documented and given to the agent as binding context. It does not renegotiate tenant separation in every task, it implements it. A decision that only exists in someone’s head is one the agent does not know.

04

Scoped agent task

One task, one clear cut. The agent changes the part of the system the task is about, not the whole of it. Several tasks run in parallel, each in its own workspace with its own clone of the repository, so two agents do not overwrite each other’s work.

05

Implementation

This is where the agent works: analysing the existing code, implementing, testing, refactoring and documenting. The tools are Claude Code, Codex and Grok. The agent is fast, but it decides nothing that was not already decided in steps two and three.

06

Automated checks

Tests and linting run in CI on dedicated runners. They are a condition for merging, not a suggestion. If a check fails, the task goes back to the agent, not on to review.

07

Review

Every change goes through a multi-stage review by Codex and CodeRabbit before it is merged. The question is no longer whether it runs, the checks have answered that. The question is whether it fits the system that is supposed to still exist a year from now. If a review finds something, the task goes back to the agent that implemented it.

08

Integration

Changes are merged through pull requests, more than 2,000 in five months. That keeps every change traceable: its reason, its checks and its review. Delivery runs through GitHub Actions, with a staging environment of its own next to production.

The difference

What sets this apart from vibe coding.

No uncontrolled prompts. The agent gets a specification and an architecture context, not a one-line wish. What it is meant to build is settled before it starts.

No merge because it runs locally. Code working on one machine is not a criterion. The criteria are passing checks in CI and a passed review.

No architecture by accident. An agent working without rules makes its own decisions in every task, and after a hundred tasks the system has a hundred architectures. Here a person makes them, once, and writes them down.

No speed without quality assurance. The speed comes from the agent. The discipline comes from the process, and it is not up for negotiation when things are urgent.

The split is unambiguous. Product decisions, architecture, prioritization and the rules everything is checked against stay with people. The agents speed up analysis, implementation, tests, refactoring and documentation, and they check each other: code is written with Claude Code, Codex and Grok, and checked by CI and, in several stages, by Codex and CodeRabbit. If you want to bring this into an existing team, the approach is described under introducing Claude Code, and what a codebase needs first is in the article Is your code too old for coding agents?

Want to build a product with agentic engineering?

Without leaving architecture and quality to chance. Let us talk about your plans in 30 minutes.

Book a call

Where it stands

Built. In progress. Vision.

Built

A working platform

Customer management, jobs, quotes and invoices, open items, site documentation and reporting, plus an LLM integration with Gemini and Mistral. The production environment is set up, and the first product website is live.

In progress

The pilot

Ten to fifteen businesses have committed to take part. Product, usability and feature areas are being developed further for the pilot.

Vision

Not built yet

Voice control, a phone assistant, an open API even in the smallest plan and, later, further agentic features. Planned, not in production.

The outcome so far: in five months a product idea has become a working SaaS platform, built by three developers across more than 2,000 pull requests. A sprawling domain has been translated into a structured product model. Agentic engineering was not an experiment next to the real development, it was the development model itself. The technical foundation for the pilot and for further AI features is in place.

What is missing here is missing on purpose: savings, usage figures or satisfaction scores. They will only exist after the pilot has run in production, and until then there are none on this page.

Technologies used

An ordinary stack. The process is what is unusual.

Development and review
  • Claude Code
  • Codex
  • Grok
  • CodeRabbit
Language models in the product
  • Gemini
  • Mistral
Application
  • Ruby
  • Go
  • PostgreSQL
  • Redis
  • Elasticsearch
  • Sidekiq
Delivery
  • Git
  • GitHub Actions
  • Dedicated CI runners
  • Docker
  • Staging environment

Frequently asked

What clients ask before deciding.

Isn’t this just vibe coding?

No. In vibe coding, the prompt decides what ends up in the system. Here every agent task starts from a specification, the agent works against documented architecture rules, and only what has passed the automated checks and a multi-stage review by Codex and CodeRabbit gets merged. Code that runs locally is not a criterion.

What does the agent decide, and what do I decide?

Product decisions, architecture, prioritization and the rules everything is checked against stay with people. The agents speed up analysis, implementation, tests, refactoring and documentation, and they do the review. They get scoped tasks, not the whole system.

How fast was it?

Five months passed from the product idea to a working platform. In that time more than 2,000 pull requests went through the flow, from three developers who each work with coding agents.

Are there results from the pilot yet?

Not yet. The platform works, the production environment is set up, and ten to fifteen businesses have committed to the pilot. Reliable figures on savings or usage will only exist after the pilot. Until then, there are none on this page.

Why is the product not named?

It is a venture of its own within a founding team and has not yet launched. For this page, the product name matters less than how it is built.