Technology: AI engineering

AI-driven software development: agents write, test and review the code; a person gates every release.

We build software with agents that write, test and review code against written rules, under the same discipline we ask of any engineer: one branch per change, checks that run before anything ships, a record of every change and a named person who approves each release. This website and the system that runs our marketing are built this way, and so is the configuration of the platforms our solutions run on.

What we build with it

What we build with AI-driven software development.

Written rules before code.

Every project carries its standing rules in one file the agents read first: naming, copy, security, what never leaves the repository. A defect that gets through becomes a new rule the same day.

One branch per change.

An agent pulls the main line, works on one branch for one change, and never commits to main, rewrites history or pushes over someone else’s work. Two people and several agents work the same code without treading on each other.

Checks before anything ships.

Unit tests, the build, and check scripts that sweep the rendered product: the layout at nine sizes, the copy rules, the search metadata, the contrast, every link. A finding stops the change; the summary line goes into the pull request.

Review by agent, release by a person.

An agent reviews the diff against the rules and the record; a preview goes to a named person, who answers with one digit. Nothing reaches production without that answer, and a release runs from the main line only.

A record of every change.

Each change leaves a dated note beside the code: the request in the requester’s words, what changed and where, the checks verbatim, the screenshots. Six months later the why is still there.

Defects become checks.

When a person finds a defect the checks missed, the fix ships with a new probe that would have caught it, so the same defect cannot come back unnoticed and is found by a person only once.

Platform configuration, the same way.

A CMS row, a CRM field, a journey in a marketing platform or a scheduled job is read before it is written, written once through its own interface, read back and logged, with the same gate as a code release.

No credential in the work.

No credential value in a file, a message or a log. Agents read secrets by name from the environment or an encrypted store at run time, and the rule is checked by a script, not assumed.

In practice

This website is built this way: every change is a branch, a build, the check scripts, a preview reviewed by a named person, a pull request and a dated record, and the same flow carries the agents that run our marketing.

The technology

What AI-driven development means here.

Agents write, test and review the code. People set the rules, read the previews and approve the releases. That division is the whole method. Everything else on this page is the discipline that makes it safe: written rules the agents read before they start, one branch per change, checks that run before anything ships, a person at the gate, and a record of every change. It is how this website is built, how the system that runs our marketing is built, and how we configure the platforms our solutions run on.

The loop: rules, change, checks, gate, record.

  1. Rules. Every project carries its standing rules in one file, in plain words, with the date and the request that produced each one. The agents read it first. A rule that is not written is a rule that will be broken.
  2. Change. An agent pulls the main line, branches for one change and makes it: the code, the data, the test for the new behaviour. It never commits to the main line, never rewrites history, never pushes over someone else’s work. Two people and several agents work the same code at once this way.
  3. Checks. The unit tests, the build and a set of check scripts that sweep the rendered product. A finding stops the change. The check’s summary line, verbatim, goes into the pull request.
  4. Gate. An agent reviews the diff against the rules and the record; a preview of the change goes to a named person, who answers with one digit. Nothing reaches production without that answer, and the release is made from the main line, never from a branch.
  5. Record. A dated note beside the code: the request in the requester’s own words, what changed and where, the checks verbatim, the screenshots, what was left undone and why. Six months later the why is still there.

What the agents do, and what they do not.

They read the codebase and the rules, propose a design, write the code and the tests, run the checks, read the results and fix what they find, draft the record and open the pull request. They review each other’s work against the same rules. They do not decide what the product should be, they do not merge, and they do not release. Those are a person’s decisions, and the method keeps them so: a credential an agent can use is scoped to what it needs, a release command runs only from an approved state, and the approval itself is a message from a named person that the record keeps.

The checks, and how a defect becomes one.

The checks are code, kept beside the product, and they sweep what a person would otherwise inspect by eye: the layout at nine viewport sizes, from a phone to a wide desktop; the copy rules (every heading ends with a period, no placeholder text, no vendor named where the page is not about that vendor); the search metadata on every route; colour contrast on every band; every related link answering; the share image of every page. Each check prints one summary line and exits with a code the pipeline reads.

When a person finds a defect the checks missed, the fix ships with a new probe that would have caught it. An eyebrow that wrapped on a wide screen became a rule that no eyebrow wraps from 700 pixels up. A hero whose figure sat below its heading on a phone became a rule that the heading comes first under 700 pixels. A menu whose lead group sat on the right became a rule that the group we want read first is leftmost. The checks are the memory of every defect a person has had to find, and the reason each is found only once.

Configuration is code, under the same rules.

A CMS row, a CRM field, a journey in a marketing platform, a scheduled job: each is production as much as a line of code, because a customer, a prospect or a colleague sees what it does. So each is changed the same way. Read the current state before writing, in the same turn. Write once, through the platform’s own interface, never by hand in a screen when an interface exists. Read it back and compare. Log the change with who asked and who approved. A change to a shared surface, a template, a routing rule, a gateway, is presumed to reach every page or every run it touches until the after-check proves otherwise. “Backwards compatible” is a result of that check, never an opinion.

What it changes for a client.

Speed, first: an agent does not wait for a sprint to start, and a change that is small stays small. The capability page on agentic solution development carries the measured figures for our own build, with the method beside them. Then confidence: the record and the checks mean a client can read why every line exists and see that it was tested before it reached them. Then portability: the rules, the checks and the records are files in the client’s own repository, so another team, or another set of agents, can pick the work up without a handover meeting.

What a client gets.

  • A repository with the code, the data, the tests and the check scripts, and the standing rules in one file at its root.
  • A record for every change, in plain words, with the request, the diff’s reach, the checks verbatim and the screenshots.
  • A release flow in which a named person approves every production change from a preview, and the release runs from the main line only.
  • The same flow for the configuration of the platforms the software runs on, with a read-back on every write.
  • No credential value in any file, message or log; secrets read by name from an encrypted store at run time.

Where it is not the right choice.

A one-off script nobody will run twice does not need the record. A prototype meant to be thrown away does not need the gate. We will say so, build it the quick way, and label it so that nobody mistakes it for the product.

Related
Build

Web application development.

Enterprise and consumer web applications, from product strategy and design through backend, data and operations.

Operate

Deployment, operations and maintenance.

Automated deployment, CI/CD, configuration management, monitoring, support and maintenance.

Operate

Quality assurance and testing.

Manual, regression, performance, API, security and web and mobile testing.

Operate

Test automation.

Functional, regression, performance, mobile and API test automation inside the CI pipeline.

Advise

Agile release planning.

Product strategy, personas, wireframes, journey and story mapping, and estimation before the first sprint.

Capability

Custom Software Development.

The capability these pages belong to: how we build custom software, and when we do not.

Other AI engineering. Agent development All technologies

Where we're not the right answer

We'll tell you if we're a fit. If we're not, we'll tell you that too.

  • Your current stack works and nobody wants to change it
  • You want licences resold at a discount and nothing else
  • Your internal team owns the operating model and is not handing it over
  • You want hours of configuration work and nothing run for you: that is on our services pages, and it is not a managed solution
How we start

Most of our best clients come to us with a feeling, not a plan.

"Something isn't working." "We're outgrowing our tools." "We're afraid to make the wrong move." No-Risk Discovery is a short, practical conversation that gets you clarity before you commit to anything big. We'll tell you if we're a fit. If we're not, we'll tell you that too.