Technology: AI engineering

Agent development: agents that act in the systems you already run, with a person at every gate.

An agent reads what is happening across a company’s applications, decides the next step against written rules and does it through those applications’ own interfaces, asking a person only where money, risk or compliance is at stake. We build agents this way for our own marketing function and for clients, on whatever platforms they already run. This page says what one is made of, how we test it and what you get.

What we build with it

What we build with agent development.

Models, chosen per job.

The model is a component, not the product. We pick a provider and a size per task by measured error rate and cost, keep the choice behind one interface, and change it when a better or cheaper one appears, without touching the rest.

Tools and connectors.

An agent acts through gateway functions that own each platform’s interface: read, write, read back. Credentials live in an encrypted store no person or prompt can read. Any vendor in a category whose interfaces carry the use case.

Memory and state.

What the agent knows between runs lives in records a person can open: the operating model, the state of each workflow, the run log. Nothing that matters is kept only inside a conversation.

Skills, written down.

Each workflow is a written procedure the agent follows: the reads before the writes, the decision, the gate, what to do when it fails. A person can read it, and a change to it is a versioned release.

Evaluation.

Every skill ships with test cases that run before a change ships. The agent’s outcomes are measured against the numbers the business already reads, so a drift in quality is a line in a report, not a surprise.

Guardrails and approvals.

Anything that costs money, carries risk or needs taste waits for a person’s one-digit answer. Caps are enforced in code beneath the gate, so an approval given in error still cannot spend past the envelope.

Observability.

One row per action in the run log: who asked, from which surface, who approved, what the interface returned, what it cost. A daily self-check tests every connection and restores a regression in the same run.

Cost control.

Each run carries its model and tool cost onto the outcome it produced, so the all-in cost per lead, per piece or per order is a monthly number the owner reads, with the agents’ own bill inside it.

In practice

Our own marketing runs on agents built this way: they plan, produce, publish and measure a week of content under a monthly spending envelope enforced in code, and a person gives three one-digit decisions a week.

The technology

What an agent is, in our terms.

An agent is a program that reads what is happening across a company’s applications, decides the next step against written rules, and does it through those applications’ own interfaces, asking a person only where a person should decide. The model inside it is one component. The rest, the rules it works to, the tools it may use, the record of what it did and the gate a person holds, is what makes it something a business can run.

We build agents for our own operation and for clients. Our marketing function runs on them today: a set of agents plans the month’s themes, proposes the week’s work, produces and publishes the pieces, reads the results back into one table and asks a person for three one-digit decisions a week. The same pattern runs the solutions we deliver for operationally complex businesses, where the agents read the CRM, the help desk, the accounting system and the inventory a client already runs.

The control plane pattern.

Every agent we build sits in a control plane above the applications, never inside one of them. The applications of record stay where they are; the client already owns them. The control plane holds the operating model (what good looks like, the rules, the thresholds), the skills that run each workflow, the gates and their approvers, the run log and the conversation.

It does three things a configured application never does. It observes across systems, joining what each one knows into one picture. It decides against the operating model: score, plan, pick, hold. It acts through the applications’ interfaces in the right order, then reads back what it did and compares the result to what it intended. A person meets the agent on one surface: a conversation pane on the screen they are already working in, or the team chat on their phone when a gate needs an answer. Talk or click on the same screen; the same gates and the same audit trail everywhere.

The parts, and why each one is there.

  • Models. Chosen per task by measured error rate and cost, held behind one interface so the provider or the size can change without the rest of the system noticing. We name model providers by category, never as our identity: the agent is ours, the model is a part.
  • Tools and connectors. Gateway functions own each platform’s interface. Every write is preceded by a read in the same turn and followed by a read-back; a read-only probe is never taken as proof that a write path works.
  • Memory and state. What the agent must know between runs lives in records a person can open: the operating model, the state of each workflow, the run log. Nothing that matters is kept only inside a conversation.
  • Skills. Each workflow is a written procedure: the reads, the decision, the writes, the gate, what to do on failure. A person can read it, and a change to it is a versioned release with its own tests.
  • Evaluation. Test cases per skill, run before a change ships; outcomes measured against the numbers the business already reads, so a drift in quality shows up as a line in a report rather than a surprise.
  • Guardrails and approvals. Anything that spends money, carries risk or needs taste waits for a person’s answer. Caps are enforced in code beneath the gate, so an approval given in error still cannot spend past the envelope.
  • Observability. One row per action in the run log: who asked, from which surface and screen, who approved, what the interface returned and what it cost. A daily self-check tests every connection and restores a regression the same run.
  • Cost control. Each run carries its model and tool cost onto the outcome it produced, so the all-in cost per lead, per piece or per order is a monthly number with the agents’ own bill inside it.

How an agent is built.

The operating model comes first, in writing: the outcomes, the rules, the thresholds, who approves what. Without it there is nothing to decide against, and an agent with nothing to decide against is a script with a larger bill. Then the gateways to the client’s platforms, each proven with one real write and its read-back. Then the skills, one workflow at a time, each with its tests and its gate. Then the routines that start the work, on a schedule or on an event, each with a scope, a cadence and a defined behaviour when it fails. The run log and the report exist from the first skill, because the first week’s numbers are the baseline for every later claim.

The surface comes last and is designed early: which screen the person is on when the agent needs them, what the pane shows, how the gate reads on a phone, what the one-digit answers mean. A gate a person cannot answer in ten seconds is a gate they will stop answering.

How we test and measure an agent.

Three layers. The skill’s own tests, run before any release, cover the decisions it makes and the shapes of what it writes. The release check reads the production surface the change touches before and after the change, with a marker only the new release carries, and fails on any field that was present before and is gone. The run log then measures the agent in production: how many runs, how many gates raised, how many answered, how long each took, what each cost and what it produced. A defect a person finds that the tests missed becomes a new test in the same change, so the same defect cannot come back unnoticed.

Where the models and frameworks come in.

We are platform agnostic about the agent’s own parts as much as about the client’s applications. A model provider, an agent framework, a vector store or an evaluation harness is chosen for the job and sits behind an interface we own. We will say which ones we used and why, as information, and we will change them when a better or cheaper one appears. None of them is the solution. The solution is the operating model, the skills, the gates and the record, which are the client’s and outlive any vendor’s release cycle.

What a client gets.

  • The agents and their skills as versioned code in a repository the client can read, with the tests beside them.
  • The gateway functions to the client’s platforms, with credentials in an encrypted store that no person or prompt can read.
  • The gates, with the thresholds and the approvers configured, and the surface they are answered on.
  • The run log and a monthly report in plain language: what ran, what it produced, what it cost, what waited on a person.
  • The operating documentation: the rules, the routines, what each failure behaviour is.
  • Either a managed service, where we run and improve the agents for a monthly price, or a handover to the client’s own team.

When an agent is not the answer.

When the work does not cross systems and nobody needs to approve anything, a configured application will do, and we will say so. When what is wanted is an agent with no gate, we are not the right partner; we build the gates in.

Related
Build

AI development.

Custom AI models, natural language processing, computer vision, automation and the strategy before them.

Integrate

Backend and API development.

Custom backends, API development and integration, enterprise and cloud services, and backend testing.

Integrate

AI, ML and data science.

Use-case discovery, data modelling and augmentation, machine learning and deep learning on your data.

Operate

Deployment, operations and maintenance.

Automated deployment, CI/CD, configuration management, monitoring, support and maintenance.

Capability

Custom Software Development.

The capability these pages belong to: how we build custom software, and when we do not.

Other AI engineering. AI-driven software development All technologies

Where we're not the right answer

We'll tell you if we're a fit. If we're not, we'll tell you that too.

  • Your current stack works and nobody wants to change it
  • You want licences resold at a discount and nothing else
  • Your internal team owns the operating model and is not handing it over
  • You want hours of configuration work and nothing run for you: that is on our services pages, and it is not a managed solution
How we start

Most of our best clients come to us with a feeling, not a plan.

"Something isn't working." "We're outgrowing our tools." "We're afraid to make the wrong move." No-Risk Discovery is a short, practical conversation that gets you clarity before you commit to anything big. We'll tell you if we're a fit. If we're not, we'll tell you that too.