AI in engineering companies · Agents

Building production software with AI coding agents

I have written before about leading the build of a production platform with AI, and about how it removed the translation layer between an architect's thinking and the code that gets written. That article was about what changed. This one is about how it actually works day to day: what AI coding agents are good at, where they fail, and what I have learned to do about it.

The platform is Fienti, a connected vehicle platform with four user surfaces, over-the-air firmware updates with integrity checking and rollback, multi-region data handling, and extensive automated security testing. I built it for a client launching a new business, as the architect guiding every part of the work. AI coding agents did a great deal of the typing. They did not do the engineering.


What the agents actually do

A coding agent is the clearest example of the loop I described in What an agent actually is, and isn't. Given a task, it reads the relevant code, makes a change, runs the tests or the build, reads the result, and tries again until the task is done or it is stuck.

In practice, I use agents for:

  • implementing a component to a design I have specified precisely
  • writing tests, including awkward edge cases I would otherwise skip
  • tracing how data flows through unfamiliar parts of a codebase
  • refactoring, once I have decided the target structure
  • building simulations to test an algorithm before it touches real hardware
  • first drafts of documentation, kept alongside the code

What they have in common is that the decision about what should exist has already been made. The agent's job is to make it exist, correctly, and prove it.


What worked

Holding the architecture myself. The single most important practice. Every significant decision about structure, data ownership, security boundaries and failure handling was mine, written down before an agent touched the code. Agents are excellent at filling in a well-defined design. Left to invent the design, they produce something plausible, locally sensible, and incoherent as a whole.

Small, bounded tasks. An agent given a large, vague task wanders. Given a specific change, with a clear definition of done and the files it should touch, it is fast and accurate. Breaking work down well is where the architect's skill shows.

Tests as the contract. Agents are far more reliable when there is an objective way to tell whether they succeeded. Tests written, or at least specified, before the implementation give the agent a target and give me evidence. Without them, "it compiles" quietly becomes the definition of done.

Checking with a different model. I routinely have a second, different AI model review work produced by the first, particularly for security-sensitive code. Different models make different mistakes. Where they disagree is exactly where I look hardest. It does not replace my review, but it improves what reaches me.

Independent testing. AI-assisted code went through the same security testing as any other code: thousands of automated tests with industry-standard security tools, run by different AI models from the ones that wrote the code. That is not optional.


What did not work, or needed watching

Confident wrong code. Agents produce code that looks right, reads well and passes a casual review, and is subtly wrong: a race condition, an unchecked edge case, a security check applied in the wrong place. The better the code looks, the more careful the review needs to be.

Drift from the design. Over a long session, an agent can gradually depart from the agreed structure, solving each local problem in a way that erodes the whole. Regular review against the design, and short tasks, are the defence.

Invented dependencies. Agents sometimes suggest software libraries that do not exist, or that exist under slightly different names. This is more than an inconvenience. Attackers have started publishing malicious packages under names AI tools are known to invent. I cover this in The AI supply chain, in the security series.

Testing the wrong thing. An agent asked to make the tests pass will sometimes change the tests. Tests need to be protected, reviewed and treated as the specification, not as another file the agent can edit.

Looking busy. An agent that is stuck will keep trying. Without limits on time, retries and cost, it can spend a surprising amount achieving nothing. I discuss the economics in The cost and infrastructure of agents.


What this means for engineering teams

The lesson I draw is not that teams are unnecessary. It is that the valuable skills shift.

Architecture becomes more valuable, not less. The quality of what agents produce depends directly on the quality of the design they are given. An organisation that invests in AI coding tools without investing in architectural clarity will simply produce more code faster, with the same problems.

Review becomes the core skill. Most of an experienced engineer's time with coding agents is spent specifying and reviewing. Reviewing AI output well is a distinct skill: assuming nothing, checking the edges, and asking what the code does that it was not asked to do.

Testing is no longer optional. Agents make writing tests cheap, and they need tests to work well. There is no longer any excuse for untested code.

The same principle as elsewhere. As I described in AI drafts, engineers decide, the AI drafts and the engineer decides. Code is no different. Someone must own every line that ships.


Four things worth taking seriously

For engineering leaders: AI coding agents multiply whatever discipline you already have. With clear architecture and good testing, they are transformative. Without them, they multiply the mess.

For senior engineers: your value is moving towards design, specification and review. Invest in those, and use agents to take the typing off your desk.

For security teams: treat AI-written code exactly as you treat human-written code. Review it, test it, and check every dependency it introduces.

For anyone adopting coding agents: start with tests. They are the contract that makes agents reliable, and the evidence that makes their work trustworthy.


I would be interested to hear how your teams are using coding agents, and whether review has become the bottleneck.


Further reading in this series

Catherine Ives-Yim

Catherine Ives-Yim

Chartered Engineer and independent technical adviser, with a lifetime at the bleeding edge of embedded systems, connected products, data platforms and AI-assisted engineering, who has advised clients across the UK, Europe, the Middle East, the Far East, North America and Africa. Based in Leeds.