Writing · AI and technology leadership

Five things I was sure of

A passing test suite is where most teams stop checking agent-written code. In one small system it missed all five of the claims that mattered, and in another, two all-green rounds of penetration testing missed seven vulnerabilities.

This summer I built a noise-monitoring system, unpaid, as community work. A Class 1 sound level meter records to an SD card; the system decodes the meter's proprietary binary files, stores each measurement, calculates the statistics that noise assessments rely on, and produces reports and exports. It runs on two small computers at different sites that copy their data to each other, so that either can fail without losing anything. Most of the code was written by AI coding agents working to my plans, and its output is evidence that goes to a local authority.

That last fact set the rule everything else followed: a plausible but wrong number is the worst possible failure, worse than a crash or a blank page. A crash gets noticed. A blank gets questioned. A confident wrong figure in a report gets relied on.

The project keeps a document called the validation brief. It lists every significant claim the system makes, ranks them by how much damage a wrong one would do, says how confident I am in each and where it could still be wrong, and gives a reviewer the commands to try to break each one. By its fourth revision it also recorded five claims I had made with confidence and later had to retract. I left them in, because how they were caught is the most useful thing in the document.

None of the five was caught by a failing test.

1. "Fifty-four files log at a 1.25-second period"

Some of the meter's files decoded to a slightly odd logging interval: 1.25 seconds instead of one. It was consistent across all fifty-four of them, which felt like confirmation. It was the opposite. 1.25 is ten divided by eight, and that was the clue: those files used eight-byte records, and the parser was reading them as ten-byte ones. Every value in them had been decoded from misaligned data.

The consistency was the bug. Fifty-four files agreeing with each other proved only that they shared the same misreading.

What catches it: whenever results are suspiciously consistent, ask what else would produce exactly that pattern. Then add an independent check that does not share the assumption. The suite now confirms the record size three separate ways, including whether the record count matches the elapsed time.

2. "The display never shows a wrong end time"

The system shows when each measurement ended. A guard checked the meter's own end time against the recorded duration and rejected anything more than five minutes out. I wrote that the display would therefore never show a wrong figure.

It would have. Behind the guard sat an older piece of code that, when the meter's value was rejected, worked an end time out by arithmetic. A measurement legitimately paused for more than five minutes would have been rejected and then replaced with a confident, wrong time. I had reasoned carefully about the part I built and not at all about the part I inherited. An independent review caught it.

The fix was not a better threshold. A pause can be any length, so no threshold is safe. The arithmetic was deleted. The display now shows the meter's own value or nothing.

What catches it: fail closed. When in doubt, show nothing. And review the inherited path behind any guard, not just the guard.

3. "It fails closed on conflicting signals"

The validation brief described a function that decides the record size as comparing two independent signals and refusing to guess when they disagreed. The code returned as soon as the first signal gave an answer, and never looked at the second.

The document and the code were written together, in the same session, and still disagreed.

What catches it: describing intended behaviour is not evidence of it. Every claim in a design document that matters should have a test that would fail if the claim were false.

Assessments link to individual measurements. Those links used the measurement's position in the day, so recovering a file that had previously failed to import shifted every later position and silently re-pointed existing links at a different measurement. I fixed it by linking on the measurement's file name instead.

I fixed it on the read path only. Writes still used the position. Reads and writes were now resolving on different keys, which is worse than either alone: the review reproduced a link appearing twice and another silently moving to a neighbouring measurement. I had even named the write path as my main worry in the handover notes, and shipped anyway.

What catches it: stating a risk is not the same as clearing it. Test the write path as well as the read path, and audit the live data before tightening a constraint. A half-applied fix is worse than none: it removes the smell and leaves the bug.

5. "The client sends the stable key"

The fix to the assessment link needed the browser to send the file name along with the position. The validation brief said it did. It did not. The script that made that edit had stopped on an error before writing the file, so the change never reached disk, and the field had been posting as empty ever since.

At the same time, the screen that assigns measurements to an assessment was broken from end to end. All 221 checks in the suite passed, because the tests spoke to the data layer and never once called the screen's own route.

What catches it: look at the result, not the intention. Read the file back. Call the endpoint once. And put at least one test at the boundary a user actually touches.

And one more, from a different system

Earlier this year I built a multi-tenant platform with a web portal and a connected mobile app, again largely with coding agents, and security was the part that mattered most. Its record tells the same story at a larger scale.

The first full security pass found one critical issue and nine high ones. The critical one was a development convenience left switched on in production that let a maintenance function run any database command it was given. Almost every other finding had the same shape: a shortcut that made development easier. An ownership check left out, so any user could act on any customer's device. A rate limit left at its test value. A debug log that printed user profiles. Firmware sent to devices without checking it was the firmware we built. None of these is exotic. Each is the shortest route to a working demo.

After those were fixed, two rounds of automated penetration testing came back clean on consecutive days: every script passed, no exploitable vulnerabilities. Two days later I asked a model from a different company to review the code adversarially. It found seven more, two of them high severity. One was a role check that let every role through, with a comment beside it saying permission was enforced by a script in the browser. A browser script is not a security boundary; any logged-in user could call that interface directly.

The automated tests were not wrong. They confirmed that the defences we had designed still worked. They could not find a defence that had never been written. Reading the code with intent to break it did.

What the five have in common

Each of the five was a confident statement resting on one unexamined assumption: that consistency meant correctness, that the code I inherited behaved like the code I wrote, that the document matched the code, that a fix on one path was a fix on both, that a change I made had landed. Every one of them sat behind a passing test suite.

Agents make this worse in a specific way. They are very good at producing work that looks finished, and very good at summarising it confidently. Their summaries describe what they intended. So do mine, it turns out. The habits that caught these five are old and unglamorous, and cost almost nothing.

  • Make the test prove it can fail. The regression test for a sync defect in this project comes with an instruction: delete one named line and expect this test to fail. If it does not fail, it was never testing anything.
  • Test against something that did not write the code. The meter maker's own export software became a reference; comparing against it found six export defects the suite had passed.
  • Use real data. The suite runs against 527 real files from the meter's own cards, which is where the 1.25-second oddity came from.
  • Have something else attack it. A second AI system, from a different company, was briefed to review the work adversarially and not to extend it courtesy. Independent review rounds found a database corruption path, a live security hole, a wrong statistical channel and a sync defect that would have lost data.
  • Rank your claims by the damage a wrong one would do, and say where each could still be wrong. It tells a reviewer where to look first.

The suite now runs 673 checks. That is not the number I would point to. The number I would point to is five: the claims that a passing suite did not catch, written down where the next person, or the next agent, will read them.


Which of your system's claims rests on "I changed it" rather than "I saw it change"?

The rules that must never break are a different problem, covered in Write the constitution before the code. If you want an independent view of agent-built software before you rely on it, that is a conversation I have often.

© 2026 Catherine Ives-Yim. All rights reserved.

Catherine Ives-Yim

Catherine Ives-Yim

Chartered Engineer and independent technical adviser, with a lifetime at the bleeding edge of embedded systems, connected products, data platforms and AI-assisted engineering, who has advised clients across the UK, Europe, the Middle East, the Far East, North America and Africa. Based in Leeds.