AI in engineering companies · The five levels of AI value

From documents to data: building an engineering model you own

In the first article in this series I described five levels of AI value in an engineering company. The second level, structure, is the one I most often find missing. Organisations try AI search over their documents, find it useful, and then jump straight to ambitions about AI agents that diagnose and act. The step in between, which is where most of the lasting value is created, gets skipped.

This article is about that step: what it means to turn documents and records into structured engineering data, why it matters more than it sounds, and how to start without turning it into a five-year information management programme.


What search cannot answer

AI search, the first level, is good at questions like "what is the procedure for replacing this controller?" or "have we seen this symptom before?" It finds passages and summarises them.

It is poor at questions like these:

  • Which products in the field contain this component, and which of those are running the firmware version affected by this vulnerability?
  • If this supplier changes a part, which products, test plans and certificates are affected?
  • Which requirements in this product have no test evidence against them?
  • Of the returns in the last six months, how many share a production batch?

Those questions do not have an answer sitting in a paragraph somewhere. The answer exists only when you connect facts that are scattered across bills of materials, build records, firmware release notes, test reports, service tickets and requirements documents. Search retrieves text. These questions need relationships.


What "structure" means in practice

Structure means identifying the things that matter in your engineering world and recording how they relate to each other. In a connected product company, those things are typically:

  • products, variants and individual units
  • components, suppliers and production batches
  • firmware and software versions, and the third-party software inside them
  • faults, symptoms, causes and fixes
  • procedures, requirements, tests and the evidence that they passed

and the links between them: this unit was built with this batch, runs this firmware, reported this symptom, was fixed by this procedure, which relates to this requirement.

None of this is a new idea. Engineers have always wanted it, and product lifecycle systems exist to hold some of it. What has changed is the cost of building it. Much of the information was only ever written as free text: in service notes, test reports, emails, meeting minutes and release notes. Extracting structured facts from that text used to need people to read every record. A language model can now do the reading, at a volume no team could match.


How AI does the work

The process has four parts, and all four matter.

Extraction. AI reads each document or record and pulls out the facts: this ticket concerns this product and serial number, reports this symptom, mentions this firmware version, and was closed with this fix.

Normalisation. The same thing is described in many ways. Part numbers are abbreviated, components have nicknames, customers describe symptoms rather than faults, and firmware versions are written three different ways. AI is good at recognising that these refer to the same thing, and mapping them to one agreed name.

Linking. Facts are connected to the records they relate to, so that a fault report links to the unit, the unit to its build record, and the build record to its components and firmware.

Validation. This is the part most often neglected. AI extraction is good but not perfect, and errors that enter a structured model travel silently into every answer built on it. Each extracted fact should carry its source and a measure of confidence, uncertain extractions should go to a person, and a sample of the rest should be checked regularly. The standard should be proportionate to the use: a model used for regulatory evidence needs a far higher bar than one used to spot trends.


Why this matters now

Three things make level 2 more valuable today than it was even a year ago.

Regulation is demanding traceability. The EU Cyber Resilience Act requires manufacturers of connected products to report actively exploited vulnerabilities, an obligation that has applied since September 2026, with the full set of obligations following in December 2027. Meeting it means knowing which software components are in which products, and which products in the field are affected by a given vulnerability. Battery passports for e-bike, e-scooter and other batteries arrive in February 2027. The EU Data Act already requires connected products to be designed so that users can access their data. All of these assume you can answer structured questions about your own products quickly and with evidence.

It is the foundation for everything above it. Connecting AI to live systems at level 3 needs identifiers that match across systems. Drafting engineering documents at level 4 needs reliable facts to draft from. Detecting emerging faults at level 5 needs faults, units, batches and versions linked together. Without level 2, each of those becomes a fragile one-off.

It is an asset the company owns. An AI model is increasingly a replaceable component. A validated, structured model of your products and their history is not. It belongs to you, it outlives whichever AI built it, and it becomes more valuable every month it is kept current. It should be held in an open, exportable form, so that it never depends on one supplier.


What goes wrong

Designing the perfect model first. The fastest way to fail is to spend six months defining every entity and relationship before extracting anything. Start with the few that answer a real question, and extend as you go.

Nobody owns it. A structured engineering model needs an owner in engineering, not only in IT. Without one, definitions drift, errors go uncorrected, and trust erodes.

Treating it as a one-off. Extracting structure from years of historical records is a project. Keeping it current is a process. New records should be captured in structured form at source where possible, with AI filling the gaps rather than repeating the historical exercise every year.

Silent errors. Without source references and validation, a wrong extraction looks exactly like a right one. That is acceptable for a trend chart and unacceptable for a compliance file.


Where it should run

Extraction is high-volume, repetitive work across the organisation's most complete body of information. That makes it a strong candidate for local or UK-hosted models, for two reasons I discuss in Where should your AI live?: the cost per task is close to zero once the hardware is in place, and the material is often sensitive. The frontier cloud models can be kept for the difficult cases the local model is not confident about.


How to start

Pick one question you cannot answer today and would value answering. For a connected product company, a good one is: "Which units in the field have this component and this firmware version?"

Build only the structure needed to answer it. Extract it from the records you have, validate a sample, and put the answer in front of the people who need it. Then add the next question.

Within a few iterations you will have the core of an engineering model, a clear view of where your records are weakest, and a foundation the rest of the ladder can stand on.


Four things worth taking seriously

For engineering leaders: the questions your organisation cannot answer quickly are a map of where structure is missing. Write them down.

For compliance and quality teams: the regulatory obligations arriving now assume traceability you may not have. AI makes building it affordable. It does not make it optional.

For anyone commissioning AI work: ask what you will own at the end. A structured, validated, exportable model of your products is worth more than any tool.

For engineers: the knowledge you record in service notes and test reports is more valuable than it looks. The more consistently it is written, the more of it can be put to work.


I would be interested to hear which question your organisation would most like to be able to answer, and cannot today.


Further reading in this series

Catherine Ives-Yim

Catherine Ives-Yim

Chartered Engineer and independent technical adviser, with a lifetime at the bleeding edge of embedded systems, connected products, data platforms and AI-assisted engineering, who has advised clients across the UK, Europe, the Middle East, the Far East, North America and Africa. Based in Leeds.