AI in engineering companies · The five levels of AI value

Where should your AI live? Cloud, UK-hosted or in the building

Over the past year, AI has become the most productive tool I have ever used as an engineer. I have used it to design, build, test and review systems that would previously have needed a team. That experience has left me with a question I now hear from almost every engineering organisation I talk to, usually phrased as a worry rather than a question: "Can we actually use this on our own work?"

Behind that worry is a real decision that most organisations are making badly, or not making at all. Where should the AI run? On a general cloud service, on infrastructure hosted in the UK, or on hardware in your own building?

The usual answers are either "just use the best model", or "we can't, it's confidential". Both are wrong more often than they are right. This article sets out the trade-offs as I see them from the engineering side, and where I think the useful opportunities are.


Three places your AI can live

It helps to be precise, because the terms get used loosely.

General cloud. A frontier model accessed through the provider's own service or a major cloud platform. This gives you the most capable models available, with the least effort. You can often choose a UK or European region and a zero-retention arrangement, which means your prompts and data are processed in that region and not kept.

UK-hosted open models. A model whose weights are publicly available, running on infrastructure in the UK that you, or someone working for you, control. No third-party AI provider sees your data. The price is capability: the best open models are good, and improving quickly, but they still trail the frontier on the hardest reasoning and coding tasks.

In the building. The same open models running on your own hardware, on your premises, potentially with no network connection at all. Nothing leaves. This is the most private option and the most demanding to set up, run and keep current.

None of these is the right answer. Each is the right answer to a different question.


The distinction almost everyone gets wrong

I want to be careful here, because this is where most procurement conversations go wrong.

Data residency is not the same as jurisdiction. Choosing a UK region from a US provider means your data is processed in the UK. It does not change which legal system can compel that provider to produce it. If what worries you is where your data physically sits, residency solves it, cheaply, and a competent supplier will already offer it. If what worries you is who could be legally obliged to hand it over, residency does not address that at all.

Model provenance is not the same as jurisdiction either. Model weights are files. Running an open model that was trained in another country, on hardware in Yorkshire, creates no legal relationship with whoever trained it, and no data leaves the premises. Conversely, calling a European provider's API still sends your data to a third party. A European model hosted on a US-owned cloud is still under US jurisdiction.

These distinctions matter because they change what you actually need to buy. In my experience, most organisations that say they need their AI to avoid particular countries actually need their data to stay in the UK and not be retained. That is the easiest requirement to meet and the least expensive. A smaller number genuinely need to be outside foreign jurisdiction, and a smaller number still, mostly defence-adjacent, have a real requirement about who trained the model. Working out which of these you actually have is the first and most valuable step.


The trade-offs, honestly stated

For each option there is a cost, and the cost is not always money.

Capability. The frontier models are measurably better on complex, multi-step engineering work: reasoning across a large codebase, holding a whole system design in view, catching subtle interactions. Open models close the gap every few months, but for the hardest problems the gap is still real. Choosing a more private option means accepting a less capable assistant, and you should know how much less for your specific task before you decide.

Cost and effort. Cloud is cheap to start and scales with use. Local hardware is a capital purchase that has to be sized correctly. The memory in the machine decides which models will fit, and the memory bandwidth decides how fast they run. Those are different things, and buying the wrong one is an expensive mistake. A capable workstation is excellent for one engineer, or a handful of automated agents. It is not a service for a whole company. Serving many users at once is a different engineering problem, with different hardware.

Operation. A cloud model is maintained by someone else. A local model is yours to update, secure, monitor and support. That is real ongoing work, and it belongs in the business case.

Connectivity. This one is often overlooked. Some engineering happens where there is no reliable network: on a factory floor, in a test cell, on a vehicle, in the field. An offline model is sometimes the only one that can be where the work is.

Evidence. If you claim that your data never leaves the UK, or never touches a foreign provider, you have to be able to prove it for the whole chain. That means where inference runs, who owns the hosting company, where the embeddings and search indexes are stored, where the logs and telemetry go, and what the development tools themselves send out while the system is being built. A partly substantiated claim is worse than an honest statement that you use a US provider in a UK region, because someone can hold you to it.


Stop choosing one option for the whole organisation

The most useful shift I can suggest is this: decide where each piece of work runs, not where "the AI" runs.

Most engineering organisations have a mix of material. Some is public or low-risk: general design questions, standard components, published regulations. Some is commercially sensitive: product roadmaps, costings, customer data. Some is the crown jewels: firmware source, proprietary algorithms, test data from unreleased products.

It makes little sense to route all of that through the most private option and accept the weakest model for everything. It makes even less sense to route all of it through a general cloud service and hope. A sensible arrangement uses the frontier model for the work that can go there, and keeps the sensitive material on UK-hosted or local models, with a clear rule about which is which. The rule is the important part. Without one, people make the decision themselves, one prompt at a time, and usually in favour of convenience.


Blending local and cloud: cost and ownership

Sensitivity is not the only reason to blend. The two other reasons, which I think are underrated, are cost and ownership.

Cost behaves differently in each place. Cloud AI is paid by use. That is ideal when use is occasional, and it is where the best models are. It becomes expensive in a way people rarely anticipate once AI is doing sustained work rather than answering questions. An AI agent working through a codebase, or processing months of test logs, reads the same material again and again, and every pass is billed. Local hardware is the opposite: a fixed cost up front, then close to nothing for each additional task. It pays for itself only if it is kept busy. A capable machine that sits idle most of the week is an expensive ornament.

That points to a natural division of labour. High-volume, routine work goes local: classifying fault reports, summarising logs, indexing documents so they can be searched, first-pass reviews, generating test cases. It is repetitive, the local models do it well, and the marginal cost is negligible. The difficult, occasional work goes to the frontier: the design question that needs real reasoning, the subtle defect that the local model could not explain. A good arrangement lets the local model handle what it can and pass on what it cannot, so the expensive capability is used only where it earns its cost.

Ownership is about continuity as much as control. When you use a cloud model, you rent a capability that someone else can change. Models are updated, retired and repriced on the provider's schedule, not yours. For most work that is fine. For engineering work that has to be repeatable, it can be a real problem. If an AI-assisted analysis forms part of your safety case or your compliance evidence, you may need to show which tool produced it and to get the same result again next year. A model you run yourself does not change unless you change it.

The more important point is what you actually own. The model is increasingly the least valuable part of the arrangement. It is replaceable, and better ones keep arriving. What is valuable, and specific to you, is everything around it: the organised and indexed body of your own technical knowledge, the rules for what goes where, the prompts and workflows that encode how your engineers work, and the record of what the AI was asked and what it answered. That is the part to own and to keep in-house. Build it so that models can be plugged in and swapped out, local or cloud, and you are never locked into one supplier, never exposed to one price rise, and never starting again when the next model arrives.

A blend is also a way to reduce what leaves the building at all. A local model can prepare work before it goes anywhere: extract only the relevant section, remove names and identifiers, and summarise rather than send the raw data. I would not rely on that alone for genuinely sensitive material, because automatic redaction is imperfect. As a way of reducing exposure for everything else, it is effective.

The honest cost of blending is complexity. Two sets of systems to run, a routing rule to maintain, and a need to measure whether the split is actually saving money. For a small organisation with light use, a well-configured cloud service may be all that is needed. The blend earns its keep when volume is high, when the material is sensitive, or when the work has to be reproducible, and in engineering, at least one of those is usually true.


Where the useful opportunities are in engineering

This is the part that gets lost when the conversation is only about risk. There is a great deal of valuable engineering work where AI now makes a real difference, and much of it involves exactly the material organisations are most nervous about sharing. That is why the choice of where it runs matters so much.

Understanding code nobody fully understands any more. Many products run on firmware and software whose original engineers have moved on. AI is remarkably good at reading a large, poorly documented codebase and explaining what it does, mapping the dependencies, generating tests and highlighting what looks wrong. The code is usually the organisation's most sensitive asset, which makes it a natural case for a local or UK-hosted model.

Security and compliance evidence. Connected products now face serious regulatory obligations, from vulnerability reporting for connected products in the EU to data-access requirements and, for batteries, digital product passports. Much of that work is careful, repetitive analysis: building an inventory of the software components, checking them against known vulnerabilities, tracing a requirement to the evidence that it has been met. AI accelerates it considerably. The inputs, again, are things most manufacturers will not upload casually.

Making sense of operational data. Telemetry from connected products, test rig logs, maintenance and warranty records, quality reports: organisations collect far more of this than they analyse. AI can find the fault patterns, correlate failures with conditions, and turn a pile of returns data into an argument about where the money is going. Fleet data often contains personal information, which brings data protection into the decision about where it is processed.

Design review and cross-checking. Checking a design against a datasheet, a standard, or last year's failure reports is slow, skilled work that people postpone. An AI assistant that has read all of it can do a first pass in minutes and flag what an engineer should look at. It does not replace the engineer's judgement. It points that judgement at the right places.

Beyond software. None of this is confined to software engineering. The same applies to manufacturing process data, non-conformance reports, maintenance logs, environmental monitoring, and test laboratories. I have used exactly this approach to analyse years of aircraft movement and noise data. Anywhere there is technical data and specialist documents that people do not have time to read properly, there is an opportunity.

In each case, what makes the work valuable is also what makes it sensitive. That is not a reason to avoid it. It is a reason to choose deliberately where it runs.


Four things worth taking seriously

For engineering leaders: the question is not whether to use AI on your own work. Your competitors already are. The question is which work, on which model, under which rule. Write the rule down.

For anyone paying for AI: own the knowledge, the rules and the record, and treat the model as a replaceable component. Put the high-volume work where it is cheap and keep the frontier for the problems that need it.

For anyone buying AI services: ask the supplier to distinguish residency, jurisdiction and provenance, and to evidence each claim end to end. If they cannot, the claim is marketing.

For engineers: the most valuable use of these tools is not writing code faster. It is applying your judgement to far more of the system than you previously had time to examine. That is where the real gain is, and it does not depend on which model you use.


The models themselves are becoming a commodity. What is not a commodity is knowing what your data actually requires, which model is good enough for each task, and how to build the whole arrangement so that you can stand behind it. That is an engineering problem, and it deserves to be treated as one.

I would be interested to hear how your organisation is approaching this, and whether you have found the balance, or are still waiting for someone to decide.


Further reading in this series

Catherine Ives-Yim

Catherine Ives-Yim

Chartered Engineer and independent technical adviser, with a lifetime at the bleeding edge of embedded systems, connected products, data platforms and AI-assisted engineering, who has advised clients across the UK, Europe, the Middle East, the Far East, North America and Africa. Based in Leeds.