AI in engineering companies · Coding agents

Keeping your code local and private

For many engineering companies, source code is the most valuable thing they own. Proprietary firmware, control algorithms, years of accumulated product knowledge expressed in software. It is also exactly what a coding agent needs to read in order to be useful.

That creates a genuine tension. The more of your code an agent can see, the better it works. The more of your code it sends to a provider, the more of your most valuable asset leaves the building. This final article in the series on coding agents is about managing that tension: understanding what leaves, deciding what may, and keeping what matters local and private.


What a coding agent sends

When you use a cloud-based coding agent, what it reads is sent to the model provider as part of each request. That typically includes:

  • the files the agent opens, in whole or in part
  • search results from across the codebase
  • output from commands it runs, such as build logs and test results
  • your instructions and the project's standing instructions
  • the history of the session

Depending on the tool, it may also send telemetry about usage, and some tools index the codebase on the provider's side to make searching faster.

None of this is hidden. But it is easy to forget that an agent asked to fix a small bug may read, and send, a good deal more than the file containing it. Secrets that happen to be in the working environment, such as keys in configuration files or credentials in command output, can be sent too.


Understand the terms

The first step is to know what your provider does with what it receives. The questions that matter:

  • Is your code used to train or improve models, and can you opt out? Terms often differ between consumer, business and enterprise plans.
  • How long is it retained, and is there a zero-retention option?
  • Where is it processed, and can you choose a UK or European region?
  • Who can access it within the provider?
  • Under which jurisdiction does the provider operate? As I explained in Where should your AI live?, data residency and legal jurisdiction are not the same thing.

For most organisations, a business or enterprise arrangement with no training on your data, limited or zero retention and a suitable processing region is an acceptable basis for most code. The question is which code falls outside "most".


Classify your code

Not all code needs the same protection. A simple classification helps:

Open. Code that is public, or would do no harm if it were. Any tier is fine.

Internal. Normal commercial code. Cloud models under suitable business terms are usually acceptable.

Crown jewels. Proprietary algorithms, core firmware, security mechanisms, code under export control or covered by customer confidentiality obligations. This should stay on UK-hosted or local models, or not be exposed to AI at all.

Once the classification exists, it can be enforced: separate repositories, tool configurations that exclude sensitive paths, and different model tiers for different projects, as described in Local, cloud and frontier: getting the most from your budget.


Keeping it local

For the crown jewels, local models make AI assistance possible without anything leaving the building.

Local models for sensitive work. Open coding models running on your own hardware can explain, document, test and modify code with no external connection. They are less capable than the frontier on the hardest problems, but for much of the work, understanding, documenting and testing, they are entirely adequate.

Local indexes. Searching and indexing the codebase can be done locally, so the index, which is effectively an organised copy of all your code, never leaves.

Hybrid working. Sensitive modules handled locally, the rest of the code with cloud models. The boundary needs to be clear, and enforced by configuration, not by engineers remembering.

Verify it really is local. A local model does not make a whole toolchain local. Extensions, plugins, telemetry and the agent tool itself may still connect outside. Check network traffic, and where it matters, run sensitive work in an environment without external access.


Practical hygiene

Whatever tier you use, some practices protect code and secrets:

  • No secrets in repositories. Keep keys and credentials in a secrets manager, never in files an agent can read. Scan for them automatically.
  • Exclude sensitive paths. Most agent tools support ignore rules. Use them for secrets, customer data, and anything classified above the tool's tier.
  • Limit what the agent can reach. Run coding agents with access only to the repository they are working on, not to everything on the engineer's machine, as described in The security of AI-written code.
  • Control network access. An agent that can reach arbitrary internet addresses can send data to them, including if misdirected by content it reads, as described in When the data gives the orders.
  • Review tool settings centrally. Individual engineers should not each decide what their tools send. Set organisation-wide configurations.
  • Mind the logs. Agent session logs contain the code they read. Protect them accordingly, as described in How data leaks out of AI systems.

The evidence burden

If you tell a customer, a partner or a regulator that their code, or your own, never leaves the UK or never reaches a particular provider, you need to be able to prove it for the whole chain: the model, the agent tool, its plugins, indexes, logs and telemetry. As I noted in Where should your AI live?, a claim that is only partly true is worse than an honest description, because you can be held to it. Make the claim only where you can evidence it.


Closing the series

Across this series, I have argued that coding agents are the most powerful level of abstraction engineers have yet been given; that they are strong on well-defined work and weak on architecture; that they need experienced engineers to direct and review them; that the code they write needs the same security discipline as any other; that tokens are the unit of their cost and deserve managing; that local, cloud and frontier models each have their place; and that the organisation's code deserves the same care as any other valuable asset.

Used that way, coding agents let experienced engineers build more, faster and better than ever before, while keeping what matters private and secure.


Four things worth taking seriously

For engineering leaders: classify your code, and decide which classes may be read by which tier of model.

For security teams: treat coding agent configuration as a security control, set centrally, and verify what actually leaves the network.

For engineers: keep secrets out of reach, use ignore rules, and use local models for the code that matters most.

For everyone: your code is your product's knowledge. Use AI on it, deliberately, and know where it goes.


I would be interested to hear whether your organisation knows exactly what its coding tools send, and to whom.


Further reading in this series

Catherine Ives-Yim

Catherine Ives-Yim

Chartered Engineer and independent technical adviser, with a lifetime at the bleeding edge of embedded systems, connected products, data platforms and AI-assisted engineering, who has advised clients across the UK, Europe, the Middle East, the Far East, North America and Africa. Based in Leeds.