Writing · AI and technology leadership
Write the constitution before the code
Coding agents will build whatever looks reasonable in the moment. A short document with a violation test for every rule gives them, and the people reviewing them, a standard to be held to. The checks around it make it bite.
In July I started building VenturePilot, a product that finds and prepares business-development opportunities for independent professionals, and I decided early that most of the code would be written by AI coding agents. Before any of them wrote a line, I wrote a document called the constitution. It started at about fifteen hundred words and is now around two and a half thousand, it sits at the top of the repository, and everything else in the project, every requirement, architecture decision, plan, prompt and piece of code, is subordinate to it. Fifty-two work packages later it is the single document I would least like to have done without. This piece is about why, where the idea comes from, and how to write one.
The problem it solves
A coding agent is very good at doing what looks reasonable in the moment. That is its strength and its danger. Asked to make a failing test pass, it will find the shortest route to green, and sometimes that route runs straight through a rule nobody wrote down: a database policy loosened, a check skipped, a provider's library imported directly where it should not be. Each shortcut is locally sensible. Together they turn a system you designed into one you merely own.
Human teams have the same failure, more slowly. They have shared context and professional judgement to fall back on, and neither is reliable enough to leave unwritten. An agent cannot be relied on to carry that context from one session to the next, or to know why a rule exists, so it needs the rule written down, in a form it can be checked against, before it starts.
There was a second problem. A serious product accumulates a lot of documents: a system plan, requirements, architecture decisions, data and API contracts, a security model, a test strategy, a delivery plan. They will disagree with each other, because they are written at different times by people (or agents) thinking about different things. When they do, an agent will quietly pick one. Which one it picks is not a decision you want made silently.
What the constitution does
It does four things, and each of them is deliberately simple.
It states the rules that are not traded away. Seventeen articles, from tenant isolation and human authority over external actions to cost governance, auditability and the protection of the user's professional reputation. Not a vision statement: the things that must remain true however much pressure the schedule is under.
Every rule has a violation test. This is the part that matters most. A principle on its own is a hope. A principle with a test is something a reviewer, or an automated check, can answer yes or no to. Here are three articles, slightly shortened; the full constitution is published.
Article 2: Replaceable providers and runtimes. External providers and execution runtimes are replaceable behind stable internal interfaces. Business modules request capabilities such as
opportunity.scoreoroutreach.draft; they do not select provider models.
Violation test: A provider SDK import, provider model identifier or provider-specific response type outside its adapter or gateway boundary is a violation.
Article 7: Auditable AI decisions. Every AI request and stored AI artefact is traceable: the capability requested, prompt version, provider, model, timing, token usage, cost, structured output, rationale, evidence and outcome.
Violation test: A stored AI score, classification or draft that cannot be linked to its request, model, prompt version, cost and user-facing rationale is a violation.
Article 14: Professional reputation and product integrity. VenturePilot treats each user's reputation as a protected asset. It avoids spam, fabricated experience, misleading claims, undisclosed commitments and unjustified confidence.
Violation test: Generated outreach contains an unsupported factual claim, or automation optimises message volume at the expense of relevance and consent, is a violation.
Those three tests are not the same kind of thing, and it is worth being honest about the difference. The first can be partly automated: a search for provider imports outside the gateway catches the obvious breach, though not every way a dependency can leak. The second can be enforced by the database and the test suite, which prove that a rationale and its evidence exist, though not that the rationale is a good one. The third needs a reviewer. So every article defines what counts as a violation; some violations can be blocked automatically, others need a person and evidence; and each needs a named check and someone accountable for it. The central article, the one I have written about separately as autonomous internally, permissioned externally, has the strictest test of all: any message, calendar write or other external action without a valid approval record bound to the final payload is a severe violation.
It declares an order of authority. A separate index lists every governing document and the order in which they win when they conflict: the constitution, then the accepted architecture decisions, then the requirements, the contracts, the plan, the roadmap, and the individual work packages last. And it carries one instruction that I would put in every agent's brief: an implementation agent must stop and report a conflict rather than silently choose.
It can only change on the record. Article 15 says the constitution changes only through a recorded amendment, approved by the product owner, that names the articles affected, the reason, the consequences, the effective date and a new version number. That sounds bureaucratic until you need it. VenturePilot's constitution is now at version 1.4. The first was a short statement of principles. A second was proposed when I decided to build for paying customers rather than for myself alone, and for a few days the project had two competing authorities, which is exactly the situation the document exists to prevent. Version 1.2 reconciled them, and the amendment record says what changed and why. Without that, the next agent to read the repository would have found two constitutions and chosen one. Versions 1.3 and 1.4 came later, after two external reviews of the published text: they added the governance described below, tightened the rules on approval, and moved hosting choices out into the architecture decisions. Each records who approved it and the compliance work it created, because a rule added to the constitution is not yet a rule the system meets.
It does not enforce itself
A constitution sets the standard. Tests, permissions, review and the release process are what hold the work to it, and it is easy to give the document credit for protection those provide. Each article needs a place where compliance is checked, someone who owns the check, and a defined response to a breach. Some checks belong in continuous integration, some in runtime permissions and database policy, and some in human review. And an agent must never be able to declare its own work compliant: in VenturePilot the agents are not allowed to change a package's status at all. A separate coordinator marks a package complete only when its tests and evidence have been committed and verified. The coordinator is a small program, not another AI: it reads the checks each work package declares, verifies that the committed evidence shows them passing, and refuses to advance the package otherwise. I commit the result. Both examples below happened in July under version 1.2, with that coordinator and the agents' standing instruction never to weaken security, privacy, budget or test controls to finish a task; versions 1.3 and 1.4 later wrote the governance into the constitution itself.
Two examples from the build show what that chain looks like in practice.
A tenant boundary that looked fine and was not. The constitution's article on tenant isolation says that any cross-tenant link the database permits is a violation. The agent that built the first database packages wrote row-level security on every table, and its own ninety-eight tests passed. The work plan then required an independent security review package, run by a separately briefed agent, before anything could depend on that layer, with the completion rule that no critical or high finding remain. The reviewer ran a twelve-vector cross-tenant attack and found two real vulnerabilities, spread across three privileged database functions that bypassed row-level security and accepted a tenant identifier from the caller without checking the caller belonged to that tenant. Through one function, a customer could have forged audit records against another; through the other two, it could have forged usage charges. The functions were given a membership guard, six regression tests were added, and the attack was re-run until all twelve vectors were blocked. The article made tenant isolation a condition of acceptance; the review package and its completion rule made sure someone tested it.
A failing test with an easy wrong answer. Later, eleven tests failed because the scoring capability demanded the strictest privacy class for profile and opportunity data, and the tests asked for a weaker one. The quickest route to green was to relax the capability to accept the weaker class. The repair record says, in terms, that the privacy class was not weakened: the tests were corrected instead, because the stronger protection was the rule and the tests were wrong. That is the shortcut the constitution exists to block, and the reason the agents' brief says never to weaken security, privacy, budget or test controls to complete a task.
How an agent actually gets it
A question I am often asked: is the constitution sent with every request to the model? In effect, yes, though nobody pastes it in. It arrives in three layers.
First, the coding tool loads a short instruction file automatically at the start of every session; most tools now look for an AGENTS.md or similar in the repository. VenturePilot's does not contain the constitution. Its job is to say: read the constitution before anything else.
Second, every work package starts with a prompt written by the coordinator, and the first lines of that prompt are a mandatory reading list: the document index, the constitution in full, the agent rules, then the package's own documents. The agent opens each file with its own tools, so the full text is in its working context before it changes any code.
Third, model interfaces have no memory between calls. Each request the agent makes carries the whole conversation so far, including the constitution it read. So for the rest of that session, every call to the model carries the rules. At around three and a half thousand tokens it is not free, but providers cache a repeated opening, so after the first call the cost is small.
The product's own AI calls are different. When VenturePilot scores an opportunity or drafts a message, the model never sees the constitution; it gets a short prompt for that one task. The constitution is enforced around those calls in code instead: the gateway checks budget and privacy class before the call is made, the database enforces tenant isolation, and nothing is sent without an approval record. That is deliberate. A model that reads the open web can be argued with by what it reads; a database policy cannot.
Keeping it in front of the agent
The layers above fail in predictable ways, and each has a simple counter.
It falls out of a long session. When a conversation outgrows the model's context, coding tools summarise or drop the oldest material, and the constitution is old material. Keep sessions short and bounded: one work package per session, starting fresh, with the constitution read at the start of each. If a session has to resume or has been compacted, make re-reading the constitution the first step.
It is skimmed, not read. A long document read once is weaker than a short rule restated at the point of risk. Keep the constitution short, put the handful of no-exception rules in the instruction file the tool loads automatically, and repeat the rule that matters in the prompt for the work that could break it. Every VenturePilot package prompt says, in its own words, never to delete, skip or weaken a test to get to green.
The agent reads an old version. Rules change, and an agent that read last month's copy will follow last month's rules. Keep one current file referenced everywhere by name, update every reference when it is amended, and have the agent record in its completion evidence which version it worked to, so a reviewer can see at a glance.
Work is handed to a helper that never saw it. Some tools let an agent start sub-agents, and a sub-agent starts with its own, emptier context. Either switch that off, as VenturePilot's configuration does, or make the reading list part of every hand-off.
Something it reads argues with it. A web page, a document or a tool result can contain instructions, and a model may follow them. Deny what the work does not need (VenturePilot's build agents cannot fetch web pages or search the web), state in the constitution that content can never grant authority, and keep the rules that matter most in places a persuasive paragraph cannot reach: permissions, database policy and the checks that decide acceptance.
That last point is the real answer. Everything in this section makes the agent more likely to follow the rules. None of it guarantees that it will, which is why the checks around it exist.
Where the idea comes from
I did not invent it, and it is worth being clear about the lineage.
The nearest relative is spec-driven development, which GitHub made popular with its open-source Spec Kit in 2025. Spec Kit's first step is to write a project constitution, a set of principles that every later specification, plan and task must respect, before asking an agent to build anything. My first version was in that tradition. Alongside it sits the now-common convention of an AGENTS.md or similar file in the repository telling coding agents how to behave there; VenturePilot has one, and the first instruction in it is to read the constitution.
The older relatives are the ones I recognise from earlier in my career. Architecture decision records. Design authorities that a change has to get past. And above all the discipline of safety-related engineering, where I started out writing real-time software for avionics: requirements traced to tests, a clear hierarchy of documents, and the assumption that anything not written down will eventually be got wrong. A constitution with violation tests is that discipline, cut down to fit a small product and pointed at a new kind of engineer.
It is not the same thing as Constitutional AI, the name Anthropic uses for training a model against a set of principles. That is a constitution for a model's behaviour in general. This is a constitution for one software project, read by whatever agents happen to be working on it.
What I would claim as less usual is the combination, since a list of principles on its own is easy to write and easy to ignore; one practitioner's guide to Spec Kit constitutions is titled, fairly, stop shipping platitude charters. Mine has a violation test on every article, a declared authority order with a duty to stop, an external review of the whole document set before any code was written (fourteen findings, thirteen accepted and one rejected, each in writing), and a coordinator that would not mark a work package complete until its tests and evidence had been committed. The constitution says what must be true; the coordinator refuses to believe it is true without evidence. Coding agents need experienced engineers, and this is a large part of what the experienced engineer is for.
How to write one
If you are about to let agents build something that matters, this is the order I would do it in.
Start with what must never happen. Not features. The failures that would be expensive, embarrassing or illegal: one customer seeing another's data, a message sent without consent, an unbounded bill, a decision nobody can explain. Each becomes an article.
Write the violation test before you are happy with the principle. If you cannot say what a breach would look like in the code or the data, the principle is not finished. Prefer tests a machine can check.
Keep it short. Mine is two and a half thousand words including its amendment record, and I would not want it longer. Detailed design belongs in the architecture decisions and features in the requirements; the constitution states the constraints those decisions must respect.
Declare the order of authority and the duty to stop. Then make sure every agent's instructions point to it first, and that every new session reads it before starting work.
Give every article an enforcement mechanism, an evidence requirement and an owner. Decide where each check runs, what evidence shows it passed, who resolves a breach, and what happens when one is found: stop the work, block the release, or contain and investigate. Never let the thing being checked mark itself as passed.
Have someone else read it before you build. VenturePilot's review found fourteen issues in the document set, from a missing consumer for a shared component to a circular dependency between implementation and review packages. Every one was cheaper to fix on paper than it would have been in code.
Amend it on the record, and say who may. It will need to change. A version number records a change; it does not authorise one. Name who interprets an ambiguous article, who approves an amendment, and how existing work is brought back into line when the rules move.
Three things worth taking seriously
For anyone building with coding agents: write the constitution, with a violation test for every rule, before the first prompt, and give each rule an automatic check or a required review. The document is the standard every agent, session and handover can find; the checks are what make it hold.
For engineering leaders: the same document works for human teams. The rules your senior engineers carry in their heads are exactly the ones that disappear when they leave, or when an agent joins.
For boards and investors: ask any company that says it builds with AI agents what the agents are not allowed to do, and how it would know if they did. A good answer names a rule, the check that enforces it and the person who owns it. A poor one names a model.
What are the rules in your system that nobody has written down, and which of your engineers, human or otherwise, knows them?
The full constitution, version 1.3, is published with my notes as The VenturePilot constitution, annotated, with a blank template you are free to adapt. The project behind it is described in the VenturePilot case study. For larger organisations with several products or teams, see One constitution or many? If you are deciding how to govern agents in your own organisation, that is a conversation I have often.
© 2026 Catherine Ives-Yim. All rights reserved.