Case study

VenturePilot: an AI product built by coding agents to a written constitution

An AI system that prepares business-development opportunities and never speaks in its owner's name without approval. In private build.

A product of my own: an AI system that finds, researches and prepares business-development opportunities for independent professionals, and never contacts anyone without its owner approving the exact message. It was built by coding agents working to a written constitution and a governed work plan, with me as architect and reviewer. It is in private build, not yet in service, and this page says so.

52work packages, each completed with committed evidence before the next could start
950+automated tests across database, API, worker and web app
25database migrations, every tenant table isolated by row-level security
43,000lines of strict TypeScript across three apps and seven packages
1rule nothing may break: autonomous internally, permissioned externally
0messages it can send without a human approving the exact words

The problem

Doing the work and finding the work are mutually exclusive. When an independent professional is inside a contract or an engagement, the pipeline dries up; when the work ends, the scramble begins. The tools built to help were built for salespeople: a CRM filled in by hand, job boards that surface volume rather than fit, outreach tools that send without thinking. I wanted a system that keeps the pipeline warm in the background, and I wanted it built so that it could never damage the reputation it exists to serve.

What it does

The owner sets a brief: goals, sectors, roles, geography, exclusions, research depth and a spending limit. VenturePilot watches approved sources for candidates that fit, follows permitted evidence trails to related organisations, people, roles and events, and scores each candidate for fit, risk, commercial value and effort, showing the evidence, the unknowns and the reasons behind every score. For the best candidates it drafts a first message grounded in evidence from the owner's actual work, and flags or blocks any claim it cannot support. Then it stops and waits. The owner approves, edits or rejects, and every decision teaches the next cycle.

It serves a portfolio career rather than a single job search: consultancy, contracts, fractional and advisory roles, board seats, speaking and commercial partnerships, tracked side by side.

The control model

The system's constitution is the document every later decision has to obey, and each of its articles carries a test for what would count as a violation. Its central line is autonomous internally, permissioned externally. Inside its walls the system may discover, research, score and draft within its budget. Nothing leaves it, no email, no calendar entry, no application, without explicit, attributable, logged human approval of the exact payload. Approvals are single-use and time-bounded, and a material edit voids them. Budgets are checked before every model call, not discovered afterwards. Suppression lists and send-rate limits live in the database, not in a prompt. Every decision is on the record. I wrote about why this matters for any agent that acts in someone's name in Autonomous internally, permissioned externally.

How it was built

This is the part most relevant to clients. I wrote the constitution, the requirements, the architecture decisions and a dependency graph of 52 work packages, each with the files it owns, the checks it must pass and the conditions under which it must stop. An external review of the document set was carried out before any code was written, and its findings were dispositioned in writing. Coding agents then built the system one package at a time, through a coordinator that only allowed a package to start when its dependencies were complete and only marked it complete when its evidence (test runs, produced artefacts, the commit) had been committed and verified. The agents were forbidden to weaken row-level security, approvals, privacy, budgets or tests to get a package through, and told to stop and report rather than improvise when blocked.

The point of that machinery is that model confidence is not delivery. An agent will report success on work that does not pass; the coordinator does not believe it until the evidence says so. That is the same discipline I apply to engineering teams, made mechanical.

Architecture

A strict TypeScript monorepo: a Next.js web app, a Fastify API and a background worker over shared domain, application, contract, database, AI-gateway, integration and observability packages. PostgreSQL on Supabase is the system of record, with authentication, row-level security on every tenant table and a transactional outbox for jobs. All model calls go through an AI gateway that records cost, enforces per-tenant budgets and can swap providers; a deterministic adapter makes the AI paths testable. Self-service trial tenants are capped and metered against trial credits. The private-beta build has no external-write capability at all: approval is exercised end to end, but nothing is sent.

Where it stands

The private-beta work plan is complete and tested. The next stage is bringing it up against real infrastructure, running it for my own practice and with a small number of invited design partners, and then writing honestly about what it does when it meets real sources. A design that has not yet met the world is a design; I will update this page when it has.

Where this helps you

If you are deciding how far to let AI agents act for your organisation, building a product with coding agents and wondering how to keep control of quality, or evaluating a company that claims to have done either, this is the kind of work I can help with. See the AI Opportunity Workshop, fractional technical leadership and technology due diligence, or book a 30-minute conversation.

Tell me what you are trying to decide, or where you feel stuck.

I will suggest the smallest useful first step. Based in Leeds, working in the UK and internationally, on site or remote.

Not ready to talk? Try the free AI Ladder Check. It takes about 3 minutes.