I have spent twenty years running product organisations end to end, and most recently I have been helping companies work out how AI agents (software that can act on its own, as well as answer questions) can be built safely into the way they operate. In practice, that means working out what is worth automating, testing it against real data instead of a tidy demo, and putting the safeguards in place that let it run unattended, without anyone having to hold their breath.
I have spent twenty years running product organisations, and more recently I have been building working prototypes with AI coding tools, alongside the engineers who build production systems. That combination is the point. I am as comfortable talking an executive team through a new operating model as I am sitting with an engineer to work out why an AI agent has started giving the wrong answer (often because it has been asked to hold too much in its head at once, which is known, a little grandly, as its 'context window' filling up).
I have restructured two companies around a new way of working. IPOINT moved from a services-led business to a product-led one, and Soho House's member platform was rebuilt from a booking tool into a content-led product, with full P&L ownership throughout. The method is much the same each time: prototype the workflow, test it with real users, and then build the governance underneath it, so that other people can keep building safely after I have gone.
Before product came engineering: an MSc in Computer Science and a BEng in Electronic & Computer Engineering, both from the University of Birmingham. Before that, I spent three years teaching in elementary and junior high schools in rural Japan, including special needs classes, which is where my interest in working across languages, regulations and very different management cultures began.
Each piece of work stands on its own and can be commissioned separately. Taken together, they trace the route from a handful of people using an AI assistant well to AI being part of how the company runs, day to day.
Two to three weeks are spent breaking the organisation's work into its parts: the purpose behind it, the decisions involved, the information it needs, how it is carried out and what tends to go wrong. The candidates are then ranked against evidence, instead of against whoever is most enthusiastic about them. The result is one chosen piece of work, a baseline to measure progress against, and a short written case for why it goes first.
Over four to six weeks, a working pilot is built on your own data, in your own environment. It is then tested against examples nobody on the team has seen before, to check that the system has learned the task and has not simply memorised the answers. It ends with a working system and a clear recommendation on whether to proceed.
Autonomy is handed over gradually. Each AI agent gets its own restricted identity, so it can only touch what it needs to (much as a new employee gets limited system access on their first day), and its work is checked against a proper set of test cases before it is trusted with more. This stage also designs for the one risk that more secure hosting cannot fix: that something an agent reads, such as an email, a document or a web page, steers what it does next. Every workflow that comes out of this stage has a named owner and a written statement of the remaining risk, which somebody has read and signed.
The groundwork for the first workflow (how information is gathered and fed in, how the parts of the system talk to one another, how it is evaluated and how it is kept under control) is turned into something reusable: a platform the rest of the organisation can build on. The effect compounds, so by the time a fourth workflow comes along, it tends to cost a fraction of the first.
Most organisations arrive at AI in much the same way. A few people start using an assistant on their own initiative, one or two get rather good at it, and licences are bought more widely. A dashboard shows adoption climbing, and everyone takes some comfort from that. Eighteen months later, the company works exactly as it always did, because the work itself never changed, only the tool people used to do it. What follows is a method for changing the work.
Adopting AI tools changes what people use. Redesigning the work around AI changes what a company can do, and that is the part that shows up in the numbers.
Here, the AI produces a first draft and the person edits it, so it becomes less a blank page and more a thinking partner to argue with. Everyone in the organisation should reach a confident L1 fairly quickly.
Whole tasks are handed over and reviewed afterwards, from within a saved workspace that already has the right context loaded into it. It's here that most of the model's mistakes tend to get caught.
At this level, people design multi-step workflows that run with proper checkpoints built in, build reusable components that others go on to use, and mentor those still working their way up from L1.
Two quite different questions tend to get asked together, and answered as if they were the same. Where does it run decides who can see the data, and what it all costs. What may it touch decides what the system is allowed near in the first place. Once the two are separated and set against each other, the right deployment option for a piece of work tends to become clear, and it is rarely the same answer twice. Click a cell to see why.
| In-suite AI | Enterprise assistant | Own agents, vendor API | Own agents, own tenant | Open-weight, self-hosted | |
|---|---|---|---|---|---|
| Public / low-sensitivity | Proceed | Proceed | Proceed | Proceed | Proceed |
| General internal | Proceed | Proceed | Proceed | Proceed | Conditional |
| Confidential business | Conditional | Conditional | Proceed | Proceed | Avoid |
| Employee personal data | Avoid | Conditional | Conditional | Proceed | Avoid |
| Customer personal data | Avoid | Avoid | Avoid | Conditional | Avoid |
Moving right along this table only changes who can see the data and what it all costs. It changes nothing about who could steer the model through something it reads, what it is allowed to do, or who is accountable when it gets something wrong. Those are decisions about architecture and management, not hosting, and they are what the next diagram is about.
It is worth being clear about one thing: a language model cannot reliably tell a real instruction from its user apart from one hidden in something it is reading, such as a stray line of text on a web page or an email dressed up as something else. The major AI labs say so themselves, and none of them describe it as solved. Where the model runs changes who can see the data and what it all costs. It changes nothing about this risk.
What works is containment, sometimes called the Rule of Two. An AI session can reasonably run unattended if it has no more than two of three properties: access to private data, exposure to content it did not write, and the ability to act or communicate with the outside world. Give it all three at once, unsupervised, and you have built exactly the thing a security team loses sleep over.
No meaningful exposure from this combination, so not the interesting case. Tick a box above.
A deterministic core, with models kept to the edges. Language models are used to interpret, draft and triage. Calculations, transactions and records of decisions are kept in deterministic code, which behaves the same way every time and so can be replayed and audited. That is what makes a system built partly on unpredictable AI explainable to a risk function, when the question is asked (and it always is).
Decisions become precedent. Every approved judgement is saved as a reusable precedent, so the system becomes more capable through use, as well as through further development.
No vendor becomes the system of record. Context, skills, evaluation sets and traces are all kept in formats the organisation owns. The choice of model therefore stays a decision that can be revisited for each workflow, on the evidence, instead of being locked in from the start.
What matters is hours given back, lead times cut and error rates reduced. Seats issued and prompts run measure activity, not value. If the value cannot be named and shown against a proper baseline, the work has not been done yet.
A prototype that runs on the organisation's own data, built in weeks, tends to settle debates that a hundred slides never could, not least because it is put in front of the people who will have to live with the result.
The destination is described; only the next step is committed to. Each phase is a real decision gate, and the first one is kept cheap enough that stopping there would cost almost nothing.
AI does the assembly work and the first pass; the decisions with real consequences are still made by people. For that reason, confidence thresholds and escalation routes are designed in from day one, instead of being added after something has gone wrong.
"Where a model runs decides who can see the data and what it all costs. How the system is designed decides what can go wrong, and that is the part that gets my attention."The sentence I ask a board to remember
Product strategy for enterprise AI platforms across financial services, travel and retail, from discovery through to production pilots. Designed and ran the four-week pilot model: map the workflow, prototype with real users, measure before scaling. Led discovery and a four-week production pilot for a luxury travel company: an AI agent that cut quote preparation from over two hours to under two minutes, with 100% accuracy.
Go-to-market for a new AI-enabled B2B SaaS platform unifying compliance and sustainability reporting, reporting to the CPO. Redesigned the product organisation and shifted the company from a services-led delivery model to a product-led one, building the pricing model to support it.
Full P&L ownership of a new-to-market digital product; rebuilt the member app from a transactional booking utility into a content-led platform. Owned four functions end to end (Product, Design & UX, Data and Platform) and authored the six-year plan the business adopted as its roadmap (revenue $2M → $55M, members 29k → 328k).
Built product safety strategy across Bumble and Badoo, translating policy intent into a concrete product and team structure, and aligned executive sponsors across London and Barcelona.
Owned Incognito Mode, browsing-data controls and Chrome's privacy settings architecture: the privacy and permissions gate every other feature had to clear before shipping, working closely with Legal, PR and Product Marketing.
Built and led a team of four product managers across web, iOS and Android; led Babbel's shift from a single consumer app into a multi-product, B2B platform business.
Both were designed from the ground up and built with an AI coding tool: I wrote the requirements, and the tool wrote the code.
Questionnaire questions are matched against a library of approved answers, and a draft response is produced for each one, with its sources. Questions that need a specialist are routed to a named expert, and every answer is tracked through to approval and export.
Built on the principle that work should be asynchronous by default. It is designed to follow Slack and email conversations and suggest a meeting only when ambiguity, confusion and frustration reach a set level, then find a free slot and send the invitations.
The same method, sized for a young builder: from understanding how AI works to putting something real in front of real users, with privacy and safety built in from the first week. It draws on three years of teaching in Japanese schools as well as twenty years in product.
The best way to get in touch is by email. I'm always happy to talk through a specific workflow, a rollout that's stalled somewhere along the way, or a question a security team has been asked with no deadline attached to it.
I read and reply to email myself, usually within a day.