AI has changed the economics of producing delivery artefacts. A status report, a business case draft, a tender response or a page of website copy that took a day now takes minutes. Most of the attention has gone to that speed while most of the risk sits somewhere else entirely: an ungoverned AI pipeline ships errors at production speed.
The failure modes are specific and they are not hypothetical. A generated claim reads plausibly but cannot be traced to a source and an evaluator or auditor falsifies it later. A figure that sits under a non-disclosure obligation leaks into a public document in a slightly reworded form. A decision the organisation settled months ago gets quietly reversed, because the tool had no way to know it was settled. None of these are tooling problems and none of them are fixed by a better prompt. They are governance problems and project delivery already owns the disciplines that fix them.
The pattern below treats AI output the way a well-run program treats any other delivery product: baseline what is true before work starts, gate the output against a defined exit standard, protect what must not leak, keep a decision record that prevents regression and hold a human consent step in front of anything irreversible.
Establish the source of truth before anything generates
Generation without a baseline produces confident drift. The single highest-value artefact in an AI-assisted pipeline is a verified evidence base that sits upstream of every output: a register of what the organisation can actually claim, with each claim traced to a source document.
For delivery organisations this usually means a claims or evidence register covering outcomes, dates, figures and attributions, verified against primary sources before anything propagates. Where a claim has been contested and resolved, record the settled wording at its exact strength and treat that wording as fixed. “Delivered to the re-baselined schedule the executive approved” and “delivered ahead of schedule” may look interchangeable, but only one of them survives a reference check and an AI tool cannot know which unless the register tells it.
This is requirements baselining applied to content. The test is the same one you would apply to a requirements document: if two people generated the same artefact independently from the register, would the claims match?
Gate the output, not the tool
Attempting to control AI by restricting which tools people use misses where the risk actually lives. The risk is in what ships, so the control belongs at the point of release.
A release gate for AI output works like any quality gate: a defined exit standard, checked the same way every time, with nothing shipping on a read-through alone. In practice the standard has three layers:
Banned content, zero tolerance. Claims that were corrected and must never reappear, wordings stronger than the evidence supports and terminology the organisation has ruled out. These are checked mechanically, because a human reviewer skims and a mechanical check does not.
Fragile content, justified per instance. Phrases that are sometimes legitimate and sometimes overclaim - “on time”, “without incident”, “no issues raised” - where every surviving instance needs a recorded justification against its source.
Structural integrity. Whatever the medium’s equivalent of a build is: links resolve, references exist, formats validate. An artefact that fails mechanically was never reviewed at all.
The gate runs before every release, not once. AI regenerates content freely, which means an error corrected last week can reappear this week unless the check that caught it runs again.
Protect what must not leak
Confidentiality failures in AI output rarely look like a pasted secret. They look like a paraphrase: a contract figure restated to the nearest thousand, a client identified by an unmistakable description rather than a name, a phrase lifted from the client’s own marketing that any search engine resolves in one query.
Two controls close this. The first is an approved-forms rule: for every sensitive fact, define the form in which it may appear publicly - a percentage rather than a contract value, a sector descriptor rather than an organisation name - and permit nothing outside those forms. The second is an independent sweep, run after the main work is finished and by a different process than the one that produced it, across everything that ships: body copy, metadata, image descriptions, code comments and the built output as well as the source. The separation matters for the same reason evaluation panels separate scoring from moderation - the process that wrote the content is the process least likely to notice what it exposed.
Keep a decision record that prevents regression
The most distinctive failure mode of AI-assisted work is confident reversal of settled decisions. A human team member remembers why the anchor names were frozen or why a claim was softened. A generation tool starts fresh every session and will helpfully “improve” its way straight back into a resolved problem.
The fix is a decision record written for a reader with no memory: what was decided, why it was decided and an explicit instruction that the decision wins over any future suggestion to the contrary. Where a past incident produced a rule, record the incident with the rule, because a rule without its reason gets argued away. This is change control and decision logging applied to instructions rather than scope. It is the difference between an AI pipeline that compounds quality and one that oscillates.
Hold the consent gate in front of anything irreversible
Everything above can and should run without a human in the loop. The step that must not is release. Publishing, sending, deleting and deploying are held behind an explicit human decision, taken with the full change in front of the approver, after the gates have passed - not instead of them.
The sequencing matters. A human approving ungated output is reviewing too much to review anything properly. A human approving gated output is making one decision: the checks passed, the diff is what I expect, ship it. That is a decision a senior reviewer can make well in thirty seconds, which is why the pattern holds up at production speed.
What this looks like assembled
Run end to end, the pipeline is short: verify the evidence base, generate against it, gate the output mechanically, sweep independently for what must not leak, obtain human consent on the exact change, release and confirm the release landed. Each step maps to a discipline delivery managers already practise - baselining, quality gates, probity separation, change control and delegated approval - which is the point. Organisations do not need a new governance framework for AI. They need their existing delivery governance pointed at a new class of output, run at the same standard they would demand of any supplier producing work at this volume.
The measure of success is the same as for any governed delivery: what shipped is traceable, what was confidential stayed confidential, what was settled stayed settled and the whole run is repeatable by someone who was not there.
Vantage Meridian designs and runs governed delivery - the artefact standards, gates and decision records that make fast-moving work defensible, including work produced with AI. See PMO stand-up and uplift.
If you are dealing with a situation like the one described in this article, send an enquiry. An initial conversation costs nothing.