Shipping AI into a product that has to pass a security review.
An AI feature is easy to demo and hard to defend. The gap between the two is where most products fail their first external assessment. Here is the order of work that avoids it.
Almost every AI feature we are asked to review was built in the same order: get the model responding, wire it into the product, demo it, then think about security when a customer asks for a SOC 2 report or an enterprise buyer sends a questionnaire.
That order is the problem. By the time the questionnaire arrives, the decisions that determine whether you pass have already been made — where data crosses your boundary, what the provider retains, how retrieved context is scoped, and whether model output is ever trusted.
We built Hai-Sensei, an AI coaching platform that sits in on meetings, transcribes them and turns what was said into coaching insights. Meeting transcripts are about as sensitive as workplace data gets. It passed its third-party security review on the first attempt, and the penetration test came back with nothing. None of that was luck, and none of it was added at the end.
The model is not the attack surface
The most common misconception is that securing an AI product means securing the model. In practice the model is a third-party API call, and almost every finding a reviewer raises is about the system around it.
- What data leaves your boundary to reach the provider, and is it the minimum needed?
- What does the provider retain, for how long, and can you evidence that contractually?
- Can content retrieved for one tenant ever reach another tenant's prompt?
- Is model output ever used as input to something that acts — a query, a tool call, a rendered page?
- Can a user reconstruct training or reference data they should not see?
Those are the questions in a serious assessment. Note that none of them are about the model itself.
Work the data boundary first
Before any prompt is written, the question to answer is what actually has to leave. For a transcription product the honest answer is a lot — you cannot summarise a meeting without sending the meeting somewhere. That makes minimisation and retention the whole game.
Redact before you send, not after
Strip or tokenise what the model does not need to do its job. Names and identifiers can frequently be replaced with stable pseudonyms before the call and mapped back in your own boundary afterwards. The model summarises Participant A; your application renders the real name from a mapping that never left your database.
Pin the retention position in writing
A reviewer will ask what the provider does with your data. Answering from memory is a finding. Answering with the data-processing addendum, the retention setting you have configured, and a screenshot of that configuration is not. Make this an artefact you maintain, not something you reconstruct under pressure.
Log the metadata, not the payload
Observability on AI features drifts toward logging full prompts and completions, because that is what helps you debug. It also quietly creates a second copy of every sensitive document, in a log store with weaker access controls and longer retention than your primary database. Log token counts, latency, model version, a content hash and a correlation id. Put the payload behind an explicit, time-boxed, audited debug mode.
Treat model output as untrusted input
This is the single rule that prevents the largest class of AI vulnerabilities, and it is routinely broken because model output looks trustworthy.
If a model can be influenced by content it reads — and in a transcription or retrieval product it always can — then its output is attacker-influenced. Anything that consumes it needs the same treatment you would give a form field from the public internet.
- Rendering output as HTML without sanitising it is a cross-site scripting vector.
- Passing output into a database query, a shell, or a file path is an injection vector.
- Letting output select which tool to call, with which arguments, is privilege escalation unless the tool layer re-authorises against the user's own permissions.
- Returning output that contains a link the user is invited to click is a phishing vector.
The mitigation is structural, not a cleverer prompt. Constrain output to a schema and validate it. Authorise tool calls server-side against the session identity rather than against whatever the model asked for. Never let a prompt instruction be the only thing standing between a user and data they are not entitled to.
Prompt injection is an authorisation problem
Prompt injection gets discussed as if it were a content-filtering challenge. It is better understood as an authorisation one.
You will not reliably stop a model from following instructions embedded in content it reads. What you can do is ensure that following them achieves nothing: if the model has no capability the current user does not already have, a successful injection produces a rude summary rather than a breach.
That means the AI layer runs with the user's permissions, not with a service account that can read everything. It is the same principle as not running your web tier as root, and it is skipped about as often.
Retrieval is where tenant isolation breaks
Retrieval-augmented generation introduces a failure mode that does not exist in a plain API call: a query returns the nearest vectors, and nearest is not the same as permitted.
Metadata filtering is the usual answer and it fails open. One code path that builds a query without the tenant filter — a background job, a migration script, a new endpoint written in a hurry — and the system returns another customer's content, scored and formatted and delivered with total confidence.
- Partition at storage time: a namespace or collection per tenant, so a missing filter returns nothing rather than everything.
- Resolve tenant identity server-side from the session. Never accept it as a client-supplied parameter.
- Write the isolation test first, and make it part of CI rather than part of the pen test.
The order that actually passes
The reason Hai-Sensei cleared its review first time is not that the security work was especially exotic. It is that it happened in the right order. Concretely:
- Decide the data boundary before writing the first prompt. What leaves, what is redacted, what is retained.
- Build authorisation into the AI layer at the start, scoped to the user rather than a service account.
- Constrain and validate model output as a schema from the first version, not as a later hardening pass.
- Partition retrieval per tenant on day one, when there is one tenant and it costs nothing.
- Keep payloads out of logs from the first commit, because retrofitting that means purging history.
- Book the external test before you think you are ready. The findings are cheaper when there is less built on top of them.
Every one of those is close to free at the beginning of a project and expensive six months in. That asymmetry is the entire argument for doing security first rather than last, and it is why we treat OWASP practice as a starting condition on AI work rather than a review gate.
What reviewers actually ask for
If you are preparing for a first external assessment on an AI product, the artefacts that shorten it most are mundane:
- A data-flow diagram showing exactly where data crosses your boundary and to whom.
- The provider's data-processing terms, plus evidence of your retention configuration.
- Your threat model for prompt injection, with the authorisation controls that make it survivable.
- Evidence of tenant isolation testing in CI, not just an assertion that it is isolated.
- An access-control matrix for the AI features specifically, not inherited from the rest of the app.
If you are building one
AI features are easy to demo and hard to defend, and the distance between those two states is mostly decisions made in the first fortnight. If you are putting AI into a product that will eventually face an enterprise buyer, a regulator or a penetration test, talk to us before you build it rather than after. We have shipped AI into healthcare, coaching and triage products and taken them through external review.
Frequently asked questions
What is different about securing an AI product?
The model is not the risk. The risk is everything around it: what data leaves your boundary to reach the provider, how long that provider retains it, whether retrieved context can cross tenants, and whether model output is ever treated as trusted input by something downstream.
Does using OpenAI mean our data trains their models?
On the commercial API, business data submitted through the API is not used to train the models by default, and zero-retention arrangements are available for qualifying workloads. What matters for a review is that you can show the contractual position in writing and that your configuration actually matches it.
What is prompt injection and why do reviewers ask about it?
Prompt injection is when untrusted content the model reads — a transcript, a document, a web page — contains instructions the model then follows. Reviewers ask because it turns any content ingestion path into a potential privilege-escalation path if model output can trigger actions.
How do you keep tenants isolated in a vector database?
Filter at query time and partition at storage time. Metadata filtering alone fails open if a query is ever constructed without the filter, so the safer pattern is a separate namespace or collection per tenant with the tenant identity resolved server-side from the session, never passed in by the client.
How long does a third-party security review usually take?
Expect two to six weeks for the assessment itself, plus remediation and retest. The variable is not the testing; it is how much has to be rebuilt once the findings come in. Products that were designed against OWASP guidance from the first commit tend to pass without a structural change.