Skip to content

Where your documents actually go

“Your data stays yours” is on every vendor site including, apparently, this one. It means nothing without saying which parts move and where they stop. So here is the whole picture, layer by layer.

Your source documents

Stay where they already live — SharePoint, your file server, your DMS.We read them to index. We do not become the system of record, and if the project ends you are not holding a migration problem.

The search index

Hosted wherever you decide, including entirely inside your own tenancy.This contains passages of your documents in both text and vector form. Treat it with the same sensitivity as the originals, because effectively it is a copy.

The embedding step

Either a hosted API or a model running on your own infrastructure.This is a real decision. Hosted is cheaper and usually better quality; self-hosted means passage text never leaves your environment. We will give you the honest trade-off rather than a default.

The answering model

Commercial API, or an open-weight model you run yourself.Whichever you pick, it sees the retrieved passages and the question. That is the point at which content leaves, if it leaves at all.

Logs and evaluation data

Your infrastructure.Question logs are quietly sensitive — what staff ask reveals a lot. People forget to think about these and they are often the least protected part of the system.

Training on your content

The question everyone asks, usually phrased as “will our data be used to train the model”. Two things are worth separating.

First: in a retrieval system, nothing is trained on your content by design. Your documents are retrieved and passed in at question time. The model is not modified.

Second, and separately: whether the model provider retains what passes through their API and uses it for their own training. That is governed by which provider and which tier of their service you are on, not by anything in the architecture. The enterprise tiers of the major providers contractually do not train on API traffic. The consumer tiers of some products do.

We will tell you exactly which provider and which tier your build uses, in writing, before you sign anything. If that answer is ever vague from any vendor, push until it is not.

Australian hosting

The index, the application and the logs can all run in Australian regions. That covers most of what people mean when they ask about data sovereignty, because it is where the bulk of your content sits at rest.

The answering model is the part that may not be, depending on provider and region availability. If everything must remain onshore, that constrains which models you can use and usually costs some answer quality. It is a legitimate requirement and we would rather design for it up front than discover it at go-live.

What we do not have

Yes AI holds no SOC 2 report and no ISO 27001 certification. We are a small Australian firm and we have not been through those audits.

If your procurement process requires either, we are not the right supplier for that piece of work and we will say so on the first call rather than three weeks into a questionnaire. Some clients build with us and host inside their own already-certified environment, which is a sensible way through it.

We would rather lose the job than imply a certification we do not hold.

What a build looks like