Custom-trained AI, inside your perimeter.Your knowledge, made answerable — without any of it leaving the building.
Every general-purpose AI tool you subscribe to knows everything about the world and nothing about your business. Training on your own material inverts that. Internally we just call it your business brain: it answers from what your organisation actually knows, shows you the document it took the answer from, and runs inside a boundary you control.
Four layers, one perimeter.
It is not a policy promise that nobody sends data outside. It is an architecture in which there is nowhere outside to send it.
Because the model and the corpus both sit inside your estate, that knowledge is structurally incapable of reaching a third-party provider. The distinction matters the first time someone asks you to evidence it.
Contracts, procedures, precedents, correspondence, historic decisions — indexed, permissioned and version-controlled. Curated and de-duplicated before anything is trained, because a corpus full of superseded documents produces a model that confidently cites the wrong version.
Every response traced to the source document and passage it came from. Chunking is tuned to the shape of your documents — clause boundaries in contracts, sections in policy, turns in correspondence — rather than a fixed token count.
Trained on how your organisation actually writes, decides and operates — including how it says no. Supervised fine-tuning on your corpus, then preference alignment so the answers match your house register rather than a generic assistant's.
Role-based access, an immutable audit trail, and human sign-off wherever a decision carries consequence. Below a defined confidence floor the answer goes to a named person before it is used.
How the training actually works.
Six disciplines, in the order they happen. Skipping any one of them is how a model that demos well fails in month four.
Curation & de-duplication
Superseded drafts, duplicate scans and documents nobody would stand behind are removed before training. The single biggest quality lever, and the least glamorous.
Corpus hygieneSupervised fine-tuning
Training on your material so the model knows your terminology, your document structures and your precedents rather than the general internet's version of them.
SFT on your corpusPreference alignment
DPO or RLHF so the model answers the way your organisation answers — including the refusals. A model that never declines is a liability in a regulated setting.
DPO / RLHFParameter-efficient adaptation
LoRA and QLoRA where a full fine-tune is not warranted; full-weight training on multi-GPU infrastructure where it is. The choice is made on evidence, not fashion.
LoRA · QLoRA · multi-GPURegression checking
Every run is checked for catastrophic forgetting against a frozen baseline before it goes near production. Gaining your domain must not cost general competence.
Held-out evaluationDeployment & retraining
Containerised into your estate, then retrained as your material changes. Input distributions shift; a model trained once is a model decaying quietly.
On-prem · tenancy · air-gappedThree places it can run.
Which one is right depends on how sensitive the data is and what your information governance will accept — and we will tell you when an air-gapped build is unnecessary.
- On-premise — in your own data centre, on hardware you own, under your physical control
- Private tenancy — your own isolated cloud environment, your billing, your access controls
- Fully air-gapped — no outbound network route whatsoever, for the material that genuinely warrants it
- UK data centres and UK-incorporated processors are the default in all three
- Non-UK hosting is named explicitly and consented in writing, never assumed
- Model weights and corpus never leave the boundary you have approved
Common questions.
Can a model really be trained on our own documents?
+
Yes. We curate and de-duplicate your material, fine-tune on it, then align the model to how your organisation actually answers — inside a boundary you control. Every answer cites the source document it came from.
What stops it inventing things?
+
Grounding and a confidence floor. Answers are retrieved from your corpus and traced to source, so if something is not in your material the model says so rather than producing something plausible. Below a defined confidence threshold the question goes to a named person instead.
What happens when our documents change?
+
The index updates continuously and the model is retrained on a schedule agreed with you. We measure input drift rather than waiting for someone to notice the answers have got worse — that is part of the Managed AI retainer.
How is this different from using ChatGPT with our files attached?
+
Three ways: the material never leaves your estate, the retrieval is engineered for your document structure rather than generic, and every answer is bound to its source so it can be audited afterwards. The third one is usually what a regulator asks about.
We can help.
Tell us what's hard, expensive, or taking too long — and we will help you identify, optimise and deploy AI for innovation, efficiency and growth.