We build what we recommend.A permanent core, extended by a vetted network.
A permanent core of UK-based LLM engineers, extended by a vetted network of specialists brought in by discipline — training infrastructure, retrieval, evaluation, security. You get the engineer who has done the specific thing before, not whoever was free that week. Everyone who works on your system writes code on it; we do not run an account layer between you and the people building the thing.
Six disciplines.
Between them the team has shipped production AI systems across healthcare, retail, fintech, media, SaaS and productivity.
Model training that survives contact with your data
Curation and de-duplication before anything is trained, supervised fine-tuning on your corpus, then preference alignment so the model answers the way your organisation actually answers. Parameter-efficient adaptation where a full fine-tune is not warranted; full-weight training on multi-GPU infrastructure where it is. Every run is checked for catastrophic forgetting against a frozen baseline before it goes anywhere near production.
- Training-data curation & de-duplication
- Supervised fine-tuning
- Preference alignment · DPO / RLHF
- LoRA & QLoRA adaptation
- Multi-GPU training infrastructure
- Catastrophic-forgetting checks
Retrieval that cites, structurally
Chunking tuned to the shape of your documents rather than a fixed token count — clause boundaries in contracts, sections in policy, turns in correspondence. Hybrid retrieval combines sparse BM25 with dense embeddings, then a cross-encoder reranks before anything reaches the model. Every passage that informs an answer stays bound to its source, so the citation is a property of the pipeline rather than something the model was asked to produce.
- Structure-aware chunking
- Hybrid sparse + dense retrieval
- Cross-encoder reranking
- Source binding & citation integrity
- Precision / recall @ k
- Permission-aware indexes
Agents you can replay
Multi-step agents built as explicit state machines rather than free-running loops, with typed tool schemas, bounded retries and hard budget ceilings. Every run writes a full trace — inputs, tool calls, intermediate state, the decision point and who approved it. An agent that cannot be replayed cannot be audited, and an agent that cannot be audited will not survive a complaint.
- Typed tool schemas
- Explicit state machines
- Bounded retries & budgets
- Memory & context assembly
- Deterministic replay logs
Held-out sets built from your real inputs
Evaluation is the deliverable most AI work quietly skips. We build a held-out set from your actual traffic and score groundedness, citation accuracy, refusal behaviour and out-of-corpus handling. Regression suites run on every model, prompt or index change, so a quality move is caught in CI rather than by a customer four months later.
- Groundedness & faithfulness scoring
- Citation accuracy
- Refusal & out-of-corpus behaviour
- Regression suites in CI
- Human review sampling
Deployed into your estate, not ours
Containerised builds that run on-premise, in your own private tenancy, or fully air-gapped with no outbound route. Quantisation and batching tuned to a stated latency budget rather than left to default, and the whole thing reproducible from source.
- Docker · reproducible builds
- GPU serving & quantisation
- Cloudflare Workers · AWS · GCP · Fly.io
- PostgreSQL · Supabase · D1 · R2
- Latency budgets & autoscaling
A model is not a product
Most of what makes an AI system usable is ordinary engineering done well — typed interfaces, real tests, sensible state management and an interface someone will actually use. TypeScript and Python throughout, test coverage before handover, source transferred on completion.
- Next.js · React · Astro · Preact
- Node.js & Python
- REST · WebSockets · queues
- Stripe · Auth.js · OAuth
- Test coverage before handover
Selected production builds.
HIPAA-compliant transcription
Regulated audio through speaker diarisation and multi-format export, behind scoped OAuth.
Deepgram · Fly.io · Supabase · Next.jsIn-store AI avatar
Camera detection driving a real-time greeting on a shop-floor screen over a WebSocket stream.
HeyGen · Python · Docker · ReactPayment platform rebuild
More than 500 tests, which surfaced around twenty billing defects in discount-stacking logic that had been live.
React · Supabase · Stripe Connect · VitestSecure cloud storage MVP
Auth, nested folders, in-browser file viewing and usage-based billing — shipped in seven days.
Next.js · GCP · Stripe · ShadcnThree-stage image pipeline
Segment, inpaint, composite — a chained pipeline rather than a single prompt and hope.
Stable Diffusion · Next.jsPresentation SaaS prototype
CSV-driven deck population with native .pptx export.
Next.js · ReactWho actually does the work.
A permanent core plus a vetted network, engaged by discipline. It is the model that lets us hold genuine depth in six specialisms rather than shallow coverage of twenty.
- A permanent UK-based core who hold the architecture and the client relationship
- Specialists engaged per discipline — training infrastructure, retrieval, evaluation, security
- Everyone who works on your system writes code on it. No account layer.
- A named accountable owner per workflow, documented and reportable
- Senior content and audience leadership inside Google and Adobe; delivery for Microsoft
- Growth and monetisation across global publishers, B2B and B2C
- Production systems shipped across healthcare, retail, fintech and media
- Where we do not have the depth, we say so and decline
Common questions.
How big is the team?
+
A permanent core, extended by a vetted network engaged per discipline — which is why you get someone who has done your specific problem before rather than whoever was free.
Who will actually build our system?
+
Named engineers, introduced during Discovery, who write the code themselves. There is no account layer between you and the people building the thing, and the named accountable owner for each workflow is documented.
What if the work needs a specialism you don't have?
+
We tell you, and either bring in a specialist from the network or decline. Declining is the more common outcome and it is the reason the work we do take on holds up.
We can help.
Tell us what's hard, expensive, or taking too long — and we will help you identify, optimise and deploy AI for innovation, efficiency and growth.