AI on your sensitive data
The documents, records and code you could not send to a cloud provider, in reach of an assistant your people use.
A model that stays put
A version you pin and control, with an audit trail of which version produced which output.
A cost you can forecast
A machine you own and running costs you know in advance. The bill no longer moves with every token.
Evidence you can show
Access controls, logs and evaluation results, ready for the regulator, auditor or client who asks.
On-Premise AI
Put AI to work on data you cannot send to the cloud. Private LLMs on hardware you own, starting with whether you need one at all.
Your best AI use cases sit on your most sensitive data

Client contracts, patient records, source code, deal documents. The work AI would help with most is often the work you cannot hand to a cloud provider.
So it stalls. The pilot runs on dummy data, the policy says no, or nobody can say for certain what is allowed.
Private deployment runs the model on a machine you own, inside your own network. The prompts, the documents and the answers stay there.
What you end up with
Is private AI right for you?
Keep AI in-house for one of five reasons.
If none applies, cloud is usually the better answer.
A contract says the data cannot leave
The strongest reason, and the easiest to verify.
A client contract, a security classification or a supply chain requirement. If a document says your data stays inside your estate, the decision is made. The only question left is how.
We will ask you which document says so.
Data protection law
UK GDPR and the Data Protection Act 2018.
Lawful basis, data minimisation, international transfers, and an impact assessment that has to stand up if anyone asks. Personal data in a prompt is still personal data, processed by a processor you may not have assessed.
UK-region cloud with no training on your data and no retention answers much of this by contract, without hardware.
Your sector regulator
Where data sits, how decisions are evidenced, what an auditor will ask.
Financial services, health and care, policing, defence and life sciences each have their own expectations. They rarely say "no cloud". What they ask for is easier to satisfy when you control the environment.
Intellectual property you cannot risk
Designs, source code, formulations, research, deal documents, case strategy.
None of that is personal data, so data protection law is not the constraint. The constraint is that it is the business.
Enterprise terms say the provider will not train on your data, and that is worth having. It does not change where the data sits, who at the provider could reach it, or who could be compelled to produce it. Some boards accept that. Some do not, and that is a legitimate commercial judgement.
For a trade secret, UK protection depends partly on the holder taking reasonable steps to keep it secret, and where it is processed forms part of that picture. It is a question for your own counsel. We raise it; we do not advise on it.
Model version control and governance
The least understood reason, and often the most important.
A provider can change the model behind an API, which quietly breaks your change control. Where a model informs a regulated decision, or sits in a validated environment, you may need a version you pin and control.
Private deployment gives you that, with an audit trail of which version produced which output.
When cloud is still the better answer
- No contract restricts the data. Often "our data cannot go to the cloud" is an assumption that no contract makes. UK-region cloud with no training on your data and no retention solves much of it.
- Your volumes are modest. Under about two million tokens a day, cloud is almost certainly cheaper.
- Your work needs the hardest reasoning. Open-weight models handle extraction, classification, summarisation and document search well. The frontier cloud models still lead on the hardest reasoning tasks.
- Your regulator allows cloud. Neither the NHS nor the FCA requires AI to run on your own hardware.
What you give up
Open-weight models still trail the frontier models on the hardest reasoning tasks. You also take on model updates, security patching, monitoring and capacity, and someone has to own them. Self-hosting does not remove vendor dependency. It exchanges it for hardware dependency.
Is running AI on your own hardware cheaper?
Usually not. Cloud inference has got much cheaper.
Frontier model prices have fallen by more than 80 per cent since early 2023. The median flagship model now costs around £5 per million blended tokens at current exchange rates.
What has risen is consumption. Agents, document search and reasoning models that bill for their own thinking use far more tokens per task. So bills rise while unit prices fall, and they rise unpredictably.
The real cost case is forecastability. On the right workload, a variable, consumption-driven bill becomes a fixed asset with a known cost.
- Under 2Mtokens a dayCloud is almost certainly cheaper. Deploy privately only if compliance requires it.
- Around 5Mtokens a day, sustainedA workstation-class machine reaches parity in roughly 18 to 24 months. At 12 months, cloud is still cheaper.
- Around 50Mtokens a day, sustainedEnterprise hardware pays back in 18 months to three years at genuine 24/7 use. At 12 months, cloud is still cheaper.
Third-party 2026 total cost of ownership analysis, adjusted for UK electricity, with 36-month straight-line depreciation. Your figures will differ, which is what the assessment establishes. Figures checked September 2026.
What it costs to run in the UK
UK businesses pay between 22p and 30p per kWh on standard fixed contracts. A desktop appliance costs a few hundred pounds a year to run.
| Deployment | Power | Per year |
|---|---|---|
| Desktop appliance, working hours | 240W, 10 hours a day | around £170 |
| Desktop appliance, always on | 240W, 24/7 | around £600 |
| Enterprise GPU node with cooling | around 9.6kW, 24/7 | around £20,400 |
Why American break-even figures understate UK costs
Most break-even analyses you will read are American, and American electricity is about a third of the UK price. An enterprise node costs roughly £13,000 a year more to run here, about £39,000 across a three-year life. Add that back to any American figure before you believe it.
What does it run on?
A desk, a power socket and a network port.
You probably do not need a server room. A class of desktop machine now runs substantial models locally, with no rack, no data centre and no facilities work. You can prove whether private AI works for a few thousand pounds of hardware you own outright.
| Machine | UK price, ex VAT | Memory | Bandwidth | Speed |
|---|---|---|---|---|
| NVIDIA DGX Spark and other GB10 machines | £3,750 to £4,250 | 128GB | 273 GB/s | 43 to 59 tokens/sec on a 120B model |
| Mac Studio, M5 Max | about £4,250 (128GB) | up to 128GB | 614 GB/s | not yet benchmarked |
| Mac Studio, M5 Ultra | from about £4,600 | 96 to 512GB | 1.2 TB/s | not yet benchmarked |
| AMD Ryzen AI Max+ 395 mini PCs | £2,350 to £2,500 | up to 128GB | around 256 GB/s | lower, and thermally dependent |
UK retail prices ex VAT; workstation-brand AMD machines cost more. Speed is single-user decode and falls as context grows. The 512GB Mac Studio ships from late October. Prices and figures move quickly. Checked September 2026.
Concurrency sets the limit
These machines hold the model in a fixed pool of memory with nothing to spill into. Every simultaneous user takes a share, so capacity runs out on concurrent users long before model size matters. Memory bandwidth sets the speed. One machine serves a team comfortably. It will not serve an enterprise. Working out which you are is the first thing the assessment does.
What does your regulator allow?
Neither the NHS nor the FCA requires AI to run on your own hardware. Private deployment earns its place for specific, narrower reasons.
Does the NHS allow AI in the cloud?
Yes. NHS England policy makes public cloud the default.
Public cloud is to be considered before any other option. Data must be hosted in the UK or a territory the UK Government considers adequate. Hosting anywhere else needs an international data transfer agreement, a transfer risk assessment, and SIRO and executive approval. Your SIRO, DPO and Caldicott Guardian have to be satisfied either way, and the level of control scales with the data classification.
If a supplier tells you the NHS requires on-premise AI, they are wrong. Your IG lead will know it.
When on-premise is still the right answer
Clinical safety
The strongest argument in health, and the least often made. DCB0129 and DCB0160 assess the clinical risk of the system as deployed. A model that changes underneath you invalidates the hazard analysis you signed off. A pinned version, with an audit trail, keeps the safety case intact.
Data gravity
Imaging and genomics generate volumes that make routine egress impractical, whatever the policy.
Connectivity
Community teams, ambulances and some estates cannot rely on a connection at the point of need.
A data sharing agreement
Research and third-party data often carry terms stricter than national policy. That is contract, and it is the first thing to check.
The standards are moving
NHS England ran a national consultation on DCB0129 and DCB0160 during 2026, with artificial intelligence explicitly in scope. No final changes have been published. If you are designing clinical AI now, expect the bar to move and build the evidence trail to match.
A practical point about trusts
Most NHS organisations are trying to reduce the infrastructure they run. A proposal that starts with a server room starts with an argument. A single machine on a desk is a very different conversation, and a very different business case.
Do FCA rules require AI to run on our own infrastructure?
No. The FCA and PRA are technology-neutral.
Both take a principles-based approach, and the FCA has said it does not currently plan to introduce AI-specific rules. Cloud is used extensively across financial services, and nothing prohibits it. What applies is the regime you already operate under.
What actually applies
Model risk management
The PRA's SS1/23 sets model risk management principles for banks: inventory, validation, monitoring and change control. Firms have told the regulator that validation is hard to scale to generative and agentic AI. It gets harder when the model can change without your knowledge.
Senior manager accountability
Under SMCR a named individual is accountable. The FCA has asked directly how that works when AI performs functions that were previously under human oversight. That person needs to evidence what the system did and why.
Operational resilience
The UK Critical Third Parties regime brings major providers into scope, and Parliament's Treasury Committee has recommended designating major AI and cloud providers. Concentration, substitutability and exit planning are live supervisory questions.
Consumer Duty
Where AI touches a customer outcome, the outcome is what gets tested.
Where private deployment earns its place
Model version control
If a model informs a credit, pricing, suitability or financial crime decision, the version behind the output is part of the evidence. A provider changing it breaks your validation and your change control. A pinned version with an audit trail makes model risk management tractable.
Confidential material
Live deal documents, client mandates, positions, unpublished research. No regulation points at it, and no regulator will decide it for you. Plenty of firms decide it does not belong on a third party's infrastructure.
Sometimes it is a judgement call
Not every reason is written down. Some firms are content that the contract covers them and still keep certain material off third-party infrastructure. That is a legitimate risk appetite position, and it is worth saying out loud.
We make sure it is a decision. If a constraint sits underneath the discomfort, the assessment finds it and you have a clear case. If not, you will know that too, and can take the finding to your risk committee.
Concentration is now a supervisory question
The Critical Third Parties regime exists because regulators are concerned about the financial system depending on a few providers. That is no reason to leave the cloud. It is a reason to have an answer if a provider is unavailable, changes terms or is designated. Some capability that does not depend on one provider is one of the available answers.
Policing, justice and defence
These carry their own constraints, including material classified above what a commercial cloud contract covers, and vetting for the people who handle it. Those cases are usually clearer than health.
How it works, and what it costs
Three steps, each bought on its own.
Assessment
Find out whether private deployment is right for your data, and what it would cost.
- A workload and token volume profile, from your actual usage
- A concurrency figure: how many people need an answer at the same time
- A data classification: what genuinely cannot leave
- A model recommendation, benchmarked on your own task
- A three-year cost comparison against staying in the cloud
The recommendation may be to stay in the cloud. The fee is the same whichever way it goes.
From £4,000from 4 effort days · timescale scoped to your estatePilot
Prove it on one real workload, on a machine you own.
- Deployment onto your hardware, integrated with the systems your people use
- Access controls and logging
- An evaluation harness, so you can prove it still performs after a change
- The assurance and audit evidence a regulator, auditor or client will ask for
Typically an appliance on your own premises. You keep the hardware whatever you decide next.
From £10,000from 10 effort days · timescale scoped to your workloadRun and Hold
Keep it current, secure and evidenced.
- Model updates and version pinning
- Security patching
- Performance and drift monitoring
- Periodic re-evaluation
- The audit trail behind all of it
Engagements are fixed-price against a stated effort estimate. Effort is what we commit to deliver in, not a day count we bill against. All prices exclude VAT, charged at the prevailing rate.
What we need from you A named owner with time to commit, any existing usage data, and twenty to fifty real examples of the task, with a view on what a good answer looks like.
Then the application Our Ignite Studio practice builds what your people use, against your private model. Proof of concept in two to three weeks, priced once we know what we are building.
Hardware Typically £2,000 to £7,000 for a pilot, excluding VAT, confirmed with a live UK reseller quote. You buy it directly or through a supplier you choose, and we take no margin on it.
Running costs Stated separately from our fee: power, cooling, hosting, support contracts and any commercial model licences.
Public sector Priced from our published SFIA rate card on the Digital Marketplace.
Several workloads, more than one site or a larger estate? Contact us to discuss scope and price.
What we do not do
Data centre design
We do not design data centre facilities, rack-scale GPU clusters, power or cooling. Desktop and small-server deployment we do ourselves.
Legal advice or certification
We do not give legal advice, and we do not certify.
Clinical safety sign-off
DCB0129 and DCB0160 need a qualified clinical safety officer. We produce the hazard analysis, evidence and documentation. A clinical safety officer signs it.
Regulatory interpretation
We build the technical controls and the evidence your risk, compliance and model validation teams need. Interpretation belongs with your compliance function or counsel.
Common questions
How do we run AI without sending data to the cloud?
You run the model on hardware you own, inside your own network. A private LLM deployment keeps the model weights, the prompts and the outputs on your infrastructure. That is a desktop appliance for a team or server hardware for an organisation, and the assessment establishes which. It can also be hybrid, with sensitive workloads running privately and everything else on cloud models.
Can we use AI if our data cannot leave our network?
Yes. That is exactly what this is for. The first question is whether the restriction is as broad as people assume, because often only a subset of data is genuinely restricted. The assessment classifies it, and the design usually splits workloads.
Is on-premise AI cheaper than the cloud?
Usually not. Frontier model prices have fallen by more than 80 per cent since early 2023, so cloud inference is much cheaper than most people assume. Private deployment tends to win on contractual and regulatory grounds, or at sustained volumes above roughly ten million tokens a day. The assessment gives you the figures for your own workload.
Are open models good enough?
For most business workloads, yes. Extraction, classification, summarisation and search across your own documents are all well within reach. The frontier models still lead on the hardest reasoning tasks.
We are worried about our data going to a cloud AI provider. Is on-premise the only answer?
No, and it is worth checking before you spend anything. UK-region cloud deployment with contractual guarantees on training and retention resolves a great many data concerns. The assessment establishes whether your situation genuinely requires local inference, or whether a contractual answer is available.
Does the FCA require AI to be hosted on our own systems?
No. The FCA and PRA are technology-neutral, and the FCA has said it does not currently plan AI-specific rules. Private deployment usually earns its place through model risk management. If a model informs a credit, pricing, suitability or financial crime decision, the version behind the output is part of the evidence. A provider changing it underneath you breaks the validation you carried out.
We are an NHS organisation. Do we have to keep AI on our own infrastructure?
No. NHS England policy makes public cloud the default for patient data, provided hosting is in the UK or an adequate territory and your SIRO, DPO and Caldicott Guardian are satisfied. On-premise usually earns its place in health through clinical safety. DCB0129 and DCB0160 assess the risk of the system as deployed, and a model that changes underneath you undermines the hazard analysis you signed off.
Our concern is our own IP, not personal data. Does that change anything?
It changes the argument. Data protection law is not your constraint, so no regulation decides it for you. The question is commercial: is the residual risk of your designs, code, research or deal documents sitting on a third party's infrastructure acceptable to your board? No-training terms address part of it. They do not address where the data sits or who could be compelled to produce it. For trade secrets, ask your lawyers whether processing arrangements form part of the reasonable steps you are expected to take to keep them secret.
What hardware do we need to run AI in house?
Less than most people expect. A pilot usually runs on a single desktop appliance costing between £2,000 and £7,000, such as an NVIDIA DGX Spark, a Mac Studio or an AMD Ryzen AI Max machine. No rack, no server room. What decides the choice is how many people need an answer at the same time, and how fast the machine's memory is. The assessment produces that number and the specification that follows. We do not sell hardware and take no margin on it.
How many people can one of these machines serve?
A team. These machines hold the model in a fixed pool of memory, and each simultaneous user takes a share of it. A handful of concurrent users is comfortable. Hundreds is a different architecture. The assessment works out which you need before you buy anything.
Can this work alongside our existing cloud AI?
Usually that is the right design. Sensitive workloads run privately, and everything else uses cloud models where they are cheaper and stronger. Splitting by data classification is almost always the better answer.
Who governs the models once they are running?
That is the part most deployments miss. Model versions, change control, evaluation evidence and audit logging all need an owner and a routine. It pairs with our AI Governance work.
On-premise AI in your sector
- Healthcare patient data, DSPT and clinical system integration
- Financial services model risk and operational resilience
- Policing and justice classified and sensitive material
- Local government resident data and public scrutiny
- Enterprise client contracts that restrict data processing


Find out whether your AI should stay in-house
Book a free discovery call. We will tell you whether private deployment fits your data, and what the assessment would involve.
Book Your Free Discovery CallNot ready for a call? Take the free AI readiness checklist. Ten questions, five minutes, scored instantly.
