On-Prem, Private Cloud, or a Box Under the Desk: What Running AI In-House Actually Costs
A cost-first look at the three ways to keep AI off public APIs — a physical appliance, a private cloud VPC, or prosumer hardware — and how to work out which one is genuinely cheapest for your real workload.
1. The question everyone asks is the wrong one
“Is on-prem cheaper than cloud?” is a bad question. It has no answer, any more than “is buying a van cheaper than hiring one?” does. It depends entirely on how many miles you drive.
In our companion piece we argued that a proprietary AI appliance is usually a poor deal for a small business — a Formula 1 car sold to a plumbing firm. That was about fit. This is about money: even when on-prem is the right shape, the sums only work under conditions most buyers never check.
The real question is total cost of ownership for your token volume at your utilisation. Everything else is a sales deck.
2. The three ways to keep AI off public APIs
If you want open-weight models running on your own data without pushing it through a public multi-tenant API, you have three broad routes:
- A physical appliance or GPU server — hardware you buy and rack on-site. Maximum control and containment, and every operational burden lands on you.
- A private cloud VPC — open-weight models deployed inside a locked-down instance on AWS, Azure, or GCP. You pay for the compute you actually consume, with no hardware to babysit.
- Prosumer local hardware — a high-spec workstation or a small cluster of consumer GPUs running Ollama or vLLM. Cheap to buy, no vendor lock-in, but you own the ops entirely.
All three keep data off public APIs. They fail on cost differently.
3. The cost drivers people forget
The sticker price of a box is the number everybody fixates on, and the least interesting. Here are the drivers that quietly decide whether owned hardware wins or bleeds you:
The drivers that break an on-prem business case
- Capex and depreciation. GPUs are a depreciating asset, not a fixed cost. Whatever you pay, spread it over the honest working life of the kit — three years is a common assumption — because next year’s hardware will do more per pound.
- Redundancy, or the cost of not having it. A single box has one power supply and one set of fans. If it dies, your AI capability goes dark until a part arrives. Real availability means a second node — roughly doubling the hardware line.
- MLOps and maintenance labour. Someone has to patch the stack, re-quantise models as new open-weight releases land, and keep the inference engine current — a salaried human, or a chunk of one, every month. The most under-counted number in every on-prem pitch.
- Utilisation. The big one, and it deserves its own section. A box you cannot keep busy is money idling; an owned GPU costs the same at 5% load or 95%.
4. Utilisation is the whole game
Cloud compute is rented by the second: when your workload is quiet, you stop paying. Owned hardware is the opposite — you pay the full capital cost up front and the meter never stops, busy or idle. So the only fair comparison is cost per useful token, not cost per hour.
Imagine — purely for the sake of an example, plug in your own figures — a workstation that costs a fixed amount per day once you spread its purchase over three years and add power and a share of someone’s time. Run it flat out serving a real queue of work and that fixed cost divides across a huge number of tokens: the per-token cost is tiny. Let the same box handle a few requests an hour and the identical cost divides across almost nothing — each token becomes absurdly expensive.
Same hardware, same invoice, wildly different economics — decided entirely by how full the pipe is. An idle owned GPU is the most expensive way to run AI ever invented.
5. The crossover logic
Put those drivers together and a simple shape emerges.
- Below some sustained token volume, cloud wins. You are not busy enough to justify a capital purchase, and pay-as-you-go means you never pay for idle silicon. Bursty, unpredictable, or still-experimental workloads live here.
- Above that volume, owned hardware can win — but only if utilisation stays high and someone maintains it. High, steady, predictable demand is what amortises capex into a low per-token cost.
Where the two lines cross depends on your volume, your utilisation, hardware prices at purchase, and the fully-loaded cost of the person keeping it alive. It differs for every organisation and moves every time prices or model efficiency shift.
| Cost driver | Physical appliance / GPU server | Private cloud VPC | Prosumer hardware |
|---|---|---|---|
| Upfront capex | High | None | Low |
| Depreciation risk | You carry it | None | You carry it |
| Redundancy | You buy and build it | Built into the platform | Largely absent |
| Maintenance labour | You own it | Shared with provider | You own it |
| Idle cost | Full cost, always | Near zero | Full cost, always |
| Best fit | High, steady, sovereign demand | Variable or unproven demand | Small teams, low ops appetite |
6. The honest caveats
A few things the crossover model tends to gloss over:
- Data residency can override the maths entirely. If a compliance perimeter genuinely prohibits data leaving your walls — as in regulated manufacturing, healthcare, or finance — on-prem may be the only lawful option, and cost becomes secondary. That is a real reason to own hardware; “it feels cheaper” usually is not.
- Every estimate is an estimate. Any ROI figure you have been quoted rests on assumed volumes, prices, and utilisation — and cloud rates move too. Treat it as a model to interrogate, not a fact; change the utilisation assumption and the whole case can invert.
None of this is a reason to freeze — just to do the arithmetic before signing a purchase order.
7. Summary: model it before you buy it
There is no universal answer, and anyone who gives you one is selling something. The decision turns on four of your own numbers: sustained token volume, realistic utilisation, current prices, and the fully-loaded cost of maintenance. Rules of thumb:
- Low or bursty volume, small team, no dedicated ops? Cloud VPC, almost always.
- High, steady, predictable volume with someone to maintain it? Owned hardware can genuinely win — if you keep it busy.
- Local control on a budget, and you can live with the ops? Prosumer hardware beats a premium appliance for most.
- A hard data-residency mandate? That may decide it regardless of cost.
The gap between “on-prem sounds cheaper” and “on-prem is cheaper for us” is exactly one spreadsheet wide. Put your real figures into a model, flex the utilisation assumption, and see where the lines cross before committing capital — because it is far cheaper to be wrong in a spreadsheet than wrong in a supply closet.
Ready to scale your operations securely?
We specialize in constructing non-invasive middleware and automating manual workflows without disrupting your core records. Let's trace your bottlenecks and outline a practical feasibility roadmap.