A private AI data center is a room, a power chain, a cooling plant and a network that together let an organisation run large models on hardware it owns, with no third-party AI provider in the request path. Most conversations about one start with GPUs. The ones that go well start with the electrical supply, because a single 8-GPU server now draws more power than the typical enterprise rack was built to deliver. Everything below comes from the vendor documentation and industry survey data we read on 30 September 2026, and it is the checklist we work through before a client orders any compute.
What a private AI data center means in practice
The phrase covers three different arrangements, and the right one depends less on budget than on where the data is allowed to go and who is going to be woken at three in the morning. All three keep the model and the prompts off shared public AI services. They differ in who owns the building and who holds the keys.
| Arrangement | Who owns the facility | What you still own | Fits when |
|---|---|---|---|
| Your own room | You | Everything: power, cooling, security, hardware, operations | Data may not leave your premises, or you already run a facility |
| Colocation cage | A colocation provider | The hardware, the network inside the cage, operations | You want owned hardware without building a room |
| Dedicated hosted cluster | A provider | The model, the data and the application | Single-tenant hardware is enough and speed matters more than ownership |
Owned facilities are not a fringe choice. The Uptime Institute Global Data Center Survey 2025 reports that on-premises data centers “remain foundational, with 45% of IT workloads still residing in corporate facilities”. What is new is what AI hardware asks of those facilities.
Power is the first constraint, not the GPU count
NVIDIA publishes the power figures for its integrated systems, and they are the right place to start. The DGX H100 and H200 user guide lists system power usage of “10.2 kW max” for one 8-GPU system. The DGX B200 user guide lists “14.3 kW max”. Both use six 3.3 kW power supplies.
| System | GPUs and total GPU memory | Max power | Rack space | Max weight | Operating temperature |
|---|---|---|---|---|---|
| DGX H100 | 8 x H100, 640 GB | 10.2 kW | 8U | 130.45 kg | 5°C to 30°C |
| DGX H200 | 8 x H200, 1,128 GB | 10.2 kW | 8U | 130.45 kg | 5°C to 30°C |
| DGX B200 | 8 x B200, 1,440 GB | 14.3 kW | 10U | 142.4 kg | 10°C to 35°C |
Now set those numbers against what existing rooms were built for. The Uptime survey reports that “the average of the modal densities in our total sample reaches almost 9 kW”, and that “More than 80% of the operators responding to our survey say their facility has no racks above 30 kW”. A single DGX H100 at 10.2 kW is already above the typical rack. Four DGX B200 systems, which is 32 GPUs, add up to 57.2 kW at maximum draw, which is a density most of the surveyed facilities have never delivered to one rack.
- Ask for the supply, not the space. The first question to a facilities team or a colocation provider is how many kilowatts can be delivered to one rack, on how many independent feeds.
- Plan for maximum, not typical. The published figures are maximums. Training runs sit close to them for days at a time.
- Count the redundancy. The DGX H100 power supplies are configured for 4+2 redundancy and the DGX B200 for 5+1. That only helps if the room gives them two independent feeds.
- Budget the rest of the chain. A UPS sized for the load, automatic transfer switching and a generator are what keep a training job alive through a mains disruption.
Size the hardware from the model, not the other way round
The second question is what you intend to run, because that decides how much GPU memory you need, and GPU memory decides the number of systems. A useful first pass is arithmetic: a model stored at 16-bit precision needs two bytes per parameter just for its weights, and one stored at 8-bit needs one. That is before the working memory each concurrent request needs, so treat the figures below as a floor rather than a target.
| Model size | Weights at 16-bit | Weights at 8-bit | Fits in one 640 GB system | Fits in one 1,440 GB system |
|---|---|---|---|---|
| 8 billion parameters | about 16 GB | about 8 GB | Yes, on a single GPU | Yes |
| 70 billion parameters | about 140 GB | about 70 GB | Yes, with room for concurrent requests | Yes |
| 405 billion parameters | about 810 GB | about 405 GB | Only at 8-bit | Yes, at 16-bit |
Two practical conclusions follow. First, most organisations do not need a cluster to serve a model; a 70 billion parameter model fits comfortably in one 8-GPU system, and a smaller one fits on far less. Second, the case for several systems joined by a high-speed fabric is training and fine-tuning, or serving many large models at once. NVIDIA lists eight ConnectX-7 cards per DGX system at up to 400 Gb/s for exactly that job. If your plan is inference on one model, a fabric is cost without a return.
This is also the point to be honest about utilisation. Owned hardware is cheapest when it is busy. A cluster that runs a nightly batch and sits idle for sixteen hours a day is an expensive way to keep data private, and it is worth pricing a dedicated hosted cluster against it before signing a purchase order.
Cooling, floor and fire: the facility checks
Every watt a GPU draws leaves as heat, so the cooling plant has to be sized to the same maximum as the power chain. The operating ranges NVIDIA publishes are the envelope the room must hold at the air intake: 5°C to 30°C for the DGX H100 and H200, 10°C to 35°C for the DGX B200. Hot-aisle or cold-aisle containment is what keeps that intake temperature stable when the systems are at full draw.
- Floor loading. At up to 142.4 kg per system, four DGX B200 systems weigh 569.6 kg before the rack, the switches and the storage are added. Check a raised floor’s rating before delivery day, not on it.
- Delivery path. Doors, lifts and corridor turns between the loading bay and the room. An 8U or 10U system on its pallet does not bend.
- Fire suppression. Clean-agent suppression and very early smoke detection protect hardware that water would destroy.
- Out-of-band management. A separate management network, isolated from production, so the platform can be reached and repaired when the main network is the thing that failed.
None of these are exotic. They are the ordinary disciplines of a data center applied at a density most rooms have not seen, and they are where a project either stays on schedule or waits three months for an electrician.
The software that turns GPUs into a private model API
Once the hardware is running, the goal is usually an internal API that applications can call the same way they would call a public one. Open-source serving engines make that straightforward. The vLLM online serving documentation states that it supports the OpenAI Completions, Chat Completions, Responses and Embeddings APIs, among others, which means existing client code can point at your own server by changing a base URL.
1# Serve a model across all eight GPUs of one system2vllm serve meta-llama/Llama-3.1-70B-Instruct --tensor-parallel-size 83 4# Call it from inside the network, with the same request shape as a public API5curl http://ai.internal:8000/v1/chat/completions \6 -H "Content-Type: application/json" \7 -d '{"model": "meta-llama/Llama-3.1-70B-Instruct", "messages": [{"role": "user", "content": "Summarise this contract clause."}]}'The serving engine is the easy layer. The work that makes it a production service is everything around it: authentication in front of the endpoint, per-team quotas, logging of who asked what for audit, dashboards for GPU utilisation, temperature and power draw, and a deployment process that can swap a model without an outage. The same discipline we describe for securing a Node.js application in production applies to the gateway in front of the model, and the retrieval layer that feeds it your documents is covered in our guide to RAG knowledge bases.
Read the model licence before you plan around it
Owning the hardware does not mean owning the model. Open-weight models come with licences, and they differ. The Llama 3.1 Community License, with a version release date of 23 July 2024, carries an additional commercial term that most organisations will never trigger and a few must plan for.
“If, on the Llama 3.1 version release date, the monthly active users of the products or services made available by or for Licensee, or Licensee’s affiliates, is greater than 700 million monthly active users in the preceding calendar month, you must request a license from Meta.”
Licences differ from one model family to the next. Put the licence of every candidate model in front of your legal team at the same time as the hardware quote, because swapping models later is cheap on the hardware and can be expensive in a contract.
When a private AI data center is the wrong answer
We build these, and we also talk clients out of them. The case is strong when data genuinely may not leave the organisation, when utilisation will be high, or when the model itself is the asset. It is weak when the workload is occasional, when nobody on staff will own the platform after handover, or when the real requirement is a contractual guarantee that a provider will not train on your data, which is a procurement question rather than a construction project.
- Write down where the data is allowed to go, in the words of whoever owns that decision. This settles the arrangement before anything else.
- Name the models you will run and do the memory arithmetic. This settles the number of systems.
- Get the deliverable kilowatts per rack from facilities or the colocation provider, on how many feeds. This settles the room.
- Estimate utilisation honestly and price the dedicated hosted option against owned hardware at that utilisation.
- Name the team that runs it on day two. If there is no name, budget for a managed service contract from day one.
For a full example of the owned-room path, from an empty room to a running model behind a private API, our private AI datacenter case study describes a cluster of NVIDIA DGX systems delivering 32 Blackwell-class GPUs, the power and cooling built around it, and the custom model we deployed and maintain on it. If you are weighing the same decision and want the five questions answered before anyone orders hardware, talk to us. Other systems we have built are on the work page, and if the end goal is an agent rather than a model API, our guide to choosing an AI agent development company covers that layer.
What is a private AI data center?
A facility, owned or leased, where an organisation runs AI models on dedicated hardware with no third-party AI provider in the request path. It can be your own room, a colocation cage holding hardware you own, or a single-tenant cluster hosted by a provider. The first two put power, cooling and operations in your hands.
How much power does an AI server need?
NVIDIA lists 10.2 kW maximum for one 8-GPU DGX H100 or H200 system and 14.3 kW maximum for one DGX B200. The Uptime Institute’s 2025 survey puts the average typical rack density at almost 9 kW, so a single system can exceed what an existing rack was designed to deliver.
How many GPUs do I need to run a 70 billion parameter model privately?
At 16-bit precision the weights alone need about 140 GB, which fits in one 8-GPU system with 640 GB of GPU memory and leaves room for concurrent requests. At 8-bit the weights need about 70 GB. Several systems are usually justified by training, fine-tuning or serving many large models, not by serving one.
Can applications call a private model the same way they call a public API?
Yes. Serving engines such as vLLM implement OpenAI-compatible endpoints, including Chat Completions and Embeddings, so existing client code can point at your own server by changing the base URL. Authentication, quotas and audit logging in front of that endpoint are still yours to build.
Is an on-premise AI data center cheaper than the cloud?
Only when the hardware is busy. Owned systems cost the same whether they run all day or once a night, so estimate utilisation honestly and price a dedicated hosted cluster at that utilisation before buying. Data residency, not cost, is usually the stronger reason to own.
Need help building this?
Let our team build it for you.
Dude Lemon builds production-grade web apps, APIs, and cloud infrastructure. Get a free consultation and project proposal within 48 hours.
Start a project