AI Integration

Private AI Data Center: What It Takes Before the First GPU Arrives

A private AI data center is a power and cooling project that happens to end in a model API. One 8-GPU server draws more than the typical enterprise rack is built for. Here is the sizing arithmetic, the facility checks, the serving stack and the licence line to read, all from primary sources.

A private AI data center is a room, a power chain, a cooling plant and a network that together let an organisation run large models on hardware it owns, with no third-party AI provider in the request path. Most conversations about one start with GPUs. The ones that go well start with the electrical supply, because a single 8-GPU server now draws more power than the typical enterprise rack was built to deliver. Everything below comes from the vendor documentation and industry survey data we read on 30 September 2026, and it is the checklist we work through before a client orders any compute.

The GPUs are the easiest part of a private AI data center to buy. The power, the cooling, the floor and the people who will keep it running are the parts that decide whether it works.

What a private AI data center means in practice

The phrase covers three different arrangements, and the right one depends less on budget than on where the data is allowed to go and who is going to be woken at three in the morning. All three keep the model and the prompts off shared public AI services. They differ in who owns the building and who holds the keys.

ArrangementWho owns the facilityWhat you still ownFits when
Your own roomYouEverything: power, cooling, security, hardware, operationsData may not leave your premises, or you already run a facility
Colocation cageA colocation providerThe hardware, the network inside the cage, operationsYou want owned hardware without building a room
Dedicated hosted clusterA providerThe model, the data and the applicationSingle-tenant hardware is enough and speed matters more than ownership
The three shapes a private AI data center takes. The rest of this article is about the first two, where the facility questions are yours.

Owned facilities are not a fringe choice. The Uptime Institute Global Data Center Survey 2025 reports that on-premises data centers “remain foundational, with 45% of IT workloads still residing in corporate facilities”. What is new is what AI hardware asks of those facilities.

Power is the first constraint, not the GPU count

NVIDIA publishes the power figures for its integrated systems, and they are the right place to start. The DGX H100 and H200 user guide lists system power usage of “10.2 kW max” for one 8-GPU system. The DGX B200 user guide lists “14.3 kW max”. Both use six 3.3 kW power supplies.

SystemGPUs and total GPU memoryMax powerRack spaceMax weightOperating temperature
DGX H1008 x H100, 640 GB10.2 kW8U130.45 kg5°C to 30°C
DGX H2008 x H200, 1,128 GB10.2 kW8U130.45 kg5°C to 30°C
DGX B2008 x B200, 1,440 GB14.3 kW10U142.4 kg10°C to 35°C
Quoted from NVIDIA’s DGX H100/H200 and DGX B200 user guides, read 30 September 2026. Each system is one 8-GPU server.

Now set those numbers against what existing rooms were built for. The Uptime survey reports that “the average of the modal densities in our total sample reaches almost 9 kW”, and that “More than 80% of the operators responding to our survey say their facility has no racks above 30 kW”. A single DGX H100 at 10.2 kW is already above the typical rack. Four DGX B200 systems, which is 32 GPUs, add up to 57.2 kW at maximum draw, which is a density most of the surveyed facilities have never delivered to one rack.

  • Ask for the supply, not the space. The first question to a facilities team or a colocation provider is how many kilowatts can be delivered to one rack, on how many independent feeds.
  • Plan for maximum, not typical. The published figures are maximums. Training runs sit close to them for days at a time.
  • Count the redundancy. The DGX H100 power supplies are configured for 4+2 redundancy and the DGX B200 for 5+1. That only helps if the room gives them two independent feeds.
  • Budget the rest of the chain. A UPS sized for the load, automatic transfer switching and a generator are what keep a training job alive through a mains disruption.

Size the hardware from the model, not the other way round

The second question is what you intend to run, because that decides how much GPU memory you need, and GPU memory decides the number of systems. A useful first pass is arithmetic: a model stored at 16-bit precision needs two bytes per parameter just for its weights, and one stored at 8-bit needs one. That is before the working memory each concurrent request needs, so treat the figures below as a floor rather than a target.

Model sizeWeights at 16-bitWeights at 8-bitFits in one 640 GB systemFits in one 1,440 GB system
8 billion parametersabout 16 GBabout 8 GBYes, on a single GPUYes
70 billion parametersabout 140 GBabout 70 GBYes, with room for concurrent requestsYes
405 billion parametersabout 810 GBabout 405 GBOnly at 8-bitYes, at 16-bit
Weights only: parameter count multiplied by bytes per parameter. Concurrent requests need additional memory on top, so leave generous headroom.

Two practical conclusions follow. First, most organisations do not need a cluster to serve a model; a 70 billion parameter model fits comfortably in one 8-GPU system, and a smaller one fits on far less. Second, the case for several systems joined by a high-speed fabric is training and fine-tuning, or serving many large models at once. NVIDIA lists eight ConnectX-7 cards per DGX system at up to 400 Gb/s for exactly that job. If your plan is inference on one model, a fabric is cost without a return.

This is also the point to be honest about utilisation. Owned hardware is cheapest when it is busy. A cluster that runs a nightly batch and sits idle for sixteen hours a day is an expensive way to keep data private, and it is worth pricing a dedicated hosted cluster against it before signing a purchase order.

Cooling, floor and fire: the facility checks

Every watt a GPU draws leaves as heat, so the cooling plant has to be sized to the same maximum as the power chain. The operating ranges NVIDIA publishes are the envelope the room must hold at the air intake: 5°C to 30°C for the DGX H100 and H200, 10°C to 35°C for the DGX B200. Hot-aisle or cold-aisle containment is what keeps that intake temperature stable when the systems are at full draw.

  • Floor loading. At up to 142.4 kg per system, four DGX B200 systems weigh 569.6 kg before the rack, the switches and the storage are added. Check a raised floor’s rating before delivery day, not on it.
  • Delivery path. Doors, lifts and corridor turns between the loading bay and the room. An 8U or 10U system on its pallet does not bend.
  • Fire suppression. Clean-agent suppression and very early smoke detection protect hardware that water would destroy.
  • Out-of-band management. A separate management network, isolated from production, so the platform can be reached and repaired when the main network is the thing that failed.

None of these are exotic. They are the ordinary disciplines of a data center applied at a density most rooms have not seen, and they are where a project either stays on schedule or waits three months for an electrician.

The software that turns GPUs into a private model API

Once the hardware is running, the goal is usually an internal API that applications can call the same way they would call a public one. Open-source serving engines make that straightforward. The vLLM online serving documentation states that it supports the OpenAI Completions, Chat Completions, Responses and Embeddings APIs, among others, which means existing client code can point at your own server by changing a base URL.

bash
1# Serve a model across all eight GPUs of one system
2vllm serve meta-llama/Llama-3.1-70B-Instruct --tensor-parallel-size 8
3
4# Call it from inside the network, with the same request shape as a public API
5curl http://ai.internal:8000/v1/chat/completions \
6 -H "Content-Type: application/json" \
7 -d '{"model": "meta-llama/Llama-3.1-70B-Instruct", "messages": [{"role": "user", "content": "Summarise this contract clause."}]}'

The serving engine is the easy layer. The work that makes it a production service is everything around it: authentication in front of the endpoint, per-team quotas, logging of who asked what for audit, dashboards for GPU utilisation, temperature and power draw, and a deployment process that can swap a model without an outage. The same discipline we describe for securing a Node.js application in production applies to the gateway in front of the model, and the retrieval layer that feeds it your documents is covered in our guide to RAG knowledge bases.

Read the model licence before you plan around it

Owning the hardware does not mean owning the model. Open-weight models come with licences, and they differ. The Llama 3.1 Community License, with a version release date of 23 July 2024, carries an additional commercial term that most organisations will never trigger and a few must plan for.

“If, on the Llama 3.1 version release date, the monthly active users of the products or services made available by or for Licensee, or Licensee’s affiliates, is greater than 700 million monthly active users in the preceding calendar month, you must request a license from Meta.”

Llama 3.1 Community License Agreement, clause 2, read 30 September 2026

Licences differ from one model family to the next. Put the licence of every candidate model in front of your legal team at the same time as the hardware quote, because swapping models later is cheap on the hardware and can be expensive in a contract.

When a private AI data center is the wrong answer

We build these, and we also talk clients out of them. The case is strong when data genuinely may not leave the organisation, when utilisation will be high, or when the model itself is the asset. It is weak when the workload is occasional, when nobody on staff will own the platform after handover, or when the real requirement is a contractual guarantee that a provider will not train on your data, which is a procurement question rather than a construction project.

  • Write down where the data is allowed to go, in the words of whoever owns that decision. This settles the arrangement before anything else.
  • Name the models you will run and do the memory arithmetic. This settles the number of systems.
  • Get the deliverable kilowatts per rack from facilities or the colocation provider, on how many feeds. This settles the room.
  • Estimate utilisation honestly and price the dedicated hosted option against owned hardware at that utilisation.
  • Name the team that runs it on day two. If there is no name, budget for a managed service contract from day one.

For a full example of the owned-room path, from an empty room to a running model behind a private API, our private AI datacenter case study describes a cluster of NVIDIA DGX systems delivering 32 Blackwell-class GPUs, the power and cooling built around it, and the custom model we deployed and maintain on it. If you are weighing the same decision and want the five questions answered before anyone orders hardware, talk to us. Other systems we have built are on the work page, and if the end goal is an agent rather than a model API, our guide to choosing an AI agent development company covers that layer.

What is a private AI data center?

A facility, owned or leased, where an organisation runs AI models on dedicated hardware with no third-party AI provider in the request path. It can be your own room, a colocation cage holding hardware you own, or a single-tenant cluster hosted by a provider. The first two put power, cooling and operations in your hands.

How much power does an AI server need?

NVIDIA lists 10.2 kW maximum for one 8-GPU DGX H100 or H200 system and 14.3 kW maximum for one DGX B200. The Uptime Institute’s 2025 survey puts the average typical rack density at almost 9 kW, so a single system can exceed what an existing rack was designed to deliver.

How many GPUs do I need to run a 70 billion parameter model privately?

At 16-bit precision the weights alone need about 140 GB, which fits in one 8-GPU system with 640 GB of GPU memory and leaves room for concurrent requests. At 8-bit the weights need about 70 GB. Several systems are usually justified by training, fine-tuning or serving many large models, not by serving one.

Can applications call a private model the same way they call a public API?

Yes. Serving engines such as vLLM implement OpenAI-compatible endpoints, including Chat Completions and Embeddings, so existing client code can point at your own server by changing the base URL. Authentication, quotas and audit logging in front of that endpoint are still yours to build.

Is an on-premise AI data center cheaper than the cloud?

Only when the hardware is busy. Owned systems cost the same whether they run all day or once a night, so estimate utilisation honestly and price a dedicated hosted cluster at that utilisation before buying. Data residency, not cost, is usually the stronger reason to own.

Need help building this?

Let our team build it for you.

Dude Lemon builds production-grade web apps, APIs, and cloud infrastructure. Get a free consultation and project proposal within 48 hours.

Start a project