BlackForge Data Centers

Site strategy by data center type

AI Training vs. Inference Data Centers: How Site Needs Differ

AI training and AI inference ask different things of a site. Training clusters need very large, dense blocks of power and cooling but can sit far from users, while inference serves live requests and is pulled toward metro and near-metro locations with strong fiber.1 Both use high-density GPU racks, so the real difference for land is scale, latency tolerance and how flexible the location can be.12

Last reviewed · 8 min read · BlackForge Data Centers

Key takeaways

  • Training is a batch job: it tolerates distance from users, so it can follow cheap, available power to remote campuses such as the 1.2 GW Stargate site in Abilene, Texas.3
  • Inference answers user requests in real time, so latency pulls it toward metros, cloud regions and interconnection hubs.14
  • McKinsey projects inference will overtake training as the dominant AI data center workload by 2030, which shifts demand toward network-rich locations.1
  • Current GPU racks draw roughly 120–132 kW each, about ten times a general-purpose CPU rack, for training and inference alike.52
  • Large training clusters can swing their power draw quickly, and grid planners now treat that behavior as a reliability issue.67

01Two different jobs: training and inference

Training is the process of building a model: thousands to hundreds of thousands of accelerators work on one job for weeks or months, exchanging results constantly. Inference is the process of using a trained model: answering a chatbot prompt, ranking search results, generating code or summarizing a document, usually while a person or application waits for the answer.

Those two jobs put different demands on land. A training run cares about how many accelerators can be wired together in one place, how much power and cooling the site can deliver, and how soon. It does not care much where the end user is. Inference cares about response time and reliability for users in a region, so it behaves more like a cloud or colocation workload than a factory.1

The balance between them is changing. McKinsey projects that inference will overtake training as the dominant AI workload by 2030, and that inference build-outs are concentrating in metro and near-metro sites with low latency and strong connectivity, while training favors large, high-density campuses.1 The IEA estimates that data centers used about 415 TWh of electricity in 2024, around 1.5% of global consumption, with the United States accounting for 45% of it.8 How that demand splits between training and inference shapes which kinds of land are in demand.

For a broader view of AI campuses, see our guide to AI data center site requirements. This guide focuses on how the two workloads pull site selection in different directions.

02Latency: where inference is tied and training is free

Latency is the time a signal takes to travel between two points and back. Light in optical fiber travels at roughly 200,000 km per second, which works out to about 5 microseconds of one-way delay per kilometer of fiber, or about 5 ms for every 1,000 km.9 Real routes are longer than straight lines, and every router and switch adds delay on top of that floor.

For a training run, those milliseconds matter inside the cluster, not between the cluster and users. The accelerators synchronize constantly, so the facility is designed for very short, very fast links between racks and buildings. The finished model is then copied to inference sites, and the training campus itself can sit hundreds of miles from the nearest large city.

Inference is the opposite. Each user request makes a round trip, and interactive applications feel slow when network delay stacks on top of model processing time. JLL’s 2026 outlook notes that inference demand requires geographic distribution to reduce latency and serve users, which it expects to drive regional deployments and systems at the edge.4 Our guide to latency requirements by workload goes deeper into how many milliseconds different applications can tolerate.

Fig. 1Training vs. inference: what each asks of a site

Power first

Training

  • Very large, contiguous power blocks
  • Tolerates distance from users
  • Dense GPU racks with liquid cooling
  • Fast, synchronized load swings
  • Follows available power and speed

Network first

Inference

  • Smaller, more numerous deployments
  • Latency-sensitive, near users
  • Dense GPU racks with liquid cooling
  • Load tracks user demand over the day
  • Follows metros, fiber and cloud regions
General patterns, not rules; many campuses host both workloads. Based on McKinsey and JLL;14 load-swing behavior from PNNL.6

03Scale: how much power each workload concentrates

Training rewards concentration. The more accelerators that can be linked in one tightly coupled cluster, the larger the model that can be trained in a given time. That is why training campuses are measured in hundreds of megawatts to gigawatts. The Stargate campus in Abilene, Texas, developed by Crusoe and leased to Oracle for OpenAI workloads, is planned at about 1.2 GW across roughly 4 million square feet.3

Inference rewards distribution. A given volume of inference can be split across many sites, each sized to the demand in its region, and each needing a reliable connection to users and to cloud networks. Individual inference deployments are therefore usually smaller than frontier training clusters, although large cloud regions can add up to very large totals. McKinsey’s analysis expects inference capacity to concentrate in metro and near-metro areas where low latency, connectivity and energy efficiency come together.1

The practical consequence is that the same acreage can be valued very differently. A 1,000-acre tract next to a high-voltage line in a remote county may suit a training campus well and an inference deployment poorly. A 40-acre parcel near a metro fiber hub may be the reverse. Our guide to what a gigawatt-scale AI campus requires covers the very large end.

Fig. 2Signals of scale in AI data centers

Global data center electricity use, 20248
415 TWh
U.S. share of that consumption8
45%
Planned power at the Abilene Stargate campus3
1.2 GW
Global figures from the IEA; Abilene from DCD reporting on the Crusoe campus.83

04Power density and cooling: more alike than different

On rack density, training and inference have converged. Both increasingly run on the same class of accelerator systems. Supermicro’s GB200 NVL72 rack configuration lists about 132 kW of total rack power with an in-rack coolant distribution unit.5 SemiAnalysis puts the NVL72 form factor at about 120 kW per rack, compared with up to 12 kW for a general-purpose CPU rack and about 40 kW for air-cooled H100 racks.2

SemiAnalysis also notes that moving well beyond 40 kW per rack is the main reason liquid cooling is required for these systems, and that many existing facilities cannot support such densities even with direct-to-chip cooling.2 For land, that means both workloads now need the mechanical yard space, piping and heat rejection that liquid cooling implies. See air vs. liquid cooling site implications and high-density GPU racks and site power.

Fig. 3Rack power by system type

  • General-purpose CPU rackup to 12
  • Air-cooled H100 rack~40
  • GB200 NVL36x2 (per rack)66
  • GB200 NVL72~120
  • Supermicro GB200 NVL72 spec132

kW per rack

Approximate figures from SemiAnalysis and a Supermicro product specification; actual draw depends on configuration.25

Where they differ is how much of that density is concentrated in one place. A training hall may fill entire buildings with these racks on a single fabric. An inference deployment may place a few rows of them inside a metro colocation facility, where the constraint is often the building’s existing power and cooling rather than the land.

05Load behavior and the grid

Training clusters do not draw power like a typical data center. Because thousands of accelerators compute and then synchronize in step, the facility’s demand can rise and fall together. A PNNL report on large loads states that AI training and inference facilities can inject large active power swings concentrated at specific frequencies over extended periods, which existing grid planning practices do not account for.6

Grid operators have responded. In May 2026 NERC issued a Level 3 Essential Action alert on large computational loads, its highest alert level, citing sudden customer-initiated load reductions and significant oscillations, and requiring registered entities to take specific actions.7 For a site, this shows up as more detailed utility questions about ride-through, ramp rates and on-site storage or controls, especially for training-heavy campuses. Our guide to NERC reliability and large loads explains the standards side.

Inference load tends to follow user activity through the day, so it looks more like a conventional cloud load. It is still large and continuous, but it is less likely to produce synchronized swings across a whole campus.

06Location flexibility and multi-site training

Training is becoming more flexible about location, not less. Google’s Gemini technical report states that Gemini Ultra was trained on TPUv4 accelerators across multiple data centers, combining 4,096-chip SuperPods over Google’s intra-cluster and inter-cluster network, and that its network latency and bandwidth were sufficient for synchronous training.10

Microsoft has gone further. It describes its Fairwater sites in Wisconsin and Atlanta as an “AI superfactory,” connected by a dedicated AI wide area network so that the sites, roughly 700 miles apart, work on large training jobs together.11 The implication for land is that a training campus no longer has to hold an entire cluster on one parcel. It does need very high-capacity fiber to its sibling sites, which is a fiber route diversity and latency question as much as a power question.

Inference has a different kind of flexibility. Requests can be routed to whichever site has capacity, so operators can place less time-sensitive work in cheaper, power-rich locations and keep latency-sensitive work near users. JLL ranks speed to power as the primary site selection criterion, followed by community support, latency and proximity to customers, and notes that grid connection waits in primary markets average more than four years.4 That tension is the core of the near-metro vs. remote power-rich sites decision.

07Screening a site for training, inference or both

Most land can be screened first by asking which workload it fits naturally, then testing whether it can flex to the other. The answer changes who the likely buyer or tenant is, what they will pay for, and which risks they will scrutinize first.

Fig. 4Power scale vs. network proximity

Tens of MW ← Available power → Hundreds of MW+

Training campus

Large power, remote; needs long-haul fiber to sibling sites.

Mixed AI campus

Scarce and priced to match; suits both workloads.

Weak fit

Small power, far from users; limited AI use.

Inference node

Modest power near users; colocation or regional cloud.

Remote ← Proximity to users and fiber hubs → Metro

A way to frame which AI workload a site suits; not a rule, and many sites end up mixed.
  1. 01Size the power: how many MW can the site plausibly receive, and on what schedule? Training needs large blocks; inference can start smaller.
  2. 02Measure the network: distance and route to the nearest metro, cloud region and internet exchange, and how many independent fiber paths exist.
  3. 03Check cooling: space and water or dry-cooling options for liquid-cooled racks at 100 kW-plus.2
  4. 04Ask about load behavior: what the utility will require on ramp rates, ride-through and oscillation controls.67
  5. 05Decide the story: a site pitched as a training campus is judged on power and speed; one pitched for inference is judged on latency and connectivity.

If you are weighing which story a parcel can support, we can get a site reviewed against both.

Common questions

Do AI training data centers need to be near cities?

Generally no. Training is a batch process, so its latency needs are inside the cluster rather than to users, which is why large training campuses such as Abilene, Texas, sit away from major metros.3 They do need high-capacity fiber, especially if they share jobs with other sites.11

Why does AI inference need to be close to users?

Every inference request makes a round trip over the network, and fiber adds about 5 ms of one-way delay per 1,000 km before equipment delays.9 Interactive applications are sensitive to that delay, so operators distribute inference capacity regionally.4

Is inference going to be bigger than training?

McKinsey projects that inference will overtake training as the dominant AI data center workload by 2030.1 Projections like this depend on efficiency gains and adoption, so treat them as directional.

Do inference data centers use less power per rack than training?

Not necessarily. Both increasingly run on the same liquid-cooled accelerator racks, which draw roughly 120–132 kW each.52 The difference is usually how many racks are concentrated in one location.

Can one site host both training and inference?

Yes. Large cloud and AI campuses often do, and operators can route less time-sensitive work to power-rich sites.4 A site near a metro with large power available is the most flexible, and usually the most expensive.

Notes

  1. 1.McKinsey & Company, “The future of AI workloads,” 2025. mckinsey.com
  2. 2.SemiAnalysis, “GB200 Hardware Architecture and Component Supply Chain,” 2024. newsletter.semianalysis.com
  3. 3.Data Center Dynamics, “Crusoe tops out final building at OpenAI Stargate data center campus in Abilene, Texas,” 2025. datacenterdynamics.com
  4. 4.JLL, “2026 Global Data Center Outlook,” 2026. jll.com
  5. 5.Supermicro, “SuperServer SRS-GB200-NVL72,” n.d. supermicro.com
  6. 6.Pacific Northwest National Laboratory, “A Methodology to Evaluate the Grid Reliability Impact of Oscillations Induced by Large Loads (PNNL-39459),” 2026. pnnl.gov
  7. 7.Midwest Reliability Organization, “NERC Issues Critical Reliability Alert as Emerging Large Loads Challenge the Grid,” 2026. mro.net
  8. 8.International Energy Agency, “Energy and AI: Executive summary,” 2025. iea.org
  9. 9.CommScope, “Latency in Optical Fiber Systems (white paper),” n.d. commscope.com
  10. 10.Gemini Team, Google, “Gemini: A Family of Highly Capable Multimodal Models,” 2023. arxiv.org
  11. 11.Microsoft, “From Wisconsin to Atlanta: Microsoft connects datacenters to build its first AI superfactory,” 2025. news.microsoft.com

Have a site in mind?

Get a straight answer on your land.

Send a parcel number, an address, a map pin or a target load. We’ll tell you what it can support and what it would take.

Start a conversation →

This guide is general information about data center site selection. It is not engineering, legal, tax or investment advice. Requirements vary by state, utility and county, so confirm the specifics for any site with the relevant authorities and advisors.

Related guides

More in Site strategy by data center type