The transition from GPU-centric planning to capacity-aware planning should be staged. The goal is to improve decision quality before the organization commits to the largest facilities and energy investments.

Many AI infrastructure plans still begin with a model-size estimate, convert that estimate into GPU demand, and then move directly into procurement. That approach may work for a limited cloud experiment. It is inadequate for a production platform that requires firm capacity, predictable service levels, controlled data placement, and a multiyear operating model.

That does not mean every enterprise AI rack will operate at that density. It means average rack-density assumptions from general-purpose virtualization environments can no longer be used safely for AI planning.

AI training and distributed inference depend on high-throughput, low-loss network fabrics, storage bandwidth, metadata performance, and predictable data movement. Moving a large cluster farther from constrained metro areas may improve access to power while increasing WAN costs, replication time, recovery complexity, user latency, or exposure to a small number of long-haul paths.

  • Power identified: A utility, developer, or provider believes capacity may be available.
  • Power studied: The load and its grid impact have entered the appropriate study process.
  • Power contracted: Commercial terms, upgrade obligations, and delivery conditions are documented.
  • Power energized: The physical service is available to the site.
  • Power qualified: The service has passed commissioning, resilience, and workload acceptance tests.

Annual renewable-energy matching or corporate averages may not explain the local impact of a specific site. Site-level energy, water, emissions, and community metrics are needed for defensible decisions.

Rack Density and Cooling Readiness

A credible AI power plan should distinguish at least five milestones:

The scale of the wider market explains why. The U.S. Department of Energy’s 2024 data-center energy report estimated that data centers consumed about 4.4 percent of U.S. electricity in 2023 and could account for approximately 6.7 to 12 percent by 2028. In June 2026, the Federal Energy Regulatory Commission directed all six regional grid operators under its jurisdiction to justify or reform how data centers and other large loads connect to the transmission system.

The target state is not a facilities project with an AI label. It is a joint operating model across business sponsors, application owners, AI platform teams, infrastructure, network engineering, data-center operations, finance, procurement, sustainability, legal, utilities, and local stakeholders.

  • Facility water loops or an approved alternative heat-rejection design
  • Coolant distribution units and appropriate redundancy
  • Manifolds, quick disconnects, filtration, water chemistry, and leak detection
  • Warm-water supply and return temperature designs that match the equipment
  • Controls integration across the building management system and data-center infrastructure management platform
  • Maintenance procedures for mixed air-cooled and liquid-cooled components
  • Spare parts, trained technicians, and vendor escalation paths
  • Commissioning under realistic heat load rather than an empty-room inspection

Most enterprises already own many of the required data sources. The problem is that they are separated across IT and facilities systems.

A board does not need to design a cooling loop. It does need confidence that the AI investment depends on a complete, owned, and validated capacity chain.

Colocation Versus Private Facilities

The practical mistake is to treat GPUs as the unit of capacity and everything else as supporting infrastructure. In production, the reverse is often true. Accelerators are one component inside a chain that includes rack power, cooling, network, storage, utility service, on-site generation, permits, environmental constraints, operating capability, and community acceptance.

Decision factor Private facility Colocation Public cloud or managed AI capacity
Time to first capacity Slow when major power or cooling upgrades are required Potentially faster when qualified high-density capacity already exists Often fastest for initial capacity, subject to region, quota, and service availability
Design control Highest Shared with provider and contract boundaries Lowest physical control, highest service abstraction
Rack-density customization Strong when the site can be redesigned Depends on suite, provider standard, coolant model, and available power blocks Provider-defined
Capital profile High capital commitment and long asset life Contracted recurring commitment with buildout charges Consumption and commitment-based operating expense
Permitting and community exposure Direct organizational responsibility Mostly provider-managed, but customer demand still influences expansion Provider-managed, though regional constraints still affect availability
Data gravity Strong for local enterprise data Strong when connected to enterprise campuses and carrier ecosystems Strong when data already resides in the same cloud region
Operational responsibility Full facilities and platform stack Platform stack plus provider relationship and shared facility processes Platform, data, application, cost, and service governance
Exit complexity Physical assets and specialized facility investment Contract terms, migration windows, cross-connects, and data movement Data egress, service dependencies, reserved commitments, and refactoring

Liquid cooling crosses facilities, controls, operations, maintenance, water treatment, safety, and vendor-support boundaries. The ownership model must be explicit.

A power-advantaged site can still be the wrong AI site.

On-Site Generation and Microgrids


TL;DR

Colocation is not automatically a shortcut. A provider may have building-level megawatts available while lacking the exact rack density, liquid-cooling configuration, network fabric, or delivery window the AI design requires. Contract language should identify the supported kW per rack, cooling method, delivery temperature, redundancy model, maintenance rights, expansion blocks, power curtailment terms, metering, service credits, and responsibilities at every handoff.

Reserved GPU capacity without qualified network and data paths is stranded capacity.

  • Fuel availability, pipeline capacity, or renewable resource variability
  • Air-quality, noise, water, and land-use permits
  • Emissions controls and sustainability commitments
  • Black-start, islanding, synchronization, and protection requirements
  • Maintenance staffing and spare-parts strategy
  • Battery duration and degradation
  • Utility coordination and export restrictions
  • Community concerns about local environmental impact

As a rough planning illustration, 50 racks designed for 100 kilowatts each create a 5-megawatt rack-level IT envelope. That does not mean a 5-megawatt utility service is sufficient. The site still needs to account for cooling and electrical overhead, network and storage systems, redundancy, maintenance states, growth headroom, and any limits placed on the interconnection.

On-site generation may reduce grid dependency while increasing air-quality, fuel, noise, land-use, and community obligations.

Network Capacity and Data Gravity

AI power has become a business capacity decision because the physical system now influences strategy, location, schedule, cost, sustainability, and public approval.

This is why AI infrastructure plans need a capacity dependency map rather than a single procurement schedule.

The strongest CIO approach is to manage that chain as a capacity portfolio. Classify workloads, translate demand into physical requirements, qualify sites and providers, separate reserved capacity from usable capacity, and expose the limiting gate in every executive review.

Only the final milestone should be counted as usable production capacity.

  • Inside the cluster: GPU fabric, east-west bandwidth, congestion control, and topology.
  • Inside the site: Storage, ingestion, management, backup, observability, and tenant segmentation.
  • Between sites: WAN capacity, cloud interconnects, replication, user access, recovery paths, and data-sovereignty boundaries.

A credible community and environmental plan should address:

Water, Environmental, and Community Constraints

Bloom Energy’s 2026 annual survey reported that utility respondents expected time-to-power to take roughly 1.5 to 2 years longer on average than hyperscalers and colocation providers expected. The same report described widening expectation gaps in Northern Virginia, the Bay Area, and Atlanta. This is vendor-sponsored survey evidence rather than a universal utility benchmark, but the planning implication is still important: the party buying compute and the party delivering power may be working from materially different dates.

A useful executive metric is not “GPUs under contract.” It is “qualified AI capacity available by service class and date.” That metric should report the limiting gate, not hide it.

Cooling choices affect both power and water. The Lawrence Berkeley National Laboratory estimated direct data-center water consumption of approximately 66 billion liters in 2023, with hyperscale and colocation facilities representing most of that total. Its report also emphasized that indirect water consumption varies with the regional electricity mix, which means a low-water cooling design can still carry a material water footprint through power generation.

The next AI capacity review should not begin with a GPU count. It should begin with a workload service class, a megawatt and cooling envelope, a site-qualified delivery date, and evidence that the organization has permission to operate at the intended scale.

  • Who pays for grid, road, water, and public-service upgrades
  • How residential and small-business ratepayers are protected
  • Expected water withdrawal and consumption under normal and peak conditions
  • Noise, emissions, lighting, backup-generation testing, and construction traffic
  • Local jobs, tax base, training, and other durable community benefits
  • Transparency on project phases, power sources, expansion limits, and mitigation commitments
  • How complaints, incidents, and environmental data will be reported after the site opens

For a traditional enterprise data center, power was often treated as a stable facility constraint. For large AI deployments, it is becoming a location and sequencing decision.

Workload Placement Based on Energy Availability

The practical CIO model is to treat AI capacity as a chain of constrained gates. The lowest available gate sets the real capacity ceiling. A business may have thousands of GPUs under contract and still have less deployable capacity than expected because power delivery, liquid-cooling readiness, network buildout, or permitting is late.

Workload class Power and placement characteristics Likely placement pattern
Real-time customer inference Low latency, high availability, limited curtailment tolerance Near users or application services, often across multiple regions or sites
Regulated private inference Strong data gravity, controlled trust boundary, predictable demand Qualified private facility or colocation with firm capacity and governed data locality
Large training and fine-tuning High power density, high network demand, often schedulable with checkpointing Power-advantaged site or cloud region with strong fabric and data-staging capability
Batch embedding and indexing Flexible scheduling, significant data movement, restartable Site or time window selected for capacity, price, and energy conditions
Development and experimentation Bursty, uncertain demand, lower utilization Cloud, shared enterprise platform, or smaller private pools
Business-critical agents Moderate compute but strong dependency on data, tools, identity, and uptime Close to governed systems of record with resilient inference capacity

Before approving a major AI infrastructure commitment, CEOs and CIOs should be able to answer these questions:

CEOs and CIOs should therefore make AI workload placement, data-center sourcing, sustainability, and community engagement part of the same capacity decision. The question is not simply, “How many GPUs can we buy?” It is, “How much reliable AI work can we operate, where can we operate it, and what physical and social permissions must be in place first?”

Build a CIO-Level AI Capacity Ledger

This model also creates better workload-placement decisions. Latency-sensitive inference can remain close to users and data. Flexible training and batch workloads can move toward power-advantaged locations or time windows. Regulated workloads can remain inside qualified private boundaries. Cloud and colocation can be used deliberately rather than as emergency overflow after a private site misses its delivery date.

The placement decision should therefore evaluate three network layers:

Data gravity adds another constraint. Sensitive datasets may already reside in a private facility, specific cloud, regional data platform, or regulated geography. The energy-optimal location may not be the data-optimal location. Moving the data can consume more time, money, and operational risk than moving the workload.

Several mistakes repeatedly make AI capacity look larger than it is.

External References

Similar Posts