
TL;DR
Colocation is not automatically a shortcut. A provider may have building-level megawatts available while lacking the exact rack density, liquid-cooling configuration, network fabric, or delivery window the AI design requires. Contract language should identify the supported kW per rack, cooling method, delivery temperature, redundancy model, maintenance rights, expansion blocks, power curtailment terms, metering, service credits, and responsibilities at every handoff.
Reserved GPU capacity without qualified network and data paths is stranded capacity.
- Fuel availability, pipeline capacity, or renewable resource variability
- Air-quality, noise, water, and land-use permits
- Emissions controls and sustainability commitments
- Black-start, islanding, synchronization, and protection requirements
- Maintenance staffing and spare-parts strategy
- Battery duration and degradation
- Utility coordination and export restrictions
- Community concerns about local environmental impact
As a rough planning illustration, 50 racks designed for 100 kilowatts each create a 5-megawatt rack-level IT envelope. That does not mean a 5-megawatt utility service is sufficient. The site still needs to account for cooling and electrical overhead, network and storage systems, redundancy, maintenance states, growth headroom, and any limits placed on the interconnection.
On-site generation may reduce grid dependency while increasing air-quality, fuel, noise, land-use, and community obligations.
Network Capacity and Data Gravity
AI power has become a business capacity decision because the physical system now influences strategy, location, schedule, cost, sustainability, and public approval.
This is why AI infrastructure plans need a capacity dependency map rather than a single procurement schedule.
The strongest CIO approach is to manage that chain as a capacity portfolio. Classify workloads, translate demand into physical requirements, qualify sites and providers, separate reserved capacity from usable capacity, and expose the limiting gate in every executive review.
Only the final milestone should be counted as usable production capacity.
- Inside the cluster: GPU fabric, east-west bandwidth, congestion control, and topology.
- Inside the site: Storage, ingestion, management, backup, observability, and tenant segmentation.
- Between sites: WAN capacity, cloud interconnects, replication, user access, recovery paths, and data-sovereignty boundaries.
A credible community and environmental plan should address:
Water, Environmental, and Community Constraints
Bloom Energy’s 2026 annual survey reported that utility respondents expected time-to-power to take roughly 1.5 to 2 years longer on average than hyperscalers and colocation providers expected. The same report described widening expectation gaps in Northern Virginia, the Bay Area, and Atlanta. This is vendor-sponsored survey evidence rather than a universal utility benchmark, but the planning implication is still important: the party buying compute and the party delivering power may be working from materially different dates.
A useful executive metric is not “GPUs under contract.” It is “qualified AI capacity available by service class and date.” That metric should report the limiting gate, not hide it.
Cooling choices affect both power and water. The Lawrence Berkeley National Laboratory estimated direct data-center water consumption of approximately 66 billion liters in 2023, with hyperscale and colocation facilities representing most of that total. Its report also emphasized that indirect water consumption varies with the regional electricity mix, which means a low-water cooling design can still carry a material water footprint through power generation.
The next AI capacity review should not begin with a GPU count. It should begin with a workload service class, a megawatt and cooling envelope, a site-qualified delivery date, and evidence that the organization has permission to operate at the intended scale.
- Who pays for grid, road, water, and public-service upgrades
- How residential and small-business ratepayers are protected
- Expected water withdrawal and consumption under normal and peak conditions
- Noise, emissions, lighting, backup-generation testing, and construction traffic
- Local jobs, tax base, training, and other durable community benefits
- Transparency on project phases, power sources, expansion limits, and mitigation commitments
- How complaints, incidents, and environmental data will be reported after the site opens
For a traditional enterprise data center, power was often treated as a stable facility constraint. For large AI deployments, it is becoming a location and sequencing decision.
Workload Placement Based on Energy Availability
The practical CIO model is to treat AI capacity as a chain of constrained gates. The lowest available gate sets the real capacity ceiling. A business may have thousands of GPUs under contract and still have less deployable capacity than expected because power delivery, liquid-cooling readiness, network buildout, or permitting is late.
| Workload class | Power and placement characteristics | Likely placement pattern |
|---|---|---|
| Real-time customer inference | Low latency, high availability, limited curtailment tolerance | Near users or application services, often across multiple regions or sites |
| Regulated private inference | Strong data gravity, controlled trust boundary, predictable demand | Qualified private facility or colocation with firm capacity and governed data locality |
| Large training and fine-tuning | High power density, high network demand, often schedulable with checkpointing | Power-advantaged site or cloud region with strong fabric and data-staging capability |
| Batch embedding and indexing | Flexible scheduling, significant data movement, restartable | Site or time window selected for capacity, price, and energy conditions |
| Development and experimentation | Bursty, uncertain demand, lower utilization | Cloud, shared enterprise platform, or smaller private pools |
| Business-critical agents | Moderate compute but strong dependency on data, tools, identity, and uptime | Close to governed systems of record with resilient inference capacity |
Before approving a major AI infrastructure commitment, CEOs and CIOs should be able to answer these questions:
CEOs and CIOs should therefore make AI workload placement, data-center sourcing, sustainability, and community engagement part of the same capacity decision. The question is not simply, “How many GPUs can we buy?” It is, “How much reliable AI work can we operate, where can we operate it, and what physical and social permissions must be in place first?”
Build a CIO-Level AI Capacity Ledger
This model also creates better workload-placement decisions. Latency-sensitive inference can remain close to users and data. Flexible training and batch workloads can move toward power-advantaged locations or time windows. Regulated workloads can remain inside qualified private boundaries. Cloud and colocation can be used deliberately rather than as emergency overflow after a private site misses its delivery date.

The placement decision should therefore evaluate three network layers:
Data gravity adds another constraint. Sensitive datasets may already reside in a private facility, specific cloud, regional data platform, or regulated geography. The energy-optimal location may not be the data-optimal location. Moving the data can consume more time, money, and operational risk than moving the workload.
Several mistakes repeatedly make AI capacity look larger than it is.
