Getty Images
The business cost of underutilized data center capacity
Underutilized data center capacity poses a costly challenge for organizations, as infrastructure often lacks resources to optimize workload efficiency and performance.
Should data center operators invest aggressively in new infrastructure to support growing cloud and AI workloads? Many existing facilities have significant underused capacity. The contradiction arises from treating installed infrastructure and usable capacity as the same thing.
Underutilization is not only about servers sitting idle. For instance, a facility can have open rack space but lack the power needed for the target workload. Compute resources can also exist without applications being able to use them effectively because of limitations elsewhere in the infrastructure.
Siraj Aziz, an analyst for cloud and data center at Omdia, a division of Informa TechTarget, explained that the amount of installed infrastructure is only one part of the equation. Productive capacity depends on whether systems and workloads can use those resources to meet performance and business requirements. The more important question is: If the capacity already exists, why can't organizations put more workloads on it, and what does that mismatch cost the business?
Available data center capacity is becoming unusable
The amount of equipment or floor space available does not determine how much additional workload a data center can support. Bruce Bateman, chief analyst for semiconductors at Omdia, described this challenge in the context of upgrading legacy facilities.
He explained that adding newer compute also requires more storage, higher-capacity routers and internet connectivity. Upgrading for higher-density AI infrastructure will also require changes to electrical distribution and cooling systems. In some cases, retrofitting liquid cooling can introduce new plumbing, permitting and facility requirements.
Organizations may intentionally provision more capacity than normal demand requires to handle workload peaks and maintain application performance. Application architecture and orchestration can also limit the efficiency of allocating available resources.
Siraj noted that this becomes harder as workloads change. Dynamic and agentic AI workloads complicate capacity planning because their processing requirements and hardware needs can vary over time. Visibility in network, compute and AI infrastructure remains fragmented across systems.
AI amplifies these infrastructure dependencies with large training clusters requiring coupled accelerators, high-bandwidth networking, substantial storage throughput, high rack power density and advanced cooling.
Inference has different resource requirements depending on the model, concurrency and latency requirements. For instance, a facility capable of supporting traditional enterprise applications cannot be assumed to have equivalent capacity for all AI workloads.
For the same reason, low CPU or GPU utilization does not indicate excess compute capacity. The processor may be waiting on memory, storage, networking or another part of the workload pipeline. Capacity must be evaluated across the system.
Stranded capacity costs the business
The cost of stranded capacity goes well beyond underutilized servers.
Data center operators have already invested capital in servers, racks, storage, networking, UPS systems, electrical distribution, cooling equipment and the facility itself. Many operating costs also continue regardless of workload, including maintenance, staffing, connectivity, monitoring and lifecycle management.
A bottleneck can also strand investment elsewhere in the facility. As a result, a single constrained resource can reduce the productive value of several other assets that the organization has paid for.
Freeing that capacity can add another layer of cost. Bateman said that an operator considering modernization must assess the total retrofit cost required to support the target workload.
This makes opportunity cost important. A colocation provider may have empty racks yet still be unable to accommodate the power density or configuration a customer requires. Those racks exist physically, but they do not become revenue-generating inventory. The enterprise can face the same problem because its existing systems cannot meet the workload's requirements.
AI workloads make this investment decision more consequential. Strong demand for AI infrastructure does not mean every legacy facility should be converted into a high-density AI data center. Bateman cautioned that upgrading compute alone does not drive utilization.
His broader point was that an operator should determine who will use the capacity and for what workload. Existing facilities may continue to generate value from conventional workloads, and selective upgrades could support other demand profiles, such as regional or private inference. A full conversion only makes economic sense when the expected workload justifies the full cost of making that capacity usable.
Measure productive capacity, not maximum utilization
Traditional utilization metrics remain valuable for managing a data center, but they do not indicate whether the infrastructure is generating useful work cost-effectively.
The most useful measures depend on what the infrastructure is expected to achieve. In conventional environments, the operator might consider jobs completed within performance requirements, cost per workload and energy consumed per unit of work. But AI infrastructure introduces additional measures such as workload throughput, concurrency, request latency and token throughput.
Siraj emphasized this system-level view of agentic AI, highlighting energy efficiency, infrastructure adaptability and total cost of ownership (TCO) as metrics that can provide a more comprehensive understanding of asset productivity.
Bateman made a similar point from the power side, emphasizing a focus on tokens per kilowatt for advanced AI infrastructure. He explained how much AI output an organization can achieve within a constrained energy budget. Current AI benchmarking evaluates several measures together, including token throughput and latency.
The principle is to measure useful workload output per constrained resource. In a power-constrained facility, that means productive workload per available megawatt. For a colocation operator, revenue or contribution margin per sellable kilowatt matters more than server utilization. The denominator varies with the workload and business model.
This does not mean driving every asset in the facility to 100% utilization. Some unused capacity provides necessary headroom for failures, maintenance, workload spikes, redundancy and future growth.
Before adding more infrastructure, leaders should identify the workload, determine which resource constrains it, calculate the total investment required to remove that constraint, and evaluate the resulting TCO and business output. Productive workload capacity is the output that determines its business value.
Abhishek Jadhav is a technology journalist covering AI infrastructure, semiconductors and advanced computing systems.