Tag: training

  • Why Running AI Models, Not Building Them, Will Drive Data Center Growth

    Why Running AI Models, Not Building Them, Will Drive Data Center Growth

    The Next Generation of AI Data Centers Explained | GMI Cloud

    For the past few years, the biggest data centers on Earth have been built for one purpose: training AI models. These facilities, packed with tens of thousands of GPUs, run for weeks at a time to teach models like GPT-4 how to generate text or images. But that era is ending. By 2026–2028, the industry consensus is that running AI models a process called inference will surpass training as the dominant driver of data center demand. This shift isn’t just a change in workload; it’s a fundamental transformation in how data centers are designed, powered, and located.

    Inference is what happens when you ask ChatGPT a question and get an answer. It’s the always-on, millisecond-sensitive process that powers every AI assistant, recommendation engine, and autonomous agent. Unlike training, which is a massive, one-time burst of compute, inference is a continuous, scaling workload that grows with every new user and every new model. As AI moves from a niche experiment to a mainstream utility, inference is becoming the new cloud workload—and it’s reshaping the data center industry from the ground up.

    The Shift from Training to Inference

    To understand why inference will dominate, you need to know the difference between the two phases of AI compute. Training is like building a rocket: you pour enormous resources into a single, intense project that lasts months. Inference is like launching the rocket every time a user asks a question—it’s the ongoing operation that keeps the service alive.

    Training workloads are batch-processed and can be run in centralized, high-density facilities. They’re tolerant of downtime and latency—if a training run pauses for an hour, no one notices. Inference, on the other hand, is latency-sensitive. When you ask Siri for the weather, you expect an answer in under a second. That means inference servers need to be geographically distributed, closer to the user, to minimize delay.

    NVIDIA has already reported that inference accounts for about 40% of its data center revenue, and it’s growing faster than training. Microsoft, Google, and Amazon are all building out regional edge data centers specifically for inference. The shift is not speculative; it’s happening right now.

    The Numbers Behind the Shift

    The growth projections are staggering. McKinsey estimates that AI-related data center capacity will grow from about 10 gigawatts (GW) in 2024 to 50–60 GW by 2030, with inference driving the majority of that growth. Goldman Sachs projects that data center power demand will increase by 165% by 2030, again with AI inference as the leading contributor.

    What’s driving this? Token generation—the unit of output for AI models—is growing at 3 to 5 times annually across major providers like OpenAI, Anthropic, and Google. As more applications integrate AI, from coding assistants to customer service chatbots, the volume of inference requests skyrockets. And each request consumes compute power, which translates directly to data center demand.

    Why Inference Is Structurally Different

    Inference isn’t just a smaller version of training; it’s a different beast altogether. Training clusters run at near-100% utilization for weeks, making them ideal for a few massive, centralized facilities. Inference, however, has variable utilization—peak during business hours, low at night. That variability requires over-provisioning and new scheduling techniques to handle the load efficiently.

    Hardware is also diverging. Training is dominated by NVIDIA’s H100 and B200 GPUs, but inference is increasingly using specialized chips like Google’s TPU, AWS’s Inferentia, and Groq’s LPU, which are optimized for low latency and high throughput per watt. Software is evolving too, with frameworks like vLLM and TensorRT-LLM that optimize models for inference, sometimes at the cost of making hardware obsolete faster than in the training era.

    The Rise of Agentic AI

    One of the most explosive drivers of inference demand is the shift toward agentic AI—autonomous agents that don’t just answer a single question but perform a series of tasks. Imagine an AI assistant that books a flight, reserves a hotel, and schedules meetings. Each of those steps requires multiple inference calls, multiplying demand by 10 to 100 times per user interaction.

    For example, a simple chatbot might make one inference call per query. An agentic system could make dozens, each with its own latency requirement. This is why companies like OpenAI and Google are investing heavily in agentic frameworks—they know that each agent multiplies the compute needed, and thus the revenue.

    Multimodal Models and Context Windows

    Text-only models were just the beginning. Multimodal models that generate images, audio, and video are far more compute-intensive at inference time. Video generation, for instance, is 100 to 1000 times more expensive per token than text. As these models become mainstream, they’ll add a massive new layer of demand.

    Another factor is the growing size of context windows. Modern models can now process over 1 million tokens in a single request—like reading a whole book before answering a question. The compute needed for inference grows quadratically with context length, meaning that a 1M-token context is not just 10 times more expensive than a 100K-token one; it’s 100 times more. As users demand longer, more nuanced interactions, the cost per request climbs.

    Power and Infrastructure Implications

    Inference workloads have lower power density per rack than training, but they require higher reliability and lower latency. That’s pushing data center design toward regional edge locations. AWS Local Zones and Azure Edge Zones are prime examples—smaller facilities distributed across cities, designed to bring compute closer to users.

    Power procurement is also shifting. Training facilities are the classic “megaprojects”—500 MW or more, built in remote areas with cheap land and power. Inference, by contrast, needs power where people are. That means a distributed portfolio of 50–200 MW sites across many regions. This creates new challenges for grid capacity and reliability, but also opportunities for integration with local renewable energy sources.

    The Economic Logic of Inference

    Training is a capital expense—you build it once and amortize the cost. Inference is a recurring operating expense—you pay per token, per request. That makes it a more predictable revenue stream for cloud providers and a persistent cost for enterprises. The unit economics of inference are improving about 2x per year, but demand is growing faster than efficiency gains. So even as each query becomes cheaper, total spending keeps rising.

    This is why hyperscalers are pouring $200 billion combined into AI infrastructure through 2026, even as skeptics question the near-term returns. They’re betting that inference will become the new cloud workload—the base of a multi-trillion-dollar industry.

    The Skeptic’s View: Is It a Bubble?

    Not everyone is convinced. Some analysts, like Sequoia’s David Cahn, have raised the “$600 billion question”: if inference revenue doesn’t materialize fast enough, the massive capex could be a bubble. If AI adoption plateaus or monetization fails, inference demand could disappoint.

    But the counterpoint is strong: even if consumer AI plateaus, enterprise and government adoption—in coding, healthcare, defense—provides a floor. Companies are already paying for AI copilots that boost productivity, and the ROI is measurable in some sectors. The question isn’t whether inference will grow, but how fast and how sustainably.

    The Energy and Sustainability Angle

    Inference’s distributed nature means power is needed where people live and work, not just in remote deserts. This creates tension with the current data center siting model, which often favors cheap land and abundant power over proximity to users. As cities compete for edge data centers, they’ll need to balance local power demands with sustainability goals.

    The good news is that inference workloads are often more flexible than training—they can be spread out and even shifted between locations based on grid conditions. This opens the door for smart load balancing that can reduce strain on the grid and integrate more renewable energy.

    Looking Ahead

    The era of inference is already here, and it will only accelerate. As AI becomes embedded in every software product, from spreadsheets to medical diagnostics, the demand for running models will dwarf the demand for training them. Data centers will evolve from massive, remote campuses into a web of distributed, edge facilities that bring compute to the user.

    For anyone planning the next decade of infrastructure, the message is clear: the future is not about building the biggest AI model; it’s about running it billions of times a day, reliably, cheaply, and fast. That’s the new reality of data center demand.

    The shift from training to inference is a fundamental change in the data center industry. It’s not just about new hardware or software—it’s about rethinking where data centers are built, how they’re powered, and how they serve the always-on, latency-sensitive demands of AI applications. As inference becomes the primary driver of demand, the winners will be those who can build the most efficient, distributed, and reliable infrastructure.

    Summary

    • Inference is overtaking training as the dominant AI compute workload, with NVIDIA reporting ~40% of data center revenue from inference and growing.
    • Data center capacity is projected to grow from ~10 GW in 2024 to 50–60 GW by 2030, driven largely by inference.
    • Inference is latency-sensitive and requires distributed edge data centers, unlike training’s centralized, batch-processed facilities.
    • Agentic AI and multimodal models multiply inference demand by 10–100x per user interaction.
    • Power procurement shifts from 500MW+ megaprojects to 50–200MW distributed portfolios, closer to users.

    FAQ

    Q: What is the difference between training and inference?
    A: Training is the process of building an AI model, using huge amounts of compute over weeks or months. Inference is the process of running that model to generate outputs, like answering a question or generating an image. Training is a one-time cost, while inference is continuous and scales with usage.

    Q: Why will inference drive more data center demand than training?
    A: Because inference is an always-on workload that grows with every user and every new model. Training, while compute-intensive, is finite and happens less frequently. As AI adoption grows, the number of inference requests multiplies, requiring more data center capacity.

    Q: How does inference affect data center design?
    A: Inference requires low latency, so data centers need to be distributed closer to users. This means more edge data centers in urban areas, with lower power density per rack but higher reliability requirements. It’s a shift from a few massive facilities to many smaller ones.

    Q: What is agentic AI and why does it increase inference demand?
    A: Agentic AI refers to autonomous agents that perform multiple steps to accomplish a task, like booking a trip. Each step involves an inference call, so a single user interaction can trigger 10-100x more compute than a simple chatbot query.

    Q: Is the growth in inference demand a bubble?
    A: Some analysts worry that AI revenue won’t justify the massive investment, but enterprise and government adoption provides a floor. Even if consumer AI plateaus, business use cases like coding and healthcare are expanding, so inference demand is likely to keep growing, though the pace is uncertain.