The end of AI is "high-quality electricity"! Under the surge of 945 terawatt-hours of demand, "power system stability" is taking over as the top priority for AI infrastructure.

date
14:57 06/08/2026
avatar
GMT Eight
The unstable power demand from artificial intelligence is harming its own data centers. Equipment failures indicate that these infrastructures, worth billions of dollars, have numerous cost and reliability issues.
The immense power demand of terawatt-scale AI training/inference clusters is already well known. However, what is less understood is how the rapid fluctuations in power demand from data centers can damage critical equipment in these massive infrastructures. In the rapidly developing AI infrastructure, high-capacity batteries, generators, cooling devices, and other key systems are under such immense pressure that they frequently fail or reach their end of life prematurely. As the boom in artificial intelligence accelerates, these technical issues mean additional costs and previously unanticipated reliability risks for data center power equipment; even a few minutes of lost uptime can significantly impact the profits data center developers can earn. Meanwhile, investors and lenders are already uneasy about the staggering expenditures, reaching hundreds of billions of dollars, related to single hyperscale cloud computing enterprises. There's growing concern in the market that the depreciation rate of these data centers' underlying hardware infrastructures may far exceed consensus expectations from prior analyses. Recent developments in the narrative that "the intense power demand from AI is damaging its own data centers" suggest that the valuation anchor for the AI infrastructure construction frenzy is shifting from simply counting the "number of GPUs/TPUs" to considering power quality, dynamic stability, lifespan of power chain equipment, and effective uptime. The true hidden bottleneck of AI data centers appears to have shifted from "Can enough power be obtained?" to "Can they withstand millisecond-level, urban-scale power fluctuations?": the simultaneous activation of hundreds of thousands of AI GPUs or dozens of NVIDIA Vera-Rubin and AMD Helios high-power AI racks can repeatedly impact gas turbines, generators, batteries, transformers, UPS systems, and cooling systems, causing equipment to fracture, wear abnormally, and face early retirement. If failures result in downtime and supply interruptions for computing power infrastructure, the lost cost is not merely that of replacing parts, but the expensive GPU assets are unable to generate profits, and these dynamic loads may also threaten the stability of external grids through voltage, frequency, and subsynchronous oscillations. This is why the NERC (North American Electric Reliability Corporation) has deemed the rapid load changes of large data centers as significant reliability risks. The end of AI truly is power is becoming an increasingly hot narrative! Megawatt-level racks are reshaping the investment landscape of data center equipment. The saying "The end of AI is power" is evolving further into "The end of AI is stable and controllable high-quality power"energy storage, UPS, high-voltage direct current, power electronics, microgrids, and advanced cooling equipment will gain higher value, and data center project valuations must account for previously underestimated hidden costs such as accelerated depreciation of equipment, downtime losses, and reliability upgrades. The ultimate constraint for AI is the ability to deliver high-quality usable power that is stable, continuous, and responsive with extreme dynamic speed to chips. Traditional CPU rack power densities are relatively low, and load variations are gentle, while systems like Blackwell couple dozens of GPUs, CPUs, NVSwitches, and high-speed network cards synchronized: the GB200 NVL72 has a full load power of about 120 kW, while the GB300 NVL72 can reach up to 142 kW; when training tasks enter collective communication, checkpoint saves, or inference batch switches, the whole rack of GPUs simultaneously raises or lowers power in milliseconds, creating immense surges, harmonics, and heat loads. NVIDIA expects to support IT racks of one megawatt and above starting in 2027; traditional 54-volt DC distribution will face physical limits in busbar size, transmission losses, and conversion stages. Therefore, each increase in AI cluster computing power requires not only buying more GPUs but also simultaneous expansion of medium voltage transformers, switchgear, busbars, rectifiers, UPS, battery storage, supercapacitors, backup generators, and cooling liquid distribution units; power conversion, distribution stability, and heat dissipation systems are transitioning from auxiliary facilities to core production materials that determine whether GPUs can power on. The strong demand for rack-level AI computing power from global enterprises for Vera Rubin and AMD Helios further indicates that the next generation of competition has shifted from "chip-level performance" to a full rack collaborative design involving "computingpower supplynetworkliquid cooling." The Vera Rubin NVL72 incorporates power smoothing mechanisms and provides approximately six times the local energy buffer of Blackwell racks, using batteries or capacitors to instantaneously supply energy during load surges and absorb excess energy during drops, avoiding direct transmission of GPU power spikes to transformers and the grid; the 800-volt DC architecture reduces current under the same power, minimizing the volume of copper, cabling, and multi-stage AC/DC conversion losses. AMD Helios integrates 72 MI455X GPUs, 31TB HBM4, centralized power management bus, and liquid cooling manifolds into a rack-level system. This is why the higher the AI computing density, the more rapidly the value and technical thresholds of power equipment increase; the structural beneficiaries are expanding from simple power generation capacity to transformers and switching equipment, high-voltage DC power, UPS and energy storage, power quality management, liquid cooling, and microgrids, while data centers lacking dynamic load management may see capital returns eroded by accelerated equipment depreciation and downtime losses. Wall Street financial giant Goldman Sachs estimates that global data center power demand will grow crazily by 220% by 2030 compared to 2023, equivalent to adding a country among the worlds top ten power consumers. The IEA (International Energy Agency) projected that global data center electricity usage would rise from around 415 terawatt-hours in 2024 to about 945 terawatt-hours by 2030, averaging an annual growth rate of around 15%, which would account for nearly 3% of global electricity consumption; in particular, AI-accelerated server electricity use is expected to grow at an annual rate of about 30%, with U.S. data center energy consumption increasing by about 240 terawatt-hours compared to 2024, a growth rate of about 130%, contributing nearly half of the U.S.'s additional electricity demand by 2030. Millisecond power fluctuations, AI is rapidly depleting its own infrastructure. Amber Villagas-Williamson, a senior consultant at the Uptime Institute, which provides standards and reliability consulting for energy suppliers and data centers, stated, "AI does indeed generate very abnormal power demands. It's like running a car engine at super high RPM continuously, which will wear out the engine faster than operating at a constant speed." Data centers have existed for decades, consuming vast amounts of electricity to ensure everything from your favorite streaming shows to online grocery orders runs smoothly. However, the underlying computing infrastructure designed for AI has notable differences due to the sheer scale of its power demands and the magnitude of fluctuations involved. The load, equivalent to that of factories, towns, or even cities, can appear or disappear within seconds, creating repeated shocks that make it challenging for connected equipment to endure. Shannon Miller, founder and president of Mainspring Energy Inc., which develops microgrid projects for industrial and data center clients, noted that the power consumption of a 1-gigawatt data center is equivalent to a city like Boston, and half of that load could be toggling on and off every few seconds. Some AI parks planned in Texas and the Midwest are even larger, with average power consumption nearly matching that of New York City. As shown in the graph, sales data for AI chips indicate a surge in data center power demand projections64% of the newly added AI electricity consumption from 2022 to 2033 will come from the U.S. Note: the BloombergNEF model estimates the implied power demand of AI chips, including computing, networking, and cooling infrastructures. When AI data centers train new models, they put especially immense pressure on the power supply systemsthis process activates all GPUs simultaneously. Just like a swarm of bees or a school of fish suddenly changing direction in the digital world, hundreds of thousands of GPUs can simultaneously increase or decrease power in milliseconds. Drew Baglino, a former Tesla executive who founded Heron Power Electronics Co., noted that the power consumption of AI facilities can sometimes soar to 50% above design capacity, stating, "Therefore, a 1-gigawatt facility could consume 1.5 gigawatts of electricity in an extremely short moment." The company is developing equipment to manage power fluctuations in coordination with NVIDIA's next-generation servers expected to consume even more energy, launching in 2027. Most equipment is not designed to handle such drastic power fluctuations. Jon Pirella, CEO of energy storage developer Terraflow Energy, likened it to shifting from sixth to first gear in a Ferrari; he shared, "You can't switch that quickly." This article is based on interviews with over thirty electricity experts from the U.S. and Europe, including representatives from power generation companies and other energy suppliers, data center developers, grid operators, utility companies, investors, standard makers, insurance companies, and regulatory agencies. Nearly all respondents indicated that the physical stress on these facilities has become very evident. Several respondents reported that small natural gas internal combustion engines used to power data centers have experienced broken crankshafts. One individual indicated that gas turbines in the xAI Colossus computing facility in Memphis, Tennessee, had cracked. This individual noted that batteries were subsequently installed in the system to help smooth out power fluctuations and alleviate pressure on the rotating turbines. xAIs parent company, SpaceX, did not immediately respond to media requests for comment. Andrew Cunningham, CEO of GeoPura Ltd., stated that even much smaller data centers in the UK have experienced turbine cracking. The company is providing hydrogen for fuel cells at some locations to smooth out power flow. Jennifer Scanlon, CEO of UL Solutions Inc., which tests and certifies new technologies, explained that cracks or wear in equipment can lead to arcingcurrent jumping between conductorspotentially damaging AI chips. A complete set of equipment, including batteries, capacitors, transformers, and flywheels, can help stabilize power flow. However, multiple respondents indicated that newly constructed data centers have not adequately adopted these technologies in the race to build AI computing power. According to sources at Uptime Institute and others collaborating with operators, some batteries installed for this purpose have had to be replaced within months or even weeks due to excessive stress. Villagas-Williamson from Uptime Institute stated that this issue has manifested in data centers globally, from the Middle East and Africa to Europe and the U.S. Downtime rates erode cash flow, and the reliability of AI computing power infrastructure is beginning to emerge as a financial risk. These problems have led to some AI computing facilities experiencing project delays or operational restrictions, thereby reducing revenue. Chris James, CEO of Joulent Inc., noted that to ensure the AI park in West Texas, with a planned capacity of 2.67 gigawatts, meets Microsoft's requirement for 99.999% reliability, additional time has been allocated in the engineering plan. The company is developing this facility in partnership with energy giant Chevron. This means that the power supply will be delayed from the originally planned start in 2027 to 2028. Jason Hoffman, chief strategy officer at data center construction and operations company Switch, stated that if key equipment fails prematurely, "the financial consequences are primarily not the cost of replacing a pump, breaker, or some power component, but rather the value lost from expensive computing power that cannot generate revenue due to downtime." The revenue loss from downtime varies widely; estimates range from a few thousand dollars to several hundred thousand dollars per minute, depending on the facility type and the workloads running. A person involved in financing such facilities stated that the fundamental assumption for data center construction is that once operational, they will run 365 days a year, around the clock. However, in reality, some facilities achieve only about 80% uptime; if this issue remains unresolved, investors in some projects may face impacts within the next 12 to 24 months. Any reliability issues will only further exacerbate market concerns about whether the hundreds of billions of dollars invested in AI will generate corresponding returns. Another critical piece of equipment in data centersthe GPU racks themselveshas also been subject to scrutiny regarding the speed of depreciation: does this industry truly have the capability to achieve the promised profitability? These reliability issues may disrupt broader power systems. The extensive network consisting of high-voltage transmission lines, transformers, and power plants requires constant calibration, yet this work is becoming increasingly difficult each year due to aging equipment, rising power demands, and extreme weather conditions. Intermittent wind and CECEP Solar Energy generation are expanding, often causing power supply to fluctuate dramatically within the hour, which has already become a source of instability for the grid. AI data centers could potentially amplify such fluctuations significantly. Srijan Roy, an electrical power quality expert at Schneider Electric based in Nashville, Tennessee, stated, "These loads are highly dynamic, or fluctuate dramatically, which can destabilize the grid; left unaddressed, this could lead to blackouts or power interruptions." He expressed particular concern that data centers might induce subsynchronous oscillations in the power flow, which could damage equipment connected to other parts of the grid. Roy stated, "This has already raised significant concerns among global utilities." In the past two years, the highest regulatory body responsible for establishing U.S. electricity reliability standardsthe North American Electric Reliability Corporationhas repeatedly issued warnings and alerts, asserting that data centers are among the greatest risks facing grid stability. According to a report published in September, the North American Electric Reliability Corporation evaluated over 33 gigawatts of operational data centers in the U.S. and found that about three-quarters of their load models "do not adequately reflect the dynamic behavior of data centers." Earlier this year, the organization issued a rare Level 3 alert urging large data centers to address these urgent risks and submit responses by August 3. From "false calculations" to energy storage buffering, the AI computing industry chain is beginning to completely reconstruct the power supply architecture. From upstream to downstream in the industrial chain, the AI industry has recognized these issues and is actively researching solutions. Dion Harris, senior director of large-scale infrastructure solutions at NVIDIA, stated that when developing the Blackwell GPU, NVIDIA began collaborating more closely with power experts. Blackwell is set to launch in 2024 and is now widely deployed in data centers. Harris noted, "We are not just manufacturing chips and processors"; we are also using these products in NVIDIA's own data centers. NVIDIA is working to make the deployment process smoother, "including in the construction, design, and engineering phases of data centers, as well as in the power supply segment." Data center users have adopted some technologies to smooth out power fluctuations of AI workloads, such as running auxiliary computationsessentially ineffective mathematical operations unrelated to the training processto maintain GPU stability. However, during surges in power demand, this approach has been criticized for wasting electricity. Martha Simko-Davis, a project manager at the U.S. Department of Energy's Rocky Mountain National Laboratory, stated that last year, the lab established a test platform near Denver, Colorado, representing the Energy Department to explore how to safely integrate AI into the power grid. She manages related projects within the Office of Electricity. This test site is equipped with GPUs and on-site power generation equipment, allowing developers to determine whether their systems can withstand the fluctuations of AI loads. One power supplier indicated that they would use this facility to test batteries, software, and other devices to mitigate oscillations that could damage both data centers and the grid. Simko-Davis stated, "We now have the opportunity to get this right."