Zheshang: "Cost per unit task completion" will become the core competitive focus of AI cloud; bullish on long-term opportunities brought by full-stack optimization and scale advantages.
The bank is bullish on the long-term investment opportunities brought by full-stack optimization and scale advantages amid robust demand for AI cloud.
Zheshang released a research report stating that with the continuous expansion of large model inference scale and technological iteration, "cost per unit task completion" will become the core competitive focus of the AI cloud business. Full-stack advantages and scale advantages constitute the two major core competitiveness of the AI cloud business. The bank is bullish on long-term investment opportunities brought by full-stack optimization and scale advantages under the high-prosperity demand for AI cloud. As the industry deepens, the core of AI cloud competition in the future will focus on "unit task cost" competition. Hyperscalers and Neoclouds, two types of vendors, rely on their respective differentiated foundation advantages to continuously optimize cluster utilization, reduce unit Token costs, and amplify the value of computing power assets and profit elasticity.
Zheshang's main views are as follows:
Core competitiveness of cloud business: full-stack advantages + scale advantages
Full-stack technology fully taps the efficiency of every unit of computing power, while scale advantages improve cluster utilization and minimize idle computing power. Large CSPs will rely on horizontal scale dividends and vertical full-stack integration capabilities to maintain industry leadership; while Neoclouds are expected to build their own differentiated competitive barriers by virtue of the specialized architecture of pure AI clusters and the efficiency gains brought by homogeneous traffic from third-party model inference.
Core competitiveness of cloud business 1: full-stack advantages
Core of full-stack advantage 1: from "local optimum" to "system optimum." Insufficient GPU utilization in computing clusters usually stems from the significant "bottleneck effect" in computing clusters. Through full-link global coordination, the overall system optimum can be achieved, losses in each link can be reduced, and comprehensive efficiency that cannot be achieved by local optimization can be obtained. Google's full-stack solution and the practice of collaborative tuning between Nvidia and Nebius both reflect the performance leap brought by global optimization.
Core of full-stack advantage 2: providing customers with higher added value. The bank divides AI cloud services from the bottom to the top into four major layers: physical infrastructure, managed infrastructure, managed inference services, and agent optimization layer. Full-stack capabilities drive AI infrastructure to upgrade toward high-value-added segments. On the one hand, global technology optimization improves resource utilization; on the other hand, it helps vendors achieve diversification and dispersion of customer structure.
Compare Nebius and CoreWeave from a full-stack perspective; the two are leading emerging AI cloud vendors in Europe and the United States and are deeply tied to Nvidia. Both focus on AI computing power workloads, but their full-stack evolution paths have diverged significantly: CoreWeave is mainly at the second layer (managed multi-tenant infrastructure), while Nebius is at the third layer (managed inference TokenFactory), with lower customer concentration and more complete full-stack technology.
Core competitiveness of cloud business 2: scale advantages
Core of scale advantage 1: eliminating traffic fluctuations through the law of large numbers and improving cluster utilization. After the scale of computing power and business traffic expands, massive inference traffic can smooth tidal load fluctuations, improve GPU cluster utilization, and compress marginal inference costs. Large-scale clusters do not eliminate traffic fluctuations, but rely on the law of large numbers to hedge risks. Technologies such as GPU partitioning, continuous batching, Prefill/Decode decoupling, and resource reuse all need to rely on large-scale tasks to fully release performance, and optimization effects are limited in small-scale traffic scenarios.
Core of scale advantage 2: decline in marginal cost (TCO). 1) Lower procurement costs. Large-scale vendors rely on procurement bargaining power and ODM direct procurement models to hardware procurement costs; 2) Continuously lock in electricity, and at the level of power energy consumption, the expansion of computing power scale and improvement in utilization dilute fixed energy consumption such as power supply and idle heat dissipation, driving PUE downward; 3) Dilute fixed costs such as team labor. From 2023 to 2025, the share of personnel costs in Google Cloud revenue continued to decline, and profit margins improved accordingly.
How to understand the respective competitive advantages of Hyperscalers and Neoclouds?
1.
Core competitiveness of CSPs. 1) Horizontal scale advantage. In 2025, the cloud business revenue of the top three CSPs all reached the tens of billions of dollars level, with AWS and Azure even achieving hundreds of billions of dollars in revenue; Neocloud vendors Nebius and CoreWeave had revenue of only $530 million and $5.13 billion, respectively. 2) Vertical full-stack integration advantage. Large cloud vendors are still actively self-developed chips on top of their existing cloud business layouts, continuously building solid competitive barriers with complete full-stack capabilities.
2.
Differentiated competitiveness of Neoclouds. Although CSPs occupy the main market share in the AI cloud track, in 2025 Neoclouds represented by Nebius and CoreWeave still achieved rapid growth, with revenue growth rates of 351% and 168%, respectively. The high growth of Neoclouds mainly stems from the current market background of a large supply-demand gap in computing power, their flexible organizational mechanism advantages, and the mutual checks and balances among multiple industry roles. In the future, Neoclouds are expected to build their own differentiated competitive capabilities by relying on the low TCO advantages brought by pure AI clusters, combined with the homogeneous traffic efficiency advantages of third-party model inference.
Risk warnings: AI technology iteration falling short of expectations; fluctuations in the computing power supply-demand pattern; intensifying industry competition; fluctuations in customer demand; policy and compliance risks.
Related Articles

Guotai Haitong: The baijiu industry still has structural opportunities; prefer leading players at each price point.

HK Stock Market Move | China United Network Communications (00762) rose over 4% in early trading. China United Network Communications releases AIDN innovative architecture, accelerating intelligent network upgrades.

HK Stock Market Move | BAIC MOTOR (01958) rose over 9% at one point this morning as expectations for auto industry consolidation heat up; the company turned from profit to loss year-on-year in the first half.
Guotai Haitong: The baijiu industry still has structural opportunities; prefer leading players at each price point.

HK Stock Market Move | China United Network Communications (00762) rose over 4% in early trading. China United Network Communications releases AIDN innovative architecture, accelerating intelligent network upgrades.

HK Stock Market Move | BAIC MOTOR (01958) rose over 9% at one point this morning as expectations for auto industry consolidation heat up; the company turned from profit to loss year-on-year in the first half.

RECOMMEND





