CMSC: In-depth tracking of the domestic computing power chip chain: Ascend 960 progress and NPO exceed expectations.

date
10:50 21/09/2026
avatar
GMT Eight
Recommend paying attention to related targets in foundry, GPU/ASIC, network chips and optical chips, packaging and testing, equipment, components, storage, and other related segments.
CMSC released a research report stating that the development pace of Ascend 960 has been moved forward, and domestic computing power is accelerating toward system-level collaboration across "compute, interconnect, storage, and management." It is recommended to focus on investment opportunities in the AI computing power industry chain and the semiconductor self-controllable track. At this conference, the Ascend 960 timeline exceeded expectations, Hi-ONE mass production is advancing, and the 4096-card super node and million-card cluster plans are expected to simultaneously drive demand in areas such as computing power chips, advanced packaging and HBM, high-speed optical interconnect/NPO, network switching, and AI storage. It is recommended to focus on targets in related areas such as foundry, GPU/ASIC, network chips and optical chips, packaging and testing, equipment, components, and storage. CMSC's main views are as follows: Huawei will focus on AI infrastructure in the future, building the silicon-based black soil of the intelligent world Huawei's AI strategy mainly focuses on seven aspects: 1) The core of Huawei's AI strategy is computing power, and it insists on hardware monetization; 2) "Super node + cluster" to build China's solid computing power foundation and create a new choice for the world; 3) Build an open-source and open computing power ecosystem, support native training of mainstream large models, and actively embrace "hundreds of models and thousands of forms"; 4) Pangu large models are used for the intelligence of Huawei products to enhance the competitiveness of its own products; 5) Provide flexible offline + online computing power solutions to accelerate industry intelligence; 6) Innovate diverse computing power, promote AI into devices and vehicles, making intelligence ubiquitous; 7) Build a new generation of communication networks around "using computing" to deliver intelligence to every person, home, and enterprise. The Lingqu series' 11-chip family was unveiled, the readiness time point for Ascend 960 was moved earlier, and DT was advanced to 27Q1 Based on Lingqu UnifiedBus, super nodes and super node clusters are built, and 11 key chips have been developed, including: 1) Compute: general computing processors (Kunpeng), intelligent computing processors (Ascend); 2) Interconnect: Lingqu interconnect chips, NPO high-density optical engines, network interface card chips; 3) Storage: intelligent storage chips, hard disk control chips, storage control chips; 4) Management: Lingqu management chips, device management control chips. Regarding Ascend chips: 1) Ascend 960DT: FP8 computing power 2PFLOPS, memory access 288GB, bandwidth 9.6TB/s, interconnect bandwidth 2.2TB/s. Originally planned to be ready in 27Q4, now advanced to 27Q1; 2) Ascend 960PR: FP8 computing power 2PFLOPS, memory access 192GB, bandwidth 2.4TB/s. Originally planned to be ready in 27Q4, now advanced to 27Q3; 3) Ascend will maintain a one-generation-per-year evolution pace. The 970 and 980 chips will be launched in 2028 and 2029. The 970 is expected to have FP8 computing power of 3.6PFLOPS, memory access 288GB, and bandwidth 14.4TB/s. The 980 is expected to have FP8 computing power of 7.2PFLOPS, memory access 384GB, and bandwidth 38.4TB/s. The Lingqu architecture is paired with the Hi-ONE optical product series, and the 960 super node will become the first product to adopt NPO At the conference, the industry's first mass-produced optical engine was also released, using "Lingqu UnifiedBus + Hi-ONE" to build a scalable super node interconnect system, pioneering the SiN-SOI silicon photonics hybrid process, the industry's first NPO with a built-in light source, and the industry's largest transmission capacity of 7.2Tbps. The industry's first super node using NPO, the 960 super node, was released. A single super node is at the 4096-card scale, can provide a maximum of 8EFP8 computing power and up to 1PB of HBM capacity, and uses 5,500 Hi-ONEs, replacing the 48,000 800G optical modules originally required, reducing power consumption by more than 550 kilowatts. The Kunpeng series is upgraded again, and the AI memory storage cluster and Agentic super node cluster help 10-trillion models The Kunpeng super node has been fully upgraded. Based on Lingqu all-optical networking, the Kunpeng super node supports a maximum of 4096 nodes and forms a unified memory pool of up to 256TB, which can significantly accelerate Agent performance. Based on Lingqu, the AI memory storage OceanStor M900 cluster is built. Huawei has proposed a multi-layer KV cache architecture and launched Huawei's L3.5-layer PB-level KV cache based on Lingqu UnifiedBus and supporting "one-hop direct connection." Based on Lingqu UnifiedBus unified interconnect, multiple interconnect protocols are unified into the Lingqu protocol, greatly reducing protocol conversion overhead and enabling equal interconnection among subsystems such as Ascend super nodes, Kunpeng super nodes, and KV cache clusters. Through two-layer CLOS four-plane networking, the maximum cluster scale can reach 512,000 cards. Combined with multi-track topology technology, it can support a maximum of 1 million-card Ascend super node clusters. Risk warnings: terminal demand falling short of expectations; macroeconomic environment impact; semiconductor domestic substitution progress falling short of expectations; intensifying industry competition, etc.