HAITONG INT'L: The price war for AI models escalates, and value rapidly shifts towards the cloud and infrastructure.
Continue to be optimistic about investment in cloud and inference infrastructure.
HAITONG INT'L released a research report stating that DeepSeek V4-Flash has further reduced the usage cost of cutting-edge models, with its API priced at just $0.14 per million input tokens and $0.28 per million output tokens. While maintaining strong comprehensive capabilities, it has driven the unit task cost down to an industry low. The investment outlook remains positive on cloud and inference infrastructure. Caution is advised regarding the model layer, as price wars will further weaken the scarcity and pricing power of individual APIs; cloud platforms will benefit from the growth in model invocation and the demand for multi-model deployments; on the infrastructure side, attention should be paid to the inference chip, server, network, and data center industry chain.
The main points from HAITONG INT'L are as follows:
The competition among models is shifting from parameter scale to post-training and system engineering.
V4-Flash has not achieved its capability leap by expanding parameters; rather, it relies on post-training and the Harness framework to strengthen task planning, tool invocation, and error recovery. In the future, the actual effectiveness of models will increasingly depend on the overall combination of model + Harness + tools + data, making it more challenging to establish barriers relying solely on the capabilities of the base model.
The pricing power of the model layer will remain under pressure.
As medium-strength models can complete most standardized tasks at extremely low prices, enterprises will more frequently switch between different models, and high-end closed-source models can only maintain a premium in complex reasoning, high-reliability code, and critical tasks. Model vendors may enter a phase of increased input and continuously declining token prices.
Low-cost models will accelerate the penetration of Agent and AI applications.
Agents typically require multiple rounds of reasoning, tool invocation, and failure retries; model price reductions will significantly improve the unit economic model for automating functions such as AI programming and marketing in middle and back offices. Competitive focuses for application companies will return to customer interfaces, industry data, and workflow closures, with products that can take responsibility for the final task outcomes holding more pricing power.
DeepSeek V4 and Kimi K3 are driving the overseas model market into a new round of price competition.
Support for the open-weight ecosystem overseas continues to warm up, and Chinese models are accelerating their entry into the model selection range of overseas developers and enterprises due to their near-state-of-the-art capabilities, open deployment, and significant cost advantages. Kimi K3 is pushing open model capabilities further closer to overseas closed-source flagships, while DeepSeek V4-Flash has driven the unit task cost down to an industry low; coupled with OpenAI's recent proactive reduction of prices for GPT-5.6 Terra and Luna, overseas model vendors may shift from merely competing on capabilities to balancing capabilities, price, and ecosystem, while model API prices and gross margins will still face pressure.
Price reductions in models will drive growth in cloud usage through the Jevons Paradox.
Although high-efficiency models reduce the computational consumption per task, lower prices will also expand the user base, invocation frequency, and economically deployable scenarios, accelerating the growth of token consumption and inference demand. As enterprises access more open-source and closed-source models simultaneously, cloud platforms will become the core entry point for unified access to computing power, model deployment, and call management. The cheaper the models are and the more widespread the applications, the more evident the benefits to cloud revenue may become.
Risk warnings: 1) AI demand may not meet expectations; 2) Geopolitical environment may disrupt the supply chain; 3) Construction of AI data centers may slow down.
Related Articles

FE HORIZON (03360) will distribute an interim dividend of HKD 0.25 per share on September 29.

FE HORIZON (03360) announced its interim results, with profit attributable to shareholders of 2.222 billion yuan, an increase of 2.68% year-on-year.

CATHAY PAC AIR (00293) will distribute an interim dividend of HKD 0.26 per share on October 8.
FE HORIZON (03360) will distribute an interim dividend of HKD 0.25 per share on September 29.

FE HORIZON (03360) announced its interim results, with profit attributable to shareholders of 2.222 billion yuan, an increase of 2.68% year-on-year.

CATHAY PAC AIR (00293) will distribute an interim dividend of HKD 0.26 per share on October 8.

RECOMMEND





