UNISOUND(09678) releases its new-generation high-performance flagship model U2 Flash

date
08:29 15/09/2026
avatar
GMT Eight
Unisound (09678) announced that the Company has officially released U2 Flash, a new-generation high-density intelligent model built through reinforcement post-training based on the U2 general-purpose foundation model.
UNISOUND(09678) announces that the Company has officially released U2 Flash, a new-generation high-density intelligent model built on reinforced post-training of the U2 general-purpose foundation model. Compared with the previous-generation U2 model, U2-Flash achieves significant improvements in task completion quality and overall efficiency: on the DeepSWE v1.1 benchmark, which represents coding capability, U2-Flash doubled its score versus the previous-generation model, surpassing models such as GLM5.3-Flash and DeepSeek-V4-Pro-0813 with a score of 64.6; its TerminalBench3.0 score rose sharply to 24.3, exceeding trillion-parameter-class models such as K3; and its SWE-Bench Pro score reached 61.6, an increase of 10.5 points over the previous generation. At the same time, the model achieves comprehensive optimization in end-to-end execution efficiency and inference cost, reducing the number of Agent task iteration steps by 20%30%, shortening the task execution cycle by 35%, and cutting Token consumption by 20%30%. U2-Flash adopts a sparse Mixture-of-Experts (MoE) architecture, with approximately 266B total parameters and approximately 10B activated parameters per inference. Its average time to first token is controlled within 3 seconds, and its peak output throughput reaches up to 300 Tokens/s. The model unifies capabilities across multiple domainsincluding coding, agents, mathematical reasoning, and instruction followinginto a single set of weights, achieving main-force-level task completion quality with an activation scale far smaller than comparable dense models. The model features implicit thinking and continuous-state reasoning mechanisms, and encapsulates reasoning intensity into a four-level controllable interface. At the high-intensity level, it outputs a complete and readable reasoning chain, ensuring that the reasoning process is verifiable and traceable. In addition, U2-Flash has completed systematic adaptation to mainstream domestic computing power platforms, providing more flexible computing options for large-scale deployment. In terms of training methodology, U2-Flash establishes an autonomous closed-loop mechanism in which the model deeply participates in its own training, marking the Company's first step in the direction of recursive self-improvement (RSI): the model can participate in task generation, trajectory analysis, and error-correction resampling, autonomously building a high-quality SWE task set of nearly 100,000 tasks; through innovations such as asynchronous Agent RL, online policy distillation from multiple self-training teacher models, and adaptive dynamic task generation, the number of effective training trajectories increases by approximately 60%, while the number of training steps decreases by approximately 55%, and autonomous inspection and repair of the training system can be achieved. All autonomous adjustments are carried out within manually set sandbox environments and validation standards, with complete, rollback-capable records retained. The release of U2-Flash marks the Company's large-model technology moving from "human-led model training" to a new stage in which "the model participates in its own evolution." Through the recursive self-improvement mechanism, the model can continuously self-verify, self-correct, and self-enhance in real tasks, laying the foundation for long-term compounding improvements in model capabilities. In the future, the Company will adhere to a technology roadmap prioritizing validation, full-process traceability, and clear boundaries. Under the premise of ensuring safety and controllability, the Company will steadily expand the scope of model participation in its own capability growth, promote the large-scale implementation of high-density intelligence in more real production scenarios, and create long-term value for shareholders.