Lates News
Luo Fuli, head of Xiaomi's MiMo team, posted on X, publicly disclosing for the first time the reinforcement learning training progress of Xiaomi's new model MiMo-V2.6. Luo Fuli said that for nearly the past six months, the team has been researching the question of how far reinforcement learning can still be scaled. Currently, MiMo-V2.6 is in the process of reinforcement learning training, and the team is scaling up in three directions, including computing resources, the training environment and task system, as well as the reward evaluation mechanism.
Latest
4 m ago

