Article
Xiaomi first disclosed MiMo-V3 a new architecture: prefill compute for million-token inputs reduced 80%Beating AI news reported that Xiaomi's MiMo team disclosed the next-generation MiMo-V3 will adopt the HySparse2 architecture. It mainly addresses the problem of Agent becoming increasingly expensive.Agent sends only very short instructions each time, but may retrieve large amounts of content from webpages, terminals, and tools. The model must first read all this new content; the longer the context, the more compute this step consumes,KV Cache and the more GPU memory it occupies.HySparse2 first splits the model into a front half and a back half. The front half processes the input, while the back half continues inference and generation. The context information needed by the back half can be generated directly from the results already computed by the front half, so prefill does not need to run the back half in full again. This idea comes from Microsoft's Research Institute's 2024 proposal of YOCO; its name is “You Only Cache Once.” The core idea is to let the back half share information already computed by the front half, storing less cache and performing less duplicate computation.
It also redesigns sparse attention. Ordinary full attention must examine the entire context each time, while HySparse2 allows only a small number of layers to do so. These layers first select the most relevant 1024 items token from the long context, and subsequent layers directly reuse these results while always retaining the most recent 128 items token. The previous-generation HySparse grouped every 64 items token into a group before selecting; as long as one group contained important information, the entire group had to be retained. It now selects token individually. With the same amount of computation, more genuinely relevant content can be retained. In separate tests of this change, the scores on two long-text retrieval benchmarks increased by 6.57 points and 8.14 points, respectively.
The paper uses a model with a total of 80000000000 parameters and approximately 3000000000 activated parameters per step for comparison. The 49-layer model needs to run only the first prefill layers during the 25 stage. When the input reaches 1000000 token, compared with the hybrid sliding-window attention used by the MiMo-V2 series,prefill compute is reduced to approximately 1/5,KV Cache from 12. 09GB to 2. 69GB.
Original link https://m.theblockbeats.info/flash/368764