InclusionAI releases Ling-3.1-flash, with a clear increase in parameter scale and per-token activation scale compared with the previous generation. The official demonstration showed it completing long-duration programming tasks, with plans to make context and an open-source model available later.
OpenRouter The model page indicates that the anonymous model Space Bunny Alpha September 23 launched in preview, supporting million-Token context, multimodal input, and tool calling. Both input and output are currently free.
Xiaomi's MiMo team publicly stated that the next generation MiMo-V3 will adopt the HySparse2 architecture to reduce long-context Agent prefill computation and KV Cache usage. Tests in the article state that with a million-token input, prefill compute is reduced to approximately the original 1/5。