Article

Dai Jifeng’s NaiveAI open-sources its first model:AI participated in 6 days of 151 rounds of optimization

Beating AI flash news, NaiveAI, founded by Tsinghua University associate professor Dai Jifeng, released Naive-N0.5-Flash and made the model weights available. The model was adapted from Xiaomi MiMo-V2.5 Base and open-sourced under the MIT protocol.

NaiveAI retained the original MoE architecture, with a focus on changing the attention mechanism. The team replaced global attention with sliding-window attention and DeepSeek sparse attention, then continued training 3250000000000 tokens. The adapted model natively supports a 1000000-token context, reducing the computational cost of processing extremely long text.AI also directly participated in the research and development. It wrote code, ran experiments, analyzed results, and continued optimization, while researchers set the direction and made key decisions. Using the inference system NaiveRT as an example, the team conducted 6 days of 151 rounds of optimization experiments, of which 63 rounds were adopted.

The officially reported peak inference speed was 2122 tok/s, but the test conditions were unusual: using 8 GPU, with thinking mode disabled, excluding input-processing time, and taking the best result from 41 requests at 1 seconds. In standard mode, the official figure was approximately 50 tok/s per user.

Source: https://m.theblockbeats.info/flash/369256