NaiveAI has introduced the open-weight Naive-N0.5-Flash model on the Hugging Face platform. This model features a Mixture-of-Experts (MoE) architecture designed for software development tasks and artificial intelligence research. Information about the model was published in mid or late September 2026.
The official product name is NaiveAI / Naive-N0.5-Flash. According to technical documentation, the total number of parameters is approximately 309 billion, with 15.5 billion weights being active. The model is distributed under the MIT license and includes inference code, as well as native support for a one-million-token context window.
The context length is achieved through a hybrid combination of the Sliding-Window Attention (SWA) mechanism and lightweight DeepSeek Sparse Attention (DSA), which utilizes grouped attention. The architecture is described as having a predominant SWA–DSA ratio of approximately 5:1, with no fully attentive layers in the stack. The model is based on the open-source MiMo-V2.5 model, followed by fine-tuning and post-training, rather than full pre-training from scratch.
In terms of serving, NaiveAI emphasizes NaiveRT—an inference stack developed using AI-oriented research. This stack integrates mega-kernel fusion, Programmatic Dependent Launch, and speculative decoding. According to the vendor, processing speed reaches approximately 50 tokens per second per user in Standard mode and up to 2000 tokens per second in Ultra-Fast mode.
Tags on Hugging Face highlight that the model is suitable for code text generation, long-context handling, and AI research tasks. This aligns with the stated concept that fine-tuning and post-training on a strong open base can push boundaries for coding agents. For developers comparing Chinese MoE releases this week, Naive-N0.5-Flash is a specific example of a 309B / 15.5B MoE with a million-token hybrid context under the MIT license, representing an architectural announcement rather than a cost or product launch news.
