NaiveAI open-sources 309B model whose hybrid attention architecture AI designed itself
Naive-N0.5-Flash, a 309-billion-parameter open-weight model that Beijing startup NaiveAI released on September 27, 2026, has no full-attention layers anywhere in its 48-layer stack. NaiveAI's own account of how the model got built is the sharper story: the company says its AI systems, not its human engineers, explored and designed that architecture.
What
Naive-N0.5-Flash is a mixture-of-experts model with 309 billion total parameters and 15.5 billion active parameters, built on Xiaomi's open MiMo-V2.5 base and released under the MIT license, per NaiveAI's model card. The 48-layer network mixes 39 sliding-window-attention layers with 9 DeepSeek Sparse Attention layers, a layout the card describes as predominantly 5-to-1, and none of the layers use full attention. It still supports a native 1-million-token context window. NaiveAI's hosted API prices the model at $0.10 per million input tokens, $0.40 per million output tokens, and $0.01 per million cached-token reads, and the company's NaiveRT inference stack claims up to 2,000 tokens per second in its fastest mode, against 50 tokens per second per user in standard mode.
On its technical blog, NaiveAI describes how the model was made: "AI explored and designed its hybrid attention architecture while optimizing its training, inference, and deployment systems," with human researchers limited to setting direction, defining constraints, and making final calls, according to the post. The company says the AI-centered research loop runs on close to 10 million sandboxed compute environments a week, with 100,000 active concurrently at peak. After the architecture change, the model went through 3.25 trillion tokens of multi-stage training at its native 1-million-token context length: 50 billion tokens of indexer warmup, 3 trillion tokens of sparse-attention training, and 200 billion tokens of learning-rate decay, per the model card. Deployment needs FP8-capable Nvidia GPUs, since the weights alone occupy roughly 315 gigabytes before accounting for inference memory, the card adds.
Sliding-window and sparse attention replace the layers that make long context expensive
Full-attention layers are the part of a transformer that gets costlier as context grows, since every token attends to every other token. Naive-N0.5-Flash removes them entirely, using a 128-token sliding window for most layers and a DeepSeek Sparse Attention indexer that selects the top 2,048 tokens for the rest, per the model card. That is a bet that a fully local-or-sparse network can hold a million-token context without paying the decoding overhead of even a handful of global-attention layers. If independent testing bears it out, it gives any lab serving long-context coding agents a concrete architectural template to copy, not just a benchmark score to chase. It also matters that NaiveAI says this design came out of an AI-run exploration process rather than a human architecture team, since that is the harder claim to verify and the one with the bigger implications if it holds: a lab replacing its own research staff's design work with model output, at a scale (nearly 10 million weekly sandboxes) large enough to plausibly search a wider design space than a small human team could in the same time.
NaiveAI ties that same claim to a longer-term goal. Its blog says Naive-N0.5-Flash was trained specifically to participate in AI research and development itself, which the company frames as a step toward what it calls recursive self-improvement, per the post. NaiveAI does not define a measurable threshold for when that loop would count as self-sustaining, and nothing in the release ties the label to a specific capability test. Readers evaluating the model should treat the RSI framing as the company's own characterization of its research process, not as an independently verified milestone.
Context and reactions
NaiveAI's own benchmark charts put Naive-N0.5-Flash at 73.6 on SWE-bench Pro and 86.7 on Terminal-Bench 2.1, alongside 32.4 on Agents' Last Exam and 17.5 on ProgramBench's Almost@1 metric, according to NaiveAI's technical blog. On the AI-R&D side, the same post reports 37.5 on PostTrainBench and 73.7 percent on MLE-bench-30, a benchmark built around automated machine-learning research tasks rather than coding. All of those figures are self-reported and have not been reproduced by an outside evaluator; NaiveAI's own model card lists the specific closed and open models each score was checked against, which at least makes the comparison auditable once someone reruns it. Pandaily's write-up of the release flagged, without confirming, that some other outlets describe NaiveAI as affiliated with Tsinghua University through its founders, and said its own coverage would stick to the NaiveAI product name rather than repeat that claim, per Pandaily's report. NaiveAI has not addressed the affiliation claim on its own channels.
What to watch next
Independent benchmark runs are the near-term test, since every coding and AI-R&D score NaiveAI has published so far is self-reported. NaiveAI has not said whether outside researchers will get any access to verify the sandbox-scale and AI-autonomy claims behind the model, or whether a larger sibling model built the same way is coming next. Also worth tracking: whether rival open-weight labs in China, several of which NaiveAI cites directly as benchmark comparisons on its own blog, respond with their own AI-authored architecture claims, or push back on the framing.
Sources
- Naive-N0.5-Flash: Building Frontier AI with AI: NaiveAI technical blog, primary
- NaiveAI/Naive-N0.5-Flash model card: Hugging Face, primary
- NaiveAI Open-Weights Naive-N0.5-Flash: 309B MoE, Hybrid SWA-DSA, 1M Context, MIT: Pandaily, secondary
