2026-10-11 17:15 UTC

NaiveAI claims its MIT-licensed Naive-N0.5-Flash — a 309B-A15.5B sparse MoE with native 1M-token context via hybrid SWA/DSA and no full-attention layers, served by its AI-optimized NaiveRT stack at up to 2,000 tokens/s — delivers frontier-comparable coding and AI-R&D capability at open weights; independent benchmarking and self-hosted adoption would establish it as a credible local coding model.

state: watchingheat: highuncertainty: highnovelscott: mediumopen-weight-model-release sparse-moe local-inference long-contextNaiveAI
Surfaced 2026-09-29T09:12:47Z — Release: open-weight 309B MoE with 15.5B active parameters built for coding and AI R&D, native 1M-token context through hybrid SWA/DSA (39 S — The velocity-spike trigger was a static-p90 crossing on a ±2-upvote wiggle (137→135), not fresh movement: the thread is flat at 135 pts / 36 comments with 0.0 pts/h at 38h, and the only comment churn is a deleted one-liner plus name-mockery. The magnitude-valve spread flag overstates reach — its second 'platform' is an echo reconstruction of the same first-party release, so periphery is still one community; no independent eval, stack support, API, or identity news has landed, leaving the case's meaning unchanged: a genuine but entirely self-reported release whose verification window remains open and realistically multi-day.

What is this?

The supplied results confirm a company called Naïve — a YC S25 startup founded by Sean Dorje and Dennis Zax (Palo Alto) selling 'autonomous AI employees' infrastructure that packages payments, incorporation, email/phone, and compute behind one API — with 30,000+ developer customers and a $28.5M Series A (Aug 2026) earmarked for frontier research on token-efficiency in agent loops. The same record carries a March 2026 exposé (verified via curl+grep on public assets) showing Naive's platform was a rebrand of the MIT-licensed 41K-star 'Paperclip' agent framework with attribution stripped. None of the snippets mention the Naive-N0.5-Flash model, its 309B-A15.5B all-sparse MoE architecture, hybrid SWA/DSA 1M context, or the NaiveRT serving stack — the release itself rests solely on the case's own evidence, and whether the model-building 'NaiveAI' is the same entity as this infra company is plausible (name, timing, stated frontier-research ambitions) but not established by these results. One landscape note: the snippets show native 1M-token context is already standard at the closed frontier (Gemini 3.5 Flash at 1M; Zhipu's GLM-5.3 at 1.3M as an agentic coding model), so the distinctive claim is doing it all-sparse at open weights, not the context length.

Why it matters to Scott

A 309B-A15.5B all-sparse MoE claiming frontier-comparable coding with zero full-attention layers at 1M context is a clean, resolvable test of Scott's attention-budget/context-rot doctrine — if sparse-only attention sustains agent-grade long-context work it changes the substrate his compaction discipline assumes, and if it degrades the way his context-rot canon predicts, that's evidence for the doctrine; either outcome bears on what he argues, and independent verification would reprice the frontier slot of his Model Barbell against his LiteLLM/Ollama local tier. But evidence class is announcement-only: the grounding confirms no web corroboration of the release itself, and if the model-building NaiveAI were ever confirmed to be the same entity as the YC startup documented (in the supplied snippets) rebranding MIT-licensed Paperclip with attribution stripped, his OSS-attribution trust canon applies hard — so the watch condition is independent LocalLLaMA benchmarking and quant availability, not the launch.
ip:concept.attention-budgetip:concept.context-rotip:concept.model-barbellip:source.the-inference-field-ebookip:concept.model-plus-harness-benchmark-unitradar:concept.open-modelsradar:concept.long-contextradar:concept.local-inferenceradar:concept.moe-inferenceradar:kimi-linear-local-validationradar:qwen3-7-flash-open-weight-releaseradar:dwarf-sparse-attention-validationradar:sliding-window-attention-sinks-conversion
queries asked of Scott's wikis
  • open-weights local coding model strategy
  • sparse MoE local inference economics throughput
  • long-context sparse attention sliding-window RAG agent memory
  • coding agent harness model requirements
  • frontier-comparable benchmark claims independent verification
  • MIT license attribution OSS fork trust

Measured heat

now 0 pts/hpeak 35 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 333h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-27 20:30 (minted)⭐ origin echo-reconstructedRelease: open-weight 309B MoE with 15.5B active parameters built for coding and AI R&D, native 1M-token context through hybrid SWA/DSA (39 S
NaiveAI on github (echo) · attributed from reddit.post.1wrs58t · published time unknown
—
09-27 18:48first on r/LocalLLaMA · published · lag ?Naive-N0.5-Flash - 309B-A15.5B
nullmove
—
09-27 18:48amplified on r/LocalLLaMA 👑reddit.post.1wrs58t
nullmove
peak 144 · 38 comments · 100% of case engagement
09-27 20:20our radar first saw it · lag ?discovery anchor: reddit.post.1wrs58t—
09-29 09:09reached heat=high · lag ? · via queue+ledger——
pace: p72 vs 1188 stories at the 168h mark (now 333h old) — ahead of dhh-agents-campfire-trilingual-ports (1.0x), behind kls-conjecture-proof-credit (1.0x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditNaive-N0.5-Flash - 309B-A15.5B
LocalLLaMA
Retrieved article excerpt

Open article · Retrieved 2026-09-27T20:29:21.571902+00:00

# [NaiveAI](https://huggingface.co/NaiveAI) / [Naive-N0.5-Flash](https://huggingface.co/NaiveAI/Naive-N0.5-Flash) Like 20 Follow NaiveAI 21

[Text Generation](https://huggingface.co/models?pipeline_tag=text-generation)[naive\_n05\_flash](https://huggingface.co/models?other=naive_n05_flash)[Mixture of Experts](https://huggingface.co/models?other=moe)[code](https://huggingface.co/models?other=code)[long-context](https://huggingface.co/models?other=long-context)[ai-research](https://huggingface.co/models?other=ai-research)[conversational](https://huggingface.co/models?other=conversational)[custom\_code](https://huggingface.co/models?other=custom_code)

License: mit

[Model card](https://huggingface.co/NaiveAI/Naive-N0.5-Flash)  [Files Files and versions  

xet](https://huggingface.co/NaiveAI/Naive-N0.5-Flash/tree/main)  [Community](https://huggingface.co/NaiveAI/Naive-N0.5-Flash/discussions)

 

Copy to bucket new

 

Naive-N0.5-Flash
Naive-N0.5-Flash

**Building Frontier AI with AI**

[🏠 Homepage](https://naive.ai/) · [📰 Technical Blog](https://naive.ai/en/research/) · [💻 GitHub](https://github.com/NaiveAI-Labs/Naive-N0.5-Flash)

## Introduction

Naive-N0.5-Flash is an open-weight **309B MoE model with 15.5B active parameters**, built for **coding and AI R&D**. It supports a **native 1M-token context window** through a hybrid of Sliding-Window Attention (SWA) and lightweight DeepSeek Sparse Attention (DSA), with no full-attention layers.

### Key Features

- **Native 1M context, without full attention.** Naive-N0.5-Flash combines Sliding-Window Attention (SWA) and lightweight DeepSeek Sparse Attention (DSA) with GQA4 at a predominantly 5:1 SWA–DSA layout. The entire network remains local or sparse, with no full-attention layers.
- **AI-optimized inference up to 2,000 tokens/s.** NaiveRT, our inference system for Naive-N0.5-Flash, was built and optimized through AI-centered R&D. It combines mega-kernel fusion, Programmatic Dependent Launch (PDL), and speculative decoding, delivering 50 tokens/s per user in Standard mode and up to 2,000 tokens/s in Ultrafast mode. See the NaiveRT case study in the [technical blog](https://naive.ai/en/research/) for the implementation and optimization process.
- **Open weights and API.** Model weights and inference code are released under the MIT license. API access will also be provided, with pricing set at $0.10 / $0.40 / $0.01 per million tokens for input, output, and cache reads, respectively.

## Model Architecture

| Property | Specification |
| --- | --- |
| Architecture | Mixture-of-Experts (MoE) |
| Total parameters | 309B |
| Active parameters | 15.5B |
| Context length | Native 1M tokens |
| Transformer layers | 48 |
| Attention-layer composition | 39 SWA layers + 9 DSA layers |
| Attention mechanism | Hybrid SWA–DSA |
| SWA window | 128 tokens |
| DSA token selection | Top 2,048 tokens for backbone attention |
| DSA KV groups | 4 (GQA4) |
| Indexer query heads | 16 |

### Hybrid SWA–DSA Attention

Naive-N0.5-Flash builds on the open-weight MiMo-V2.5 base model, which has a simple architecture with strong foundational capabilities in world knowledge and deep research. Most layers use Sliding-Window Attention (SWA), whose per-token decoding cost does not grow with context length, while a small number of global-attention layers preserve long-range information. At million-token context lengths, however, these global-attention layers account for much of the decoding overhead.

Naive-N0.5-Flash replaces the global-attention layers with DeepSeek Sparse Attention (DSA). A lightweight indexer scores the full history, while the backbone computes attention only over a selected subset of tokens. Although the indexer still scans the full history and the full KV cache is retained, sparse attention substantially reduces attention computation and memory access. Adapting the model to this new attention structure was one objective of continued pretraining.

[Hybrid SWA–DSA architecture showing the network stack and the DSA attention module, including the 16-head indexer and top-2,048 token selection.](https://huggingface.co/NaiveAI/Naive-N0.5-Flash/blob/main/assets/hybrid-swa-dsa-architecture.svg)

*Figure 1. The hybrid attention stack and DSA module.*

The network consists of **eight six-layer modules**. A standard module contains five SWA layers followed by one DSA layer, with the first layer of the first module also replaced by DSA. SWA uses a **128-token window**, while DSA selects the **top 2,048 tokens** for backbone attention. Both attention types incorporate sink bias.

Unlike the original MLA-based DSA implementation, Naive-N0.5-Flash replaces MLA with grouped-query attention (GQA) using **four KV groups**. For the architecture design process and indexer efficiency comparison, see model architecture in the [technical blog](https://naive.ai/en/research/).

### Training Overview

Following the architectural changes, Naive-N0.5-Flash completed 3.25T tokens of multi-stage training with a native 1M-token context window: 50B tokens of Indexer Warmup, 3T tokens of Sparse Attention Training, and 200B tokens of Learning Rate Decay. This process adapted the model to its new sparse attention architecture while substantially improving its AI R&D and coding capabilities. See the [technical blog](https://naive.ai/en/research/) for training details.

## Evaluation Results

[Coding benchmarks comparing Naive-N0.5-Flash with other models across seven software engineering and agentic tasks.](https://huggingface.co/NaiveAI/Naive-N0.5-Flash/blob/main/assets/coding-benchmarks.pdf)

*Figure 2. Coding and agentic task results. Naive-N0.5-Flash is highlighted in yellow.*

[AI R&D benchmarks covering PostTrainBench, MLE-bench-30, PaperBench, SOL-ExecBench, NanoChat AutoResearch, and NanoGPT SpeedRun.](https://huggingface.co/NaiveAI/Naive-N0.5-Flash/blob/main/assets/ai-rd-benchmarks.pdf)

*Figure 3. AI research and systems optimization results. Metric directions are indicated in the figure.*

Evaluation setup and metric notes

**Evaluation setup.** Unless otherwise noted, our evaluations of Naive-N0.5-Flash use Claude Code 2.1.207 with a 1M-token context window, temperature 1.0, and top-p 0.95. The harness exposes only basic file I/O and Bash tools.

**Sources for reported benchmark scores are as follows:**

- **GLM-5.3 and GLM-5.3-Flash:** [GLM-5.3 blog](https://z.ai/blog/glm-5.3) and [GLM-5.3-Flash blog](https://z.ai/blog/glm-5.3-flash), respectively.
- **Kimi-K3:** [Kimi-K3 model page](https://huggingface.co/moonshotai/Kimi-K3).
- **Qwen-3.8-Max:** [Qwen-3.8-Max blog](https://qwen.ai/blog?id=qwen3.8).
- **Hy4-preview:** [Hy4-preview model page](https://huggingface.co/tencent/Hy4-preview).
- **DeepSeek-V4.1-Flash:** [DeepSeek-V4.1-Flash model page](https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash).
- **Step-5-preview:** [Step-5-Preview-BF16 model page](https://huggingface.co/TypeSafeAI/Step-5-Preview-BF16).
- **Fable-5 (w/ fallback):** [GLM-5.3 blog](https://z.ai/blog/glm-5.3).
- **SWE-Bench Pro:** GPT-5.6-Sol, Opus-5, and Opus-5.5 scores are drawn from the [GPT-5.6 blog](https://openai.com/index/gpt-5-6/) and the [Claude Opus 5.5 System Card](https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba5024219938094694bd/Claude%20Opus%205.5%20System%20Card.pdf).
- **DeepSWE v1.1:** The Muse-Spark-1.3 score comes from its [Muse-Spark-1.3 blog](https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology). GPT-5.6-Sol and Opus-5 scores come from the [DeepSWE v1.1 leaderboard](https://deepswe.datacurve.ai/). The Opus-5.5 score comes from the [Claude Opus 5.5 System Card](https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba5024219938094694bd/Claude%20Opus%205.5%20System%20Card.pdf).
- **Terminal-Bench 2.1:** Muse-Spark-1.3, GPT-5.6-Sol, and Opus-5 scores come from the [Muse-Spark-1.3 blog](https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology). The GPT-6-Astra score comes from the [Terminal-Bench 2.1 leaderboard](https://www.tbench.ai/?version=2.1).
- **ALE-CLI:** GPT-5.6-Sol, GPT-6-Astra, Muse-Spark-1.3, Opus-5, and Opus-5.5 scores come from the [ALE-CLI leaderboard](https://agents-last-exam.org/leaderboard).
- **FrontierSWE v1:** We calculate the Dominance score using the competing systems’ results as of August 23, 2026.
- **ProgramBench:** We report the Almost@1 score. GPT-5.6-Sol and Opus-5 scores come from the [ProgramBench leaderboard](https://programbench.com/).
- **MLE-bench-30:** Gemini-3.5-Flash, Gemini-3.6-Flash, Grok-4.5, and GPT-5.6-Luna scores come from the [Gemini 3.6 Flash model card](https://deepmind.google/models/model-cards/gemini-3-6-flash/). Following its evaluation protocol, we report the average position score of Naive-N0.5-Flash.
- **PaperBench:** MiniMax M3, Opus-4.7, GPT-5.5, and Gemini-3.1-Pro scores come from the [MiniMax M3 model page](https://huggingface.co/MiniMaxAI/MiniMax-M3).
- **SOL-ExecBench, NanoChat AutoResearch, and NanoGPT SpeedRun:** Naive-N0.5-Flash scores were obtained using our in-house AutoResearch harness. Recursive Superintelligence Inc. scores come from its [research article](https://www.recursive.com/articles/first-steps-toward-automated-ai-research). Following the SOL-ExecBench update, we use its updated score from the [leaderboard](https://research.nvidia.com/benchmarks/sol-execbench).

## Deployment

Naive-N0.5-Flash supports FP8 mixed-precision inference. For general use, we recommend setting the sampling parameters to `temperature=1.0` and `top_p=0.95`.

### Quick Start with Transformers

Naive-N0.5-Flash requires FP8-capable NVIDIA GPUs. The model weights occupy approximately 315 GB; allow additional GPU memory for inference.

```
pip install "transformers[torch,kernels]>=5.17.0"
```

```
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "NaiveAI/Naive-N0.5-Flash-FP8"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    trust_remote_code=True,
    dtype="auto",
    device_map="auto",
)

inputs = tokenizer.apply_chat_template(
    [{"role": "user", "content": "Hello!"}],
    add_generation_prompt=True,
    return_dict=True,
    return_tensors="pt",
).to(model.device)

output = model.generate(**inputs, max_new_tokens=2048)
response = tokenizer.decode(
    output[0, inputs["input_ids"].shape[1]:],
    skip_special_tokens=True,
)
print(response)
```

## License

Naive-N0.5-Flash is released under the MIT License.

## Citation

If you find Naive-N0.5-Flash useful in your research or work, please cite:

```
@misc{naiveai2026naiven05flash,
  title  = {Naive-N0.5-Flash: Building Frontier AI with AI},
  author = {{NaiveAI Team}},
  year   = {2026},
  url    = {https://naive.ai/en/research/}
}
```

## Acknowledgments

Naive-N0.5-Flash builds on the work of the open-source community and gives back to it. We thank the Xiaomi MiMo team for making their MiMo-V2.5 base model publicly available, the DeepSeek team for their work on DeepSeek Sparse Attention (DSA), and the SGLang team and community for their open-source inference infrastructure.

## Contact

For questions, feedback, or collaboration, please contact us at [email protected] or follow us on X at @naiveailab. You can also find our open-source projects and model releases on [GitHub](https://github.com/naiveai-labs) and [Hugging Face](https://huggingface.co/NaiveAI).

Downloads last month
:   -
nullmove14338
🟧 echo.github ⭐Release: open-weight 309B MoE with 15.5B active parameters built for coding and AI R&D, native 1M-token context through hybrid SWA/DSA (39 SNaiveAI——

Interpretation history

Decision trace