2026-10-11 17:13 UTC

ParaRNN creator bugkira claims the released PyTorch/Triton library parallelizes nonlinear RNN training through Newton-based associative scans with roughly 200-fold wall-clock speedups over sequential unrolling, potentially removing a major training bottleneck for recurrent alternatives to transformers.

state: resolvedheat: lowuncertainty: mediumnovelscott: lowrecurrent-models triton alternative-architecturesbugkiraParaRNN

What is this?

ParaRNN is an Apple-released framework (ICLR 2026 oral, arXiv 2510.21450) that parallelizes training of nonlinear RNNs by recasting the recurrence as a single system of nonlinear equations solved with Newton's iterations plus custom parallel reductions, claiming up to 665x speedup over sequential unrolling and the first 7B-parameter classical RNNs (ParaLSTM/ParaGRU) matching similarly sized Transformers and Mamba2 on perplexity, with the codebase open-sourced. The radar's case, however, concerns a same-named library announced by independent creator bugkira (Daniil Sereda) — a PyTorch/Triton implementation covering sLSTM, RWKV-7, CfC, and Titans cells with a claimed ~200x speedup — and the supplied web results say nothing about that person or artifact, leaving unresolved whether it is Apple's codebase, an independent reimplementation, or merely same-named kin (the case's FlashRNN naming mismatch stands). The method class itself carries independent peer-reviewed support per the case record (a NeurIPS 2026 spotlight using the same DEER/Newton fixed-point scheme reporting >100x speedups), so the technique looks real while the bugkira-specific ~200x figure and artifact identity rest only on reconstructed echo testimony.

Why it matters to Scott

The new NeurIPS corroboration validates the DEER/Newton method class, not bugkira's artifact, and the supplied canon carries no live position on nonlinear-RNN training or recurrent-architecture viability — the only genuine tie is dev:project.crypto, his archived 2023–24 TensorFlow/Keras sequence-model research, which parallelized nonlinear-RNN training would once have served directly but no longer bears on. This remains lineage on the radar's recurrence thread (follow-on to radar:concept.recurrent-models alongside tupoi, Complex-KDA, and RWKV-8) rather than something that would change what Scott builds or argues; the evidence-class and speedup-claim matches are the generic 'illustrates his framework' pattern, not news. Relevance stays low unless the ~200x artifact claim is independently verified or a canon page emerges holding an alternative-architectures position this would confirm or extend.
dev:project.cryptoradar:concept.recurrent-modelsradar:concept.alternative-architecturesradar:concept.triton-kernelsradar:concept.reproducibilityradar:tupoi-constant-memory-llmradar:complex-kda-recurrent-releaseradar:rwkv8-m1-low-memory-training
queries asked of Scott's wikis
  • triton custom kernel projects GPU fusion work
  • RWKV Mamba SSM alternative architecture positions
  • local inference constant memory streaming long context cost
  • associative scan parallel prefix scan numerical iterations
  • early sequence modeling LSTM RNN work history
  • evaluating speedup claims reproducing open source benchmark repos

Measured heat

now 0 pts/hpeak 13 pts/hcomments 0/hpeers p50momentum: steady2 platformsage 857h
points/hour across evidence · reading as of 2026-10-05 18:22:53.291200+11:00 · deterministic, not a model opinion

How the heat travelled

08-30 14:00⭐ origin echo-reconstructedThe earliest primary artifact is the author's initial repository commit: “Initial ParaRNN: cells, Newton+scan, fused Triton, Autograd generi
Daniil Sereda on github (echo) · attributed from reddit.post.1w9oi7l
—
09-07 10:31first on r/MachineLearning · published · +188.5hParaRNN: Parallel Training for Non-Linear RNNs (sLSTM, RWKV-7, CfC, Titans) via Triton Newton Scans [P]
bugkira
—
09-07 10:31amplified on r/MachineLearningreddit.post.1w9oi7l
bugkira
peak 2 · 1 comments · 2% of case engagement
10-01 13:12amplified on r/MachineLearning 👑reddit.post.1wuz2s4
DangerousFunny1371
peak 145 · 18 comments · 98% of case engagement
09-08 13:20our radar first saw it · +215.3hdiscovery anchor: reddit.post.1w9oi7l—

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditParaRNN: Parallel Training for Non-Linear RNNs (sLSTM, RWKV-7, CfC, Titans) via Triton Newton Scans [P]
MachineLearning
bugkira21
🟧 echo.github ⭐The earliest primary artifact is the author's initial repository commit: “Initial ParaRNN: cells, Newton+scan, fused Triton, Autograd generiDaniil Sereda——
🟠 redditParallel-in-Time Training of Recurrent Neural Networks for Dynamical Systems Reconstruction [R]
MachineLearning
DangerousFunny137115618

Interpretation history

Decision trace