2026-10-11 16:38 UTC

Kaarelson claims the released LingBot-World 2.0 1.3B implementation runs at 16 FPS on one RTX 5090, potentially enabling interactive world-model experimentation on a single consumer GPU.

state: corroboratedheat: lowuncertainty: mediumnovelscott: lowworld-models local-inference interactive-simulationkaarelkaarelson

What is this?

LingBot-World 2.0 is an open-source interactive video world model from Robbyant (Ant Group's embodied-AI unit), released July 9, 2026 with a 14B primary checkpoint; a 1.3B 'Small' variant explicitly designed for single-GPU real-time inference was released Sept 10, 2026, which resolves the case's earlier weights-availability ambiguity. Builder kaarelson claims his optimized implementation runs that 1.3B at 16 FPS / 832×464 on one RTX 5090 (vs ~6 FPS stock) — claimed 2.7x over Robbyant's stack, 2.5x over SGLang Diffusion, 1.9x over NVIDIA FlashDreams — via an fp16 decoder, SageAttention instead of FlashAttention, and minor custom kernels; code is open-sourced and Linux-only. Notably, the HN thread and r/StableDiffusion crosspost carrying these claims are now directly observable in search results (previously echo-reconstructed), but no independent reproduction or repo-level verification of the numbers exists. Coverage context: every frame-rate figure in this niche is vendor-reported with no independent leaderboard, LingBot's own 720p/60 FPS headline depends on a deployment stack Robbyant won't release, Amap's separate ABot-World-0 (5B, Apache 2.0) reports 720p up to 16 FPS on one 5090 with its default config at 13.3 FPS, and a second independent builder (lucidml_lover) has separately demonstrated a custom-trained 1B realtime world model on one 5090 — so the pattern has two independent data points, neither validating Kaarelson's specific benchmark.

Why it matters to Scott

Territorial adjacency, not a position-level intersection: the released 1.3B open world model and Kaarelson's fp16-decoder/SageAttention/custom-kernel path are a concrete instance of the hardware-aware local inference practice his gamepc model zoo lives on, and the case's vendor-reported-with-no-reproduction epistemics run straight through his evidence-class-ladder discipline — but nothing here challenges, extends, or newly arrives at a position his canon holds, the 16 FPS claims remain claimant-reported and 5090-bound while his card is 3090-class. LingBot's 60 FPS headline riding an unreleased stack is a clean illustration of his claim-grading canon, yet illustration is not bearing and there is no dated-receipts angle, so it stays his kind of topic rather than news for him.
dev:project.gamepcdev:concept.hardware-aware-local-inferenceip:concept.evidence-class-ladderradar:single-5090-playable-world-modelradar:concept.world-modelsradar:concept.inference-optimizationradar:lingbot-video-action-world-modelradar:concept.benchmark-integrity
queries asked of Scott's wikis
  • local video inference model zoo consumer GPU benchmarks
  • world models as interactive agent environments simulation
  • inference optimization fp16 quantization attention kernel speedups
  • real-time AI-generated playable game worlds local-first
  • verifying unverified performance claims open-source releases

Measured heat

now 0 pts/hpeak 10 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 549h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-18 19:45 (minted)⭐ origin echo-reconstructedThe linked implementation is presented on Show HN as running LingBot-World 2.0's 1.3B model at 16 FPS on one RTX 5090.
kaarelkaarelson on github (echo) · attributed from hn.story.49758853 · published time unknown
—
09-18 19:12first on hacker news · published · lag ?Show HN: LingBot-World 2.0 (1.3B) running at 16 FPS on one RTX 5090
kaarelson
—
10-06 20:38first on r/LocalLLaMA · published · lag ?Local AI World Model Part 2 - Deep NN to turn Images into Playable Characters, with prompt switching mid rollout
lucidml_lover
—
09-18 19:12amplified on hacker newshn.story.49758853
kaarelson
peak 4 · 2 comments · 10% of case engagement
10-06 20:38amplified on r/LocalLLaMA 👑reddit.post.1wzdm08
lucidml_lover
peak 67 · 27 comments · 90% of case engagement
09-18 19:20our radar first saw it · lag ?discovery anchor: hn.story.49758853—
pace: p32 vs 1032 stories at the 336h mark (now 549h old) — ahead of addom-local-coding-harness (1.5x), behind agentsec-static-config-auditing (0.8x)

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: LingBot-World 2.0 (1.3B) running at 16 FPS on one RTX 5090kaarelson42
🟧 echo.github ⭐The linked implementation is presented on Show HN as running LingBot-World 2.0's 1.3B model at 16 FPS on one RTX 5090.kaarelkaarelson——
🟠 redditLocal AI World Model Part 2 - Deep NN to turn Images into Playable Characters, with prompt switching mid rollout
LocalLLaMA
lucidml_lover6727

Interpretation history

Decision trace