Kaarelson claims the released LingBot-World 2.0 1.3B implementation runs at 16 FPS on one RTX 5090, potentially enabling interactive world-model experimentation on a single consumer GPU.
state: corroboratedheat: lowuncertainty: mediumnovelscott: lowworld-models local-inference interactive-simulationkaarelkaarelson
What is this?
LingBot-World 2.0 is an open-source interactive video world model from Robbyant (Ant Group's embodied-AI unit), released July 9, 2026 with a 14B primary checkpoint; a 1.3B 'Small' variant explicitly designed for single-GPU real-time inference was released Sept 10, 2026, which resolves the case's earlier weights-availability ambiguity. Builder kaarelson claims his optimized implementation runs that 1.3B at 16 FPS / 832×464 on one RTX 5090 (vs ~6 FPS stock) — claimed 2.7x over Robbyant's stack, 2.5x over SGLang Diffusion, 1.9x over NVIDIA FlashDreams — via an fp16 decoder, SageAttention instead of FlashAttention, and minor custom kernels; code is open-sourced and Linux-only. Notably, the HN thread and r/StableDiffusion crosspost carrying these claims are now directly observable in search results (previously echo-reconstructed), but no independent reproduction or repo-level verification of the numbers exists. Coverage context: every frame-rate figure in this niche is vendor-reported with no independent leaderboard, LingBot's own 720p/60 FPS headline depends on a deployment stack Robbyant won't release, Amap's separate ABot-World-0 (5B, Apache 2.0) reports 720p up to 16 FPS on one 5090 with its default config at 13.3 FPS, and a second independent builder (lucidml_lover) has separately demonstrated a custom-trained 1B realtime world model on one 5090 — so the pattern has two independent data points, neither validating Kaarelson's specific benchmark.
Why it matters to Scott
Territorial adjacency, not a position-level intersection: the released 1.3B open world model and Kaarelson's fp16-decoder/SageAttention/custom-kernel path are a concrete instance of the hardware-aware local inference practice his gamepc model zoo lives on, and the case's vendor-reported-with-no-reproduction epistemics run straight through his evidence-class-ladder discipline — but nothing here challenges, extends, or newly arrives at a position his canon holds, the 16 FPS claims remain claimant-reported and 5090-bound while his card is 3090-class. LingBot's 60 FPS headline riding an unreleased stack is a clean illustration of his claim-grading canon, yet illustration is not bearing and there is no dated-receipts angle, so it stays his kind of topic rather than news for him.
dev:project.gamepcdev:concept.hardware-aware-local-inferenceip:concept.evidence-class-ladderradar:single-5090-playable-world-modelradar:concept.world-modelsradar:concept.inference-optimizationradar:lingbot-video-action-world-modelradar:concept.benchmark-integrity
queries asked of Scott's wikis
- local video inference model zoo consumer GPU benchmarks
- world models as interactive agent environments simulation
- inference optimization fp16 quantization attention kernel speedups
- real-time AI-generated playable game worlds local-first
- verifying unverified performance claims open-source releases
Measured heat
now 0 pts/hpeak 10 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 549h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion
How the heat travelled
pace: p32 vs 1032 stories at the 336h mark (now 549h old) — ahead of addom-local-coding-harness (1.5x), behind agentsec-static-config-auditing (0.8x)
Evidence (3) — ⭐ canonical anchor
Interpretation history
2026-10-11T03:29:17Z
The velocity spike on the corroborating reddit post was transient (6.3× baseline) and has cooled to 0 pts/h; peer percentile is 11% and the case is 22 days old with no cross-platform spread. The feasibility thesis — single-consumer-GPU realtime world models — remains corroborated by two independent builders (Kaarelson's LingBot-World 1.3B claim and lucidml_lover's custom 1B demo), but neither validates the other's specific numbers: Kaarelson's 16 FPS benchmark is still claimant-reported with no independent reproduction, and lucidml_lover's code release is unconfirmed. No new implementations, reproductions, or weight releases.
2026-10-07T00:25:06Z
grounded: novel/low — Territorial adjacency, not a position-level intersection: the released 1.3B open world model and Kaarelson's fp16-decoder/SageAttention/custom-kernel path are a
2026-10-07T00:17:06Z
First independent corroboration: lucidml_lover's separate realtime text-guided 1B world model on one RTX 5090 confirms the consumer-GPU feasibility pattern this case tracks, so the case no longer rests on a single unverified claim. It does not verify Kaarelson's specific LingBot-World 16 FPS benchmark, which remains echo testimony with no independent reproduction.
2026-10-06T23:36:36Z
evidence attached: reddit.post.1wzdm08 — Independent builder's realtime 1B world model on one RTX 5090 corroborates the consumer-GPU realtime world-model feasibility the case turns on.
2026-09-19T04:23:24Z
The available discussion remains the submitter's own performance claim, not independent validation; the GitHub echo repeats that same account. The explicit resolution tradeoff narrows the claimed capability, but there is no demonstrated change in availability or reproducibility to justify elevated attention.
2026-09-18T19:51:50Z
grounded: novel/low — Scott’s gamepc model zoo establishes an adjacent interest in local video inference, but the hits show no LingBot dependency or interactive video-world-model pro
2026-09-18T19:45:04Z
case created — A linked implementation and explicit single-GPU performance claim make this a bounded local-inference episode.
Decision trace
- 10-11 14:29repriceThe velocity spike on the corroborating reddit post was transient (6.3× baseline) and has cooled to 0 pts/h; peer percentile is 11% and the case is 22 days old with no cross-platform spread. The feasi
- 10-07 21:22sensor_dirtycomment_update
- 10-07 14:23sensor_dirtyvelocity_spike
- 10-07 11:25repriceFirst independent corroboration: lucidml_lover's separate realtime text-guided 1B world model on one RTX 5090 confirms the consumer-GPU feasibility pattern this case tracks, so the case no longer
- 10-07 11:25groundTerritorial adjacency, not a position-level intersection: the released 1.3B open world model and Kaarelson's fp16-decoder/SageAttention/custom-kernel path are a concrete instance of the hardware-
- 10-07 10:36attachIndependent builder's realtime 1B world model on one RTX 5090 corroborates the consumer-GPU realtime world-model feasibility the case turns on.
- 10-07 10:28propose_attachIndependent builder's realtime 1B world model on one RTX 5090 corroborates the consumer-GPU realtime world-model feasibility the case turns on.
- 09-19 14:23repriceThe available discussion remains the submitter's own performance claim, not independent validation; the GitHub echo repeats that same account. The explicit resolution tradeoff narrows the claimed
- 09-19 14:23review_screenscreen failed: RuntimeError: litellm codex-luna exhausted: HTTPError: HTTP Error 503: Service Unavailable
- 09-19 05:51groundScott’s gamepc model zoo establishes an adjacent interest in local video inference, but the hits show no LingBot dependency or interactive video-world-model project this would change; consumer-GPU dep
- 09-19 05:45createA linked implementation and explicit single-GPU performance claim make this a bounded local-inference episode.