2026-10-11 17:11 UTC

The LingBot-Video team claims its released 13B sparse-MoE model can generate physically plausible action-conditioned robot rollouts, potentially providing an open foundation for prediction and planning in robotics.

state: expiredheat: lowuncertainty: highconvergesscott: lowworld-models physical-ai open-modelsLingBot-Video

What is this?

Robbyant released LingBot-Video as an open, Apache-2.0 sparse-MoE diffusion-transformer video model aimed at embodied AI and physical simulation rather than only creative video generation. Supplied reports claim strong self-reported RBench results and describe physically informed pretraining, but they conflict on model size—13B total/1.4B active in the case versus 30B/3B in one snippet—and do not clearly establish that LingBot-Video itself performs action-conditioned planning. Some action-conditioned world-model and closed-loop-control claims appear to concern the related LingBot-VA or LingBot-VA 2.0 systems, so the exact relationship between those models is unresolved here.

Why it matters to Scott

The Apache-licensed release weakly converges with Scott’s preference for independently operable model foundations, while its physical-rollout claims invoke his requirement that simulations be validated through world-loop closure. For now it is only another example of those positions: the supplied evidence does not establish action-conditioned planning, real-world closed-loop performance, or practical local inference, and the model specifications conflict.
ip:framework.sovereign-software-assuranceip:concept.world-loop-closuredev:concept.hardware-aware-local-inferenceradar:concept.world-modelsradar:concept.physical-airadar:concept.open-modelsradar:concept.sparse-moe
queries asked of Scott's wikis
  • world models for robot prediction and planning
  • video generation as a physical simulator
  • open-weight foundations for physical AI
  • action-conditioned rollout models
  • sparse MoE inference economics
  • learned simulators versus explicit robotics models

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditLingBot-Video: sparse-MoE video diffusion transformer (13B total, 1.4B active) post-trained as an action-conditioned world model[R]
MachineLearning
Savings-Display512361
🟧 echo.paper ⭐The original technical report presents LingBot-Video as a DiT video-pretraining system for embodied intelligence. It explicitly evaluates “MShuailei Ma et al. (project lead: Ka Leong Cheng)——

Interpretation history

Decision trace