2026-10-11 16:37 UTC

inclusionAI claims its released Ling-3.0-flash-VL adds native image and video understanding with a one-million-token context and 5.5B active parameters out of 124B total, potentially enabling long-context visual-agent workflows with sparse inference compute.

state: corroboratedheat: lowuncertainty: mediumconvergesscott: mediumopen-models multimodal-agents long-context-inferenceinclusionAI

What is this?

Ling-3.0-flash-VL is a released MIT-licensed multimodal model from inclusionAI (Ant Group's AI lab), built on the Ling-3.0-flash text model: a sparse MoE with 124B total parameters and 5.5B active per token, adding native image/video input (ViT encoder, VideoRoPE for spatiotemporal localization) aimed at agentic visual workflows. The official Hugging Face model card — the case's canonical anchor — states a context window of up to 256K tokens, which settles the earlier tension: the 1M-token framing appears only in third-party aggregator copy (Baseten claims 1M; Krater and OpenRouter list ~262K) and in ZenMux's description of the text sibling as 'native 256K context extendable up to 1M', so the 1M figure looks like laundered extension-spec marketing rather than the VL model's shipped capability. MindStudio adds architecture detail: a 5:1 ratio of Kimi Delta Attention (KDA, a technique originating in Moonshot AI's Kimi line) to gated MLA layers across 42 layers. Local deployment is forming — llama.cpp PR #29151 (unmerged) adds VL support — but no independent visual-quality or long-context validation exists yet.

Why it matters to Scott

Converges on two fronts: the spec arc is a dated receipt for Discussed Is Not Deployed — the announced 1M-token context deflated to the 256K that the official card, aggregators, and the llama.cpp PR author all actually build to, exactly the claim-ceiling gap his evidence frameworks codify — and the release hands the YouTube-video ebook's local vision-replacement question its newest concrete candidate (a 5.5B-active multimodal MoE with a forming llama.cpp path). MEDIUM not HIGH because nothing visual is validated, community impressions have it lagging Qwen3.8-Flash in its class, and the support PR is unmerged and quiet — so this queues a hardware-aware evaluation rather than changing what he builds or argues today.
ip:source.how-to-read-a-youtube-video-ebookdev:concept.hardware-aware-local-inferenceip:framework.discussed-is-not-deployedradar:off-the-shelf-vlm-video-searchradar:concept.llama-cppradar:llama-cpp-minimax-m3-visionradar:qwen38-omni-flash-releaseradar:concept.inference-economics
queries asked of Scott's wikis
  • How to Read a YouTube Video screen-capture pipeline local vision model replacement
  • local vision-language model evaluation unified memory 128GB Mac MoE
  • MoE sparse active parameters local inference economics cost per token
  • long-context KV cache memory footprint 256K tokens local runtime
  • model card context window claims vs shipped spec verification
  • llama.cpp multimodal vision model support maturity GGUF porting

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 792h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-08 16:24 (minted)⭐ origin echo-reconstructedThe quoted model card describes native image and video understanding, 124B total parameters with 5.5B activated per token, and context suppo
inclusionAI on blog (echo) · attributed from reddit.post.1wasdnn · published time unknown
—
09-08 15:57first on r/LocalLLaMA · published · lag ?inclusionAI/Ling-3.0-flash-VL · Hugging Face
jacek2023
—
09-08 15:57amplified on r/LocalLLaMA 👑reddit.post.1wasdnn
jacek2023
peak 152 · 23 comments · 72% of case engagement
09-10 04:52amplified on r/LocalLLaMAreddit.post.1wc97an
niacolhealth
peak 6 · 7 comments · 5% of case engagement
09-20 21:11amplified on r/LocalLLaMAreddit.post.1wlt33y
autonoma_2042
peak 11 · 8 comments · 8% of case engagement
09-24 08:36amplified on r/LocalLLaMAreddit.post.1wow7nt
jacek2023
peak 26 · 10 comments · 15% of case engagement
09-08 16:20our radar first saw it · lag ?discovery anchor: reddit.post.1wasdnn—
pace: p76 vs 519 stories at the 720h mark (now 792h old) — ahead of jenny-local-coding-harness (1.0x), behind claude-code-remote-attribution-injection (1.0x)

Evidence (5) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditinclusionAI/Ling-3.0-flash-VL · Hugging Face
LocalLLaMA
jacek202315223
🟧 echo.blog ⭐The quoted model card describes native image and video understanding, 124B total parameters with 5.5B activated per token, and context suppoinclusionAI——
🟠 redditSame Ling model, different long-context curve: INT4/vLLM vs Q5/llama.cpp on one Spark
LocalLLaMA
niacolhealth17
🟠 redditLing 3.0 Tiny vs. Gemma 26B-A4B MoE
LocalLLaMA
autonoma_2042118
🟠 redditmodel : add Ling 3.0 VL support by aetherbird · Pull Request #29151 · ggml-org/llama.cpp
LocalLLaMA
jacek20232610

Interpretation history

Decision trace