2026-10-11 18:02 UTC

Independent evaluations will determine whether InclusionAI’s MIT-licensed Ling-3.0-Flash delivers competitive coding and design quality with practical local inference despite its 127.5B-total, 5.1B-active sparse-MoE architecture.

state: resolvedheat: lowuncertainty: mediumconvergesscott: highling-3-flash open-models local-inferenceInclusionAI

What is this?

Ling-3.0-Flash is an InclusionAI sparse mixture-of-experts language model intended for agentic, coding, and cost-efficient inference workloads. The strongest technical snippet describes a 124.4B-parameter, 42-layer base with 512 routed experts, eight active per token, roughly 5.5B active parameters, and a 3.1B multi-token-prediction layer that brings the full checkpoint to 127.5B; other listings instead summarize it as 124B total and 5.1B active. The supplied material reports official weights, quantized local runs, and claims of improved coding, design, tool use, and harness compatibility, but it does not provide the underlying independent evaluations needed to establish competitive quality; the MIT license claim also appears only in an evidence title, while one snippet says license sources conflict.

Why it matters to Scott

Ling-3.0-Flash converges with Scott’s hardware-aware local-inference and usable-capability positions, while its unverified coding/design claims directly call for his model-plus-harness and evaluation-driven benchmarking approach. It could merit a practical harness-and-quantization test for his local model stack, but the supplied evidence does not yet establish quality, licensing certainty, or compatibility with his own hardware.
ip:concept.model-plus-harness-benchmark-unitip:concept.evaluation-driven-developmentip:concept.usable-mass-over-unusable-powerdev:concept.hardware-aware-local-inferencedev:project.gamepcradar:concept.moe-inferenceradar:concept.local-inferenceradar:concept.quantizationradar:concept.model-evaluationradar:kimi-linear-local-validation
queries asked of Scott's wikis
  • sparse MoE local inference economics
  • open-weight model licensing and sovereignty
  • coding-agent model evaluation harnesses
  • quantization tradeoffs for local agent models
  • active parameters versus memory footprint
  • design and frontend coding benchmarks

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (26) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditinclusionAI/Ling-3.0-flash weights are up on Hugging Face — MIT, BF16 plus an official FP8
LocalLLaMA
derspenti19044
🟠 redditDesign systems from code alone - Without external images, Ling-3.0-flash generated webpages across Bauhaus, Bohemian, acid design, and more—using CSS gradients, SVG paths, typography, and layout to preserve each visual language.
LocalLLaMA
AcanthisittaOk1699377
🟧 echo.other ⭐The first public weight upload was the official inclusionAI/Ling-3.0-flash Hugging Face repository. Its model card says: “We're introducing inclusionAI——
🟠 redditLing-3.0-flash MXFP4 released and running locally on one DGX Spark.
LocalLLaMA
niacolhealth5319
🟠 redditTwo flags took the official Ling-3.0-flash INT4 from 20.8 to 38.7 tok/s on one DGX Spark
LocalLLaMA
AcanthisittaOk1699398
🟠 redditinclusionAI/Ling-3.0-tiny · 8B A1.3B MoE· Hugging Face
LocalLLaMA
-Cubie-32251
🟠 redditLing 3.0 Flash on Strix Halo
LocalLLaMA
Badger-Purple2215
🟠 redditwhat the subscription still buys
OpenAI
truecakesnake172
🟠 redditA local model read its own file back before running it. Which turns of your Claude Code loop would you actually hand off?
ClaudeAI
Danare_113141
🟠 redditLing-3.0-flash quant ladder on one DGX Spark: the whole thing sits in a 32 to 40 tok/s band
LocalLLaMA
AcanthisittaOk1699217
🟠 reddit[llama.cpp PR #26608] Ling-3.0 support (unmerged)
LocalLLaMA
Public_Umpire_1099427
🟠 redditBenched a 124B on one DGX Spark for a week and published all of it — 38.7 tok/s on the fastest path he found, 2.4x DeepSeek V4 Flash on the same box
LocalLLaMA
AcanthisittaOk16991815
🟠 redditOne prompt on a local box built this dashboard front end. The data behind it is fake. Toy or tool?
artificial
Asleep-Pilot-414223
🟠 redditA 124B emitted 15,128 tokens in a single response on one DGX Spark, decode went 35.62 → 35.68 tok/s across the whole thing
LocalLLaMA
AcanthisittaOk1699321
🟠 redditLing 3.0 support merged into llama.cpp
LocalLLaMA
parepeg11547
🟠 redditnoctrex/Ling-3.0-flash-MXFP4_MOE-GGUF · Hugging Face
LocalLLaMA
pmttyji140
🟠 redditIs Ling 3 tiny underrated for its size?
LocalLLaMA
Hot_Example_44564444
🟠 redditLing-3.0 (BailingMoE3) lands in llama.cpp mainline - Quick benchmarks on Intel Arc B580
LocalLLaMA
Polaris_debi55111
🟠 redditAntLing’ve open-sourced 6 Base Model checkpoints for Ling-3.0-tiny & Ling-3.0-flash, covering pre-trained, mid-trained, and WSM-merged stages.
LocalLLaMA
AcanthisittaOk169916511
🟠 redditling 3.0 flash/tiny base models
LocalLLaMA
jacek2023744
🟠 redditLing 3.0 Tiny makes an amazing auxillery model for Hermes (Qwen 3.8 27B as the primary model)
LocalLLaMA
My_Unbiased_Opinion2725
🟠 redditLing-3.0 released all 6 base checkpoints: 2 sizes × 3 stages
LocalLLaMA
niacolhealth794
🟠 redditLing-3.0 opens six base checkpoints across three training stages
artificial
Designer_Mouse_610972
🟠 redditAntLing released a dspark draft model for Ling-3.0-flash
LocalLLaMA
Ihtien4122
🟠 redditLing-3.0-flash q5-k-l poor oneshot slop
LocalLLaMA
Thin_Pollution884308
🟠 redditOpen-weight transparency can mean more than one downloadable endpoint
OpenAI
creditme773

Interpretation history

Decision trace