2026-10-11 17:11 UTC

Independent benchmarks will determine whether Meituan’s LongCat-Flash-Lite-Sparse can deliver practical 256K-context inference on 24GB GPUs by combining sparse MoE activation with a RAM-offloaded n-gram lookup table.

state: expiredheat: lowuncertainty: highnovelscott: nonelocal-inference open-models moe-inference long-contextMeituan

What is this?

Meituan released LongCat-Flash-Lite, an open-weight, non-reasoning 68.5B-parameter mixture-of-experts model that activates roughly 2.9–4.5B parameters per token and supports up to 256K context via YaRN. Its distinguishing architecture allocates about 31.4B parameters to an n-gram embedding layer, which Meituan says improves performance and reduces MoE inference bottlenecks; it is positioned particularly for coding and agentic tool use. The supplied snippets do not independently establish practical 256K-context operation on a 24GB GPU or the claimed RAM-offloaded lookup-table configuration, so those points still require third-party benchmarks and implementation evidence.

Why it matters to Scott

No intersection found: there are no Scott wiki hits tying this architecture or its 24GB long-context inference claim to his positions or projects, and no radar hits showing the development or actor is already tracked.
queries asked of Scott's wikis
  • commodity-GPU economics for sparse MoE inference
  • RAM offloading and memory hierarchy for local LLMs
  • n-gram embeddings versus transformer compute
  • long-context reliability in coding agents
  • open-weight models for agentic tool use
  • effective context versus advertised context windows

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (4) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditMeituan just dropped LongCat-Flash-Lite-Sparse
LocalLLaMA
Gohab200116527
🟧 echo.paper ⭐The primary artifact is the Meituan LongCat technical report, whose abstract states: “To facilitate further research, we also introduce and Meituan LongCat Team——
🟠 redditUncensored Multi-Model Releases, LongCat-Flash-Lite with MTPs, Jamba2-Mini, Qwen3.5-9B-Nikusui-v1 with MTPs and Qwen3.5-27B-Nikusui-v1 with MTPs, Available in Safetensors and GGUF Formats!
LocalLLaMA
LLMFan46528
🟠 redditLongCat-Flash-Lite-Sparse Is Now Available for Download
LocalLLaMA
LLMFan469716

Interpretation history

Decision trace