Facebook Research and UW (Shao, Shen, Zettlemoyer, Koh et al.) released the official Context Language Models codebase and paper claiming models that treat their own context as a freely editable file beat SOTA context-management strategies zero-shot (e.g., +11.4% BrowseComp-Plus accuracy at โ21.5% FLOPs, gains on 12โ24-hour agent tasks) with day-1 Pi support โ external adoption or replication of the context-as-a-file approach across harnesses and models, or ContextBench, would establish CLMs as a durable research direction rather than a one-off repo.
state: watchingheat: lowuncertainty: mediumconvergesscott: highagent-memory context-management llm-architecture meta-researchRulin ShaoLuke ZettlemoyerPang Wei KohNathan LambertMike Lewis
Surfaced 2026-10-03T15:55:11Z โ Abstract: "We introduce Context Language Models (CLMs), language models that natively manage their own context. We implement this by treatin โ The attention wave has fully crested and decayed: the front-page thread is flat (~0 pts/h, 2.6th-percentile peer velocity at ~121h age) and the incremental comments re-amplify the same three critiques already on file (suffix-cache surprise, attention-cost/cache-economics concern, novelty challenge) rather than adding substance. With no adoption, replication, independent verification, or ContextBench drop, the case downshifts to a cold watch on its named establishment triggers; the magnitude-valve spread reading reflects the days-old peak, not current motion, which is why heat prices low.
What is this?
On 29 Sep 2026, Rulin Shao, Shannon Zejiang Shen, Luke Zettlemoyer, Pang Wei Koh, Nathan Lambert, Mike Lewis, Wen-tau Yih and others posted 'Context Language Models' (arXiv:2609.37725), proposing a model class that natively manages its own context by treating the context as a freely editable file โ shifting context management from external harness code to intrinsic model behavior, learnable both in-context and parametrically, and extending naturally to multi-agent systems where each agent's context is a separate file. The abstract confirms the headline claims: zero-shot CLMs built on existing models beat SOTA context-management strategies (+11.4% accuracy at โ21.5% FLOPs on BrowseComp-Plus, +5% at โ59% FLOPs on 12-hour EdgeBench, +65% greater improvement at equal compute on a 24-hour multi-repository agent-swarm task). The supplied snippets verify the paper, its author roster, and its claims, but do not corroborate the Meta Superintelligence Labs institutional framing, an official codebase, day-1 Pi support, or the existence of 'ContextBench' โ and none of the other search results show external adoption or replication of the context-as-a-file approach; the adjacent hits (ICLR 2026 agentic-context-engineering, context-engineering surveys) are same-territory work, not evidence of uptake.
Why it matters to Scott
A heavyweight Meta/UW team has independently arrived at Scott's core position โ context is a freely editable file the agent itself compacts and discards โ making this a dated-receipts opportunity for the Context Engineering ebook and a direct comparative datapoint on whether agent-authored compaction beats harness-side strategies. Its model-intrinsic-vs-harness locus claim also pressures the scaffolding-hypothesis premise that the adaptive layer must live outside the frozen model, but with replication and external adoption still absent (per the grounding), that tension stays a watch-item rather than a verdict.
ip:framework.context-engineeringip:source.context-engineering-why-building-ai-agents-feels-like-programming-on-a-vic-20-again-ebookdev:concept.agent-authored-context-compactionip:concept.scaffolding-hypothesisradar:agentic-context-management-paperradar:concept.context-managementradar:concept.agent-memoryradar:self-written-notes-reasoning-gainsradar:anthropic-claude-5-context-engineering
queries asked of Scott's wikis
- "context as a file" agent-editable memory
- harness-level context management compaction strategy
- model-intrinsic behavior vs external scaffolding division
- multi-agent context sharing memory isolation
- long-context agent compute cost FLOPs economics
- RAG vs in-context knowledge storage tradeoff
Measured heat
now 0 pts/hpeak 36 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 314h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion
How the heat travelled
pace: p75 vs 1188 stories at the 168h mark (now 314h old) โ ahead of research-agent-compression-generalization (1.0x), behind gguf-quant-filename-mismatch (1.0x)
Evidence (5) โ โญ canonical anchor
| source | object | author | score | comments |
| ๐ง hn | Context Language Models (CLMs)Retrieved article excerptOpen article ยท Retrieved 2026-10-01T00:29:49.643231+00:00 # Context Language Models (CLMs)
[Rulin Shao](https://rulinshao.github.io/)1,2,
[Shannon Zejiang Shen](https://www.szj.io/)3,
[Junjie Oscar Yin](https://oseyincs.io/)1,2,
[Yuetai Li](https://yuetl9.github.io/)1,
[Minheng Wang](https://minhengwang.github.io/)1,
[Hamish Ivison](https://ivison.id.au)1,
[Radha Poovendran](https://people.ece.uw.edu/radha/)1,
[Nathan Lambert](https://natolambert.com/)4,
[Teng Xiao](https://tengxiao1.github.io/)1,
[Mike Lewis](https://ai.meta.com/people/209431298931133/mike-lewis/)2,
[Wen-tau Yih](https://scottyih.org/)2,
[Luke Zettlemoyer](https://homes.cs.washington.edu/~lsz/)1,2,
[Pang Wei Koh](https://koh.pw/)1
1University of Washington ย 2Meta Superintelligence Labs ย 3MIT ย 4Trillium Labs
[arXiv](https://arxiv.org/abs/2609.37725)
[Twitter](https://x.com/RulinShao/status/2105282444270448647)
[Context Language Models](https://github.com/facebookresearch/context-language-models/blob/main/assets/teaser.png)
We introduce **Context Language Models (CLMs)**, language models that natively manage their own
context. We implement this by treating the **context as a file** and allowing the model to make
unrestricted updates to this file. This allows the model to learn what is most important to
maintain in context, and naturally extends to multi-agent systems where multiple agent
contexts coexist as files.
- **Zero-shot.** Building CLMs zero-shot with existing models outperforms SOTA
context-management strategies across a variety of tasks: 11.4% higher accuracy with 21.5%
fewer FLOPs on BrowseComp-Plus, 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench,
and 65% greater improvement with the same compute on a 24-hour multi-repository agent-swarm
task.
- **In-context learning.** We show that CLMs can be steered with natural-language
instructions evolved through a standard skill-optimization loop, improving held-out
accuracy by up to 35.9 points on a context-management task while reducing compute.
- **Reinforcement learning.** We also introduce an online reinforcement learning method for
CLMs, improving Qwen3.5-9B performance on BrowseComp-Plus by 47.6% while using 12% fewer
FLOPs.
## Day 1 Support
CLM for [Pi](https://github.com/earendil-works/pi):
```
pi install npm:@lolipopshock/pi-clm
```
## Getting started
Run the minimal CLM agent on any [Harbor](https://github.com/laude-institute/harbor) task:
```
pip install -e .
clm-harbor run -p <harbor-task> -a clm-minimal -m openai/<model> \
--agent-kwarg api_base=http://localhost:8000/v1
```
`clm-harbor` is the Harbor CLI with CLM available as `-a clm-minimal`. See
[`clm/clm_harness`](https://github.com/facebookresearch/context-language-models/blob/main/clm/clm_harness) for configuration and serving.
## Repository
| | |
| --- | --- |
| [`clm/clm_harness`](https://github.com/facebookresearch/context-language-models/blob/main/clm/clm_harness) | CLMs implemented in [Harbor](https://github.com/laude-institute/harbor) |
| [`clm/clm_icl`](https://github.com/facebookresearch/context-language-models/blob/main/clm/clm_icl) | In-context learning for CLMs |
| [`clm/clm_rl`](https://github.com/facebookresearch/context-language-models/blob/main/clm/clm_rl) | Reinforcement learning for CLMs |
| [`suffix_cache_reuse`](https://github.com/facebookresearch/context-language-models/blob/main/suffix_cache_reuse) | Suffix Cache Reuse for CLM efficient serving |
## Coming soon
- ContextBench
## Citation
If you find our work helpful, we would appreciate it if you could cite our paper:
```
@article{shao2026context,
title = {Context Language Models},
author = {Shao, Rulin and Shen, Shannon Zejiang and Yin, Junjie Oscar and Li, Yuetai and
Wang, Minheng and Ivison, Hamish and Poovendran, Radha and Lambert, Nathan and
Xiao, Teng and Lewis, Mike and Yih, Wen-tau and Zettlemoyer, Luke and Koh, Pang Wei},
journal = {arXiv preprint arXiv:2609.37725},
year = {2026}
}
```
## License
This project is licensed under [CC BY-NC 4.0](https://github.com/facebookresearch/context-language-models/blob/main/LICENSE). See also [NOTICE](https://github.com/facebookresearch/context-language-models/blob/main/NOTICE). | handfuloflight | 2 | 0 |
| ๐ง echo.paper โญ | Abstract: "We introduce Context Language Models (CLMs), language models that natively manage their own context. We implement this by treatin | Rulin Shao et al. (University of Washington, Meta Superintelligence Labs, MIT, Trillium Labs) โ submitted to arXiv by Rulin Shao | โ | โ |
| ๐ง hn | Context Language Models | emersonmacro | 177 | 52 |
| ๐ reddit | Y'all this is a sexy paper; context language models LocalLLaMA | Combinatorilliance | 109 | 42 |
| ๐ reddit | Context Language Models artificial | striketheviol | 3 | 1 |
Interpretation history
2026-10-08T04:51:27Z
The Reddit thread (1wyf63m) produced two sharp velocity spikes (~45x and ~39x baseline) around Oct 6 but has since cooled to 0 pts/h; the HN front-page wave crested days earlier. Both discussion bursts were appreciation and re-raises of known critiques (suffix-cache surprise, cache-hit-rate economics, attention cost, novelty challenge), not external adoption, replication, or ContextBench release. The case's durability criteria remain unmet. Sibling topics (agent-memory, context-management) run hot, but this episode's own attention has fully decayed.
2026-10-05T22:01:46Z
Meaning shift: the story now has confirmed cross-community traction โ a Reddit ML thread engaging the idea positively a week after the HN crest, moving at top-decile cohort velocity โ but the spread is discussion-shaped, not uptake-shaped: appreciation and re-raises of the same four critiques on file, zero implementations, adoption, replication, or ContextBench. Hot periphery with cold substance keeps the case a watched release; heat prices medium on the live Reddit ripple and hot sibling topics, not on the already-delivered days-old HN peak that drove the magnitude-valve reading.
2026-10-05T20:39:21Z
evidence attached: reddit.post.1wyi5nb โ shared external link with case evidence
2026-10-05T20:39:21Z
evidence attached: reddit.post.1wyf63m โ shared external link with case evidence
2026-10-03T15:27:15Z
magnitude valve eligible (multi-platform, top-decile engagement) and never alerted; deterministic escalation to deliver
2026-10-01T19:26:25Z
The release has crossed from repo announcement into genuine technical debate โ the front-page thread engages the suffix-cache finding, an attention-cost critique of self-managing context, and a novelty challenge โ but the debate still tests the authors' own claims rather than evidencing external adoption, so the durability criteria (adoption across harnesses/models, replication, ContextBench) remain unmet and the case stays a watched release episode.
2026-10-01T18:31:31Z
evidence attached: hn.story.49922437 โ The official Context Language Models paper hitting the HN front page is independent spread for the open context-as-a-file case.
2026-10-01T00:54:05Z
origin walked (opencode/cheap-glm, conf 0.92): anchor hn.story.49915539 -> echo.paper.83044bac05 by Rulin Shao et al. (University of Washington, Meta Superintelligence Labs, MIT, Trillium Labs) โ submitted to arXiv by Rulin Shao
2026-10-01T00:49:51Z
grounded: converges/high โ A heavyweight Meta/UW team has independently arrived at Scott's core position โ context is a freely editable file the agent itself compacts and discards โ makin
2026-10-01T00:40:55Z
case created โ A first-party Meta Superintelligence Labs release with heavyweight authorship, a named model class, comparative claims directly in Scott's agent-memory/context territory, and immediate ecosystem integration is a concrete release episode whose adoption trajectory is worth tracking, and it is distinct from the older corroborated agentic-context-management-paper case.
Decision trace
- 10-11 22:08review_screenjev screen: no material development (noul=0.19)
- 10-08 15:51repriceThe Reddit thread (1wyf63m) produced two sharp velocity spikes (~45x and ~39x baseline) around Oct 6 but has since cooled to 0 pts/h; the HN front-page wave crested days earlier. Both discussion burst
- 10-06 19:20sensor_dirtyvelocity_spike
- 10-06 10:22sensor_dirtyvelocity_spike
- 10-06 09:01repriceMeaning shift: the story now has confirmed cross-community traction โ a Reddit ML thread engaging the idea positively a week after the HN crest, moving at top-decile cohort velocity โ but the spread i
- 10-06 07:39attachshared external link with case evidence
- 10-06 07:39attachshared external link with case evidence
- 10-06 07:22propose_attachshared external link with case evidence
- 10-06 07:22propose_attachshared external link with case evidence
- 10-04 01:55pushAbstract: "We introduce Context Language Models (CLMs), language models that natively manage their own context. We implement this by treatin โ The attention wave has fully crested and decayed: th
- 10-04 01:27repriceThe attention wave has fully crested and decayed: the front-page thread is flat (~0 pts/h, 2.6th-percentile peer velocity at ~121h age) and the incremental comments re-amplify the same three critiques
- 10-04 01:27alert_heldAbstract: "We introduce Context Language Models (CLMs), language models that natively manage their own context. We implement this by treatin โ The attention wave has fully crested and decayed: th
- 10-04 01:27alert_routeAbstract: "We introduce Context Language Models (CLMs), language models that natively manage their own context. We implement this by treatin โ The attention wave has fully crested and decayed: th
- 10-02 23:21sensor_dirtyvelocity_spike
- 10-02 21:21sensor_dirtycomment_update
- 10-02 16:21sensor_dirtyvelocity_spike
- 10-02 13:21sensor_dirtycomment_update
- 10-02 09:22sensor_dirtycomment_update
- 10-02 06:21sensor_dirtyvelocity_spike
- 10-02 05:26repriceThe release has crossed from repo announcement into genuine technical debate โ the front-page thread engages the suffix-cache finding, an attention-cost critique of self-managing context, and a novelt
- 10-02 04:31attachThe official Context Language Models paper hitting the HN front page is independent spread for the open context-as-a-file case.
- 10-02 04:29propose_attachThe official Context Language Models paper hitting the HN front page is independent spread for the open context-as-a-file case.
- 10-01 10:54promote_anchororigin walk conf 0.92
- 10-01 10:49groundA heavyweight Meta/UW team has independently arrived at Scott's core position โ context is a freely editable file the agent itself compacts and discards โ making this a dated-receipts opportunity
- 10-01 10:40createA first-party Meta Superintelligence Labs release with heavyweight authorship, a named model class, comparative claims directly in Scott's agent-memory/context territory, and immediate ecosystem