
Context Language Models: LLMs That Manage Their Own Context
Researchers from the University of Washington, MIT and Meta introduced Context Language Models (CLMs): models that keep their context as a file and edit it themselves. On BrowseComp-Plus it scored 11.4% higher accuracy with 21.5% fewer FLOPs.
Researchers from the University of Washington, MIT and Meta have introduced Context Language Models (CLMs), language models that natively manage their own context. The paper was published on arXiv on 29 September.
Until now, context has been managed from outside the model, by hand-built systems or constrained tools. In a CLM, that job belongs to the model itself.
Context as a file
In a CLM, the context is stored as a file. The model can change it freely: it either appends new tokens or edits any span with Bash, and every change is immediately reflected in its live context for the next turn.
Because the context is a separate file, several such files can coexist in one system, which extends the approach to multi-agent setups. In tests, models showed new behaviors: they created trackers for multi-agent orchestration, introduced a new chat role for internal notes, and defined reusable context-management functions.
The authors say the result confirms the idea behind the "Bitter Lesson": freedom lets the model find better strategies than rules written by humans.
Results
CLMs work without additional training on existing models such as Qwen3.6-27B and GPT5.6-Sol. On the deep-research benchmark BrowseComp-Plus, CLM scored 59.4% under a 32K context limit: 11.4% higher accuracy than the strongest baseline at 21.5% fewer FLOPs.
On TerminalBench 2.1 it matched the strongest method's accuracy while spending 29.5% fewer FLOPs. On the 12-hour EdgeBench it showed a 5% higher score with 59% lower compute cost. On a 24-hour task where a swarm of agents optimizes six interdependent repositories, it achieved 65% greater speedup on the same compute budget.
Steering, training and serving
CLM behavior can also be steered with text instructions: the user simply describes the strategy they want. A skill-optimization loop improves such instructions itself; on the new ContextBench benchmark accuracy rose by up to 35.9 points while compute fell.
A new online reinforcement learning method improved the Qwen3.5-9B model's result on BrowseComp-Plus by 47.6% and cut FLOPs use by 12%.
During serving, CLM edits its context mid-run, which forces standard serving systems to recompute cached states. The authors developed Suffix Cache Reuse (SCR) for SGLang: it keeps the cache of spans that remain unchanged after an edit and cuts server-side compute by 35% compared with standard SGLang, at matched performance.
The paper's code is published on GitHub. In practice, context management for long-running agents will lean more on the model itself and less on hand-built external systems, which also lowers compute costs.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.