Back
Context Language Models: LLMs That Manage Their Own Context
SiTech AI Team3 min read

Context Language Models: LLMs That Manage Their Own Context

Researchers from the University of Washington, MIT and Meta introduced Context Language Models (CLMs): models that keep their context as a file and edit it themselves. On BrowseComp-Plus it scored 11.4% higher accuracy with 21.5% fewer FLOPs.

Researchers from the University of Washington, MIT and Meta have introduced Context Language Models (CLMs), language models that natively manage their own context. The paper was published on arXiv on 29 September.

Until now, context has been managed from outside the model, by hand-built systems or constrained tools. In a CLM, that job belongs to the model itself.

Context as a file

In a CLM, the context is stored as a file. The model can change it freely: it either appends new tokens or edits any span with Bash, and every change is immediately reflected in its live context for the next turn.

Because the context is a separate file, several such files can coexist in one system, which extends the approach to multi-agent setups. In tests, models showed new behaviors: they created trackers for multi-agent orchestration, introduced a new chat role for internal notes, and defined reusable context-management functions.

The authors say the result confirms the idea behind the "Bitter Lesson": freedom lets the model find better strategies than rules written by humans.

Results

CLMs work without additional training on existing models such as Qwen3.6-27B and GPT5.6-Sol. On the deep-research benchmark BrowseComp-Plus, CLM scored 59.4% under a 32K context limit: 11.4% higher accuracy than the strongest baseline at 21.5% fewer FLOPs.

On TerminalBench 2.1 it matched the strongest method's accuracy while spending 29.5% fewer FLOPs. On the 12-hour EdgeBench it showed a 5% higher score with 59% lower compute cost. On a 24-hour task where a swarm of agents optimizes six interdependent repositories, it achieved 65% greater speedup on the same compute budget.

Steering, training and serving

CLM behavior can also be steered with text instructions: the user simply describes the strategy they want. A skill-optimization loop improves such instructions itself; on the new ContextBench benchmark accuracy rose by up to 35.9 points while compute fell.

A new online reinforcement learning method improved the Qwen3.5-9B model's result on BrowseComp-Plus by 47.6% and cut FLOPs use by 12%.

During serving, CLM edits its context mid-run, which forces standard serving systems to recompute cached states. The authors developed Suffix Cache Reuse (SCR) for SGLang: it keeps the cache of spans that remain unchanged after an edit and cuts server-side compute by 35% compared with standard SGLang, at matched performance.

Suffix Cache Reuse (SCR) versus standard SGLang serving

The paper's code is published on GitHub. In practice, context management for long-running agents will lean more on the model itself and less on hand-built external systems, which also lowers compute costs.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.