Back
Why large context windows are a budget, not a workspace
SiTech AI Team3 წთ. საკითხავი

Why large context windows are a budget, not a workspace

A developer note argues that language models lose accuracy long before their advertised context limits. The practical answer: keep sessions short and hand off written specs instead of raw history.

A note published on Garrit's Notes on May 6, 2026 argues that the headline numbers attached to large language models' context windows describe capacity rather than usable working space. The author credits a video that split the window into two zones: a "smart zone", where the model stays sharp, and a "dumb zone", where attention drops off and the model starts forgetting instructions given minutes earlier. In his account, the boundary sits near 100,000 tokens.

Agents reach the limit before lunch

The gap matters most for coding agents, which burn through tokens quickly. A few file reads, a long debugging session and a sprawling test run can push a session past 100k tokens before lunch. Vendors, meanwhile, advertise windows of 200k, 1M and even 2M tokens, as if those numbers described a usable working set.

The author points to published research, including the RULER benchmark and Chroma's report on "context rot", which he says show that effective context is a fraction of the advertised figure and that performance degrades gradually as the window fills. His conclusion is that the bigger number is largely a marketing figure, and that the architectures behind it do not solve the underlying attention problem.

Compaction helps, but arrives late

Modern agents have begun to compensate. Tools such as Claude Code auto-compact: when a session grows long, the agent summarises its history and starts fresh. That helps, the note says, but compaction triggers after time has already been spent in the degraded zone, and the summary is written by a model that is itself already degraded. Better than nothing, the author adds, but not the situation he would choose.

His alternative is to open a new session and hand it a specification he wrote himself. That is a higher-signal transfer than any automated summary, because a person decides what still matters going forward. He describes it as the breadcrumb approach applied to agents: leave an artifact that the next session, or the next person, can pick up cleanly.

Artifacts instead of history

Projects such as obra/superpowers and mattpocock/skills organise entire agent workflows around small named artifacts — product requirement documents, plans, skills and sub-agent handoffs. Each one moves information out of the live session and into something the next session can read, which keeps the working context inside the zone where the model is still sharp. The author's rule is to treat the context window as a budget: assume only the first part of it is really working, and move everything else into writing.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.