
Google Researchers Tackle Test Memorization in Self-Improving AI Agents
Google Cloud AI Research and university partners propose RRSI, a method that regularizes how AI agents rewrite their own harnesses, preventing test-task memorization while cutting token use by about 30 percent.
Google Cloud AI Research and several universities have published a paper describing a method that keeps self-improving AI agents from memorizing their test tasks while reducing compute costs. The approach, called RRSI (Regularized Recursive Self-Improvement of Agent Harnesses), targets the software layer around a language model rather than the model itself.
Self-optimization leads to memorization
Modern AI agents wrap a fixed language model in a harness, a framework of prompts, workflows, tools, memory, and logic that controls what the model sees at each step. According to the paper, much of the recent progress in agents comes from work on the harness, not from new models. Newer methods automate the loop by having a language model rewrite the harness itself based on feedback from test tasks, a practical form of recursive self-improvement.
The researchers found a catch. Because the agent keeps working on the same limited set of test tasks, it ends up memorizing them. Scores on training tasks go up while gains on new, unseen tasks shrink or disappear. The search memorizes patterns that fit only one benchmark, favors candidates that score well by chance, and adds unnecessary complexity that raises test scores without making the agent better.
Shrinking budgets and a strict critic
RRSI works on both ends of the optimization loop while leaving the harness fully editable. A budget caps how many independent edits a candidate can bundle at once, and that budget shrinks over time. Early rounds allow larger rewrites, while later rounds permit only small changes that can be clearly traced to a result. The system tracks earlier attempts to avoid chasing the same failed ideas and deliberately experiments with untouched parts of the harness when progress stalls.
When picking changes, a critic reviews every proposal and rejects any that hardcode task names, solutions, or other benchmark-specific tricks. Another rule accepts higher compute costs only with a measurable performance gain, and components that no longer help get removed.
Gains on unseen tasks
The team tested RRSI on eight benchmarks spanning coding, agentic office work, and engineering design, with Claude Opus 4.8 frozen throughout. RRSI gained up to 14.1 points on trained tasks and up to 4.7 points on five unseen benchmarks, while using about 30 percent fewer tokens at runtime than the unregularized version. Overall performance never fell below the baseline on any unseen benchmark, and the biggest gain of 4.7 points came on JobBench.
Every compared method did well on training tasks, but results flipped on new ones, with two methods ending below the baseline harness. RRSI posted the smallest training gain of all variants and was the only method well above baseline on unseen tasks.
A coding harness optimized with Gemini 3.5 Flash raised the accuracy of the much weaker Gemini 3.1 Flash Lite from 11.2 to 14.6 points without modifications, suggesting the discovered mechanisms do not depend on the capability of the model used to find them. The authors note the study covers only harnesses around frozen models and does not address changing model weights. The code is available on GitHub.
Related work
Manually designed harnesses often fail to generalize. Tests on ARC-AGI-3 showed Opus 4.6 scoring 97.1 percent in a familiar environment and 0 percent in an unfamiliar one with a purpose-built harness. Nvidia recently presented SoL-Pi, where a research agent automatically rebuilds the harness of coding agents, cutting token use by up to 49 percent without a noticeable performance drop. Shortly before that, Google had agents dream about past search runs to improve their search strategy, also leaving the model unchanged.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.