Google researchers show self-improving agents memorize their test tasks. Your prompt tuning can do the same
If you keep tweaking a prompt until your five test examples pass, you may be teaching it those five examples and nothing else.
What happened
The Decoder reports on a paper from Google Cloud AI Research and several universities about agent harnesses: the prompts, workflows, tools, memory and logic wrapped around a fixed model. Newer methods let a model rewrite its own harness over and over based on test task feedback. The paper finds this leads agents to memorize their test tasks. Training scores go up while gains on new tasks shrink or vanish, partly because the search favors lucky candidates and piles on complexity that only helps the benchmark.
Their method, RRSI, adds limits. A budget caps how many edits a candidate change can bundle, and it shrinks over time so late changes are small and traceable. A critic rejects proposals that hardcode task names, solutions or benchmark tricks. Higher compute is accepted only with a measurable gain, and components that stop helping get removed.
Tested on eight benchmarks with Claude Opus 4.8 frozen, RRSI had the smallest training gain of the variants but was the only one well above baseline on unseen tasks, gaining up to 4.7 points there and up to 14.1 on training tasks. It used about 30 percent fewer tokens than the unregularized version. The code is on GitHub.
My take
You do not need a research lab to fall into this. Most AI automations get tuned the same way: run a few real examples, adjust the prompt, repeat until they look right. Then a new kind of lead or invoice shows up and the step quietly fails.
What I take from RRSI for everyday builds:
- Keep a holdout set. Tune on some examples, and judge on others the prompt has never seen.
- Change one thing at a time. Small, traceable edits beat big rewrites you cannot explain.
- Delete special cases. If a prompt names a specific client or edge case to pass a test, that is memorization.
- Make cost earn its place. A longer prompt or extra step stays only if the holdout results improve.
- Prune. Old instructions that no longer help just add tokens and confusion.
Have an AI step that works on your examples but breaks on real traffic? Tell me what it does at romielwillautomate.dev and I will suggest how to test it properly.
More posts
- Google's EmbeddingGemma 2 runs search and RAG on device in about 191MB of RAM. Private client data can stay localOct 7, 2026
- Customers' AI agents are getting blocked by human checks and bot defenses. Your client's site may be turning buyers awayOct 7, 2026
- Siena raised $17M to give support, shopping and social agents one shared memory of each customer. Your CRM should do the sameOct 7, 2026
