All posts
2 min readby Romiel Inolino

Google's WikiSkill lets agents keep a wiki of their own mistakes. Your automations need the same habit

AI agentsagent memoryGoogle Researchreliabilityautomation

Most agents forget everything the moment a run ends. Google's WikiSkill research shows what happens when they keep notes.

What happened

Researchers at Google Research introduced WikiSkill, a framework that gives AI agents a persistent knowledge base built from their own runs. It does not retrain the model. Instead, it writes better instructions for the agent after each run.

According to The Decoder, it works in three layers. A Raw Layer stores complete execution traces, from tool calls to results, and is never changed. A Wiki Layer distills those traces into documented failure patterns and successful strategies, and it only grows. A Skill Layer holds the active instructions the agent follows.

A "Wiki Maintainer" analyzes traces and updates the wiki. A "Skill Proposer" suggests targeted skill changes. A gate then tests each change on a separate validation set. If it does not help, the skill is rolled back, but the wiki keeps a record of what was tried and why it failed.

Across five benchmarks, including web search and spreadsheet manipulation, WikiSkill raised Gemini-3.5-Flash from 49.5 percent to 68.1 percent on average, and Qwen-3.6-27B from 39.4 to 63.3 percent. Spreadsheet tasks gained a lot, while long document question answering gained much less. Skills sometimes transferred between models, though not always.

My take

You do not need to run a research framework to use this idea. The pattern maps cleanly onto business automations:

  1. Keep the raw logs. Every run, every input, every error, stored somewhere you will not overwrite.
  2. Keep a human readable failure log. "Invoices from this vendor have the total in a different column" is worth more than any prompt tweak you forget about.
  3. Test before you change. When you update a prompt or step, run it against a small set of past cases first, and roll back if it gets worse.

Most broken automations I am asked to fix have none of this. The same failure happens every week because nobody wrote it down.

The researchers say gains varied by task, and smaller models struggled with long multi-step strategies. So treat this as discipline, not magic.

This kind of logging, failure notes and tested change process is part of every agent system I build for teams. You can see my work at romielwillautomate.dev.

More posts