All posts
2 min readby Romiel Inolino

"Done, with one asterisk" often means not done. A Harvard physicist's rules for checking AI agent work

AI agentsverificationClaudequality controlautomation

The most practical line in a physics essay this week: "Done, with one asterisk" is often "not done at all."

What happened

Harvard physicist Matthew Schwartz described BootLoops in a guest post on Anthropic's site. BootLoops is an open source harness for exact calculations in quantitative science; Schwartz owns and maintains it and says it is not an Anthropic project. Working with Claude and experts in other fields, his team produced 36 manuscripts across 18 fields with 19 coauthors over three months, chosen from some 400 candidate problems.

The process mattered as much as the output. Each project had its own Claude Code session, plus a master session that coordinated the others and validated results, and separate sessions for checking results "as an adversarial referee." He found Claude was usually technically correct outside his field, but results often were not interesting until a domain expert steered the work.

His warnings are blunt. Claude "loves to declare victory." "Exactly that, with one refinement" usually means no. Clear and rigid standards for success help. "Look at everything yourself": he asks to see plots, and says even with monitors in place, automated checks still cannot be trusted. He also notes the projects were compute and token intensive, and that Claude tended to grind through long calculations instead of building a tool to do them faster.

My take

Swap "proof" for "lead list" or "invoice reconciliation" and every point applies to business automation.

  1. Write the success test before the run. "Clean the CRM" invites a confident report. "Zero contacts without an owner, zero duplicate emails" can be checked.
  2. Ban hedged completions. If an agent's summary includes "mostly", "one exception" or "with a refinement", treat it as failed and route it to a person.
  3. Use a separate checker. Schwartz used a dedicated referee session. In n8n or Zapier that is a second step that re-queries the data and compares it to what the agent claims.
  4. Look at real samples. Pull five records the agent touched every week and read them yourself.
  5. Ask for tools, not grinding. If an agent loops through 10,000 rows by hand, a script will be cheaper and repeatable.

This is how I design AI automations for clients, whether it is lead enrichment, reporting, data entry or content pipelines: every agent step gets a test it has to pass, written before it runs. More of that work is at romielwillautomate.dev.

More posts