All posts
2 min readby Romiel Inolino

AI proposed, humans decided: what 700+ task logs from an AI lab say about human in the loop design

AI AgentsHuman in the LoopResearchAutomation DesignGovernance

More agent activity is not the same as more agent autonomy. A new study puts numbers on that.

What happened

The Decoder reports on a study by the team behind Atria Dawn Preview, a 744 billion parameter mixture of experts agentic model from the Shanghai Artificial Intelligence Laboratory, with researchers from Fudan University involved. The team analyzed more than 700 task logs from 56 participants, plus the logs of the agents they used.

The findings:

  • AI was used in 96.5% of reviewed tasks. The median number of agent actions per human input rose from 11 to 28.5 over four weeks, but the authors caution this is not growing autonomy.
  • Of 455 completed AI assisted tasks, 151, roughly a third, were rated infeasible without AI.
  • The most common pattern for methods and parameters was "AI proposes, human selects" at 55.4%. Humans made 85.5% of those decisions and AI made 9.2%. Humans set goals and scope in 93.4% of cases.
  • When tasks got stuck, human intervention moved work forward in 76% of cases, mostly by adding context or diagnosing problems. Full human takeovers were just 0.7%.

The authors also flag a "rubber-stamp risk": when agent chains get too long to review, humans can end up approving what they cannot check. Many participants ran agents in autonomous modes to avoid constant approvals, a boundary the Decoder notes was drawn out of convenience, not deliberate choice.

My take

This lines up with how the best business automations actually run. The agent drafts, sorts and proposes. A person picks, approves and adds context. The value is not that the human does less thinking, it is that they stop doing the typing.

Two design lessons I take for client workflows:

  1. Put the human at the decision, not at every step. Approving 40 micro actions per run trains people to click yes without reading. One clear approval with a short summary of what will happen is better.
  2. Decide the autonomy boundary on purpose. Write down which actions run on their own and which wait for a person, based on cost of a mistake, not on what feels annoying this week.

Also note the third of tasks that would not have happened at all. That is often where the real return sits: the report nobody had time to build, the follow up nobody had time to send.

More posts