Agent reliability reports land, and long-task failure rates stay high

Multiple tests showed agents still fail often on long, multi-step tasks, with reliability widely accepted as the main obstacle to adoption.

September 18, 2025 ·Source:Public reports AgentsReliabilityEvaluation

Multiple independent tests showed agents still fail often on tasks requiring dozens of steps, with errors accumulating along the chain. Reliability is widely accepted as the main obstacle to deploying agents.

Key facts

  • Failure rates stay high on long multi-step tasks
  • Errors accumulate along the execution chain
  • Reliability decays exponentially rather than linearly
  • Few effective mid-course correction mechanisms

Big Share take

Reliability is agents’ biggest weakness and least discussed, because it demos badly. An agent that succeeds ninety percent per step is down to barely ten percent over twenty steps, and that arithmetic decides what it can do.

Compiled from public reports; opinions are for reference only.

Share on Telegram ↗