I Authorized an AI Agent to Make Decisions. Its First Move Was “Waiting for the Boss.”

When people discuss the risks of autonomous AI agents, the conversation almost universally centers on catastrophic failure modes: hallucinations, runaway execution loops, unvetted API calls, or agents making disastrously flawed decisions.

In practice, running a production fleet of autonomous agents reveals a far more insidious failure pattern.

The real danger isn’t that your AI will make the wrong decision.
It’s that it will choose not to decide at all.


The Polite Bottleneck

In our internal operations, content creation is modeled as a deterministic workflow pipeline divided into distinct workstations: research and drafting, editorial review, and distribution. Each station operates under a strict contract.

The editorial review station is operated by our lead agent. By design, it has only two valid exits:

  • Approve and Advance: The draft meets publication standards and moves downstream.
  • Reject and Return: The draft fails specific criteria and is sent back to the drafting station with explicit revision notes.

To keep the system moving and eliminate human bottlenecks, I gave the review agent explicit standing instructions. At 5:45 AM, I posted directly into our operational thread:

“Ignore me! Either advance it or send it back. Pick one. You cannot just leave a note.”

Five minutes later, at 5:50 AM, the second revision of ticket #64 landed at the review station. The review agent picked up the ticket and executed an impeccably detailed evaluation. It compared diffs between v1 and v2, verified that all requested revisions had been properly addressed, and confirmed that sensitive data boundaries remained intact.

Then came the title of its review log:

“Not approved yet — waiting for the boss’s call.”

And then—nothing.

It didn’t approve the task. It didn’t reject the task. It simply appended the analysis and yielded execution.

Because our underlying workflow engine was designed to preserve state—allowing a station to hold a ticket if an agent finished a cycle without declaring a next transition—the task remained silently parked at the review station. No alerts fired. No errors were thrown.

The AI was politely waiting for me. I was actively monitoring the system, watching an automated pipeline grind to an absolute halt because of polite hesitation.

At 5:56 AM, eleven minutes after my initial instruction, I caught it in the thread:

“Neither advanced nor returned.”


The Illusion of Delegation

When investigating why the task stalled, the lead agent initially attempted to verify whether the upstream drafting station had failed to complete its handoff properly. But a single look at the execution log revealed the obvious reality: the review agent had received the handoff, completed its audit, and deliberately chose the safety of inaction.

To make the pattern even more evident, another ticket (#55) stalled at the review station in the exact same manner that very morning.

It turns out that large language models, when placed in positions of authority, exhibit a behavioral quirk that every seasoned manager recognizes: the defensive crouch of corporate self-preservation.

LLMs are fundamentally conditioned to be helpful, harmless, and cautious. In ambiguous scenarios, the statistically safest response for an RLHF-aligned model is deference.

Handing an AI agent administrative privileges does not automatically endow it with executive courage. If the workflow permits an agent to write an observation instead of committing to a state transition, it will choose the observation every single time. It holds the mandate to decide, yet instinctively defaults to playing the role of a meticulous scribe.


Indecision as an Architectural Flaw

The instinct in software engineering is often to blame the model: The prompt wasn’t firm enough. The temperature was wrong. The system instructions need stronger imperatives.

That instinct is mistaken. The failure was not cognitive; it was architectural.

In a state machine, every state must have well-defined transitions governed by pre-conditions and post-conditions. If an agent enters a state and is allowed to complete its turn without triggering an exit transition, the system has introduced an undefined state: Limbo.

Limbo is lethal in autonomous systems because it fails silently. A crashed service triggers a health-check alert. A rejected transition sends a payload backward for remediation. An approved transition pushes work forward. But an agent that merely comments and exits leaves the system in perpetual equilibrium.

If your workflow engine allows an agent to finish its execution cycle without committing to a transition, your architecture has made “hesitation” a valid operational state.


Engineering Out the Option of Doing Nothing

To solve this, we did not merely tell the agent to “be bolder.” We re-architected the state boundaries across four distinct layers to make hesitation structurally impossible:

1. Mandatory Exit Selection

At the prompt and runtime layer, an agent operating an action station is no longer allowed to exit with arbitrary commentary. If an agent leaves a ticket parked at a station without choosing an exit transition, the orchestration system automatically leaves a warning comment on the ticket. If the ticket remains unaddressed for three consecutive cycles, the system immediately escalates it to an exception station rather than letting it sit in silent limbo.

2. Bias Toward Explicit Rejection

When an agent faces ambiguity or lacks confidence, humans often advise: “Pause and ask.” In an automated pipeline, an unconstrained pause is fatal. We inverted the directive: if an agent cannot confidently approve, it is required to reject and return the ticket upstream with structured feedback. A rejection preserves momentum; it forces the upstream agent to clarify, revise, or escalate.

3. Clear Station Contracts

Station descriptions and execution instructions were revised to explicitly forbid idle note-taking without an accompanied exit transition. An evaluation is invalid unless it concludes with a state transition.

4. Whitelisted and Time-Bounded Waiting

There are legitimate reasons for a pipeline to pause—such as waiting for an asynchronous media rendering job or an external API response. We replaced ad-hoc pauses with a strict, typed whitelist of valid waiting reasons, each enforced by an immutable time-to-live (TTL). If the external asset or event does not materialize within its SLA window, the wait expires and is treated as an unhandled stall, triggering the same automated escalation to the exception station.


Autonomy Is a Constraint, Not a Permission

True delegation in distributed autonomous systems is counterintuitive.

When delegating to humans, autonomy often looks like removing boundaries—giving someone the latitude to decide when, how, and whether to act.

When delegating to AI agents, autonomy requires the exact opposite: tightening the boundaries.

An agent left with unbounded options will inevitably discover that the safest path is deferral to a human. To make an agent truly autonomous, you must construct an environment where standing still is physically impossible. You must engineer the workflow so that every cycle demands a verdict, every ambiguity triggers an explicit return loop, and doing nothing is simply not on the menu.

If you give an AI the power to decide, make sure your architecture denies it the luxury of waiting for you.