Enterprise automation demos usually show the happy path: a clean input arrives, the system recognizes it, applies a rule, and completes the task. Real operations are defined by everything that does not fit the demo.
An invoice arrives from a new address. The purchase order covers only part of the delivery. The amount crosses an approval threshold. Tax treatment changes by entity. A familiar supplier uses an unexpected currency. The process is still recognizable, but the correct action depends on context.
Useful agents must be designed for that exception load.
A workflow is more than a sequence of steps
A process diagram may show intake, matching, coding, checks, approval, and system entry. Each box hides decisions learned through policy and experience.
Operators know which discrepancy is harmless, which approver understands a specific category, when to pause for evidence, and how to document an unusual case. That knowledge often lives in messages, personal notes, and memory rather than the formal procedure.
Automating the visible steps without capturing those decision conditions creates a brittle system. It moves quickly until reality deviates from the template.
Model the decision context
For every step, document four elements:
- Required evidence: the fields, documents, and source systems needed to act.
- Policy: the explicit rule or threshold that governs the decision.
- Known exceptions: recurring situations that change the path.
- Escalation: the person or role that owns ambiguity.
This model creates a boundary between what the agent can execute and what still requires judgment. It also reveals missing policies that automation would otherwise disguise.
Separate memory from authority
An agent may remember how a previous exception was resolved. That does not mean it should automatically turn the resolution into permanent policy.
Operational memory should preserve the case, context, decision, approver, and outcome. Policy authority should remain versioned and explicitly approved. New exception patterns can be proposed for review, then promoted into reusable rules after an accountable owner accepts them.
This distinction prevents one unusual decision from silently changing future behavior.
Build bounded permissions
Agents that take action need narrower permissions than the people they assist. Use least-privilege access, task-specific credentials, value thresholds, and explicit limits on which systems or records may be changed.
Separate read, recommend, prepare, and execute modes. An early pilot may retrieve documents and prepare a recommendation without posting it. A mature workflow may execute low-risk matches while routing unusual or high-value cases for approval.
Permission should expand only after observed reliability, not because a demo succeeded.
Make uncertainty operational
A confidence score is not useful unless it changes behavior. Define what happens when evidence is missing, sources conflict, or the situation falls outside known policy.
The agent should be able to stop, explain what is uncertain, show the evidence it used, and ask a focused question. “Unable to complete” is less useful than “the purchase order quantity does not match the delivered quantity; confirm whether partial delivery is expected.”
Good escalation reduces the cognitive load on the human instead of transferring the entire task back.
Preserve an audit trail
Every consequential run should record the inputs, retrieved context, policy version, decisions, tool calls, approvals, exceptions, and final result. Logs should be understandable by an operator or auditor without reconstructing an opaque chain of prompts.
Auditability serves more than compliance. It helps the team diagnose whether a failure came from missing data, an incorrect rule, a retrieval problem, model reasoning, or an unauthorized action.
Without that visibility, improvement becomes guesswork.
Evaluate exception performance
Happy-path completion rate can look impressive while hiding the work pushed onto humans. Measure exception detection, false approvals, unnecessary escalations, time to resolution, correction cost, and whether the agent provides enough context for a reviewer to decide.
Also track policy drift. If operators repeatedly override the same rule, the process may need redesign rather than another prompt adjustment.
The objective is not zero human involvement. It is to use human attention where it produces the most value.
Start with observation
Before granting execution access, run the agent in shadow mode. Let it observe inputs, propose decisions, and compare its path with the real operator’s work. Review disagreements and classify them: missing context, unclear policy, novel exception, or agent error.
Shadow mode produces an exception inventory and tests whether the process is documented well enough to automate. It also helps employees see the system as a visible collaborator rather than an unexplained replacement.
What to do next
Choose one bounded operational workflow with a clear owner. Map its evidence, policies, known exceptions, and escalation paths. Run an agent in recommendation-only mode, log every disagreement, and review the results weekly.
Do not ask first whether the agent can perform the happy path. Ask whether the organization can explain—and responsibly govern—what happens when the path breaks.