Skip to article
All insights
AI agents3 min read

Enterprise AI agents: Which workflows are worth a pilot?

An inbox full of exceptions, scattered project knowledge, repeated requests for missing information: agents become interesting where people assemble evidence and decide what happens next. Whether they earn their place depends on the workflow. Choosing that workflow is already a significant investment decision.

At a glance

  1. 01

    Choose a clearly defined work product with an accountable team.

  2. 02

    Assess how reliably outcomes can be judged and mistakes contained.

  3. 03

    Measure the benefit after review and rework in the actual process.

Start with a work product

“Automate customer service” is too broad for a pilot. “Prepare an evidence-backed complaint file for a case handler” has a beginning, an outcome and a recipient. That definition reveals the required data, who judges quality and where responsibility changes hands. Real, anonymised cases are especially useful: what do they share, where do exceptions begin, and which information is repeatedly missing?

Anthropic distinguishes predefined workflows from agents that choose their own steps and tools. For workflow selection, that makes an agent plausible when information needs vary. Where rules are stable and variations limited, conventional automation should also be part of the comparison.

Exhibit 01

Workflow and consequences determine the scope

Two questions determine the approach.
Predictable workflow
Variable workflow
High impact
Rules + approval

Bound execution. Assign responsibility.

Assistance + approval

Prepare options. Keep people in authority.

Limited impact
Conventional automation

Defined steps and testable rules.

Pilot candidateBounded agent

Handle variation. Verify the outcome.

Conceptual selection framework: workflow variability and potential consequences shape the appropriate scope of action. Value and verifiability are assessed alongside them.

Make quality recognisable in advance

A team should recognise a good result before seeing its first agent response. For a prepared complaint file, that might mean completeness, correct account matching, supported statements and an appropriate next action. Fluent prose cannot establish those qualities. If specialists disagree about the expected outcome, the definition of the work needs attention first.

The NIST AI RMF connects operating context, evaluation and risk management. Our practical interpretation is to assemble representative cases, known exceptions and a clear rejection rule. An agent must make missing information visible. Distinguish a false statement from an unnecessary referral to manual handling: both impose costs, but in different ways.

Match authority to consequences

For an initial deployment, what happens after an output matters greatly. An internal draft can be discarded. A sent commitment, changed order or released payment creates different consequences. Divide the workflow at those boundaries. Research, recommendation, approval and execution can each require different permissions.

In the complaint example, an agent could find documents, flag inconsistencies and prepare a response. A designated person would initially retain authority over refunds. This is an illustrative scenario. Its suitability depends partly on document access respecting existing roles and on uncertain cases actually being stopped. Human review only helps when the reviewer has the time, information and responsibility to perform it.

Measure value where the work finishes

The relevant unit is a correctly completed case. Establish current handling time, clarification requests and rework, then compare the same type of case with agent assistance. Include time spent reading, checking and correcting, as well as handovers to other teams. Faster preparation can lose its advantage if verification becomes more demanding.

Prioritise a workflow where effort recurs, outcomes can be judged and mistakes remain manageable. It also needs a business owner who will decide what happens after the pilot. Ask whether released capacity has a practical destination: shorter queues, better coverage or work that currently goes undone. The first useful assignment may be small. Its result should support a credible decision about the next expansion.

Your next step

Assess a workflow for agent fit

An anonymised description of the task, expected outcome and current process is enough for an initial pilot conversation.

Discuss a pilot

Sources & further reading

  1. Anthropic · Building effective agents

    Supports the distinction between predefined workflows and dynamic agent control.

  2. NIST · AI RMF Core

    Connects operating context, evaluation and ongoing risk management. The selection questions are our editorial interpretation.