How to choose your first AI agent use case

Compare candidate tasks, define the agent's permissions and complete a pilot brief with a baseline, review process and decision criteria.

Written by
Agentic Economy
Updated
View Markdown
On this page

Choose a recurring task with a clear result, accessible information and someone who can judge the work. Then test whether an agent improves that task enough to justify its running and review costs.

This guide produces a pilot brief you can discuss with your team or a supplier. It adapts the three comparison dimensions in Microsoft's business-planning guidance: business impact, technical feasibility and user desirability. The worksheet and filled example are our practical application of that framework, not a tested deployment or a Microsoft template.

Before you begin

Bring the person who does the work and someone who can authorise access to its information. Gather examples of normal cases, exceptions and work that needed correcting. You do not need to buy a platform to make this first decision.

If you are still building a shortlist, start with practical business use cases. If the steps already follow clear rules, compare fixed automation with an agent before proceeding.

1. Turn broad ideas into specific tasks

Write down a few pieces of work that create repeat effort. For each, record the trigger, input, result and person who uses it.

Replace “automate sales” with “prepare a proposal outline from an approved service catalogue and the customer's request”. Replace “handle repairs” with “prepare a checked equipment record before a booked visit”. These are illustrative task descriptions; the second will be our worked example.

Ask the person doing the job where their next step changes. Do they have to find a missing record, reconcile conflicting information or choose which question to ask? If you cannot identify such a decision, a simpler tool may cover the need.

Check: someone familiar with the task should be able to recognise when it begins and what a satisfactory handover contains.

2. Compare impact, feasibility and desirability

Use Microsoft's three dimensions to compare candidates. Its framework suggests a scale from one to five. A short written assessment can be more useful initially if you do not have evidence to justify a precise score.

Dimension Questions to answer Evidence to bring
Business impact What problem would improve? Does it matter enough to fund? Current delay, staff effort, corrections or missed opportunities.
Technical feasibility Can the system access suitable information and operate within the required boundaries? Example inputs, system access, known failure cases and a reviewer.
User desirability Will the people involved use and support the changed process? Their account of the problem and the proposed handover.

Record unknowns rather than scoring them optimistically. A promising task with inaccessible records is not yet ready for a meaningful trial. A technically simple output that nobody needs should not win because it is easy to demonstrate.

For our illustrative repair business, appointment preparation could be a candidate because the coordinator already gathers the information, the source records exist and the technician can describe a useful handover. Whether the task consumes enough effort to justify the trial remains something to measure.

3. Define the permitted contribution

Write what the agent may read, produce and change. Name what requires a person. Confirm that the chosen system can enforce those boundaries before giving it access.

For the repair trial, begin with historical cases and a draft-only output. The agent may inspect supplied booking records and photographs. It may identify missing or contradictory details and propose a follow-up question. It may not contact customers, diagnose faults, order parts or change bookings.

That is a deliberately limited trial of preparation. Sending an information request could be tested later with separate authority. If your task requires paid services, add a scoped spending allowance and decide what information may leave the business.

Check: the operator and reviewer should agree on an example of an allowed action and an example that must stop or escalate.

4. Establish a baseline and test difficult cases

Record how the current process performs before claiming improvement. Microsoft recommends linking pilot decisions to business measures. Choose a few that describe the completed job: staff minutes including corrections, usable handovers and unresolved exceptions.

Use the same completion criteria for the existing process and the trial. Include straightforward records, incomplete inputs and conflicting evidence. A reviewer should check the underlying records rather than accept the agent's statement that the work is complete.

Anthropic's evaluation guidance recommends clear success criteria, examples drawn from actual work and tests of when actions should and should not occur. For the repair case, correctly flagging uncertain equipment identification can be a successful result.

Do not choose only the easiest bookings. The ambiguous records are where the proposed agent's flexibility needs to prove useful. Keep historical records supplied to the pilot separate from any live system it could change.

5. Complete the pilot brief

Copy the fields below and replace the example with your task. This is an illustrative brief; its proposed boundaries and decision criteria are not measured results.

Field Filled repair-business example
Task and trigger Prepare equipment details when a booked visit has incomplete information.
Output Equipment record with source references, conflicts and a proposed follow-up question.
Owner and reviewer Service coordinator owns the trial; a technician checks equipment identification.
Inputs Approved historical booking records, previous visits and supplied photographs.
Permissions Read the supplied material and produce a draft. No customer messages, purchases or booking changes.
Baseline Measure staff preparation and correction time, plus completeness of the current handover.
Cases Include complete records, missing details, conflicting records and unclear photographs.
Quality condition Correctly supported details or an explicit unresolved question; no unsupported model number presented as verified.
Value condition Reduce total staff effort while maintaining the agreed handover quality.
Review point Review the agreed case set before allowing any live action.
Stop or redesign Any unauthorised action; repeated unsupported identification; or checking effort that removes the expected benefit.

Set a review date and agree who will make the decision. For a live pilot, specify its duration, workload and cost ceiling before it starts.

6. Decide what the evidence supports

At review, separate an output-quality problem from an access problem or an unsuitable task. Each points to a different response: improve the checks, fix the information boundary or choose another use case.

You should now have one candidate, a bounded contribution and an agreed way to assess it. A successful draft-only trial supports the next controlled step. It does not establish that the agent should receive every permission needed to complete the wider business process.

If the next step needs an outside capability, read how AI agents buy services before adding a supplier to the workflow.