# AI agent adoption: demand, evidence and the measurement problem

What current AI-agent adoption evidence measures, what it misses and why accepted jobs remain the clearest operating denominator.

Source: https://aecon.ai/guides/ai-agent-adoption-measurement
Author: Agentic Economy
Published: 2026-09-13
Updated: 2026-09-14

**Chapter 2 of The Hitchhiker's Guide to the Agentic Economy**

[Report hub](/guides/hitchhikers-guide-agentic-economy-2026) · Previous: [what is the agentic economy?](/guides/what-is-the-agentic-economy) · In this chapter: [current evidence](#section-what-the-current-evidence-can-establish) · [Australian baseline](#section-an-australian-baseline) · [measurement framework](#section-agentic-economys-market-measurement-framework) · [job record](#section-the-underlying-job-record)

## Our view

Agentic activity is visible across software development, customer service, research, commerce and enterprise operations, but the market does not yet have a common measure of adoption. Current evidence mixes product availability, assistant use, pilots, self-reported scaling, tool calls, payments and completed work. Those events describe different stages of the market and cannot support one headline number.

**AI agent adoption** is the recurring use of an agentic system within a defined population and workflow. Any reported rate needs to name the population, the qualifying action and the observation window.

The job is the most useful unit of measurement. It connects what software was permitted to choose, the action it took, whether the result was accepted, the human effort that remained and whether the customer or team returned. That record remains comparable even while industry-wide measures are immature.

> **Key Takeaways**
>
> - Published studies use different denominators for agent adoption in 2026; each supports a different claim.
> - Product availability, reported use and accepted economic work are different market events.
> - Australian evidence points to early, piecemeal use and mixed realised returns among the medium and large firms consulted by the RBA.
> - Transaction, token and tool-call counts can show activity without proving independent demand, successful delivery or profit.
> - The strongest operating measure is the accepted outcome per unit of cost and human intervention.

![A circular freight interchange representing activity that must be traced to an accepted destination](/media/editorial-library/freight-curve-full.webp)

### Sources and scope

Agentic Economy compared official Australian evidence, independent evaluations, provider telemetry and survey research. Each number is kept inside its original population, event and study design. The market-measurement framework and job record are Agentic Economy analysis dated 13 September 2026. [About Agentic Economy](/about) explains the publication. Readers can submit corrections through the [contact page](/contact).

## The market is measuring different things

Agent adoption figures currently describe several different markets at once. Surveys measure what organisations say they are exploring or scaling. Product telemetry measures behaviour inside one provider's environment. Benchmarks measure performance on a defined task set. Payment networks, marketplaces and infrastructure providers count transactions, listings, calls or tokens.

These signals are all useful. They are not interchangeable. Our conclusion is that the market has clear evidence of experimentation and expanding product supply, but far less public evidence connecting agent activity to accepted work, repeat demand and durable economics.

## Why the headline numbers do not add up

One report may count organisations that use generative AI. Another may ask executives whether they are scaling agents. A provider may count sessions, tool calls or tokens across its own customers. A payment network may report transactions. A marketplace may publish listings. Each can be accurate within its method and still answer a different question.

The word *adoption* is doing too much work. It can describe:

- a person trying an AI assistant;
- a team using AI regularly inside an existing process;
- a fixed workflow calling a model;
- a system selecting tools under human approval;
- an unattended production process;
- an agent initiating a purchase;
- a supplier receiving repeat, accepted orders from software buyers.

These are not points on one universal maturity ladder. A business can have high assistant use and no delegated external action. Another can run a narrow unattended workflow without widespread employee use. A commerce platform can process agent-initiated payments without proving that the underlying suppliers gained profitable new demand.

The report therefore asks two questions of every market number:

1. **What event was counted?** A response, user, workflow, action, payment, delivery or accepted outcome?
2. **What was the denominator?** Survey respondents, provider customers, active accounts, attempts, unique buyers, transactions or completed jobs?

Without both, a percentage or total cannot support an adoption claim.

## What the current evidence can establish

The public evidence is strongest in four areas.

### Providers can document capability

Model and cloud companies can show that their products support tools, long-running execution, computer use, subagents or managed runtimes. Protocol maintainers can publish specifications and reference implementations. Commerce and payment providers can document checkout states, mandates, credentials and settlement flows.

This establishes that a capability exists or is available under stated conditions. It does not establish how many independent customers use it, how reliably it performs or whether the economics work.

### Independent evaluations can measure bounded performance

Research groups can evaluate agents on defined task suites. METR's task-completion time-horizon work measures the difficulty of software tasks that frontier systems can complete with a stated success probability. METR also cautions that its suite does not measure general real-world autonomy and that estimates above its reliable range should be interpreted carefully ([METR, "Task-Completion Time Horizons of Frontier AI Models"](https://metr.org/time-horizons/)).

This is useful capability evidence because the task, method and limitation are visible. It does not establish that an agent can perform an equivalent duration of arbitrary business work.

### Product telemetry can show behaviour inside one environment

**Product telemetry** is event data recorded inside a provider's own product environment. Providers can aggregate it to show how customers use their products. OpenAI's enterprise reporting, for example, describes a shift from assistance towards execution within its customer base ([OpenAI, "From assistance to execution"](https://openai.com/index/how-enterprises-put-ai-to-work/)). Anthropic has analysed human-agent interactions to study autonomy in its deployed products ([Anthropic, "Measuring AI agent autonomy in practice"](https://www.anthropic.com/research/measuring-agent-autonomy)).

These datasets can reveal session patterns and feature use at a scale few external researchers can observe. They remain first-party views of self-selected customers. Product telemetry may not show downstream human checking, work completed in another system or business outcomes outside the provider's boundary.

### Surveys can report organisational intent and self-described practice

Surveys and interviews can show what leaders say they are doing, buying or planning. A 2026 McKinsey survey reported that 40 per cent of respondents at large organisations said their organisations were scaling AI agents, compared with 27 per cent in the prior year; the reported share for smaller organisations was 22 per cent. The same research reported 37 per cent seeing any positive EBIT impact from AI and 6 per cent meeting McKinsey's high-performer definition ([McKinsey, "The State of AI in 2026: On the road to ROI"](https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai)).

The combination is more informative than the scaling number alone. It describes reported activity alongside an outcome gap. It remains a self-reported sample under McKinsey's definitions, not a census of agentic production or commerce.

## An Australian baseline

One useful current Australian source describes early and uneven use among the medium and large firms consulted. It does not establish a national adoption rate or an absence of activity elsewhere.

The Reserve Bank of Australia received 105 responses from medium and large firms between June and August 2025 as part of its liaison on technology and AI investment. It found that AI use was generally piecemeal, practical agentic adoption was low and realised returns were mixed, while many firms expected benefits to increase over time ([RBA, "Technology investment and AI: what are firms telling us?"](https://www.rba.gov.au/publications/bulletin/2025/nov/technology-investment-and-ai-what-are-firms-telling-us.html)).

That evidence has clear boundaries. The liaison sample is not a representative national survey. It does not provide a current denominator for small professional, trade or creative service businesses. It also measures organisational AI investment and experience more broadly than the definition of delegated work used in this report.

The useful conclusion is narrower: within the RBA liaison evidence, global product availability was ahead of practical adoption and realised returns among the participating firms. Whether the same pattern holds across Australian small businesses, sectors, regions and public bodies remains unmeasured. That gap creates two obligations.

- Global availability is not Australian adoption.
- Australian adoption is best understood through baselines around defined workflows and outcomes.

Local systems also shape what adoption can look like. PayTo, the New Payments Platform and Peppol eInvoicing can support payment authority and structured financial records. OAIC privacy guidance, ASIC expectations and ASD cyber-security guidance affect how organisations govern data and action. These are enabling and constraining institutions, not measures of demand.

<!-- ORIGINAL ANALYSIS: Agentic Economy developed the market-measurement framework and its claim-record fields for this 13 September 2026 edition. -->

## Agentic Economy's market-measurement framework

Agentic Economy separates three dimensions that market reports routinely collapse into a single adoption number.

| Dimension | Values used in this report | Why it matters |
| --- | --- | --- |
| **Observed event** | Specified, announced, documented, available, observed, invoked, paid, delivered, accepted, reconciled, repeated | States exactly how far the recorded event travelled |
| **Source origin** | Provider, customer, Agentic Economy test, third-party researcher, official body | Shows who produced or controls the record |
| **Study design** | Product telemetry, reproducible test, survey, interview, case record, audit, accounting record, census | Exposes the population, method and likely blind spots |

A protocol can be fully specified without a production implementation. A provider can document access without disclosing adoption. A private deployment can deliver accepted work without producing public evidence. Independent measurement improves source separation, but it only supports the defined sample and method.

Agentic Economy records six fields for each market claim: `claim ID · event · source origin · study design · date/version · limitation`. Together they form the reusable claim ledger behind this report.

## Why activity can be mistaken for a market

Digital infrastructure creates abundant machine-readable events. Those events are attractive because they are countable. They also need economic context.

### Tool calls

A tool call shows that an agent attempted or completed an operation. Several calls may belong to one job, including retries and failed branches. A higher call count can mean more demand, a more complicated task or a less efficient system.

### Tokens and compute

Token and compute usage show consumption of infrastructure. They do not reveal whether the output was useful, whether a customer paid for it or whether the supplier earned a margin.

### Listings and endpoints

A marketplace listing or reachable endpoint establishes discoverability or availability. It does not prove that an agent can authenticate, receive authority, pay, obtain a valid result or recover from failure.

### Transactions

A payment event can be a genuine purchase, test, subsidy, internal loop, facilitator operation, retry or abusive activity. On-chain data improves public inspectability but may still lack the identity of the principal, the promised job and the delivered outcome. Third-party x402 directories and analyses are useful because they probe endpoints and attempt to classify activity, while explicitly warning about coverage and concentration ([x402 List methodology](https://www.x402-list.com/methodology)).

### Partner and support counts

A partner list can show ecosystem interest and distribution potential. It does not establish that all partners have shipped an integration or that customers use it in production.

No metric is useless. The error is asking it to prove a later state than it records.

## The job-level measurement record

<!-- UNIQUE INSIGHT: This chapter joins technical, human and commercial events into one job-level adoption record. -->

A national dataset is not required to measure one delegated job. The underlying operating record already exists across application, human and finance systems; the analytical task is to join it.

An **accepted job** is a delivered job that a responsible person or system has validated against its stated acceptance criteria.

| Field | What to record | Why it changes the decision |
| --- | --- | --- |
| **Objective** | Job, customer and acceptance criterion | Separates activity from useful completion |
| **Agent and principal** | System, owner and responsible organisation | Establishes who acted for whom |
| **Permitted choice** | Steps or tools the system could select | Shows what was genuinely agentic |
| **Authority** | Data, action, budget, expiry and approval limits | Reveals the consequence boundary |
| **Job and events** | Stable job ID, start, stop, latency, retries and failure state | Keeps tool calls and retries inside the job they served |
| **Delivery** | Output or state change promised and returned | Connects the system to the offer |
| **Acceptance** | Validation, responsible reviewer and decision | Identifies the completed job |
| **Human effort** | Review, correction, escalation and support time | Prevents hidden labour from disappearing |
| **Economics** | Revenue or value less model, tool, payment and labour cost | Tests whether the workflow is viable |
| **Repeat use** | Return, renewal or expanded scope | Indicates durable demand |

The job-level measurement stack separates:

```text
accepted jobs / eligible jobs
first-pass accepted jobs / delivered jobs
cost / accepted job
human minutes / accepted job
```

Retries are events inside a job unless a customer has submitted a new job. Partial acceptance and rework remain separate events. As an illustration, a system that completes 95 per cent of jobs may still be unsuitable when the remaining 5 per cent can cause a severe financial, privacy or safety incident.

## The minimum market dataset

Twenty comparable jobs or four weeks of activity provides a useful first operating view of a bounded workflow. It is not a statistical threshold or a safe sample for rare, high-consequence failures. The value lies in exposing rework and cost patterns that a one-day demonstration cannot establish.

1. **Population:** teams, customers or cases eligible to use the system.
2. **Active use:** unique users or workflows that invoked it during a defined period.
3. **Delegation:** the action the system could choose and complete.
4. **Completion:** delivered and accepted outcomes, not responses alone.
5. **Reliability:** failure, retry, duplicate, correction and escalation rates.
6. **Human effort:** review and recovery time.
7. **Economics:** cost and value per accepted outcome.
8. **Retention:** repeat use by a defined cohort.
9. **Risk:** incidents and near misses by severity.
10. **Baseline:** the same measures for the prior process where available.

These fields do not create one market-wide agent adoption number. They produce comparable case records that can later support sector and national analysis.

Material differences disappear inside a single average. Results need job type, complexity and consequence, with severe incidents reported individually. A system with 19 accepted jobs from 20 can still be commercially unfit when the remaining job exposes confidential data or sends an unauthorised payment.

## The underlying job record

One row represents each eligible job, joined through a stable job ID across model, tool, payment and business systems.

```text
job_id, customer_or_team, objective, eligible_at, started_at, delivered_at,
accepted_at, first_pass_accepted, retry_count, human_minutes, model_cost,
tool_cost, payment_cost, labour_cost, revenue_or_value, incident_severity,
refund_or_rework, repeated_within_window
```

The prior process supplies the relevant comparison. Reliability and profitability require the same fields on both sides; market-level claims additionally require the vendor population, observation window, data boundary, incident record, sampling frame and definitions.

For example, a pilot with 20 eligible jobs, 18 deliveries and 15 first-pass acceptances has 75 per cent first-pass acceptance, not 90 per cent success merely because 18 outputs were produced. The five jobs requiring correction still affect labour cost and the decision to scale.

### The funding record

Join the job rows to a short decision record:

| Field | Starting record | Observed record |
| --- | --- | --- |
| **Baseline** | Eligible volume, labour minutes, delays, errors and full cost | Like-for-like result for the same job classes |
| **Investment** | Integration, security, data, training and change cost | Actual spend plus unfinished work |
| **Operating cost** | Expected model, tool, payment, support and review cost | Actual cost per delivered and accepted job |
| **Benefit** | Time, capacity, quality, revenue or risk result expected | Realised result with owner and observation window |
| **Loss boundary** | Maximum acceptable financial, privacy, safety or client consequence | Incidents and near misses reviewed individually |
| **Decision rule** | Scale, redesign and stop thresholds with a named owner | Decision, reason and next review date |

This record cannot manufacture certainty from 20 jobs. It makes the assumptions and decision visible. High-consequence work may require a larger sample, targeted stress tests, independent assurance or no delegation at all.

## What would change our view

A representative Australian dataset that separates assistance, delegated action, accepted work, full cost and repeat use would materially change this chapter. So would longitudinal records showing that self-reported scaling reliably predicts accepted outcomes or EBIT improvement. Until then, the defensible conclusion is that demand and capability are visible while comparable economic adoption remains poorly measured.

Four different market facts are routinely collapsed into “adoption”: a definition states what qualifies; documented supply states what a product claims to support; observed deployment records use in a named environment; and an accepted economic outcome joins completed work to cost, consequence and repeat demand. Evidence from an earlier state cannot prove a later one.

## Where we see the market going

Agent measurement will move from activity towards completed work. Token volume, tool calls, listings and transactions will remain useful operational signals, but the strongest providers will connect them to accepted jobs, human effort, failure, contribution and repeat demand.

That shift will expose a less uniform market than current adoption headlines suggest. Some categories will show high use and weak economics; others will show narrow use and strong value. Sector, job type and consequence will matter more than one global agent-adoption percentage.

## About the author and editorial record

Agentic Economy is the accountable publisher of this report. Joel Chan founded the publication in Perth to research the infrastructure, companies and operating choices shaping the agentic economy in Australia. This chapter is source-led analysis, not an official adoption statistic or a representative survey of Australian firms.

[Back to the report hub](/guides/hitchhikers-guide-agentic-economy-2026) · Next: [The agentic economy market map](/guides/agentic-economy-market-map).
