An AI agent is useful when it completes a defined job to an acceptable standard at a cost the business can justify. A persuasive demonstration, a quick response or an industry adoption forecast does not establish that result. This guide sets out what to measure before expanding a pilot.
Separate forecasts from observed results
In a June 25, 2025 forecast, Gartner predicted that more than 40% of agentic AI projects would be canceled by the end of 2027, citing costs, unclear business value and inadequate risk controls. That is a forecast about a future deadline, not an observed Q1 2026 failure rate.
Apply the same distinction to adoption figures and vendor case studies. Ask when the evidence was collected, which organizations or tasks it covers, what was excluded and who measured the result. An adoption percentage cannot tell you whether your own workflow will pay back.
Define what finished means
Choose one bounded job. For an order-status workflow, success might require matching the right customer and order, using the current fulfillment record, giving an accurate answer and recording the interaction. A confident answer about the wrong order fails that definition even if the ticket closes automatically.
Write the acceptance criteria before running the pilot. Separate tasks eligible for automation from requests that require a person. Keep a record of both, so narrowing the eligible set cannot silently inflate the completion rate.
| Measure | How to record it |
|---|---|
| Verified completion | Correctly completed eligible tasks divided by all eligible tasks attempted; also report how much incoming work was eligible. |
| Human intervention | Tasks requiring review, approval, correction or takeover, with the time spent on each. |
| Errors and rework | Incorrect outcomes and reopened tasks, separated by severity and cause. |
| Completion time | Elapsed time from receipt to verified outcome, including waiting and retries; keep first-response time separate. |
| Operating cost | Model, platform, integration, review, exception and maintenance costs over the same measurement window. |
Compare the same work at the same quality standard
Record a baseline using representative tasks from the current process. Keep the lead source, task mix, operating hours and acceptance criteria comparable during the pilot. Report the task count and date range beside every rate. If the workload changes, show that change rather than attributing the entire difference to AI.
For cost per verified completion, divide the full operating cost of the pilot by its verified completed tasks. Failed attempts still contribute cost. Show the implementation fee separately and state the period over which you spread it when estimating payback. Compare against a baseline that also includes human handling and rework.
If staff finish sooner but payroll and outside spending remain unchanged, describe the result as time released. Track whether that capacity is used for additional productive work. Do not report an estimated hourly value as cash saved.
Test mistakes and handoffs, not only successful examples
NIST’s Generative AI Profile describes confabulation as confidently produced false content. It recommends evaluation in conditions similar to deployment and cautions against generalizing from narrow assessments. A single model error percentage does not establish the safety of a business workflow.
Include missing records, contradictory instructions, stale information, repeated requests and unavailable integrations in the test set. Verify what the agent does when it cannot complete the job. Check that a person receives enough context to take over and that retries do not create duplicate actions.
The connection method does not replace these checks. Whether a tool uses an API, a native connector or MCP, define the permitted actions, keep access limited to the job and test the integration’s failure behavior. Review consequential actions with a person where the agreed scope requires it.
Decide when to continue, pause or expand
Set the acceptance thresholds and the review date before launch. Include quality, cost and failure limits, not just the share of work handled automatically. Pause when the system exceeds an agreed limit, investigate the cause and retest before expanding access or volume.
Use the AI project checklist and workbook to record the scope, evidence and handover requirements. If the measurement shows a worthwhile build, the AI Agents service describes how we scope a bounded agent around the tools and approvals the job needs.
