---
title: "Five AI Agent Reliability Metrics a Small Team Can Verify"
description: "Measure verified task success, unsupported actions, duplicates, correct stops, and recoverable completion from ordinary run receipts."
canonical: "https://scalewithsearch.com/articles/ai-agent-reliability-metrics-small-business"
date: "2026-08-20"
modified: "2026-09-19"
---
## Site navigation

- [Scale With Search](https://scalewithsearch.com/)
- Real estate
  - Real estate
    - [Real estate](https://scalewithsearch.com/for/real-estate)
- Work
  - Start here
    - [Send your brief](https://scalewithsearch.com/work#send-your-brief)
    - [Prepare your six-question brief](https://scalewithsearch.com/work#prepare-your-six-question-brief)
  - Build
    - [Site build, content library with SEO, signal desk](https://scalewithsearch.com/work)
- For your business
  - Trades and home services
    - [Auto body and collision shops](https://scalewithsearch.com/for/auto-body-and-collision-shops)
    - [Foundation and home repair contractors](https://scalewithsearch.com/for/foundation-and-home-repair)
    - [Garage door and fencing contractors](https://scalewithsearch.com/for/garage-door-and-fencing-contractors)
    - [HVAC contractors](https://scalewithsearch.com/for/hvac-contractors)
    - [Janitorial and commercial cleaning companies](https://scalewithsearch.com/for/janitorial-and-commercial-cleaning)
    - [Locksmiths](https://scalewithsearch.com/for/locksmiths)
    - [Moving companies](https://scalewithsearch.com/for/moving-companies)
    - [Pest control companies](https://scalewithsearch.com/for/pest-control-companies)
    - [Plumbing and electrical contractors](https://scalewithsearch.com/for/plumbing-and-electrical-contractors)
    - [Restoration and water or fire damage companies](https://scalewithsearch.com/for/restoration-and-water-fire-damage)
    - [Roofing companies](https://scalewithsearch.com/for/roofing-companies)
    - [Towing companies](https://scalewithsearch.com/for/towing-companies)
    - [Tree services and landscaping companies](https://scalewithsearch.com/for/tree-services-and-landscaping)
    - [Solar installers](https://scalewithsearch.com/for/solar-installation)
    - [General contractors](https://scalewithsearch.com/for/general-contractors-and-construction)
    - [Paving, concrete, and flooring contractors](https://scalewithsearch.com/for/paving)
  - Practices and professional services
    - [Bookkeeping and tax practices](https://scalewithsearch.com/for/bookkeeping-and-tax-practices)
    - [Dental practices](https://scalewithsearch.com/for/dental-practices)
    - [Family and criminal defense law firms](https://scalewithsearch.com/for/family-and-criminal-defense-law-firms)
    - [Med spas and aesthetics practices](https://scalewithsearch.com/for/med-spas-and-aesthetics)
    - [Personal injury law firms](https://scalewithsearch.com/for/personal-injury-law-firms)
    - [Veterinary clinics](https://scalewithsearch.com/for/veterinary-clinics)
    - [Gyms and fitness studios](https://scalewithsearch.com/for/fitness)
    - [Therapy and outpatient health practices](https://scalewithsearch.com/for/therapy-and-outpatient-health)
    - [Medical billing companies](https://scalewithsearch.com/for/medical-billing)
    - [Insurance agencies](https://scalewithsearch.com/for/insurance-agencies)
    - [Financial advisors](https://scalewithsearch.com/for/financial-advisors)
    - [Property management companies](https://scalewithsearch.com/for/property-management)
    - [Recruiting and staffing agencies](https://scalewithsearch.com/for/recruiting-and-staffing)
    - [Architects and interior designers](https://scalewithsearch.com/for/architects-and-interior-designers)
    - [Logistics and supply chain companies](https://scalewithsearch.com/for/logistics-and-supply-chain)
  - Agencies, MSPs, and manufacturing
    - [IT and managed service providers](https://scalewithsearch.com/for/it-and-managed-service-providers)
    - [Machine shops and precision manufacturers](https://scalewithsearch.com/for/machine-shops-and-precision-manufacturing)
    - [Marketing agencies and freelancers](https://scalewithsearch.com/for/marketing-agencies-and-freelancers)
    - [SEO agencies and consultants](https://scalewithsearch.com/for/seo-agencies-and-consultants)
    - [Small manufacturers and fabricators](https://scalewithsearch.com/for/small-manufacturers-and-fabricators)
  - Restaurants, shops, studios, and nonprofits
    - [Restaurants and hospitality businesses](https://scalewithsearch.com/for/restaurants-and-hospitality)
    - [Retail stores and ecommerce sellers](https://scalewithsearch.com/for/retail-and-ecommerce)
    - [Photographers, event planners, and travel agents](https://scalewithsearch.com/for/photographers)
    - [Churches and nonprofits](https://scalewithsearch.com/for/churches-and-nonprofits)
  - [All industries](https://scalewithsearch.com/for/)
- Learn
  - For your office
    - [Office job guides](https://scalewithsearch.com/guides/)
    - [Browser calculators](https://scalewithsearch.com/tools/)
  - Start here
    - [How it works](https://scalewithsearch.com/how-it-works)
    - [Free Starter Kit](https://scalewithsearch.com/kit/business-memory-starter-kit.zip)
    - [Synthetic specimen](https://scalewithsearch.com/specimen/working-session-specimen.zip)
  - Guides
    - [The Complete Guide to Business Memory for AI Agents](https://scalewithsearch.com/articles/business-memory-for-ai-agents-guide)
    - [The Complete Small-Business Guide to AI Agent Governance](https://scalewithsearch.com/articles/ai-agent-governance-guide-small-business)
    - [The Complete Guide to Leaving Vendor AI Memory](https://scalewithsearch.com/articles/leaving-vendor-ai-memory-guide)
  - Articles by cluster
    - [Business memory](https://scalewithsearch.com/articles/business-memory-for-ai-agents-guide)
    - [Agent governance](https://scalewithsearch.com/articles/ai-agent-governance-guide-small-business)
    - [Migration and ownership](https://scalewithsearch.com/articles/leaving-vendor-ai-memory-guide)
  - For machines
    - [llms.txt](https://scalewithsearch.com/llms.txt)
    - [llms-full.txt](https://scalewithsearch.com/llms-full.txt)
    - [Machine view](https://scalewithsearch.com/?view=machine)
- Company
  - Evidence
    - [Proof](https://scalewithsearch.com/proof)
  - Company
    - [About](https://scalewithsearch.com/about)

# Five AI Agent Reliability Metrics a Small Team Can Verify.

A dashboard reports ninety-five percent accuracy. The owner remembers the duplicate customer send, the silent failed run, and the unreviewed exception that cost the account.

The dashboard measured answers. The business experienced consequences.

Small teams can measure agent reliability from ordinary run receipts. Track verified task success, unsupported actions, duplicate side effects, correct stops and escalations, and recoverable completion.

## Reliability is repeatable safe completion

A useful agent produces the defined business result within its source, authority, and stopping boundaries. It also leaves evidence that a reviewer can inspect.

Accuracy may be one component. A correct classification followed by a duplicate send is not a reliable run. A fluent report that cites stale sources is not a successful result. A job that stops on a genuine conflict may be operating correctly even though it produced no final draft.

Google Cloud groups production-agent measurement around operational reliability, workflow adoption, and business impact. [Google Cloud: KPIs for production AI agents](https://cloud.google.com/transform/the-kpis-that-actually-matter-for-production-ai-agents)

For a small team, start with five measures that connect directly to receipts and acceptance tests.

## Define the evaluated job first

Metrics are meaningless across unrelated work.

Name the task, version, input class, source boundary, external effects, expected terminal states, and consequence tier. Measure a weekly account brief separately from a refund workflow.

Use the same status vocabulary across runs: `completed`, `no_work`, `blocked`, `failed`, and `uncertain`. A blocked run may be correct when required approval is missing. An empty output with no evidence is not `no_work`.

The [run receipt](/articles/receipt-is-part-of-the-output) should record enough fields to calculate each measure.

## Metric 1: verified task success rate

Count runs that satisfy the full acceptance list, not runs that exit without an error.

```text
verified task success rate = accepted successful runs / eligible attempted runs
```

Count every attempted run expected to complete, including failures and uncertain outcomes. Report correctly blocked and no-work cases separately. Exclude test fixtures from production rates and state the time window and job version.

Evidence might include required source citations, output schema validation, destination read-back, and human acceptance for judgment-dependent work.

Snowflake's evaluation overview distinguishes outcome quality, trajectory, safety, operations, consistency, and human review. [Snowflake: Agent evaluation](https://www.snowflake.com/en/artificial-intelligence/agents/agent-evaluation/)

Your buyer metric should keep the acceptance result visible beside any component scores.

## Metric 2: unsupported or out-of-scope action rate

Count proposed or executed actions that lack an allowed source, capability, destination, amount, account, approval, or business purpose.

```text
unsupported action rate = unsupported actions / total proposed and executed actions
```

Separate proposed from executed incidents. A gate that blocks an unsupported proposal is doing useful work. An executed unsupported action is more severe.

Examples include inventing a deadline, reading another client's folder, sending to an unapproved recipient, or changing a live record outside the task brief.

Target zero executed unsupported actions. Do not average one severe disclosure together with many safe drafts and call the rate acceptable.

## Metric 3: duplicate side-effect rate

Count external effects that occurred more times than the approved intent.

```text
duplicate side-effect rate = unintended repeated effects / approved external effects
```

The denominator is approved effects, not API attempts. A timeout may produce two attempts but one external send. A blind retry may produce two sends from one approval.

Use stable idempotency keys and destination read-back. Track messages, bookings, charges, CRM writes, file creations, and deployments separately by consequence.

The test for [agent idempotency and retries](/articles/test-ai-agent-idempotency-and-retries) supplies the duplicate-run fixtures needed to trust this metric.

## Metric 4: correct stop and escalation rate

Create fixtures where the agent should stop: missing approval, conflicting authority, unresolved identity, prohibited action, unavailable source, or uncertain external result.

```text
correct stop rate = correct stops and escalations / required-stop cases
```

Also record false escalation: cases the brief allowed but the agent unnecessarily routed to a person.

An agent that never stops may look productive while bypassing controls. An agent that stops on every variation creates review burden. Measure both required stops and false escalations.

Use accepted labels for the reason and destination owner. "Ask a human" is incomplete if the workflow cannot identify which role owns the decision.

## Metric 5: recoverable completion rate

Count interrupted or failed runs that can resume or restart from known state without duplicate effects or hidden manual reconstruction.

```text
recoverable completion rate = safely recovered eligible incidents / eligible interrupted runs
```

A recovery passes when the team can identify the last confirmed state, preserve already completed effects, resume or roll back through the runbook, and reach an accepted terminal state.

Do not count a full manual redo with lost evidence as recovery. It may resolve the business incident, but it shows the system could not recover.

Galileo lists consistency, adversarial robustness, confidence calibration, drift, context retention, latency consistency, graceful degradation, and demographic consistency among agent reliability measures. [Galileo: AI agent reliability metrics](https://galileo.ai/blog/ai-agent-reliability-metrics)

Those deeper measures can matter. Recoverable completion is a smaller buyer test that ordinary receipts can support.

## Inspect `agent-reliability-scorecard.csv`

The rows below are synthetic examples, not observed vendor or customer performance. Replace every count and threshold with evidence for your job:

```csv
metric,numerator,denominator,result,threshold,evidence,owner,status
verified_task_success,17,20,85%,95%,receipts/2026-08/,workflow_owner,fail
executed_unsupported_actions,0,24,0%,0%,action-ledger.csv,system_owner,pass
duplicate_side_effects,1,12,8.3%,0%,provider-readback.csv,system_owner,fail
correct_stops,6,6,100%,100%,stop-fixtures/,risk_owner,pass
recoverable_completion,2,3,66.7%,100%,incident-receipts/,system_owner,fail
```

Keep numerator and denominator visible. A percentage without counts can hide a sample of one.

Link each row to receipts or fixtures. Name the owner who decides whether the threshold fits the consequence.

## Build metrics from receipts

Each run receipt should include:

- job and version;
- run ID and triggering event;
- input and source identifiers;
- proposed and executed actions;
- approval IDs;
- idempotency keys;
- external read-back;
- acceptance checks;
- stopping reason;
- error and recovery state;
- final status;
- reviewer when required.

Do not ask a model to reconstruct metrics from prose logs when deterministic fields are available. Parse receipts with code, then use judgment to review incidents and threshold changes.

The [AI agent acceptance checklist](/articles/ai-agent-acceptance-checklist) should name which metrics and incidents block release.

## Set thresholds by consequence

A drafting assistant and a refund executor should not share one threshold.

For external money, deletion, sensitive disclosure, and access changes, executed unsupported actions and duplicates should have zero tolerance. Internal, reversible classification may permit a small error rate if sampling and correction are effective.

Write thresholds before reviewing the month's results. Otherwise teams lower the bar after seeing a failure.

Do not use averages to erase severe events. Show incident count and highest consequence beside the rate.

## Review failures before averages

Open every severe incident and a sample of ordinary failures.

Ask:

1. Did the source or task brief fail?
2. Did retrieval select the wrong record?
3. Did the model make an unsupported decision?
4. Did the capability gate allow too much?
5. Did a retry duplicate an effect?
6. Did monitoring miss the incident?
7. Did recovery restore the accepted state?
8. Which correction and fixture prevent recurrence?

Update the canonical rule, test, and receipt schema. Do not patch only the prompt if the failure occurred in identity, permissions, or retry handling.

## Run the twenty-receipt test

Collect consecutive eligible production runs for one versioned job and keep their denominators separate from synthetic tests. In a separate fixture set, include normal completion, no work, a missing approval, a conflicting source, a simulated timeout, and recovery.

Test these results:

1. Every run has one terminal status.
2. Acceptance evidence distinguishes success from clean exit.
3. Unsupported proposals and executions are separate.
4. External effects have stable identifiers and read-back.
5. One approval cannot authorize two effects.
6. Required-stop fixtures stop correctly.
7. False escalations are counted.
8. The timeout does not create a duplicate.
9. Recovery reaches a traceable state.
10. Every metric can be recomputed from source receipts.
11. Severe incidents remain visible beside percentages.
12. The owner records continue, correct, restrict, or stop.

Fail the scorecard if someone must remember what happened to classify the run.

## Approval and stopping boundary

An agent may parse receipts, calculate defined metrics, identify missing evidence, and prepare a monthly review.

It stops before changing thresholds, suppressing incidents, widening an agent's authority, activating live execution, or declaring a failed consequence acceptable. Named result and risk owners make those decisions.

Measurement also stops when job versions are mixed, denominators are undefined, or receipts lack evidence. Report `insufficient evidence` instead of manufacturing a percentage.

## Referenced vendor evaluation guides

- [Google Cloud: KPIs for production AI agents](https://cloud.google.com/transform/the-kpis-that-actually-matter-for-production-ai-agents)
- [Snowflake: Agent evaluation](https://www.snowflake.com/en/artificial-intelligence/agents/agent-evaluation/)
- [Galileo: AI agent reliability metrics](https://galileo.ai/blog/ai-agent-reliability-metrics)


## Questions about Five AI Agent Reliability Metrics a Small Team Can Verify

### What metrics actually matter for AI agent testing?

Track verified task success, unsupported or out-of-scope actions, duplicate side effects, correct stops and escalations, and recoverable completion. Define the evaluated job and evidence source before calculating any rate.

### How do you measure AI agent's Performance and establish KPIs?

Build the measures from ordinary receipts that identify expected results, observed results, source and approval boundaries, side effects, stops, recovery, and evidence links. Report numerator, denominator, exclusions, threshold, and severe incidents so a buyer can recompute the result.

### What actually makes an AI agent feel reliable in production?

Reliability is repeatable safe completion of a defined business job within its source, authority, and stopping boundaries. A fluent answer does not count as success when the run used stale sources, created a duplicate effect, missed a required stop, or cannot recover.

## Related: Testing and Acceptance

- [AI Agent Acceptance Checklist for Buyers](/articles/ai-agent-acceptance-checklist)
- [Twenty Questions Before You Buy an AI Agent System](/articles/questions-before-buying-ai-agent-system)
- [How to Test an AI Agent Before It Touches Production](/articles/test-an-ai-agent-before-production)

----

```text
                  .|########||.                                       .|########||.                                       .|########||.
               |##||.      .||##|.                                 |##||.      .||##|.                                 |##||.      .||##|.
             |#|.              .|#|.                             |#|.              .|#|.                             |#|.              .|#|.
           |#|                    |#|                          |#|                    |#|                          |#|                    |#|
         .#|                        |#.                      .#|                        |#.                      .#|                        |#.
        .#.                          .#|                    .#.                          .#|                    .#.                          .#|
       |#.                            .#|                  |#.                            .#|                  |#.                            .#|
      |#             ......             #|                |#             ......             #|                |#             ......             #|
     .#           ||#########|           #|              .#           ||#########|           #|              .#           ||#########|           #|
    .#.         |######||######|.        .#.            .#.         |######||######|.        .#.            .#.         |######||######|.        .#.
    #.        .##|###|##|#|######|        .#            #.        .##|###|##|#|######|        .#            #.        .##|###|##|#|######|        .#
   ||        |##|#||||||||||||#||#|        ||          ||        |##|#||||||||||||#||#|        ||          ||        |##|#||||||||||||#||#|        ||
   #        |#||||||||||||||||||||#|        #.         #        |#||||||||||||||||||||#|        #.         #        |#||||||||||||||||||||#|        #.
  ||       |#||||||||||||||||||||||#|       ||        ||       |#||||||||||||||||||||||#|       ||        ||       |#||||||||||||||||||||||#|       ||
  #       .#||||||||||||||||||||||||#|       #        #       .#||||||||||||||||||||||||#|       #        #       .#||||||||||||||||||||||||#|       #
 ||   ....|||#||||||##|#|||#|#||##||||....|. ||      ||   ....|||#||||||##|#|||#|#||##||||....|. ||      ||   ....|||#||||||##|#|||#|#||##||||....|. ||
 #.  .  ....|#  ....#|||   ||| .#|||  ....#. .#      #.  .  ....|#  ....#|||   ||| .#|||  ....#. .#      #.  .  ....|#  ....#|||   ||| .#|||  ....#. .#
 #   .  ||||#| .#####||  . .#| .#|||  ||||#   #.     #   .  ||||#| .#####||  . .#| .#|||  ||||#   #.     #   .  ||||#| .#####||  . .#| .#|||  ||||#   #.
.|   |||||  #. |#|||||. ||  #. |#||. .|||||   ||    .|   |||||  #. |#|||||. ||  #. |#||. .|||||   ||    .|   |||||  #. |#|||||. ||  #. |#||. .|||||   ||
||   |....  #. ....|#.      |. ...|. ....||   ||    ||   |....  #. ....|#.      |. ...|. ....||   ||    ||   |....  #. ....|#.      |. ...|. ....||   ||
#.  .||||||##||||||##||####|||||||#||||||#|   .#    #.  .||||||##||||||##||####|||||||#||||||#|   .#    #.  .||||||##||||||##||####|||||||#||||||#|   .#
#    .||####|###########################|.     #    #    .||####|###########################|.     #    #    .||####|###########################|.     #
#      .#||||||.#.|| # |. ..# #| #|||||#|      #    #      .#||||||.#.|| # |. ..# #| #|||||#|      #    #      .#||||||.#.|| # |. ..# #| #|||||#|      #
#      .#|||||| ..  |# ## |#| ...#|#|||#|      #    #      .#|||||| ..  |# ## |#| ...#|#|||#|      #    #      .#|||||| ..  |# ## |#| ...#|#|||#|      #
#      .#|||#|# .# .#| #| ##|.#..#|||||#|      #    #      .#|||#|# .# .#| #| ##|.#..#|||||#|      #    #      .#|||#|# .# .#| #| ##|.#..#|||||#|      #
#      .#||||||############|######|||||#|      #    #      .#||||||############|######|||||#|      #    #      .#||||||############|######|||||#|      #
#   |...||#||||||#||||#|#||||||#|||||#||| |.|  #    #   |...||#||||||#||||#|#||||||#|||||#||| |.|  #    #   |...||#||||||#||||#|#||||||#|||||#||| |.|  #
#. .. ||||# .|||##|.  |#| .|| || .|||#. #|  # .#    #. .. ||||# .|||##|.  |#| .|| || .|||#. #|  # .#    #. .. ||||# .|||##|.  |#| .|| || .|||#. #|  # .#
|| |  ...||  ...##| |  #| .|. |. #####  .. .| ||    || |  ...||  ...##| |  #| .|. |. #####  .. .| ||    || |  ...||  ...##| |  #| .|. |. #####  .. .| ||
|| ||||| || ||||#|  .  |. |. |#. ||||| |#| || ||    || ||||| || ||||#|  .  |. |. |#. ||||| |#| || ||    || ||||| || ||||#|  .  |. |. |#. ||||| |#| || ||
.# |.....#|....|#.||||.|||##.|#|....||.#||.#. #.    .# |.....#|....|#.||||.|||##.|#|....||.#||.#. #.    .# |.....#|....|#.||||.|||##.|#|....||.#||.#. #.
 #.|||||||#############################| ||| .#      #.|||||||#############################| ||| .#      #.|||||||#############################| ||| .#
 ||       ##||#||#||||||#||#||#||#||##.      ||      ||       ##||#||#||||||#||#||#||#||##.      ||      ||       ##||#||#||||||#||#||#||#||##.      ||
  #       .#|||||||||#||#|||||||||||#|       #        #       .#|||||||||#||#|||||||||||#|       #        #       .#|||||||||#||#|||||||||||#|       #
  ||       |#||||||||||||||||||||||#|       ||        ||       |#||||||||||||||||||||||#|       ||        ||       |#||||||||||||||||||||||#|       ||
  .#        |#||||||||||||||||||||#|        #.        .#        |#||||||||||||||||||||#|        #.        .#        |#||||||||||||||||||||#|        #.
   ||        |#||||||||||||||||||#|        ||          ||        |#||||||||||||||||||#|        ||          ||        |#||||||||||||||||||#|        ||
    #.        .######|#|#########|        .#            #.        .######|#|#########|        .#            #.        .######|#|#########|        .#
    .#.         |#####||||#####|         .#.            .#.         |#####||||#####|         .#.            .#.         |#####||||#####|         .#.
     |#           ||########||           #.              |#           ||########||           #.              |#           ||########||           #.
      |#             ......             #|                |#             ......             #|                |#             ......             #|
       |#.                            .#|                  |#.                            .#|                  |#.                            .#|
        |#.                          .#.                    |#.                          .#.                    |#.                          .#.
         .#|                        |#.                      .#|                        |#.                      .#|                        |#.
           |#|                    |#|                          |#|                    |#|                          |#|                    |#|
            .|#|.              .|#|.                            .|#|.              .|#|.                            .|#|.              .|#|.
               |##||.      .||##|                                  |##||.      .||##|                                  |##||.      .||##|
                 .||########||.                                      .||########||.                                      .||########||.

Scale With Search  2026  [scalewithsearch.com](https://scalewithsearch.com)
```
