Measurement and governance guide · RevOps automation
AI Sales Agent KPIs: Cost, Quality, Pipeline and Failure Metrics
AI may assist research organization and drafting. A human editor reviews every published page, checks material claims against the cited sources and owns the final decision. No company paid for placement in this article.
AI use policyAgent-ready brief
AI takeaways
Keep the key points here, or take a source-aware text brief into Claude, ChatGPT or another AI workspace.- 01Instrument the workflow and define every denominator before choosing dashboard metrics.
- 02Pair an accepted business outcome with a quality metric and a harm metric.
- 03Separate booked meetings, held meetings, sales acceptance, opportunities, pipeline and revenue.
- 04Include human review, correction, infrastructure, data and failure recovery in total operating cost.
- 05Use trace-stage failures, warning thresholds, stop rules and named re-entry authority.
Measure execution, quality and control, sales outcomes, and economics and risk together so faster activity cannot masquerade as better pipeline.
01 / Instrument the workflow before choosing KPIs
Instrument the workflow before choosing KPIs
- a stable run ID;
- lead, contact and account IDs;
- trigger and eligible-population rule;
- workflow, prompt and model version;
- source records and timestamps;
- proposed and executed actions;
- permission and reviewer;
- provider response and postcondition;
- final disposition;
- downstream CRM outcome;
- correction, rollback and failure category.
Define the unit of analysis
| Unit | Example question |
|---|---|
| Agent run | Did one workflow execution reach a valid terminal state? |
| Unique lead | What share of eligible people received a correct disposition? |
| Attempt | Did a specific send or call action succeed? |
| Conversation | Was the interaction meaningful and correctly handled? |
| Meeting | Was it booked, held and accepted by sales? |
| Opportunity | Did it meet the CRM admission rule? |
| Account | Did the workflow improve account-level coverage or expansion? |
02 / Use a metric dictionary
Use a metric dictionary
| Field | What to record |
|---|---|
| Business question | The decision this KPI supports |
| Unit of analysis | Run, lead, attempt, conversation, meeting, opportunity or account |
| Event source | CRM, agent log, provider, calendar or human review |
| Numerator | Exact included outcomes |
| Denominator | Exact eligible population |
| Included and excluded states | Cancellations, retries, tests, duplicates and missing data |
| Cohort and segment | Source, market, channel, owner, offer and risk tier |
| Period | Event window and reporting window |
| Attribution window | When downstream outcomes count |
| Owner and cadence | Who reviews it and how often |
| Warning and stop threshold | Action taken, not only a color |
| Diagnostic drill-down | Which failure categories explain movement |
03 / Layer 1: execution and system health
Layer 1: execution and system health
Eligible-record processing rate
Eligible records reaching a valid terminal state ÷ all eligible recordsSuccessful-action rate
Confirmed external actions ÷ attempted permitted actionsDuplicate-action rate
Duplicate external actions ÷ attempted external actionsUnclassified-outcome rate
Runs without a recognized terminal disposition ÷ completed or timed-out runsother category makes every downstream rate less trustworthy.Latency
timestamp of first valid response attempt − timestamp of eligible lead eventRetry and tool-failure metrics
- retry attempts per run;
- exhausted retries;
- tool timeouts;
- partial writes;
- provider errors;
- validation rejects;
- recovery time;
- exceptions older than the service expectation.
04 / Layer 2: quality and human control
Layer 2: quality and human control
Unsupported or materially incorrect response rate
hallucination rate. Define:Reviewed responses containing at least one unsupported or materially incorrect claim ÷ all reviewed responses- what counts as supported evidence;
- whether an irrelevant but true claim is an error;
- severity levels;
- whether the error reached a buyer;
- who adjudicates disagreement;
- how the sample was selected.
Evidence-grounding rate
Material decisions with an openable, current source record ÷ reviewed material decisionsHuman correction rate
Agent outputs changed before acceptance ÷ reviewed agent outputsOverride rate
Executed or recommended decisions reversed by a human ÷ decisions eligible for overrideEscalation precision
Correct escalations ÷ all escalationsCases that required escalation but did not receive it ÷ reviewed cases requiring escalation
Clean-handoff rate
Human-accepted handoffs containing all required fields ÷ all handoffsCRM field error rate
Incorrect, unauthorized or conflicting field writes ÷ reviewed agent-authored field writes05 / Layer 3: sales outcomes
Layer 3: sales outcomes
Eligible lead → reached → meaningful conversation → qualified → meeting booked → meeting held → sales accepted → opportunity created → closed wonMeaningful-conversation rate
Human-validated commercial conversations ÷ reached eligible leadsQualified-conversation rate
Conversations meeting the approved qualification rule ÷ meaningful conversationsMeeting-booked rate
Meetings booked ÷ named eligible populationHeld-meeting rate
Meetings completed ÷ meetings bookedSales-acceptance rate
Meetings accepted by sales as appropriate ÷ meetings heldOpportunity-creation rate
Opportunities meeting the CRM admission rule ÷ sales-accepted meetingsPipeline and revenue
- qualified pipeline created by the defined cohort;
- agent-sourced versus agent-assisted pipeline;
- opportunities progressed within the attribution window;
- closed-won value;
- time to outcome.
assisted when humans or other channels materially contributed. A clean label is better than false causal precision.06 / Layer 4: economics and risk
Layer 4: economics and risk
Total operating cost
- product subscription and usage;
- models and providers;
- data, enrichment and verification;
- phone, SMS, WhatsApp, email and carrier costs;
- integration and implementation;
- monitoring, QA and review time;
- human reply and exception work;
- remediation, CRM cleanup and rollback;
- retained human selling work;
- exit and replacement cost.
Cost per accepted meeting
Total operating cost for the cohort ÷ sales-accepted meetingsCost per created opportunity
Total operating cost for the cohort ÷ opportunities meeting the admission ruleCost per qualified pipeline dollar
Total operating cost for the cohort ÷ qualified pipeline value created within the attribution rulecost per qualified pipeline without naming its unit.Payback and attributed value
Risk and harm metrics
- suppression breaches;
- calls or messages outside approved scope;
- complaints;
- policy violations;
- unauthorized CRM changes;
- sensitive-data exposure;
- unrecorded human requests;
- failed handoffs;
- repeat contact after a stop state;
- remediation time.
07 / Diagnose failures by trace stage
Diagnose failures by trace stage
| Failure family | Example | Likely owner |
|---|---|---|
| Input | Missing consent, stale role, wrong account match | Data/RevOps |
| Evidence | Source absent, old or irrelevant | Research/content owner |
| Reasoning | Unsupported inference or wrong qualification | Agent/product owner |
| Policy | Prohibited action proposed or allowed | Governance owner |
| Tool | Timeout, invalid arguments or partial write | Engineering/integration |
| State | Duplicate, stale retry or cross-channel conflict | Orchestration owner |
| Handoff | Missing context, wrong queue or no fallback | Sales operations |
| Downstream | Meeting rejected or opportunity reversed | Sales manager/RevOps |
08 / Design the baseline and pilot
Design the baseline and pilot
- Freeze eligibility, offer, channels, qualification and CRM rules.
- Choose a current sample and record exclusions.
- Capture the previous workflow's events with the same definitions.
- Run the agent in shadow or approval mode.
- Review a defined quality sample.
- Measure accepted downstream outcomes inside a fixed window.
- Record concurrent changes such as offer, staffing or seasonality.
- Promote one permission or segment at a time.
09 / Segment AI sales agent KPIs before comparing
Segment AI sales agent KPIs before comparing them
- lead source and trigger;
- new versus existing account;
- inbound, warm follow-up and outbound motion;
- market, language and timezone;
- channel and channel sequence;
- offer, campaign and qualification rule;
- owner, queue and handoff destination;
- workflow, prompt, model and provider version;
- risk tier and permission mode;
- complete, missing and contradictory evidence.
10 / Treat attribution as a claim with evidence
Treat attribution as a claim with evidence
- the first eligible event;
- all material touches inside the attribution window;
- which person or system made the consequential decision;
- when the CRM outcome was created and accepted;
- how reopened, merged or transferred opportunities are handled;
- which concurrent campaigns are excluded or marked assisted.
11 / Run a weekly scorecard review
Run a weekly scorecard review
- Safety and customer harm: suppression, complaints, severe unsupported claims, unauthorized actions and failed human requests.
- State integrity: missing events, duplicates, uncertain provider responses, retries and rollback.
- Quality sample: evidence use, qualification, correction, override and handoff completeness.
- Sales acceptance: held meetings, acceptance, opportunity admission and rejection reasons.
- Economics: full cost, review burden and cost per accepted outcome.
- Version and mix: workflow changes, cohort shift, incidents and external campaign changes.
| Field | Example content |
|---|---|
| Review period | Exact event and outcome windows |
| Cohort | Eligibility rule and material exclusions |
| Versions | Workflow, model, provider and policy |
| Evidence | Scorecard link and reviewed case set |
| Decision | Continue, narrow, correct, promote or stop |
| Owner | Accountable person for each change |
| Re-entry test | Cases and metrics required before promotion |
12 / Set warning, stop and re-entry rules
Set warning, stop and re-entry rules
| State | Required definition |
|---|---|
| Warning | Metric, value, sample, review owner and investigation deadline |
| Approval mode | Which autonomous actions revert to review |
| Shadow Mode | Which external actions stop while proposals continue |
| Full stop | Which triggers disable the workflow or channel |
| Rollback | Which state and version are restored |
| Re-entry | Evidence required before live actions resume |
- any suppression breach;
- duplicate external action above the allowed zero or near-zero tolerance;
- severe unsupported price, legal or security claim;
- unlogged CRM mutation;
- inability to route a human request;
- missing audit events;
- provider incident that makes state uncertain.
13 / Reporting cadence
Reporting cadence
- Real time: suppression, severe policy breach, duplicate action, tool outage and failed human request.
- Daily during pilot: execution, exceptions, corrections and handoffs.
- Weekly: quality sample, qualification, held meetings, sales acceptance, cost and failure distribution.
- Monthly or cohort-close: opportunity creation, pipeline progression, revenue, payback and strategic review.
14 / Frequently asked questions
Frequently asked questions
What is the most important AI sales agent KPI?
Should we track messages and calls?
How do we measure hallucinations?
How should voice-agent latency be measured?
How do we calculate AI sales agent ROI?
When should we pause an AI sales agent?
15 / Scorecard checklist
Scorecard checklist
- the unit and event source are named;
- numerator and denominator are explicit;
- tests, retries, duplicates and missing data are handled;
- the cohort and attribution window are fixed;
- activity, acceptance and revenue are separate;
- human corrections and overrides are recorded;
- total cost includes retained human work;
- failure categories explain movement;
- warning and stop rules have owners;
- quarantined or vendor numbers are not presented as benchmarks.
Research note
Methodology
- 01The scorecard structure applies NIST lifecycle risk management and current agent-evaluation research to sales operations.
- 02Every KPI is defined with a unit, event source, numerator, denominator, cohort and attribution window.
- 03No universal conversion, latency, cost or ROI threshold is asserted; teams must derive targets from their own baseline and severity model.
Source ledger
Sources & editorial notes
- 01NIST AI Risk Management Framework
nist.gov · Primary, official or disclosed research source used for the bounded claim cited in this guide; scope and current status require rechecking.