How to measure AI voice agent outcomes
Connect conversations to operational results and evaluate your pilot with a fair comparison.
- Author
- Tigy AI team
- Published
- Updated
To measure voice agent outcomes, define the task and count confirmed results among eligible requests. Combine that measure with corrections, repeat contacts, handoffs and staff effort. Tigy AI call data supports review; bookings and sales must be verified in the responsible systems.
Define success before the pilot
Measure a voice agent by its task outcome: a correct answer, confirmed lookup, recorded appointment or lead accepted by sales. Choose a verifiable definition before the pilot and identify where the evidence lives. Ending a call is a technical event; resolving a request requires a business criterion.
Specify what goes into the calculation. For resolution, divide resolved eligible requests by all requests eligible for that conversation. Explain excluded topics to avoid a misleading comparison.
Track quality and effort
Pair the main outcome with signals that help interpret it. An ended conversation may mean someone gave up. A transfer may be the correct decision.
- Completed tasks verified in the destination system.
- Handoff reasons and tool failures.
- Repeat contact about the same issue.
- Clarity of context delivered to the team.
- Time and work needed to complete the request.
Compare similar situations
Record a baseline for the current process and select a pilot period with comparable volumes and request types. Changes in opening hours, staffing or campaigns can affect the outcome even when the agent is unchanged.
For small samples, review individual conversations and avoid turning limited observations into performance promises. Include telephony, models, integrations and human work when evaluating cost per completed task.
Use evidence to choose the next change
Classify issues as incomplete knowledge, ambiguous instructions, unavailable integrations or out-of-scope requests. Choose an adjustment, rerun tests and check whether it improved the metric without creating new errors.
Tigy's reports and call review help you observe the service. Some business outcomes, such as a completed sale or confirmed appointment, must be verified in the responsible system. Combine these sources before deciding to expand the pilot.
Combine outcomes, experience and effort
Choose a primary task outcome and protective measures. Scheduling needs confirmed reservations, duplicates and corrections. Support needs verifiable resolution, repeat contact and routing. Sales needs accepted contacts, corrected data and completed next steps.
Sample experience using clarity, repeated questions, interruptions and expectations. Human review with explicit criteria explains ratings. “Bad” is not actionable; “confirmed a reservation without a system result” identifies a fixable condition.
Include post-call effort. Short calls can shift work to another team. Measure review, correction and callbacks across the whole process. Operational benefits require total improvement without worse outcomes. Derived measures can be calculated outside Tigy by combining run evidence with the responsible system's data.
Build a comparison you can explain
Record volume, request types, operating hours and service conditions before the pilot. Compare similar slices: new campaigns, holidays and staffing changes can affect outcomes independently. Without a controlled comparison, describe the observed period and its limitations.
Compare instructions using the same scenarios, keeping tools and documents constant where possible. Change one variable and state the intended improvement. Changing voice, knowledge and transfer rules together prevents attributing the effect to each.
Report counts alongside percentages in small samples. Five successes out of six and fifty out of sixty yield equal percentages but different evidence. Review serious failures individually and preserve them as tests. Expansion requires considering severity, stability and exception handling.
Close the loop between analysis and correction
Classify failures by origin: missing knowledge, interpretation, incorrect fields, tools, telephony or subsequent process. Fix the responsible source and convert the call into a reproducible fictional-data scenario. Transcripts establish speech; external confirmation establishes operations.
Assign an owner, review deadline and return-to-pilot criteria. Content problems differ from authorization failures. Labeling everything “improve prompt” hides dependencies text cannot fix.
Reevaluate after corrections and check new effects. Confirmation can improve field accuracy while increasing duration, an acceptable cost when it reduces rework. Evaluation should explain that tradeoff rather than pushing every measure in the same direction.
Define the population before calculating a rate
A resolution rate needs a clear denominator. If an agent handles order inquiries, mixing billing questions and wrong-number contacts into one rate can obscure performance and scope limitations. First count conversations containing an eligible request, those with sufficient data and those needing service outside defined capabilities.
A simple operational formula divides completed eligible requests by all evaluated eligible requests. “Completed” needs evidence: valid retrieval, a confirmed record or a correct answer from an approved source. This is a working definition, not a universal standard. Your company may require another population, but should document it before comparing periods.
Consider a fictional set of one hundred conversations. Sixty concern a lookup the agent can perform; twenty concern another department; ten contain no identifiable request; ten end before minimum data collection. If forty-eight of the sixty eligible lookups complete, that population's resolution rate is eighty percent. Dividing forty-eight by one hundred answers a different question: the share of all contacts ending in that resolution. Both figures may help, provided they have different labels.
Beware of exclusions that artificially improve the indicator. Removing all tool failures from the denominator gives a limited view of conversation, not the complete service. To measure response quality when the API (application programming interface) works, publish that as an additional subset. Retain a measure reflecting customer experience when dependencies fail.
Define rules for conversations containing multiple intentions. A call may resolve an order inquiry while leaving an address change pending. You can measure separate tasks and the call as a whole. What matters is not choosing the most favorable unit after observing the result.
Create labels two reviewers can apply consistently
Before automating analysis, review fictional examples or records handled under your team's process manually. Use a small set of explicitly defined categories: completed, partly completed, routed, source failure, understanding failure and abandonment. A category should describe either outcome or cause; record them separately when both dimensions matter.
Have two reviewers classify a small sample without seeing each other's judgments. Investigate definitions wherever they disagree. One reviewer may interpret “routed” as a tool invocation, another as staff answering the call. That difference prevents reliable comparison. Write inclusion and exclusion examples until the criterion can be applied consistently.
Do not infer customer intent without evidence. A call ending after an answer may indicate satisfaction, interruption or abandonment. Without confirmation or sufficient system evidence, use a category reflecting uncertainty. Reports can state how many cases do not support a conclusion; hiding that uncertainty turns an estimate into a claim of certainty.
Failure analysis should connect observations to origin. A wrong answer may arise from outdated documents, a correct lookup of the wrong order, or an explanation changing the meaning of a valid response. Record input, available source and observed action. This chain helps choose a correction that actually changes the outcome.
Turn recurring failures into tests afterward. A report listing problems alone does not improve the agent. Each test needs reproducible input, expected behavior and evidence of change. Preserve the previous version for comparison when changes affect instructions, sources or tools.
Compare periods under the same service definition
An agent may appear better after a change because it receives simpler contacts. Compare distributions of reasons, times and channels alongside outcomes. If one period includes telephone calls and another only text tests, the difference does not measure prompt quality alone. The channel introduces speech recognition, turn-taking and network conditions that text cannot assess.
Use medians and higher points in the distribution to investigate duration and waiting where data allows. An average may hide a small set of very slow calls. Connect delay to task and outcome: collecting necessary booking details is different from asking an already answered question three times. Interpret time through what happened during conversation.
When calculating cost, disclose what is included. Platform usage, telephony, external services and human review may come from separate sources. Cost per call is not cost per completed task. Dividing observed cost by confirmed resolutions brings the measure closer to the objective, provided volume and success definitions remain visible.
Decisions to expand a pilot should consider significant failures even when the overall rate improves. A small number of duplicate records or incorrect confirmations may require correction first. Report observed counts and circumstances without converting a short sample into a statistical prediction for the entire operation.
Finally, select a review cadence the team can maintain. Compare consistent windows, record configuration changes and assign an owner for each recurring cause. Useful measurement ends with a decision: preserve a working behavior, fix a specific failure or gather more evidence where uncertainty remains.
