Skip to content
Tigy AI
Tigy AIVoice agentsBuild conversations for phone and webIntegrationsConnect agents to your systemsControl and reliabilityTest, monitor, and refine agents
Explore the platformSupport agentsAnswer common questions and route requestsLead qualificationUnderstand each contact's needsHow it worksGo from setup to live callsPricingFind a plan to get startedDocsLearn how to configure your agent
Areas of focus
TelecommunicationsFinancial servicesHealthcareTechnologyRetail and e-commerceMedia and entertainmentTravel and hospitality
Use cases
Customer supportLead qualificationAI receptionist
Business profiles
EnterpriseStartups
DocsBlogPricing
Platform
Voice agentsIntegrationsControl and reliability
Solutions
TelecommunicationsFinancial servicesHealthcareTechnologyRetail and e-commerceMedia and entertainmentTravel and hospitalityCustomer supportLead qualificationAI receptionistEnterpriseStartups
PricingDocsBlog
SIGN IN
Blog/Insights

How to measure AI voice agent outcomes

Connect conversations to operational results and evaluate your pilot with a fair comparison.

Author
Tigy AI team
Published
Jul 2, 2026
Updated
Oct 4, 2026
Explore customer support agentsCreate an agent
Organic forms between light and deep shadows.
Voice agent metrics

In this article

  • Define success before the pilot
  • Track quality and effort
  • Compare similar situations
  • Use evidence to choose the next change
  • Combine outcomes, experience and effort
  • Build a comparison you can explain
  • Close the loop between analysis and correction
  • Define the population before calculating a rate
  • Create labels two reviewers can apply consistently
  • Compare periods under the same service definition
In this article
  • Define success before the pilot
  • Track quality and effort
  • Compare similar situations
  • Use evidence to choose the next change
  • Combine outcomes, experience and effort
  • Build a comparison you can explain
  • Close the loop between analysis and correction
  • Define the population before calculating a rate
  • Create labels two reviewers can apply consistently
  • Compare periods under the same service definition

To measure voice agent outcomes, define the task and count confirmed results among eligible requests. Combine that measure with corrections, repeat contacts, handoffs and staff effort. Tigy AI call data supports review; bookings and sales must be verified in the responsible systems.

Key takeawayChoose an outcome metric and track errors and team effort alongside it.

Define success before the pilot

Measure a voice agent by its task outcome: a correct answer, confirmed lookup, recorded appointment or lead accepted by sales. Choose a verifiable definition before the pilot and identify where the evidence lives. Ending a call is a technical event; resolving a request requires a business criterion.

Specify what goes into the calculation. For resolution, divide resolved eligible requests by all requests eligible for that conversation. Explain excluded topics to avoid a misleading comparison.

Track quality and effort

Pair the main outcome with signals that help interpret it. An ended conversation may mean someone gave up. A transfer may be the correct decision.

  • Completed tasks verified in the destination system.
  • Handoff reasons and tool failures.
  • Repeat contact about the same issue.
  • Clarity of context delivered to the team.
  • Time and work needed to complete the request.

Compare similar situations

Record a baseline for the current process and select a pilot period with comparable volumes and request types. Changes in opening hours, staffing or campaigns can affect the outcome even when the agent is unchanged.

For small samples, review individual conversations and avoid turning limited observations into performance promises. Include telephony, models, integrations and human work when evaluating cost per completed task.

Use evidence to choose the next change

Classify issues as incomplete knowledge, ambiguous instructions, unavailable integrations or out-of-scope requests. Choose an adjustment, rerun tests and check whether it improved the metric without creating new errors.

Tigy's reports and call review help you observe the service. Some business outcomes, such as a completed sale or confirmed appointment, must be verified in the responsible system. Combine these sources before deciding to expand the pilot.

Combine outcomes, experience and effort

Choose a primary task outcome and protective measures. Scheduling needs confirmed reservations, duplicates and corrections. Support needs verifiable resolution, repeat contact and routing. Sales needs accepted contacts, corrected data and completed next steps.

Sample experience using clarity, repeated questions, interruptions and expectations. Human review with explicit criteria explains ratings. “Bad” is not actionable; “confirmed a reservation without a system result” identifies a fixable condition.

Include post-call effort. Short calls can shift work to another team. Measure review, correction and callbacks across the whole process. Operational benefits require total improvement without worse outcomes. Derived measures can be calculated outside Tigy by combining run evidence with the responsible system's data.

Build a comparison you can explain

Record volume, request types, operating hours and service conditions before the pilot. Compare similar slices: new campaigns, holidays and staffing changes can affect outcomes independently. Without a controlled comparison, describe the observed period and its limitations.

Compare instructions using the same scenarios, keeping tools and documents constant where possible. Change one variable and state the intended improvement. Changing voice, knowledge and transfer rules together prevents attributing the effect to each.

Report counts alongside percentages in small samples. Five successes out of six and fifty out of sixty yield equal percentages but different evidence. Review serious failures individually and preserve them as tests. Expansion requires considering severity, stability and exception handling.

Close the loop between analysis and correction

Classify failures by origin: missing knowledge, interpretation, incorrect fields, tools, telephony or subsequent process. Fix the responsible source and convert the call into a reproducible fictional-data scenario. Transcripts establish speech; external confirmation establishes operations.

Assign an owner, review deadline and return-to-pilot criteria. Content problems differ from authorization failures. Labeling everything “improve prompt” hides dependencies text cannot fix.

Reevaluate after corrections and check new effects. Confirmation can improve field accuracy while increasing duration, an acceptable cost when it reduces rework. Evaluation should explain that tradeoff rather than pushing every measure in the same direction.

Define the population before calculating a rate

A resolution rate needs a clear denominator. If an agent handles order inquiries, mixing billing questions and wrong-number contacts into one rate can obscure performance and scope limitations. First count conversations containing an eligible request, those with sufficient data and those needing service outside defined capabilities.

A simple operational formula divides completed eligible requests by all evaluated eligible requests. “Completed” needs evidence: valid retrieval, a confirmed record or a correct answer from an approved source. This is a working definition, not a universal standard. Your company may require another population, but should document it before comparing periods.

Consider a fictional set of one hundred conversations. Sixty concern a lookup the agent can perform; twenty concern another department; ten contain no identifiable request; ten end before minimum data collection. If forty-eight of the sixty eligible lookups complete, that population's resolution rate is eighty percent. Dividing forty-eight by one hundred answers a different question: the share of all contacts ending in that resolution. Both figures may help, provided they have different labels.

Beware of exclusions that artificially improve the indicator. Removing all tool failures from the denominator gives a limited view of conversation, not the complete service. To measure response quality when the API (application programming interface) works, publish that as an additional subset. Retain a measure reflecting customer experience when dependencies fail.

Define rules for conversations containing multiple intentions. A call may resolve an order inquiry while leaving an address change pending. You can measure separate tasks and the call as a whole. What matters is not choosing the most favorable unit after observing the result.

Create labels two reviewers can apply consistently

Before automating analysis, review fictional examples or records handled under your team's process manually. Use a small set of explicitly defined categories: completed, partly completed, routed, source failure, understanding failure and abandonment. A category should describe either outcome or cause; record them separately when both dimensions matter.

Have two reviewers classify a small sample without seeing each other's judgments. Investigate definitions wherever they disagree. One reviewer may interpret “routed” as a tool invocation, another as staff answering the call. That difference prevents reliable comparison. Write inclusion and exclusion examples until the criterion can be applied consistently.

Do not infer customer intent without evidence. A call ending after an answer may indicate satisfaction, interruption or abandonment. Without confirmation or sufficient system evidence, use a category reflecting uncertainty. Reports can state how many cases do not support a conclusion; hiding that uncertainty turns an estimate into a claim of certainty.

Failure analysis should connect observations to origin. A wrong answer may arise from outdated documents, a correct lookup of the wrong order, or an explanation changing the meaning of a valid response. Record input, available source and observed action. This chain helps choose a correction that actually changes the outcome.

Turn recurring failures into tests afterward. A report listing problems alone does not improve the agent. Each test needs reproducible input, expected behavior and evidence of change. Preserve the previous version for comparison when changes affect instructions, sources or tools.

Compare periods under the same service definition

An agent may appear better after a change because it receives simpler contacts. Compare distributions of reasons, times and channels alongside outcomes. If one period includes telephone calls and another only text tests, the difference does not measure prompt quality alone. The channel introduces speech recognition, turn-taking and network conditions that text cannot assess.

Use medians and higher points in the distribution to investigate duration and waiting where data allows. An average may hide a small set of very slow calls. Connect delay to task and outcome: collecting necessary booking details is different from asking an already answered question three times. Interpret time through what happened during conversation.

When calculating cost, disclose what is included. Platform usage, telephony, external services and human review may come from separate sources. Cost per call is not cost per completed task. Dividing observed cost by confirmed resolutions brings the measure closer to the objective, provided volume and success definitions remain visible.

Decisions to expand a pilot should consider significant failures even when the overall rate improves. A small number of duplicate records or incorrect confirmations may require correction first. Report observed counts and circumstances without converting a short sample into a statistical prediction for the entire operation.

Finally, select a review cadence the team can maintain. Compare consistent windows, record configuration changes and assign an owner for each recurring cause. Useful measurement ends with a decision: preserve a working behavior, fix a specific failure or gather more evidence where uncertainty remains.

In the Tigy documentation

  • Reports
  • Agent runs and usage
  • Billing

Make every conversation count.

Create an agent

Keep the conversation going

Diffuse light and soft shadows in an abstract composition.
Text and voice

How to test a voice agent with text and audio

Light and shadow bands with a grain texture.
Voice agent pilot

How to launch an AI voice agent pilot in your business

Soft contours over light and shadow fields.
Voice agent test checklist

Voice agent testing checklist before publishing

Soft color fields over a dark background.

How to compare voice agent versions with manual tests

Organic forms between light and deep shadows.
Cost per completed task

What does a voice agent cost? Calculate cost per completed task

Soft light veils around a textured abstract background.
Voice agent evaluation matrix

Voice agent evaluation matrix: scenarios and acceptance criteria

Organic light ribbons with a soft texture.
Voice agent latency

Voice agent latency: how to investigate slow responses

Diffuse light and soft shadows in an abstract composition.
SIP telephony

SIP telephony in Tigy: routing calls to voice agents

Tigy AI

PRODUCT

  • Voice agents
  • Integrations
  • Control and reliability
  • Demos
  • How it works
  • Pricing

Solutions

  • Telecommunications
  • Financial services
  • Healthcare
  • Technology
  • Retail and e-commerce
  • Media and entertainment
  • Travel and hospitality

Use cases

  • Customer support
  • Lead qualification
  • AI receptionist

Business profiles

  • Enterprise
  • Startups

Legal

  • Legal center
  • Terms of service
  • Privacy
  • Cookies

Resources

  • Blog
  • Documentation
  • AI documentation
  • Contact us

Social media

  • LinkedIn
  • Instagram