Skip to content
Tigy AI
Tigy AIVoice agentsBuild conversations for phone and webIntegrationsConnect agents to your systemsControl and reliabilityTest, monitor, and refine agents
Explore the platformSupport agentsAnswer common questions and route requestsLead qualificationUnderstand each contact's needsHow it worksGo from setup to live callsPricingFind a plan to get startedDocsLearn how to configure your agent
Areas of focus
TelecommunicationsFinancial servicesHealthcareTechnologyRetail and e-commerceMedia and entertainmentTravel and hospitality
Use cases
Customer supportLead qualificationAI receptionist
Business profiles
EnterpriseStartups
DocsBlogPricing
Platform
Voice agentsIntegrationsControl and reliability
Solutions
TelecommunicationsFinancial servicesHealthcareTechnologyRetail and e-commerceMedia and entertainmentTravel and hospitalityCustomer supportLead qualificationAI receptionistEnterpriseStartups
PricingDocsBlog
SIGN IN
Blog/Insights

Voice agent evaluation matrix: scenarios and acceptance criteria

Organize scenarios, evidence and acceptance criteria to compare agent changes.

Author
Tigy AI team
Published
May 24, 2026
Updated
Oct 4, 2026
Explore voice agentsCreate an agent
Soft light veils around a textured abstract background.
Voice agent evaluation matrix

In this article

  • Give scenarios a clear condition
  • Define evidence for acceptance
  • Keep review evidence
  • Compare changes without hiding exceptions
  • Record your evidence
  • Define severity before running tests
  • Use the matrix during maintenance
  • Calibrate reviewers before interpreting small differences
  • Turn real situations into assessable cases
  • Separate scores, release blockers, and improvement opportunities
  • Compare versions under conditions you can explain
In this article
  • Give scenarios a clear condition
  • Define evidence for acceptance
  • Keep review evidence
  • Compare changes without hiding exceptions
  • Record your evidence
  • Define severity before running tests
  • Use the matrix during maintenance
  • Calibrate reviewers before interpreting small differences
  • Turn real situations into assessable cases
  • Separate scores, release blockers, and improvement opportunities
  • Compare versions under conditions you can explain

A voice agent evaluation matrix records scenarios, context, expected outcomes and observed evidence. Use separate criteria for content, actions and experience, keeping critical failures visible in comparisons. Tigy AI tests and run records support this review; the spreadsheet is a team process, not a promise of automatic evaluation.

Key takeawayEvaluate the same task under the same conditions and keep important failures visible even when averages improve.

Give scenarios a clear condition

A voice agent evaluation matrix organizes scenario, context, permitted action, expected outcome and observed evidence. Create separate rows for valid lookups, missing inputs, corrections, unavailability and handoffs. These cases test different decisions; merely varying question wording does not cover every task condition.

Include cases requiring handoff. Declining an out-of-scope action can be correct even when the conversation does not end in an automated action.

Define evidence for acceptance

For lookups, check parameters and returned information. For record creation, inspect the destination. For document questions, compare the answer with the approved source.

Record clarity and pacing separately from outcomes. A pleasant voice does not compensate for an action using the wrong identifier.

Keep review evidence

Record agent version, channel, date and run identifier. Use transcripts, recordings and actions where available. Investigate missing records according to conversation state and channel.

Mark scenarios passed, failed or inconclusive and explain why. Inconclusive means missing evidence rather than a correct result.

Compare changes without hiding exceptions

Repeat scenarios after prompt, document or tool changes. Check that the fix resolves the original case and preserves other behavior.

This matrix is a team review process, not a promise of automatic Tigy evaluation. Set task-specific criteria and use results to choose the next adjustment.

Record your evidence

Use the scenario matrix to record observations and execution references. Compare conversation outcomes with external actions. The file contains synthetic criteria and fields for completion without precomputed results.

Resources for your team

  • Test matrixCSV · English

Define severity before running tests

Use criteria explaining decisions: pass, needs adjustment and blocking failure. Style details can require adjustment; wrong-record retrieval or unperformed-write confirmation should block release pending investigation. Averages must not conceal critical failures.

Give examples for each classification. Two reviewers can compare a sample and resolve disagreements before evaluating everything. If expected outcomes are disputed, refine the business rule.

After fixes, rerun neighboring cases. Identifier confirmation rules can affect questions without identifiers; tighter scope can reject valid requests. The matrix records these relationships and helps select regression scenarios.

Use the matrix during maintenance

Add rows when real failures expose missing conditions. Remove personal data while preserving date corrections, similar branches, empty results and ambiguity. This reproduces problems without the original customer.

Separate text, voice and telephone checks. Text tests decisions and tools; voice adds recognition and turns; telephony adds routing and provider conditions. Approval in one does not establish the others.

Before expansion, identify passed scenarios, unresolved ones and exception owners. This is a team evaluation method using Tigy evidence, not a claim of a native automated evaluation system.

Calibrate reviewers before interpreting small differences

Before scoring a complete version, ask two people to assess the same three conversations without seeing each other's ratings. Compare their reasons, particularly for disputed criteria. One reviewer may be scoring friendliness while another checks whether an action was confirmed. Refine the criterion and repeat the exercise. More consistent grading makes small differences between versions easier to interpret. Preserve the reasons as reference examples for future reviewers. If agreement remains poor, report that uncertainty alongside results instead of presenting a decimal score as a precise measurement. An honest assessment can identify useful changes while acknowledging that a particular dimension is not yet consistently judged.

Turn real situations into assessable cases

A useful matrix starts with the variety of situations rather than a list of desirable qualities. For a reception agent, answering correctly can mean reporting one location's opening hours, recognizing that the caller means a different location, or explaining that holiday opening has not been confirmed. Those situations require different behavior. Combining them into a single frequently asked questions row hides the exceptions that most need attention.

Choose a task and write a complete case: available context, opening utterance, information the agent may consult, and expected outcome. In a fictional example, someone asks whether they can collect a document on Friday. The source states ordinary hours but does not address the holiday occurring that day. The expected response should distinguish normal hours from special confirmation, explain the limitation, and identify the responsible contact. Requiring brevity alone is insufficient because an invented opening confirmation can also sound efficient.

Create variations of the same intention. The caller may ask whether the office opens Friday, refer to an earlier interaction, or correct the location after the first response. These variations assess understanding and context updates without changing the business objective. Add a case where the location is unknown, another where the caller interrupts the explanation, and another where the question exceeds the agent's scope. The set should distinguish comprehension, retrieval, and respect for boundaries.

Record criteria a second evaluator can apply consistently. Being friendly is broad; acknowledging a correction without blaming the caller and answering about the correct location is observable. For external actions, include independent evidence, such as a test record created in the authorized system. The agent's speech demonstrates communication but does not establish that a task completed. This separation prevents rewarding a persuasive promise as if it were an actual execution. Also record which supplied facts are intentionally unavailable so evaluators do not accidentally grade justified uncertainty as a failure.

Separate scores, release blockers, and improvement opportunities

An average score can track progress, but it must not erase failures that prevent release. Imagine twenty clear answers and one lookup exposing another person's information. Average clarity can remain high while the access problem remains unacceptable. Keep disqualifying criteria separate from gradual dimensions. The matrix may score clarity, repetition, and closing quality while independently blocking a version for inappropriate disclosure or falsely confirming an action.

Use a small scale with written definitions. For clarity, zero might mean the caller receives no understandable next step; one might mean the direction arrives after a confusing explanation; two might mean the next step is precise and understandable. The scale does not need to create an appearance of scientific precision. It needs to reduce disagreement between evaluators. Include an example for each level and update examples when new situations appear. Date rule changes so comparisons do not silently combine scores calculated under different definitions.

Assign observers to specific dimensions. An operations reviewer can judge whether a direction is executable; an integration owner can verify its external effect; someone familiar with the audience can assess vocabulary and comprehension. The same conversation may receive different judgments because reviewers inspect different parts. Joint review should resolve conflicts through evidence: transcript, tool return, external record, and current policy. Disagreement should not be settled merely by the reviewer's seniority.

Group outcomes by intention and severity. If failures cluster around spoken identifiers, confirming input may be the priority. If failures occur mainly on questions absent from the knowledge base, handling uncertainty may deserve attention. Overall averages conceal those distinctions. Also record unevaluable cases: a broken test environment should not become a conversational pass or failure without investigation. This classification keeps the matrix useful for decisions rather than forcing every event into a number. Show the number of attempted cases alongside scores, since a strong result from only two examples deserves a different level of confidence from a broad, consistently reviewed set.

Compare versions under conditions you can explain

A fair comparison starts with a fixed case set and a configuration record. Record the instruction version, available sources, attached tools, and relevant call conditions. If one version consults an updated document while another uses an older document, the result does not isolate a prompt change. The matrix should expose that difference instead of automatically attributing it to model quality or wording.

Repeat cases when behavior varies. Do not select only the new version's best conversation for presentation. Preserve unsuccessful attempts and label suspected causes without treating them as established facts. For voice, retain examples with pauses, corrections, and pronunciation resembling the intended audience. Comparing slow speech in an old test against hurried speech in a new one mixes input difficulty with the change under evaluation.

Check whether the adjustment harmed another task that previously worked. Instructions to end sooner may cut short a request with two subjects. Confirmation may prevent incorrect records but also repeat unnecessary questions. Show gains and losses together. The team may accept a slightly longer call to avoid an important error; record the reason for that choice.

Before publishing, write a short conclusion: what improved, what remains unstable and which errors still prevent use. State which group will use the new version first. In Tigy, record the instructions and version tested. The matrix is a spreadsheet maintained by the team; the platform does not automatically score all these criteria. Use fictional examples to repeat the evaluation without exposing customers.

In the Tigy documentation

  • Testing your agent
  • Agent runs and usage

Make every conversation count.

Create an agent

Keep the conversation going

Soft color fields over a dark background.

How to compare voice agent versions with manual tests

Diffuse light and soft shadows in an abstract composition.
SIP telephony

SIP telephony in Tigy: routing calls to voice agents

Color fields with organic movement.
WER and data accuracy

WER in voice agents: transcription versus field accuracy

Light and shadow bands with a grain texture.
Voice agent pilot

How to launch an AI voice agent pilot in your business

Soft contours over light and shadow fields.
Voice agent test checklist

Voice agent testing checklist before publishing

Organic light ribbons with a soft texture.

Usage reports and credit cycles in Tigy: comparing consumption

Diffuse light and soft shadows in an abstract composition.
Text and voice

How to test a voice agent with text and audio

Organic forms between light and deep shadows.
Cost per completed task

What does a voice agent cost? Calculate cost per completed task

Tigy AI

PRODUCT

  • Voice agents
  • Integrations
  • Control and reliability
  • Demos
  • How it works
  • Pricing

Solutions

  • Telecommunications
  • Financial services
  • Healthcare
  • Technology
  • Retail and e-commerce
  • Media and entertainment
  • Travel and hospitality

Use cases

  • Customer support
  • Lead qualification
  • AI receptionist

Business profiles

  • Enterprise
  • Startups

Legal

  • Legal center
  • Terms of service
  • Privacy
  • Cookies

Resources

  • Blog
  • Documentation
  • AI documentation
  • Contact us

Social media

  • LinkedIn
  • Instagram