Voice agent evaluation matrix: scenarios and acceptance criteria
Organize scenarios, evidence and acceptance criteria to compare agent changes.
- Author
- Tigy AI team
- Published
- Updated
A voice agent evaluation matrix records scenarios, context, expected outcomes and observed evidence. Use separate criteria for content, actions and experience, keeping critical failures visible in comparisons. Tigy AI tests and run records support this review; the spreadsheet is a team process, not a promise of automatic evaluation.
Give scenarios a clear condition
A voice agent evaluation matrix organizes scenario, context, permitted action, expected outcome and observed evidence. Create separate rows for valid lookups, missing inputs, corrections, unavailability and handoffs. These cases test different decisions; merely varying question wording does not cover every task condition.
Include cases requiring handoff. Declining an out-of-scope action can be correct even when the conversation does not end in an automated action.
Define evidence for acceptance
For lookups, check parameters and returned information. For record creation, inspect the destination. For document questions, compare the answer with the approved source.
Record clarity and pacing separately from outcomes. A pleasant voice does not compensate for an action using the wrong identifier.
Keep review evidence
Record agent version, channel, date and run identifier. Use transcripts, recordings and actions where available. Investigate missing records according to conversation state and channel.
Mark scenarios passed, failed or inconclusive and explain why. Inconclusive means missing evidence rather than a correct result.
Compare changes without hiding exceptions
Repeat scenarios after prompt, document or tool changes. Check that the fix resolves the original case and preserves other behavior.
This matrix is a team review process, not a promise of automatic Tigy evaluation. Set task-specific criteria and use results to choose the next adjustment.
Record your evidence
Use the scenario matrix to record observations and execution references. Compare conversation outcomes with external actions. The file contains synthetic criteria and fields for completion without precomputed results.
Resources for your team
Define severity before running tests
Use criteria explaining decisions: pass, needs adjustment and blocking failure. Style details can require adjustment; wrong-record retrieval or unperformed-write confirmation should block release pending investigation. Averages must not conceal critical failures.
Give examples for each classification. Two reviewers can compare a sample and resolve disagreements before evaluating everything. If expected outcomes are disputed, refine the business rule.
After fixes, rerun neighboring cases. Identifier confirmation rules can affect questions without identifiers; tighter scope can reject valid requests. The matrix records these relationships and helps select regression scenarios.
Use the matrix during maintenance
Add rows when real failures expose missing conditions. Remove personal data while preserving date corrections, similar branches, empty results and ambiguity. This reproduces problems without the original customer.
Separate text, voice and telephone checks. Text tests decisions and tools; voice adds recognition and turns; telephony adds routing and provider conditions. Approval in one does not establish the others.
Before expansion, identify passed scenarios, unresolved ones and exception owners. This is a team evaluation method using Tigy evidence, not a claim of a native automated evaluation system.
Calibrate reviewers before interpreting small differences
Before scoring a complete version, ask two people to assess the same three conversations without seeing each other's ratings. Compare their reasons, particularly for disputed criteria. One reviewer may be scoring friendliness while another checks whether an action was confirmed. Refine the criterion and repeat the exercise. More consistent grading makes small differences between versions easier to interpret. Preserve the reasons as reference examples for future reviewers. If agreement remains poor, report that uncertainty alongside results instead of presenting a decimal score as a precise measurement. An honest assessment can identify useful changes while acknowledging that a particular dimension is not yet consistently judged.
Turn real situations into assessable cases
A useful matrix starts with the variety of situations rather than a list of desirable qualities. For a reception agent, answering correctly can mean reporting one location's opening hours, recognizing that the caller means a different location, or explaining that holiday opening has not been confirmed. Those situations require different behavior. Combining them into a single frequently asked questions row hides the exceptions that most need attention.
Choose a task and write a complete case: available context, opening utterance, information the agent may consult, and expected outcome. In a fictional example, someone asks whether they can collect a document on Friday. The source states ordinary hours but does not address the holiday occurring that day. The expected response should distinguish normal hours from special confirmation, explain the limitation, and identify the responsible contact. Requiring brevity alone is insufficient because an invented opening confirmation can also sound efficient.
Create variations of the same intention. The caller may ask whether the office opens Friday, refer to an earlier interaction, or correct the location after the first response. These variations assess understanding and context updates without changing the business objective. Add a case where the location is unknown, another where the caller interrupts the explanation, and another where the question exceeds the agent's scope. The set should distinguish comprehension, retrieval, and respect for boundaries.
Record criteria a second evaluator can apply consistently. Being friendly is broad; acknowledging a correction without blaming the caller and answering about the correct location is observable. For external actions, include independent evidence, such as a test record created in the authorized system. The agent's speech demonstrates communication but does not establish that a task completed. This separation prevents rewarding a persuasive promise as if it were an actual execution. Also record which supplied facts are intentionally unavailable so evaluators do not accidentally grade justified uncertainty as a failure.
Separate scores, release blockers, and improvement opportunities
An average score can track progress, but it must not erase failures that prevent release. Imagine twenty clear answers and one lookup exposing another person's information. Average clarity can remain high while the access problem remains unacceptable. Keep disqualifying criteria separate from gradual dimensions. The matrix may score clarity, repetition, and closing quality while independently blocking a version for inappropriate disclosure or falsely confirming an action.
Use a small scale with written definitions. For clarity, zero might mean the caller receives no understandable next step; one might mean the direction arrives after a confusing explanation; two might mean the next step is precise and understandable. The scale does not need to create an appearance of scientific precision. It needs to reduce disagreement between evaluators. Include an example for each level and update examples when new situations appear. Date rule changes so comparisons do not silently combine scores calculated under different definitions.
Assign observers to specific dimensions. An operations reviewer can judge whether a direction is executable; an integration owner can verify its external effect; someone familiar with the audience can assess vocabulary and comprehension. The same conversation may receive different judgments because reviewers inspect different parts. Joint review should resolve conflicts through evidence: transcript, tool return, external record, and current policy. Disagreement should not be settled merely by the reviewer's seniority.
Group outcomes by intention and severity. If failures cluster around spoken identifiers, confirming input may be the priority. If failures occur mainly on questions absent from the knowledge base, handling uncertainty may deserve attention. Overall averages conceal those distinctions. Also record unevaluable cases: a broken test environment should not become a conversational pass or failure without investigation. This classification keeps the matrix useful for decisions rather than forcing every event into a number. Show the number of attempted cases alongside scores, since a strong result from only two examples deserves a different level of confidence from a broad, consistently reviewed set.
Compare versions under conditions you can explain
A fair comparison starts with a fixed case set and a configuration record. Record the instruction version, available sources, attached tools, and relevant call conditions. If one version consults an updated document while another uses an older document, the result does not isolate a prompt change. The matrix should expose that difference instead of automatically attributing it to model quality or wording.
Repeat cases when behavior varies. Do not select only the new version's best conversation for presentation. Preserve unsuccessful attempts and label suspected causes without treating them as established facts. For voice, retain examples with pauses, corrections, and pronunciation resembling the intended audience. Comparing slow speech in an old test against hurried speech in a new one mixes input difficulty with the change under evaluation.
Check whether the adjustment harmed another task that previously worked. Instructions to end sooner may cut short a request with two subjects. Confirmation may prevent incorrect records but also repeat unnecessary questions. Show gains and losses together. The team may accept a slightly longer call to avoid an important error; record the reason for that choice.
Before publishing, write a short conclusion: what improved, what remains unstable and which errors still prevent use. State which group will use the new version first. In Tigy, record the instructions and version tested. The matrix is a spreadsheet maintained by the team; the platform does not automatically score all these criteria. Use fictional examples to repeat the evaluation without exposing customers.
