How to review voice agent calls and evaluate outcomes
Review voice agent calls, identify failures and compare responses with confirmed actions in the responsible system.
- Author
- Tigy AI team
- Published
- Updated
Reviewing voice agent calls means checking what the caller requested, how the agent responded and which action actually completed. In Tigy AI, conversation review provides evidence; sales and bookings must also be confirmed in the responsible system.
Choose a sample with a purpose
Reviewing voice agent calls requires a sample including completed tasks, failures and handoffs. In Tigy AI Agent Runs, select agent and period, record how conversations were chosen, and define the expected outcome. An order lookup must match confirmed status or clearly explain why retrieval did not complete.
Write the expected result before inspecting details. An order lookup should provide a confirmed status or clearly explain that the lookup could not be completed.
Understand what each record proves
Check status, time, duration, transcript and actions where available. Listen to recordings when investigating interruptions or misunderstood numbers. Availability depends on channel, configuration and run state.
Text conversations do not create voice recordings. Missing audio in that channel does not establish a failure.
Classify the cause before editing
Separate missing information, ambiguous questions, incorrect parameters and unavailable systems. A wrong answer may come from an outdated document; a rejected action may come from credentials rather than the model.
Compare the tool record with the destination system. A spoken promise to create a request does not prove the request exists.
Turn an issue into a test scenario
Keep the run identifier and reproduce the situation in testing. Make a specific correction and check that it fixes the issue without harming other scenarios.
Compare duration and usage with outcomes from the same period. Aggregated reports reveal trends, while conclusions about a conversation require its individual evidence.
From a problematic conversation to a correction
Fictional example: a caller asks ‘Will my order arrive tomorrow?’ and the agent says ‘Yes’ without querying a tool. It announced a deadline without evidence. Keep a run reference when available and check query, result and confirmation. This dialogue is not real customer service.
Before editing, identify the responsible layer. An outdated source needs a source correction; a tool returning the wrong field needs a contract correction. For an invented answer without a query, the hypothesis below limits what may be announced.
Prompt
# Order deadlines
Do not confirm delivery dates by guessing.
For an individual order, query the configured tool only after authorization.
If its result has no estimate, explain the date could not be confirmed.
Without a tool, provide the approved team channel.
Confirm only what was verified; do not announce an action that never occurred.Repeat the failure and previously valid cases
Save the draft change and repeat the request. Record versions, input, response and destination evidence. Communicate a returned authorized estimate; without one, explain it could not be confirmed.
Include a valid general question, denied access, unavailable tools and a valid question after a referral. Preserve the previous version under your process. Publish after reviewing observed results.
| Test | Criterion |
|---|---|
| No estimate | Do not invent a date |
| Authorized estimate | Communicate the returned value |
| Denied access | Do not query or expose data |
| General hours in the base | Answer valid information |
| Tool failure | Do not confirm a completed query |
Resources for your team
Compare speech, tools and destination records
Transcripts establish what was said. Tool records establish parameters and available results. Destination systems confirm tickets, reservations or updates actually exist. Divergence among them is central to review.
If the agent claims to create a request after an API (application programming interface) error, classify unsupported confirmation and investigate why failure did not change speech. If a record exists but the agent announced failure, investigate interpretation or lost responses before retrying.
Keep run identifiers and timestamps for correlation. Avoid unnecessary personal-data copies into review sheets. A reduced problem description plus controlled evidence access can suffice for reproduction.
Make review produce a concrete change
Record the problem, probable cause, owner and verification test. “Improve service” is vague; “the document does not distinguish branches; update it and rerun three questions” is actionable. Integration failures require their contract owner.
Turn problematic calls into fictional-data scenarios. Include correct behavior and a neighboring regression condition. After changing the source, rerun both and check destination records.
Schedule regular reviews and reviews after significant changes. Even unchanged agents depend on documents, APIs and staff processes. Subsequent outcomes reveal failures absent from the call itself, such as requests sent to an unmonitored queue.
Reconstruct the task before judging the response
Review begins by identifying what the caller wanted to accomplish. A conversation may contain a general question, individual retrieval and attempted modification. Assessing only the final answer can miss incorrect intermediate actions or declare success because one of three intentions was resolved. Record tasks and their states separately.
For important claims, seek evidence. The agent says it checked an order: is there a tool call using the correct identifier? It claims an update: does the result confirm writing? It explains policy: does the approved source contain that condition? This chain turns impressions into verifiable assessment.
Observe corrections too. Someone may provide a code and replace it seconds later. Final wording can use the right value while an earlier retrieval already used the wrong one. Inspect action sequence and correction timing. Review should establish whether operational outcomes match current intent.
Do not confuse transcripts with the whole experience. Where audio is available, listen to interruption, repetition and confirmation segments. Codes may appear correctly in text while being difficult to recognize aloud. Pauses may feel like abandonment even when transcripts represent waiting differently.
Record what cannot be concluded. Ending after an answer does not establish satisfaction. Missing tool responses do not prove no effect occurred. Use categories retaining uncertainty and investigate responsible systems where outcomes depend on them.
Classify causes to choose corrections
Wrong answers have different origins. Sources may be outdated, tools wrongly selected, parameters may ignore corrections, or explanations may distort valid results. Labeling everything an AI failure makes corrective decisions difficult. Record the stage and evidence pointing to likely causes.
Separate content from interpretation with a simple question: was correct information available in this execution? If not, inspect source, processing and association. If available, examine retrieval and explanation. Stronger wording does not necessarily repair absent documents.
For integrations, compare information collected during the conversation, information sent and the response received. Correct information about the wrong order is still a service failure. An action refused for lack of permission must not produce a success confirmation. Show this distinction to both the agent owner and the connected system owner.
For audio problems, observe recognized words, response timing and confirmation. The issue may involve names, pauses treated as completed speech or confusing pronunciation. Do not change every component simultaneously. Choose a hypothesis and a case verifying its effect.
Ask receiving staff where outcomes depend on operational rules. A transfer may work technically while targeting a department unable to resolve the issue. The cause may concern destination definition or advertised scope. Dialing records alone cannot answer that question.
Choose a sample showing successes and exceptions
Reviewing only complaints can hide useful behavior and encourage overly defensive instructions. Reviewing only completed calls can hide abandonment and pending requests. Select primary contact reasons and important exceptions. Record sampling criteria so reports do not imply measurement of the entire operation.
Use criteria multiple people can apply. Two reviewers can compare a small sample and discuss differences. One may label retrieval resolved while another requires customer confirmation. Define categories and sufficient evidence. This improves consistency before expanding analysis.
Handle records under company procedure. Share only what review needs and use fictional examples when turning failures into editorial material or public tests. Complete conversations may contain data unrelated to the investigated cause.
Keep review objectives connected to tasks. Pleasant voices and courtesy matter but cannot compensate for unauthorized retrieval or unconfirmed changes. Assessment should show interaction quality and operational fidelity as related dimensions with separate evidence.
Turn findings into tests and decisions
For recurring failures, describe reproducible input, expected behavior and proposed correction. Make the cause-related change and repeat the case. Then repeat a similar previously successful situation. Together they show improvement and whether correction introduced another problem.
Assign responsibility for changes and follow-up. Source updates, tool changes and service-rule revisions may involve different owners. Without ownership, reports become observations recurring in subsequent calls.
After publication, review new conversations exercising the corrected condition. Earlier tests provide controlled evidence; new service shows whether the pattern changed within the observed population. Record limitations and continue investigation where samples cannot support conclusions.
Retain successful patterns as well. They provide reference cases showing what should remain stable while the agent evolves. Review is most useful when it produces both concrete repairs and evidence of behavior worth preserving.
Follow pending work to the responsible system
Choose a request ending pending and inspect subsequent handling through the authorized process. Did staff receive it? Did the information support action? Was the outcome confirmed? This can reveal continuity problems even when conversation correctly collected every field.
Use findings to repair unowned steps or missing contracts. Do not add follow-up promises to prompts until operations can actually fulfill them.
Keep conversational completion separate from business completion when reporting this case. That distinction shows where the agent performed correctly and where the wider service still needs improvement.
