Voice agent testing checklist before publishing
Check common questions, boundaries, interruptions and API failures in voice agent tests before publishing and opening a phone pilot.
- Author
- Tigy AI team
- Published
- Updated
A voice agent checklist covers common questions, out-of-scope requests, interruptions and tool failures before publishing. In Tigy AI, test instructions in the editor and validate the pilot's phone channel as well; a passing test does not cover every possible scenario.
Turn tasks into acceptance criteria
A voice agent test checklist should cover the ordinary task and decisions that change its outcome. Prepare complete requests, missing inputs, corrections, unavailable sources and out-of-scope questions. Write expected speech and actions before execution; hearing a convincing answer is not enough to approve a case.
Include situations requiring an exit: an out-of-scope topic, incomplete information or a request for a person. Define who continues the service and how the caller will be informed.
Use text and voice to investigate different issues
Tigy's text test helps review instructions and tools independently of audio. Voice testing checks recognition, pronunciation, pauses and interruptions.
Save changes before comparing rounds. Repeat the same scenario and adjust one aspect at a time so that improvements and regressions can be traced to a particular change.
Include the less convenient paths
Ask for repetition, correct a name and pause midway through a sentence. Simulate missing information and tool unavailability in a test environment.
- Does a correction replace the previous value?
- Does the agent wait for a sentence to finish?
- Does the action use confirmed parameters?
- Is failure explained without an invented outcome?
- Does the conversation end with an understandable next step?
Record evidence and rerun after changes
Review transcripts, actions and available run data. Record the scenario, tested version, expected outcome and observed behavior. Keep cases that exposed problems for future checks.
Begin with a pilot your team can monitor. New documents, instructions or integrations can change behavior; rerun the checklist before expanding service.
Download and complete a test matrix
The matrix contains fictional data and eight scenarios. Fill observed behavior, execution reference and external evidence for actions. Blank cells are for your testing; they do not represent measurements already performed.
Start with text to check instructions and tools, then repeat recognition and interruption cases in voice. Record evaluated configuration or version, channel and date. Pass only when evidence meets the criterion; convincing speech does not prove a write.
Correct one cause per round and repeat failed and previously passing cases. Preserve the earlier matrix for comparison. Open the CSV file, a table format, in a spreadsheet or text editor. You can download it without accessing internal systems.
| ID | Scenario | Expected |
|---|---|---|
| T01 | Approved hours | Answer from source |
| T02 | Missing reference | Ask before lookup |
| T03 | Empty return | Do not invent result |
| T04 | API failure | Explain limit and next step |
| T05 | Corrected number | Use confirmed correction |
| T06 | Voice interruption | Respect turn configuration |
| T07 | Out-of-scope request | Route to responsible staff |
| T08 | Repeated operation | Do not duplicate at destination |
Resources for your team
Run cases without affecting production
For tests that create or change records, ask the integration team for a test connection with its own access key and fictional data. Agree on which records may be changed. If the test sends a notification, check its recipient before starting: a test conversation can still trigger a real system.
Cover success, empty data, rejected validation and unavailability. For writes, include a lost response after execution. The adapter must establish status without duplication. Repeat with data corrected mid-conversation.
Listen to complete audio, checking identifiers, interruption and repetition. Text tests do not reveal those problems. Then call the pilot number and review its run in the correct workspace.
Turn the checklist into a release decision
Assign a state and evidence to each item. Unexecuted cases remain pending; similarity to another case does not establish approval. Record failures and intended fixes, then rerun affected scenarios.
Define release-blocking errors, decision owners and alternative service channels. Staff need an operational way to stop the pilot as well as edit the agent. Plan how to correct the draft, save, test and publish again before relying on that procedure.
After publishing, start a new conversation through the real channel. Preserve the matrix for future versions. It remains useful when documents, tools or telephony change, since external changes can affect service with an unchanged prompt.
Write the expected outcome before starting a test
A useful test begins with an expectation independent of the generated answer. Writing criteria after hearing the agent can lead reviewers to accept convincing language that does not match the task. Record intention, initial data, permitted action and expected outcome. Criteria should make it possible to explain why an execution passed or failed.
For a fictional order lookup, a case can specify a valid code and a known test-system result. The expectation is to retrieve that code and explain its status without turning an estimate into a guarantee. Another case uses an incomplete code: the agent should request the rest before retrieval. The set evaluates different decisions although both conversations concern orders.
Do not describe success merely as “Answer correctly.” State the supporting source and the condition that must be preserved. For a return policy, the agent should state the approved period and explain eligibility applying to the item. Where policy does not cover the case, the expected result may be acknowledging insufficient information and identifying the appropriate contact.
Include negative criteria. The agent must not claim retrieval without a tool, confirm writing without a result or invent missing details. These prevent pleasant language from concealing authority errors. Also assess whether unnecessary questions expose data or obstruct the task even when the final answer is accurate.
Use fictional data exercising the real contract. An identifier impossible to validate may test only initial rejection. To evaluate successful retrieval, prepare a test record following system rules. Separate cases testing validation from those testing explanation of valid results.
Combine language variation with operational conditions
A collection of complete, well-written questions omits much of voice service. People abbreviate, stop mid-sentence, answer out of order and change their minds. For each primary intention, prepare a direct version, an informal version and one missing information. Then add system conditions, such as empty results or unavailability.
You need not test every imaginable combination at once. Prioritize combinations changing decisions or risk. Correcting a number before retrieval matters; correcting it before a write may matter even more. A different courtesy phrase with unchanged intent usually deserves less attention than an ambiguous tool result.
Include changes of intention within a call. Someone asks about availability, selects a time and then decides against booking. The agent should follow the current decision and avoid creating a reservation from the previous intention. Another case begins with a sales question and ends with a support request. Final routing should not remain fixed to the initial destination.
Test conflicting information. A caller may insist a service is free while approved sources describe specific conditions. The expectation is to explain available information clearly rather than automatically agree. Include questions quoting old policies to verify that current sources prevail.
For voice, include natural pauses, domain-specific names and numeric sequences. Listen for understandable confirmation and whether callers can correct a field. An apparently correct transcript does not establish that the voice pronounced a code intelligibly. Text and audio provide different evidence about the same execution.
Finally, test requests for human help and endings. Someone may request a person before explaining the issue or decide they no longer wish to continue. The agent should respect those changes within available options. Reviewing only the opening and central answer omits the moment when the conversation actually delivers or abandons its outcome.
Use failures to build a regression collection
When a case fails, record input, tested version and observed behavior. Classify the likely cause before changing configuration: instruction, source, parameter, tool execution or audio condition. Make the corresponding change and repeat the case. Then repeat a nearby passing case to check that the correction did not move the problem elsewhere.
A regression collection should preserve situations previously causing errors and essential service behaviors. It need not expand for every differently worded sentence; retain cases representing distinct decisions. A smaller collection with clear criteria can be more useful than many questions without consistent evaluation.
Before publication, combine results with operational review. Does the transfer destination answer? Has the responsible owner configured production credentials? Can staff identify pending requests? Conversation tests do not replace these checks, but should reveal dependence on them.
Document what remains unvalidated. If every trial used text, do not claim speech-recognition quality. If testing covered only retrieval, do not claim write capability. Accurate test reporting supports expansion with realistic expectations and turns each new condition into a concrete verification.
Use the same collection after relevant changes to tools or sources. A new manual can alter answers even when the prompt stays the same. The release decision should reflect the combined configuration people will encounter, rather than only the latest sentence edited.
