How to test a voice agent with text and audio
Use each mode to investigate instructions, tools and audio behavior before publishing.
- Author
- Tigy AI team
- Published
- Updated
The Tigy AI editor supports agent tests through text and voice. Text helps investigate instructions and tools; audio checks recognition, pronunciation and turns. Repeat the same scenarios in both modes and verify actions in external systems. For telephony, add a call to the configured number: editor testing does not validate that connection.
Use the same scenario in both modes
Test by text to check instructions and tools, then test by voice to hear speech recognition, pronunciation, pauses and interruptions. Use the same scenario and information in both modes. Save changes, resolve editor warnings and ask the integration team to confirm test connections are being used when actions affect other systems.
A lookup scenario can include a valid number, a missing number and a correction. Following the same sequence makes differences between modes easier to interpret.
Isolate conversation logic
Send one question at a time in text testing. Inspect tool calls, parameters and how results are explained. This reduces the influence of speech recognition on the investigation.
If a lookup fails with correctly typed inputs, inspect its definition, attachment and instructions before investigating audio.
Check what only audio reveals
Start voice testing with authorized microphone access and speak naturally. Check the transcript, number reading, pauses and corrections made while the agent replies.
Listen to lengthy responses. A list that is readable in chat may be hard to follow aloud and need shorter questions.
Compare evidence, not just status
Review the run and destination system for external actions. A pleasant answer in both modes does not confirm a record was saved.
Document the scenario and tested version, fix the identified cause and repeat. End sessions when done; publish after verifying the channels you intend to use.
Keep conditions comparable
Use the same saved agent, documents and integration environment. Different data prevents isolating audio effects. Record changes and verify the workspace.
Use a known lookup reference. Inspect text parameters, then speak and correct a digit to verify updated voice parameters and understandable confirmation.
Telephone validation remains necessary before rollout. Browser success does not establish routing, provider audio or transfer behavior.
Keep tests explaining failures
Preserve fictional-data versions of similar names, pauses and interruptions causing failures. Reproduce them after correction without requiring original personal recordings.
Separate decision correctness from experience. Correct actions can involve repeated questions; comfortable calls can use wrong records.
Update scenarios after changes to models, sources, tools or channels. Unexecuted cases remain pending even when another channel passes.
Use text to investigate logic and voice to investigate conversation
Text and voice answer different questions about an agent. Text makes it easier to repeat exact wording, inspect a rule's interpretation, and observe tool usage without audio variation. Voice introduces pronunciation, speech recognition, pauses, and interruptions. An agent answering written messages correctly may still hear an identifier incorrectly or interrupt a spoken sentence. The checks complement each other.
At a fictional store, staff test an order lookup in text. The agent confirms the number, invokes the tool, and explains its response correctly. A voice test then uses the same number spoken naturally. If one digit is recognized incorrectly, the failure does not necessarily invalidate lookup instructions; it points toward recognition and confirmation in that channel. Rewriting the commercial prompt alone may not solve it. Inspect the spoken input and the identifier used for the actual lookup.
Start with text when the main question concerns an instruction, source, or state interpretation. Use direct, incomplete, and out-of-scope questions. Once behavior is understandable, move to voice to observe it during real conversation. Do not replace all voice testing with written messages simply because the results look stable. Text cannot establish whether a caller can hear or interrupt the answer, nor whether the agent correctly interprets a pause within a spoken identifier.
Keep objectives identical across channels when comparing logic. If the written scenario is simple but the spoken one includes several corrections, do not attribute differences only to audio. First repeat the same content, then introduce variations specific to speech. This order helps locate causes and prevents interpreting every difference as a model defect. Record which aspect each test establishes so later reviewers know whether an issue was checked through text, voice, or both.
Prepare expected outcomes and the environment before a session
A useful test needs an expected outcome, not merely an interesting question. For each scenario, describe what the agent may answer, which action it may perform, and which boundary it must preserve. A valid lookup may require a confirmed identifier and a response consistent with the result. An unavailable tool requires explaining inability and the defined alternative. These criteria make review less dependent on overall impressions of naturalness.
Use test endpoints and credentials for tools creating records or sending messages. Agent testing may access real systems when they are configured. Do not assume the word “test” prevents external effects. Review attached tools and supplied context before starting. For writes, inspect the corresponding system to confirm effects and prevent duplication. The test should establish intended behavior without creating unintended real customer activity.
Verify that the agent is saved, the account has permission, and workspace usage is authorized. For voice, allow browser microphone access. A startup issue may concern access, usage, or the device rather than instructions. Resolve those conditions before changing agent content in response to a session that never reached conversation. Record the startup error separately from conversational failures.
If a scenario depends on initial context, supply test values and include a missing-value variation. Handle absence according to instructions rather than inventing information. Record version, sources, and tools so repetition is comparable. Accidentally testing a different configuration can produce an incorrect conclusion about a correction. Preparation distinguishes task failures, environment failures, and version differences. It also makes it easier for another person to reproduce the scenario without guessing which values or integrations were present during the original attempt.
Include natural speech, corrections, and pauses
After a basic voice case, test how people actually speak. They may begin with a story, correct a number, interrupt an answer, or pause while remembering information. These variations need not become huge scenarios. A pause within a sentence reveals turn-ending behavior; a corrected identifier reveals whether the latest value is used; a repetition request reveals whether the needed information can be provided again.
In a fictional call, a customer dictates an order number, notices an error, and changes its final digit. Check transcription and the actual lookup. If the tool uses the first number, conversational data updating may be responsible. If recognition never captured the correction, investigate audio. A final answer may look correct for the wrong reference, making action inspection necessary rather than assessing wording alone.
Test a pause within a sentence and silence after completion. Turn-taking and inactivity controls have different functions. Tone guidance to “be patient” does not replace conversation settings. Change one control at a time where practical and repeat the same situation. If delay occurs after an external lookup, inspect integration before changing speech detection. Different causes may sound similar to a caller but need different remedies.
Include interruption of a long answer and confirmation of dates or locations. Callers should be able to correct relevant details without forced farewell. Listen to available recordings for pronunciation and clarity. Adequate transcription does not prove the spoken experience was comfortable. Voice review should examine what was heard and the operational outcome while separating audio difficulty from interpretation difficulty. This keeps corrections focused and prevents unnecessary prompt changes for problems that actually arise in timing or channel behavior.
Keep cases demonstrating corrections and boundaries
A corrected failure should become a repeatable case. Record the question, initial state, and expected outcome using fictional information where sufficient. After adjustment, test the original and a nearby variation. Fixing “What are Saturday hours?” should not make the agent apply those hours to Sunday. Regression should establish both the new answer and preservation of its conditions.
Separate content, action, and experience criteria. Answers should match sources, tools should use permitted parameters, and voice conversation should allow clear communication. A test may pass two dimensions and fail the third. This distinction helps decide whether to update documents, review tool descriptions, or change conversation settings. A single “it sounded good” judgment does not guide maintenance. Concrete evidence makes corrections explainable to another reviewer.
When testing a new version, include ordinary cases that already worked. A rule preventing unsupported confirmation may become indiscriminate refusal; instructions to ask fewer questions may remove an essential condition. Success and boundary controls show whether usefulness survived the correction. Record results with the tested version to avoid mixing evidence from different configurations. The collection should remain focused enough that it is actually used after relevant changes.
After publishing, verify a new conversation in the intended channel. Saving and publishing are distinct; an editor test does not establish that the public version is the same. For telephony, number assignment also needs verification. Final validation should follow the real path. The desired result is a small useful scenario collection revealing errors before expansion and supporting future review without relying on memory or an improvised session. Keep both the failed case and its successful repetition so the team can understand what changed and why the new outcome is considered correct.
