How to compare voice agent versions with manual tests
Separate hypotheses, scenarios and evidence to evaluate changes without altering multiple conditions together.
- Author
- Tigy AI team
- Published
- Updated
Comparing voice-agent versions means repeating known scenarios and assessing behavior before and after a change. In Tigy AI, use history as a reference, keep data and conditions comparable and inspect replies, tools and outcomes. Manual comparison is not an automatic traffic-splitting experiment.
Version comparison: an observable hypothesis and baseline
Write a verifiable hypothesis, such as confirming an ambiguous reference before lookup. Record the baseline version and change the relevant component. If documents, tools and voice change together, the test assesses their combined effect rather than one instruction. Include a case that must continue working too.
Record reference configurations and problematic examples. Different responses do not automatically establish improvement.
Repeat comparable situations
Test valid requests, missing details, corrections and system failures with matching data and conditions. Note destination changes.
Text isolates logic from audio. Spoken-understanding hypotheses also require voice tests.
Compare actions, not just wording
Inspect questions, parameters, responses and destination results. Shorter phrases help only when necessary information remains.
Record passed, failed or inconclusive scenarios. Missing external evidence cannot establish completed actions.
Publish after review
Consult history, edit the draft, save and test before publication. Previewing old versions does not automatically restore them.
This manual process assumes no automatic traffic splitting or Tigy A/B experiment feature. Review new conversations and retain scenarios for later changes.
Run the same set and look for regressions
For fictional new confirmation rules, test correct, corrected and missing references, including parameters and timing.
Include unrelated cases for regressions.
Evaluate source fidelity and operations rather than response preference alone.
Write the behavior change you expect to observe
Comparison starts with concrete hypotheses. “The agent improved” does not guide testing. “After correction, queries use confirmed codes” describes behavior observable in parameters and outcomes. State the expected change before running examples so interpretation does not follow whichever result looks favorable.
Separate problems from proposed fixes. Repeated information may indicate recognition mistakes, ambiguous questions or association failures. Instruction changes are appropriate hypotheses when investigation connects behavior to instructions or conversational structure. External delivery failures need different owners and checks.
In a fictional case, agents answer with another branch's policy. Staff may revise branch identification questions and selected sources. If both change, record that comparison evaluates a combined revision rather than attributing every effect to one new sentence. Keep the effective configuration visible to reviewers.
Define what must not deteriorate. Rules preventing wrong-record queries may begin requiring confirmation for every simple question. Review should preserve already useful cases as well as demonstrate intended corrections. Include caller effort and truthful next steps where they affect the task.
Choose a few criteria tied to outcomes and severity. Clear evaluation can explain approval and failure without depending on impressions from demonstrations. Distinguish critical failures from style preferences so polished wording cannot offset unauthorized access, wrong records or unsupported claims of completed actions.
Record what is being compared
Instruction revisions do not represent every service dependency. Documents, tools, models and routes can change during the same period. Record relevant configuration for each test to understand what supported responses. Include effective resource selection rather than assuming workspace availability means identical access.
Use identical cases and controlled data where possible. If references point to different states in second tests, responses can change without revision effects. Compare equivalent conditions or describe differences preventing direct attribution. Prepared data should represent the task while avoiding unintended changes to live records.
Avoid comparing one revision in old conversations with another in new sessions without considering history. Earlier requests can affect interpretation. Start comparable sessions and preserve conditions important to hypotheses. Cases evaluating long histories should reproduce those histories deliberately rather than acquire accidental context from prior experiments.
Separate saving, testing and publishing. Draft comparison does not establish that new calls already use approved configuration. Post-publication checks are separate and follow effective channels. Do not describe a local test as evidence of telephone routing or production behavior.
Preserve identifiers and concise notes supporting reproduction. Keep credentials and unnecessary personal information out of records. Short precise documentation is more useful than extensive outputs lacking revision association. If configuration differences remain unknown, retain that limitation in conclusions instead of claiming that wording alone caused observed changes.
Build correction and preservation cases
Collections need the failure motivating changes and requests already working. Isolated tests may show local improvement while hiding new refusals or confusion elsewhere. Choose controls relevant to the same task and a few unrelated successful paths where broader instructions could have side effects.
In a fictional example, new instructions require reference confirmation before queries. Test correct codes, corrected codes and absent codes. Add general questions not requiring references. These should still receive information without unnecessary collection. This checks that cautious behavior remains proportional to the operation.
Include empty returns, refused access and unavailable services for integrated tasks. New rules should not turn those states into success. Define caller-facing expectations and permitted external behavior. Inspect parameters and records, not only responses that sound compatible with the intended rule.
For voice, include interruption and repetition requests. Changes can improve text while producing sentences too long to hear. Comparison should observe the channel motivating revisions and its specific conditions. Browser voice checks do not establish telephone routing or transfer support.
Keep representative cases in public terminology. Vary examples where decisions change rather than creating dozens of equivalent rephrasings. The objective is finding relevant effects while maintaining a collection staff can repeat. Record expected clarification outcomes so ambiguous requests are not judged against an invented definitive answer.
Evaluate against criteria rather than preference alone
Review each case against expected outcomes. Compare information, boundaries, parameters and next steps. Responses may seem friendlier while using incorrect sources or promising actions that never occurred. Preferences can inform style decisions but should not replace operational correctness criteria.
Where possible, keep judgment independent of expectations about revisions. Authors may see intended behavior where responses remain ambiguous. Second review of critical cases helps detect this. Reviewers should use the same definitions so disagreement reveals unclear evidence or criteria rather than hidden differences in standards.
Record improvements, regressions and inconclusive outcomes. Missing evidence is not approval. If external systems do not support checking recorded changes, conclusions about completed actions remain limited even when final messages claim success. State what evidence is available and what still needs verification.
Consider severity and recovery effort. Extra questions may be appropriate for important fields. Repeated questions without decision effects can increase abandonment or rework. Evaluation should explain judgments in relation to tasks rather than treating all shorter conversations as better.
Use examples showing actual changes. Concrete cases with fictional values support discussion without distributing full personal accounts or reducing comparison to isolated ratings. Include a failed or uncertain case alongside successful examples when presenting conclusions, so reviewers can assess the practical limits of the revision before approving broader use.
Use samples for proportionate decisions
Small manual comparisons support review of executed scenarios. They do not prove universal improvement across all service. State sample size and scope alongside outcomes. Distinguish observed cases from expectations about production behavior or future demand.
Define publication conditions according to project processes. Fixing target failures and preserving essential cases may support limited scope if no critical failures remain. Authorization regressions require different priority. Do not let many easy successes offset unacceptable outcomes affecting access or external state.
If revisions fail to solve problems, return to hypotheses. Sources may be outdated or external contracts may not distinguish outcomes. Adding instructions after every attempt can make configuration harder to maintain without addressing causes. Inspect evidence before choosing another wording change.
After publication, start new conversations on real channels. Check association and relevant operational outcomes. This verifies effective revision use without assuming comparison sessions represent subsequent service. Where telephony is involved, editor checks alone do not approve number routing or transfer.
Record decisions and reasons. Staff need to know approved behavior and pending checks so expansion does not exceed evidence. Include owners for unresolved dependencies and the conditions under which conclusions should be revisited. A useful decision explains what is ready now, what remains limited and how the team will recognize whether the expected improvement continues during operation.
Preserve cases for the next revision
Comparisons leave examples and decisions supporting future reviews. Organize them by task and problem with evidence needed for repetition. Avoid depending on the tester's memory. Concise reproducible records make reviews usable when different staff take responsibility for maintenance.
Update cases when policies or integrations change, preserving revision reasons. Earlier expected outcomes may legitimately stop applying. Staff should distinguish approved changes from regressions. Record new validity conditions so tests do not enforce obsolete guidance merely because it was once correct.
Review later conversations according to scope and risk. Separate new problems from known failures. When untested conditions appear, add representative controlled-data cases and investigate causes. Preserve uncertainty where available evidence does not establish which stage failed.
Maintain operational restoration paths according to available capabilities. Agent versions do not automatically restore APIs, replaced files or external routes. Dependencies need owners and their own verification. Do not describe a written contingency plan as an implemented automatic rollback feature.
Expand evaluation from operational findings. Maintained collections prevent every decision from restarting with demonstration quality. Comparison supports concrete improvements staff can explain and reproduce. Keep established controls during expansion, and retire cases only when their underlying task or rule genuinely changes. This turns version review into continuing evidence of expected service rather than repeated preference judgments about whichever response was heard most recently.
