WER in voice agents: transcription versus field accuracy
Compare transcription and task outcomes against a human reference as separate measures.
- Author
- Tigy AI team
- Published
- Updated
WER means Word Error Rate: substitutions, deletions and insertions in a transcript divided by the word count of a reviewed reference. For voice agents, combine this measure with accuracy of fields such as identifiers, names and dates. One incorrect digit can prevent retrieval despite low WER. In Tigy AI, this guide describes manual or external evaluation rather than assuming a native WER dashboard.
Build a reviewed reference
Choose authorized test recordings or synthetic scripts read by your team. Two reviewers can compare the reference transcription and resolve disagreements. Define punctuation, spelled-out numbers and filler handling before comparison.
Calculate word error rate
WER is substitutions plus deletions plus insertions, divided by reference word count. Count errors by aligning the sequences. With ten reference words, one substitution and one deletion yield 2/10, or 20%, as a teaching example rather than a Tigy measurement.
Check task fields
Evaluate name, order code, date and phone separately when needed. One wrong address word can matter more than several filler words. Record expected and recognized values and whether conversational confirmation corrected the problem.
Test prompt-based correction
Ask the agent to confirm codes in groups and ask one question at a time. Request repetition when a name is unclear before performing a lookup. Transcription vocabulary and conversation instructions serve different purposes; check documented options in your configuration.
Use measures for diagnosis
Compare similar samples and record audio conditions. This is a manual evaluation procedure; it does not assume a native WER dashboard. Include correct actions, repetition and abandonment so a text metric does not improve at the expense of service.
Repeat the request under different conditions
Use the synthetic identifier DEMO-71 and a fictional product name. Prepare a written reference and verify what was actually spoken before evaluating transcription. Change one condition per round: quiet surroundings, background noise, pauses between digits and a corrected number. Do not caricature accents; work with voluntary participants speaking naturally and authorized to join the test.
Keep agent, version, channel and phrase comparable. Record microphone, condition and run time. Text testing reviews instructions but does not measure audio recognition. Repeat conditions to observe variation; one call does not establish an overall accuracy rate.
First compare results without changing the term dictionary. When investigating particular names, change only that setting and repeat the same audio. When used by the speech-to-text service, the dictionary helps recognize words; it does not change how the agent’s voice pronounces them.
Resources for your team
Separate transcription errors from task errors
Hypothetical example: the human reference contains ten words; transcription has one substitution, one deletion and no insertions. WER = (1 + 1 + 0) / 10 = 20%. This demonstrates the calculation, not Tigy's performance. Document normalization of numbers, punctuation and abbreviations.
A wrong order digit can break the task despite low WER. Record the confirmed identifier, value sent to the tool and destination result. A spoken correction must reach query parameters.
| Condition | Verification |
|---|---|
| Quiet surroundings | Correct transcription and field |
| Noise | Field confirmed before querying |
| Pauses between digits | Complete number |
| Correction from DEMO-17 to DEMO-71 | Query uses DEMO-71 |
| Repeat after a change | Version and condition recorded |
How do you calculate and interpret WER?
Align the transcript with a reviewed human reference using consistent rules for numbers, punctuation and abbreviations. Count substitutions, deletions and insertions. In a teaching example with 20 words, two substitutions and one insertion give WER = (2 + 0 + 1) / 20 = 15%. This illustrates calculation, not Tigy AI performance.
The denominator is the reference word count, not the transcript word count. WER can exceed 100% when there are many insertions. Read the percentage alongside specific errors and critical fields: the average alone does not establish whether the correct order was queried or conversational confirmation repaired its identifier.
Prepare references that represent the audio
Transcription evaluation compares recognized text with reviewed references. If references contain errors or describe what speakers should have said, measures stop representing recognition of the audio. Preserve actually spoken content according to the selected annotation rule. References need review because incorrect expected text can make accurate recognition appear wrong.
Define punctuation, number, abbreviation and interrupted-word handling beforehand. “Twenty one” and “21” may express equivalent information, but unnormalized comparison can count them differently. Do not change conventions afterward to favor one configuration. Record decisions so reviewers apply the same treatment across the sample.
Review ambiguous passages. When nobody can confidently hear a word, record that condition and decide consistent sample handling. Apparently precise references based on guessing create misleading conclusions. Avoid silently using task expectations to fill uncertain audio because that can hide recognition difficulty.
Use identical references in comparisons and record audio samples and tested configurations. The objective is to detect recognition differences rather than combine data changes, writing conventions and configuration into one result variation. Keep evaluation information limited to what is necessary, using controlled examples where possible. This article describes a manual or external evaluation process; it does not imply that Tigy includes a native WER dashboard or automatic scoring of every field.
Read results alongside error types
WER uses substitutions, deletions and insertions after aligning reference and recognized sequences. Percentages summarize samples, while error types help guide investigation. Comparing total word counts alone does not provide this diagnosis: equally long sequences can contain different words, and different lengths do not identify which edits are needed.
Deletions may remove important words, insertions may add unspoken content and substitutions may replace terms. In each case, inspect audio and its relationship to the task before choosing corrections. A recurring error pattern deserves concrete examples showing where it occurs rather than an assumption that all errors share one cause.
Do not select universal approval thresholds without considering service. Names, codes and numbers have different effects from supporting words. Samples with equal WER may lead to very different operational outcomes. Review whether errors change identity, intent, dates or other fields governing decisions.
Use percentages as recognition-quality signals and supplement them with case inspection. Do not automatically convert the measure into resolution, safety or satisfaction rates. Those conclusions require evidence from later conversation stages. Keep comparisons tied to stated samples and normalization rules, avoiding unsupported claims that one observed score predicts every language, channel or environment. A transcription metric supports diagnosis; it does not independently establish that an agent completed its task correctly.
Track fields through the operation using them
For critical fields, record expected, recognized, confirmed and tool-submitted values. These four stages show whether initial mistakes were corrected or later interpretation introduced incorrect values. Final-state inspection adds evidence of what the external system actually used, rather than relying entirely on submitted parameters.
In a fictional example, callers dictate five-digit references. Transcription changes the final digit, but confirmation supports correction. Initial recognition quality remains useful information; final outcomes should also record that queries used the correct reference. Preserve both measures so successful recovery does not erase the initial problem and initial errors do not imply inevitable task failure.
The transcript may also be correct while the agent chooses a different number mentioned in the conversation. Changing the speech-recognition service may not solve that problem. Check which number was confirmed, which was used in the lookup and how the instructions explain that choice.
Evaluate consequences. Unusual name pronunciation and queries to wrong records are not equivalent. Define criteria by field and action, especially values changing authorization, person, branch or date. Report critical-field results separately from general transcription scores so serious mistakes remain visible even when overall text appears accurate. Avoid declaring operational reliability from a low aggregate error rate without checking these decision-relevant values.
Choose samples representing the service
Samples containing only clean speech and prepared sentences may not represent actual calls. Include audio conditions, number expressions and common public terminology. Record differences so analysis can show where performance changes. Use controlled or appropriately authorized material and avoid collecting unrelated personal information solely to make examples sound realistic.
Separate languages and channels where they change experience. Configuration working in browser-based editor voice tests does not establish equivalent telephone outcomes. Avoid comparing differently difficult samples as though transcriber configuration were the only change. Record sample composition alongside reported scores so comparisons remain interpretable.
Include negation and correction. “Not fifteen, fifty” requires preserving current values and sentence meaning. Tests should observe recognition and subsequent use because conversations may hear the words yet act on earlier numbers. Include dates and alphanumeric references where relevant to the actual task rather than relying only on ordinary sentences.
Maintain development samples and reserved cases for generalization checks where the process allows. Adjusting to known example lists may improve those lists without demonstrating benefit on unreviewed requests. Use reserved cases after changes and preserve failures for later regression checks. Describe limited sample coverage honestly instead of presenting a small curated collection as a complete account of production performance across all callers and conditions.
Change the responsible stage and check side effects
When errors originate in recognition, examine audio and available transcriber options. When they appear later, review interpretation, confirmation and parameters. This distinction avoids replacing components unrelated to observed failures. Keep diagnosis concrete enough to show which stage differs from expected behavior.
Change one hypothesis at a time where feasible. Repeat cases motivating changes and previously successful controls. Record effective configuration and compare using identical evaluation rules. Improvements that cannot be reproduced do not support strong conclusions. If several changes are necessary together, explain that limitation instead of attributing the entire effect to one component.
Include caller effort and time until useful next steps. Confirming every term may improve some fields while making service exhausting. Choose confirmations proportional to field importance and verify that they enable actual correction. Repeating incorrect values without allowing clear correction is not a useful recovery mechanism.
Do not hide regressions in averages. Changes may improve general sentences while worsening relevant codes. Critical groups should remain visible so decisions consider impact as well as aggregate percentages. Evaluate completed tasks and unresolved cases alongside recognition measures. A lower WER can support an improvement claim about the tested transcription sample, but broader claims about service need evidence that callers obtained correct results without unacceptable additional effort or failures elsewhere in the journey.
State what evaluation proves and what remains to verify
Useful reports identify samples, periods, configuration, normalization rules and results. Include representative failures and recovery examples. Without these conditions, percentages can appear comparable while measuring different audio and tasks. Clear reporting lets another reviewer understand the practical meaning of observed differences.
Show WER, critical-field outcomes and task results as distinct dimensions. The first describes recognition, the second tracks important data and the third confirms process completion. Their relationship guides subsequent investigation. Report cases where recovery corrected initial errors and cases where correct transcription still led to incorrect actions, preserving the diagnostic distinction.
State coverage limits. If long calls or particular languages were not tested, do not extend conclusions to those cases. Define missing checks before expansion. Maintain enough evidence to repeat the evaluation while excluding unnecessary sensitive information from shared reports.
In Tigy, treat this process as manual or external evaluation tied to available material and integrations. Do not announce native measurement absent from product documentation. Diagnostic value comes from trustworthy evidence and verifiable decisions even when analysis occurs outside the platform. Revisit the sample when channels, terminology or task scope change because earlier scores may no longer represent the operating environment. Use new observed results to support updated conclusions rather than carrying forward an old percentage as a permanent property of the agent.
