Skip to content
Tigy AI
Tigy AIVoice agentsBuild conversations for phone and webIntegrationsConnect agents to your systemsControl and reliabilityTest, monitor, and refine agents
Explore the platformSupport agentsAnswer common questions and route requestsLead qualificationUnderstand each contact's needsHow it worksGo from setup to live callsPricingFind a plan to get startedDocsLearn how to configure your agent
Areas of focus
TelecommunicationsFinancial servicesHealthcareTechnologyRetail and e-commerceMedia and entertainmentTravel and hospitality
Use cases
Customer supportLead qualificationAI receptionist
Business profiles
EnterpriseStartups
DocsBlogPricing
Platform
Voice agentsIntegrationsControl and reliability
Solutions
TelecommunicationsFinancial servicesHealthcareTechnologyRetail and e-commerceMedia and entertainmentTravel and hospitalityCustomer supportLead qualificationAI receptionistEnterpriseStartups
PricingDocsBlog
SIGN IN
Blog/Insights

WER in voice agents: transcription versus field accuracy

Compare transcription and task outcomes against a human reference as separate measures.

Author
Tigy AI team
Published
May 17, 2026
Updated
Oct 4, 2026
Explore voice agentsCreate an agent
Color fields with organic movement.
WER and data accuracy

In this article

  • Build a reviewed reference
  • Calculate word error rate
  • Check task fields
  • Test prompt-based correction
  • Use measures for diagnosis
  • Repeat the request under different conditions
  • Separate transcription errors from task errors
  • How do you calculate and interpret WER?
  • Prepare references that represent the audio
  • Read results alongside error types
  • Track fields through the operation using them
  • Choose samples representing the service
  • Change the responsible stage and check side effects
  • State what evaluation proves and what remains to verify
In this article
  • Build a reviewed reference
  • Calculate word error rate
  • Check task fields
  • Test prompt-based correction
  • Use measures for diagnosis
  • Repeat the request under different conditions
  • Separate transcription errors from task errors
  • How do you calculate and interpret WER?
  • Prepare references that represent the audio
  • Read results alongside error types
  • Track fields through the operation using them
  • Choose samples representing the service
  • Change the responsible stage and check side effects
  • State what evaluation proves and what remains to verify

WER means Word Error Rate: substitutions, deletions and insertions in a transcript divided by the word count of a reviewed reference. For voice agents, combine this measure with accuracy of fields such as identifiers, names and dates. One incorrect digit can prevent retrieval despite low WER. In Tigy AI, this guide describes manual or external evaluation rather than assuming a native WER dashboard.

Key takeawayMeasure word error and field accuracy separately; confirm identifiers before taking action.

Build a reviewed reference

Choose authorized test recordings or synthetic scripts read by your team. Two reviewers can compare the reference transcription and resolve disagreements. Define punctuation, spelled-out numbers and filler handling before comparison.

Calculate word error rate

WER is substitutions plus deletions plus insertions, divided by reference word count. Count errors by aligning the sequences. With ten reference words, one substitution and one deletion yield 2/10, or 20%, as a teaching example rather than a Tigy measurement.

Check task fields

Evaluate name, order code, date and phone separately when needed. One wrong address word can matter more than several filler words. Record expected and recognized values and whether conversational confirmation corrected the problem.

Test prompt-based correction

Ask the agent to confirm codes in groups and ask one question at a time. Request repetition when a name is unclear before performing a lookup. Transcription vocabulary and conversation instructions serve different purposes; check documented options in your configuration.

Use measures for diagnosis

Compare similar samples and record audio conditions. This is a manual evaluation procedure; it does not assume a native WER dashboard. Include correct actions, repetition and abandonment so a text metric does not improve at the expense of service.

Repeat the request under different conditions

Use the synthetic identifier DEMO-71 and a fictional product name. Prepare a written reference and verify what was actually spoken before evaluating transcription. Change one condition per round: quiet surroundings, background noise, pauses between digits and a corrected number. Do not caricature accents; work with voluntary participants speaking naturally and authorized to join the test.

Keep agent, version, channel and phrase comparable. Record microphone, condition and run time. Text testing reviews instructions but does not measure audio recognition. Repeat conditions to observe variation; one call does not establish an overall accuracy rate.

First compare results without changing the term dictionary. When investigating particular names, change only that setting and repeat the same audio. When used by the speech-to-text service, the dictionary helps recognize words; it does not change how the agent’s voice pronounces them.

Resources for your team

  • Audio conditions matrixCSV · English

Separate transcription errors from task errors

Hypothetical example: the human reference contains ten words; transcription has one substitution, one deletion and no insertions. WER = (1 + 1 + 0) / 10 = 20%. This demonstrates the calculation, not Tigy's performance. Document normalization of numbers, punctuation and abbreviations.

A wrong order digit can break the task despite low WER. Record the confirmed identifier, value sent to the tool and destination result. A spoken correction must reach query parameters.

Synthetic test criteria
ConditionVerification
Quiet surroundingsCorrect transcription and field
NoiseField confirmed before querying
Pauses between digitsComplete number
Correction from DEMO-17 to DEMO-71Query uses DEMO-71
Repeat after a changeVersion and condition recorded

How do you calculate and interpret WER?

Align the transcript with a reviewed human reference using consistent rules for numbers, punctuation and abbreviations. Count substitutions, deletions and insertions. In a teaching example with 20 words, two substitutions and one insertion give WER = (2 + 0 + 1) / 20 = 15%. This illustrates calculation, not Tigy AI performance.

The denominator is the reference word count, not the transcript word count. WER can exceed 100% when there are many insertions. Read the percentage alongside specific errors and critical fields: the average alone does not establish whether the correct order was queried or conversational confirmation repaired its identifier.

Prepare references that represent the audio

Transcription evaluation compares recognized text with reviewed references. If references contain errors or describe what speakers should have said, measures stop representing recognition of the audio. Preserve actually spoken content according to the selected annotation rule. References need review because incorrect expected text can make accurate recognition appear wrong.

Define punctuation, number, abbreviation and interrupted-word handling beforehand. “Twenty one” and “21” may express equivalent information, but unnormalized comparison can count them differently. Do not change conventions afterward to favor one configuration. Record decisions so reviewers apply the same treatment across the sample.

Review ambiguous passages. When nobody can confidently hear a word, record that condition and decide consistent sample handling. Apparently precise references based on guessing create misleading conclusions. Avoid silently using task expectations to fill uncertain audio because that can hide recognition difficulty.

Use identical references in comparisons and record audio samples and tested configurations. The objective is to detect recognition differences rather than combine data changes, writing conventions and configuration into one result variation. Keep evaluation information limited to what is necessary, using controlled examples where possible. This article describes a manual or external evaluation process; it does not imply that Tigy includes a native WER dashboard or automatic scoring of every field.

Read results alongside error types

WER uses substitutions, deletions and insertions after aligning reference and recognized sequences. Percentages summarize samples, while error types help guide investigation. Comparing total word counts alone does not provide this diagnosis: equally long sequences can contain different words, and different lengths do not identify which edits are needed.

Deletions may remove important words, insertions may add unspoken content and substitutions may replace terms. In each case, inspect audio and its relationship to the task before choosing corrections. A recurring error pattern deserves concrete examples showing where it occurs rather than an assumption that all errors share one cause.

Do not select universal approval thresholds without considering service. Names, codes and numbers have different effects from supporting words. Samples with equal WER may lead to very different operational outcomes. Review whether errors change identity, intent, dates or other fields governing decisions.

Use percentages as recognition-quality signals and supplement them with case inspection. Do not automatically convert the measure into resolution, safety or satisfaction rates. Those conclusions require evidence from later conversation stages. Keep comparisons tied to stated samples and normalization rules, avoiding unsupported claims that one observed score predicts every language, channel or environment. A transcription metric supports diagnosis; it does not independently establish that an agent completed its task correctly.

Track fields through the operation using them

For critical fields, record expected, recognized, confirmed and tool-submitted values. These four stages show whether initial mistakes were corrected or later interpretation introduced incorrect values. Final-state inspection adds evidence of what the external system actually used, rather than relying entirely on submitted parameters.

In a fictional example, callers dictate five-digit references. Transcription changes the final digit, but confirmation supports correction. Initial recognition quality remains useful information; final outcomes should also record that queries used the correct reference. Preserve both measures so successful recovery does not erase the initial problem and initial errors do not imply inevitable task failure.

The transcript may also be correct while the agent chooses a different number mentioned in the conversation. Changing the speech-recognition service may not solve that problem. Check which number was confirmed, which was used in the lookup and how the instructions explain that choice.

Evaluate consequences. Unusual name pronunciation and queries to wrong records are not equivalent. Define criteria by field and action, especially values changing authorization, person, branch or date. Report critical-field results separately from general transcription scores so serious mistakes remain visible even when overall text appears accurate. Avoid declaring operational reliability from a low aggregate error rate without checking these decision-relevant values.

Choose samples representing the service

Samples containing only clean speech and prepared sentences may not represent actual calls. Include audio conditions, number expressions and common public terminology. Record differences so analysis can show where performance changes. Use controlled or appropriately authorized material and avoid collecting unrelated personal information solely to make examples sound realistic.

Separate languages and channels where they change experience. Configuration working in browser-based editor voice tests does not establish equivalent telephone outcomes. Avoid comparing differently difficult samples as though transcriber configuration were the only change. Record sample composition alongside reported scores so comparisons remain interpretable.

Include negation and correction. “Not fifteen, fifty” requires preserving current values and sentence meaning. Tests should observe recognition and subsequent use because conversations may hear the words yet act on earlier numbers. Include dates and alphanumeric references where relevant to the actual task rather than relying only on ordinary sentences.

Maintain development samples and reserved cases for generalization checks where the process allows. Adjusting to known example lists may improve those lists without demonstrating benefit on unreviewed requests. Use reserved cases after changes and preserve failures for later regression checks. Describe limited sample coverage honestly instead of presenting a small curated collection as a complete account of production performance across all callers and conditions.

Change the responsible stage and check side effects

When errors originate in recognition, examine audio and available transcriber options. When they appear later, review interpretation, confirmation and parameters. This distinction avoids replacing components unrelated to observed failures. Keep diagnosis concrete enough to show which stage differs from expected behavior.

Change one hypothesis at a time where feasible. Repeat cases motivating changes and previously successful controls. Record effective configuration and compare using identical evaluation rules. Improvements that cannot be reproduced do not support strong conclusions. If several changes are necessary together, explain that limitation instead of attributing the entire effect to one component.

Include caller effort and time until useful next steps. Confirming every term may improve some fields while making service exhausting. Choose confirmations proportional to field importance and verify that they enable actual correction. Repeating incorrect values without allowing clear correction is not a useful recovery mechanism.

Do not hide regressions in averages. Changes may improve general sentences while worsening relevant codes. Critical groups should remain visible so decisions consider impact as well as aggregate percentages. Evaluate completed tasks and unresolved cases alongside recognition measures. A lower WER can support an improvement claim about the tested transcription sample, but broader claims about service need evidence that callers obtained correct results without unacceptable additional effort or failures elsewhere in the journey.

State what evaluation proves and what remains to verify

Useful reports identify samples, periods, configuration, normalization rules and results. Include representative failures and recovery examples. Without these conditions, percentages can appear comparable while measuring different audio and tasks. Clear reporting lets another reviewer understand the practical meaning of observed differences.

Show WER, critical-field outcomes and task results as distinct dimensions. The first describes recognition, the second tracks important data and the third confirms process completion. Their relationship guides subsequent investigation. Report cases where recovery corrected initial errors and cases where correct transcription still led to incorrect actions, preserving the diagnostic distinction.

State coverage limits. If long calls or particular languages were not tested, do not extend conclusions to those cases. Define missing checks before expansion. Maintain enough evidence to repeat the evaluation while excluding unnecessary sensitive information from shared reports.

In Tigy, treat this process as manual or external evaluation tied to available material and integrations. Do not announce native measurement absent from product documentation. Diagnostic value comes from trustworthy evidence and verifiable decisions even when analysis occurs outside the platform. Revisit the sample when channels, terminology or task scope change because earlier scores may no longer represent the operating environment. Use new observed results to support updated conclusions rather than carrying forward an old percentage as a permanent property of the agent.

In the Tigy documentation

  • Transcriber
  • Testing your agent

Make every conversation count.

Create an agent

Keep the conversation going

Light and shadow bands with a grain texture.
Voice agent pilot

How to launch an AI voice agent pilot in your business

Organic light ribbons with a soft texture.
Voice agent latency

Voice agent latency: how to investigate slow responses

Organic forms between light and deep shadows.
Cost per completed task

What does a voice agent cost? Calculate cost per completed task

Soft light veils around a textured abstract background.
Voice agent evaluation matrix

Voice agent evaluation matrix: scenarios and acceptance criteria

Soft color fields over a dark background.

How to compare voice agent versions with manual tests

Diffuse light and soft shadows in an abstract composition.
SIP telephony

SIP telephony in Tigy: routing calls to voice agents

Organic forms between light and deep shadows.
Voice agent metrics

How to measure AI voice agent outcomes

Soft contours over light and shadow fields.
Voice agent test checklist

Voice agent testing checklist before publishing

Tigy AI

PRODUCT

  • Voice agents
  • Integrations
  • Control and reliability
  • Demos
  • How it works
  • Pricing

Solutions

  • Telecommunications
  • Financial services
  • Healthcare
  • Technology
  • Retail and e-commerce
  • Media and entertainment
  • Travel and hospitality

Use cases

  • Customer support
  • Lead qualification
  • AI receptionist

Business profiles

  • Enterprise
  • Startups

Legal

  • Legal center
  • Terms of service
  • Privacy
  • Cookies

Resources

  • Blog
  • Documentation
  • AI documentation
  • Contact us

Social media

  • LinkedIn
  • Instagram