Skip to content
Tigy AI
Tigy AIVoice agentsBuild conversations for phone and webIntegrationsConnect agents to your systemsControl and reliabilityTest, monitor, and refine agents
Explore the platformSupport agentsAnswer common questions and route requestsLead qualificationUnderstand each contact's needsHow it worksGo from setup to live callsPricingFind a plan to get startedDocsLearn how to configure your agent
Areas of focus
TelecommunicationsFinancial servicesHealthcareTechnologyRetail and e-commerceMedia and entertainmentTravel and hospitality
Use cases
Customer supportLead qualificationAI receptionist
Business profiles
EnterpriseStartups
DocsBlogPricing
Platform
Voice agentsIntegrationsControl and reliability
Solutions
TelecommunicationsFinancial servicesHealthcareTechnologyRetail and e-commerceMedia and entertainmentTravel and hospitalityCustomer supportLead qualificationAI receptionistEnterpriseStartups
PricingDocsBlog
SIGN IN
Blog/Guides

STT, LLM and TTS: how voice agent architecture works

Understand a voice agent's stages and locate the cause of a conversation that did not work.

Author
Tigy AI team
Published
Aug 4, 2026
Updated
Oct 4, 2026
Explore voice agentsCreate an agent
Soft color fields over a dark background.

In this article

  • STT turns speech into text
  • The LLM uses instructions and context
  • TTS turns the answer into audio
  • Locate the failure before adjusting
  • Evaluate the complete response path
  • Choose components for the service you need
  • Reconstruct the path between speech, decisions and audio
  • Understand what Tigy makes configurable
  • Test recognition using service vocabulary
  • Evaluate decisions against task contracts
  • Listen to answers as part of the outcome
  • Use the same case to locate differences
In this article
  • STT turns speech into text
  • The LLM uses instructions and context
  • TTS turns the answer into audio
  • Locate the failure before adjusting
  • Evaluate the complete response path
  • Choose components for the service you need
  • Reconstruct the path between speech, decisions and audio
  • Understand what Tigy makes configurable
  • Test recognition using service vocabulary
  • Evaluate decisions against task contracts
  • Listen to answers as part of the outcome
  • Use the same case to locate differences

A voice agent needs to listen, interpret and respond. Speech recognition turns audio into text, the language model interprets the request, and speech synthesis turns the response into audio. These stages are known as STT, LLM and TTS. Tigy manages these components while your team configures instructions, documents, tools, available voices and conversation controls.

Key takeawayTreat recognized words, generated answers and audible speech as three separate pieces of evidence.

STT turns speech into text

STT (speech-to-text) converts speech into text; ASR (automatic speech recognition) is another term for automated speech recognition. Compare audio with the transcript before evaluating the answer. Noise, names and service vocabulary can change an identifier before the model decides which tool to use.

Tigy manages the transcriber. The dictionary can supply terms relevant to recognition; test product names and abbreviations in your service conversations.

The LLM uses instructions and context

The language model interprets a question using agent instructions and available information. A knowledge base supplies documents; a configured tool can retrieve data or perform an action.

If recognition is correct but the answer is wrong, inspect sources, instructions and tool responses. A model does not automatically know your current company rules.

TTS turns the answer into audio

Speech synthesis converts the answer's text into sound. Listen to how the voice presents amounts, dates and names. Clear written sentences may need simpler phrasing to be understood over the phone.

Tigy offers selected voices and manages synthesis. The recognition dictionary is not a phonetic pronunciation control; changing STT terms should not be expected to change the voice.

Locate the failure before adjusting

For replies before a question ends, inspect turn controls. For slow lookups, examine the tool and destination. For unclear audio, compare the written answer with the available recording.

Repeat the same scenario after changing a setting. Separating stages gives each change a testable hypothesis and avoids adjustments that conceal the original cause.

Evaluate the complete response path

Time from speech completion to first audio includes recognition, decisions, network, tools and response generation. Fast model inference is not necessarily fast customer experience.

Compare knowledge-only questions with tool-dependent ones. Delays isolated to lookups suggest integration investigation. Shorter responses can reduce total duration without changing time to first audio.

Record channel, connection conditions, configured models and tools. Compare similar conditions and inspect slow cases as well as averages. Aim for predictable, understandable and source-supported interaction.

Choose components for the service you need

Evaluate transcription with representative vocabulary, interpretation with ambiguous requests and pronunciation with real phrases. Short demos do not establish performance on codes, abbreviations or interruptions.

Change one available setting at a time, such as instructions, sources, tools, selected voice or turn controls, while holding the other scenarios constant. Tigy AI manages STT and LLM; this procedure does not include selecting external providers.

Check available Tigy configuration options and evaluate the whole agent. Do not assume identical outcomes across languages and channels. Tests should represent operation and continue after dependency changes.

Reconstruct the path between speech, decisions and audio

To understand voice failures, trace information through stages. Customers speak, transcription represents the message, the model decides how to respond or invoke tools, and speech synthesis produces audio. Wrong answers can originate before decisions or afterward. Inspecting only the final sentence does not reveal origin.

Consider a fictional code. Someone says “forty-seven,” but transcription records “forty-six.” The model may correctly retrieve the value it received while the task reaches the wrong record. Explanation matches input and service still fails. Correction needs recognition and confirmation.

In another case, transcription records the right code but the model invokes modification instead of retrieval. The cause concerns decisions, descriptions or instructions rather than voice quality. In a third, text and retrieval are correct but spoken dates confuse listeners. Review needs the audible output.

This trace prevents broad changes without causes. Maintain expected speech, observed transcript, action and response text in test records. Listen to relevant audio where available. Evidence identifies the stage actually needing adjustment.

Preserve caller corrections in the trace too. They reveal whether a later value replaced the first one before any external operation occurred.

Understand what Tigy makes configurable

Current Tigy documentation describes platform-managed STT and LLM. Users do not select external providers or models for those stages. Voice uses a curated Tigy voice collection; external identifiers and arbitrary synthesis providers cannot be added. This changes practical guidance.

Instead of recommending provider combinations, investigate controls actually available. Instructions, sources, tools and conversation settings affect service. Agent dictionaries can supply recognition terms where used by the transcriber. They are not phonetic dictionaries altering synthesized pronunciation.

Do not confuse architectural knowledge with interface options. Explanations can present STT, LLM and TTS conceptually while guiding operations through managed capability. Users need to know which adjustments they can make and which problems require platform or integration investigation.

Review old instructions requesting external model selection. Current configuration should guide examples and procedures. Where controls are absent, do not turn abstract recommendations into implementation steps. Explain the layer and test behavior staff can observe.

This distinction also keeps release checks concrete. Verify selected voices and available conversation settings rather than inventing provider configuration the product does not expose.

Test recognition using service vocabulary

Generic demonstration vocabulary does not represent every domain challenge. Retail tests should include products, neighborhoods and abbreviations. Reception tests should include specialties, surnames and dates. Observe where transcription changes information affecting decisions.

Use dictionaries deliberately. Documentation describes comma-separated terms supplied as recognition support. Compare before-and-after trials with the same sentences. Do not conclude that this repairs every similar name or promise changed TTS pronunciation.

Confirm important fields conversationally. Correct recognition in one example does not remove the need to verify codes, dates and identifiers where precision matters. Confirmation should let callers identify and correct errors. Then inspect the value reaching tools.

Include sentences with pauses. Premature responses can cut data before transcription completes. Review turn-ending detection in conversation settings. Changing answer wording does not necessarily repair when the system stopped listening.

Keep audio conditions visible in findings. Quiet microphone trials and noisy telephone calls are different evidence. Repeat meaningful cases through the actual intended channel before generalizing recognition quality.

Evaluate decisions against task contracts

At the decision stage, examine alignment between intent, source and action. General questions use approved information, individual data uses authorized tools, and out-of-scope actions require recognition. Fluent language does not establish correct selection.

Test missing information and pressure for confirmation. Callers may demand certainty about estimates or insist on unavailable actions. Models should retain boundaries. External systems maintain their own authorization because prompts do not replace access controls.

For tools, inspect descriptions and parameters. Broad names can cause inappropriate selection, and unclear fields can receive wrong values. Test result interpretation too: acceptance differs from completion, forecasts from guarantees, and absent records from proof that no problem exists.

For failures, change cause-related instructions and repeat nearby cases. Do not repair unsupported confirmation by refusing every lookup. Decision quality requires distinguishing conditions rather than sounding confident or cautious in every circumstance.

Include combined intentions to verify that successful completion of one task does not grant permission for another. Retrieval may be allowed while modification still needs separate authorization.

Listen to answers as part of the outcome

Response text must work when heard. Long sentences containing several conditions can be accurate yet difficult to follow. Organize result, important condition and next action into short parts. Do not remove currency, units or estimate wording merely to speed speech.

Compare available voices with real vocabulary. Listen to names, dates, codes and waiting explanations. Pleasant opening samples do not establish full-service performance. Ask listeners to repeat heard values and check whether confirmation was understandable.

Responses should allow interruption under configuration. Callers may correct dates during speech. Review whether the agent stops, understands and resumes using the new value. Silence after interruption does not prove updating; inspect the next action.

Bring evidence across stages together. Tasks are correct when input, decisions, effects and explanations match intent and approved contracts. Component separation supports causal repair, while experience still needs validation as a complete conversation.

Retain a representative full-call case for future checks. It can expose regressions spanning several stages that isolated tests would miss.

Use the same case to locate differences

Prepare a test sentence with a known outcome and retain conditions while comparing adjustments. Changing vocabulary, channel and task simultaneously prevents attributing differences to a layer. Use text to check decisions and voice to examine input and output while preserving business facts.

Record observed effects instead of generalizing from one execution. A correctly recognized name supports that case; other names and environments still need testing. Investigation becomes more useful when it describes limits and the next relevant case.

If text succeeds and voice fails, inspect the transcript and timing before rewriting all instructions. If both fail with the same correct input, examine the shared decision or integration contract. If spoken wording alone is unclear, review output phrasing and selected voice. This sequence narrows the search while retaining the complete task as the final verification.

In the Tigy documentation

  • Transcriber
  • LLM
  • Voice

Make every conversation count.

Create an agent

Keep the conversation going

Color fields with organic movement.
AI voice agents

What is an AI voice agent and how does it work?

Organic forms between light and deep shadows.
Closing a support call

How to end a voice agent call with a clear next step

Soft contours over light and shadow fields.
IVR or AI voice agent?

IVR or AI voice agents: choosing for your customer service task

Soft light veils around a textured abstract background.
Publishing agent versions

How to save, test and publish agent versions in Tigy

Soft contours over light and shadow fields.
Conversational and generative AI

Conversational and generative AI: differences in customer service

Soft color fields over a dark background.

How to choose a voice and test your AI agent's vocabulary

Soft contours over light and shadow fields.
Preparing knowledge documents

How to prepare documents for an agent's knowledge base

Organic light fields for prompt agents.
Sales conversation practice
Instructions

# Conversation

Confirm before acting.

Example instructions

How to practice sales conversations with an AI agent in Tigy

Tigy AI

PRODUCT

  • Voice agents
  • Integrations
  • Control and reliability
  • Demos
  • How it works
  • Pricing

Solutions

  • Telecommunications
  • Financial services
  • Healthcare
  • Technology
  • Retail and e-commerce
  • Media and entertainment
  • Travel and hospitality

Use cases

  • Customer support
  • Lead qualification
  • AI receptionist

Business profiles

  • Enterprise
  • Startups

Legal

  • Legal center
  • Terms of service
  • Privacy
  • Cookies

Resources

  • Blog
  • Documentation
  • AI documentation
  • Contact us

Social media

  • LinkedIn
  • Instagram