API failures in voice agents: troubleshooting and caller responses
Distinguish errors, empty results and incomplete responses without inventing confirmations.
- Author
- Tigy AI team
- Published
- Updated
When a voice agent cannot look up an order or record a request, the problem may involve the supplied details, the connection or your business system. That connection is called an API: a way for two systems to exchange information. In Tigy AI, start with the customer’s request, check the outcome and involve the integration owner to identify the cause. A missing record, denied access and an unconfirmed outcome need different explanations.
Check where the lookup stopped
Review a conversation where the problem occurred. Check the order number supplied by the caller and whether the agent attempted the lookup. If it did not, ask the configuration owner to check that the tool is selected and its instructions explain when to use it.
If an attempt was made, the integration team should check the service address, submitted details and access authorization. Keep the conversation time and reference to support the investigation without sharing passwords or access keys.
Separate possible responses
No result may mean that the order does not exist, but it can also mean that the lookup did not finish. Ask the responsible team to distinguish these situations. Denied access calls for a permission check; delays call for a service-availability check.
Receiving a system response is not enough to confirm completion. For a booking, check whether it confirms the reservation or only receipt of the request. The caller’s explanation must reflect that difference.
Prepare a useful explanation
Explain handling for each case in instructions. Verify identifiers for missing records; explain unavailable lookups and offer the agreed continuation path.
Before resending a request that creates or changes records, check the previous attempt’s outcome. Agree with the technical team on how the system prevents duplicate requests.
Reproduce failures in testing
Use test credentials and data for nonexistent identifiers, missing fields, unavailable services and rejected operations. Check the conversation and destination record.
Fix one cause at a time and repeat the test. Changing conversational instructions does not renew an expired access key; renewing a key does not correct an incorrectly submitted order number.
When confirmation is missing, check before retrying
A request may be saved even when its confirmation never reaches the agent. Sending it again can then create two records. Claiming completion without checking would also be incorrect.
Agree with the integration team on a way to find the request by its reference and check its outcome before retrying. Meanwhile, the agent should explain that completion could not be confirmed and offer the available follow-up route. Writing this instruction does not create a system lookup.
Test a missing confirmation, a repeated request and corrected details using fictional records. Check your business system and decide when staff should take over. Repeating a lookup and repeating record creation require different precautions.
Prioritize by consequence and exposure
Not every defect deserves the same correction order. A slow opening-hours lookup and an operation creating duplicate records may affect equal numbers of calls but have different consequences. Record frequency, impact, and recoverability separately. Frequency measures exposure; impact describes what happens to callers or operations; recoverability indicates whether a dependable procedure can restore the correct state.
Use concrete examples during triage. An understandable unavailability message may let a customer continue through another channel. A false confirmation may cause someone to attend an appointment that does not exist. A return containing another person's data requires access-control review even if observed once. Priority should reflect those consequences, not only complaint counts. Silent failures may become visible in external records before they appear in support requests.
Choose how to limit the problem while it is fixed. You may need to suspend an action that changes records, reduce available tasks or send certain lookups to staff. Define the owner and the test needed to resume. After correction, repeat the situation revealing the failure. On resumption, monitor the affected task and retain the human alternative until operation is verified in real use.
Locate the boundary where the request stopped
When an agent says it could not look up an order, that statement can represent very different failures. The intention may have been misunderstood, the identifier captured incorrectly, the tool never called, or the external service returned no expected record. Investigation must separate these possibilities. Changing instructions before locating the boundary can improve the apology while leaving the original defect untouched.
Reconstruct one case with fictional data. Compare what the caller said, the submitted parameter, the returned result, and the final response. If the caller supplied order 315 but the parameter contains 350, start with capture and confirmation. If the parameter is correct but the service reports no match, inspect environment, access rules, and record existence. If the return contains a valid order but the agent says nothing was found, output interpretation or ambiguous descriptions may be responsible.
Record when the failure occurred. The connection may drop before submission or after the system saved the change. Repeating a lookup may be acceptable; repeating booking creation without checking may create two. Ask the owner to inspect the destination system before deciding. No reply does not prove that nothing happened, and receiving a reply does not establish completion.
Use state names indicating what is known. Request not sent, request rejected, outcome confirmed, and outcome unknown are more useful than a generic error. That classification guides speech and recovery. In the unknown state, the agent can explain that completion cannot be confirmed and provide a follow-up route defined by the team. It should neither convert uncertainty into success nor declare permanent failure merely to close the conversation. Preserve a correlation reference when the integration supports one, allowing staff to connect the conversation to external events without requiring callers to repeat sensitive details.
Design returns that support an honest response
An integration should return meaning rather than an opaque data block. For availability checks, distinguish found availability, unavailable slots, invalid input, and temporary service unavailability. Each state can expose only the fields needed for conversation. An empty availability response should not sometimes mean no slots and sometimes mean an outage. If those outcomes share a representation, the agent has to guess what occurred.
For actions that create or change records, agree with the system owner on the information proving completion. A confirmed booking may include a reference, date, time and location. A request received for review must remain pending. If the reply only says success, ask whether that means receipt or completion. This definition prevents incorrect promises to customers.
Ask the system owner to check details and permission before performing the action. Instructions help the agent request appropriate information but do not replace this verification. For an order lookup, the system must check who may access it and return only necessary details. Knowing a telephone number does not automatically prove the caller’s identity.
Ask the responsible team to prepare test situations: a known order, a missing order, an invalid reference and an unavailable service. Check the agent’s explanation in each case, including replies with missing information. In Tigy, the tool description must state exactly what it does and which details it needs; a tool that only receives requests must not promise to complete the service.
Recover service without multiplying the problem
Recovery must depend on the operation. For reads, another attempt may resolve a temporary outage, but unlimited retries prolong the call and increase load on an already struggling service. Agree on a bounded policy with the technical team and provide an understandable alternative. The agent can explain that lookup is unavailable and identify the responsible contact. It does not need to narrate internal codes or error traces that do not help the caller decide what to do next.
When a task creates or changes a record, first discover whether the previous attempt worked. A record may have been saved before the reply was lost. The integration team can prepare a lookup using the request reference or protection against duplicate submissions. That capability must exist and be tested in the system. If verification is unavailable, hand over an unconfirmed outcome and avoid promising that retrying is safe.
Prepare staff to continue with the request, confirmed details, attempted action and known outcome under applicable access rules. If the agent only says there was a problem, the caller may repeat everything and staff may resend an already-created request. The procedure needs to identify where to check before acting. When confirmation takes time, explain only the timing or channel actually defined by operations.
After identifying the cause, keep a fictional example to repeat during future changes. If information was missing from the system reply, test that absence. If the tool description confused the agent, test its corrected description and another task that already worked. Record the change and inspect the system record. A more courteous explanation helps the customer, but correction is ready only when the outcome is communicated accurately and the process avoids repeating the harm.
