Context windows: what longer conversations require from your agent
Understand the model's working context and test corrections, topic changes and relevant information.
- Author
- Tigy AI team
- Published
- Updated
A context window is the limited amount of information a language model can consider when generating a response. It can include instructions, messages and supplied retrieval results. For a Tigy AI voice agent, the concept helps test corrections and long conversations, but does not imply permanent memory across calls. Reviewable history, selected documents and current-response context serve different purposes.
Distinguish context, knowledge and history
Instructions, messages and information supplied through documents or tools may contribute to a model's response context. Capacity is measured in tokens, text units that do not necessarily correspond to complete words.
Knowledge bases maintain sources for retrieval; reviewable history records an execution. Neither concept alone guarantees that every earlier record enters a new call.
Keep instructions relevant to the task
Clearly define goals, necessary questions and tool usage. Remove duplicate or conflicting rules that make the next action harder to determine.
Select sources relevant to the support task. Do not copy an entire archive into instructions to try to guarantee access. Test attached-document retrieval and use tools for current information from responsible systems.
Test details corrected near the end
Prepare a scenario with an initial office choice, intervening questions and a corrected office before the final lookup. Inspect tool parameters and the agent's confirmation.
Include changed goals and resumed earlier questions. Verify the value actually used as well as conversational fluency. A coherent reply can still accompany a lookup using outdated details.
Do not turn capacity into a memory promise
Tigy maintains the agent model as a managed service. This guide requires neither provider selection nor context-window configuration in the interface. Evaluate behavior within your operation's permitted tasks and durations.
If continuity needs information from an earlier call, design an authorized external-system lookup. Define necessary fields and confirm potentially changed details. Reviewable history alone does not establish automatic retrieval into future calls.
Are context, knowledge base and history the same?
Context is material available for the current response. A knowledge base maintains retrievable documents. History records a run for staff investigation. Having a transcript does not establish that its content automatically enters the next call.
If a task depends on earlier service, define an authorized query to the system retaining the necessary data. Confirm details that may have changed. A useful test compares two calls using synthetic records and checks how the second obtained each field, instead of inferring memory from conversational fluency.
Understand what contributes to a response
Conversations rely on sources with different roles. Instructions define expected behavior. History records what was said. Documents support rules. Tools return system state. A sentence appearing recently in a conversation does not acquire authority simply through proximity. This distinction matters when plausible caller statements conflict with approved information.
In a fictional example, a caller claims a fee no longer exists. The agent should not adopt that claim as policy. It can acknowledge the question and consult an approved source. If available material does not resolve the conflict, the answer should preserve uncertainty and offer a possible next step. Repeating an unsupported statement confidently would create a new policy without authorization.
Corrections to someone's own preferences serve a different purpose. When a caller changes their desired date, the next query should use the updated value. Distinguish that change from an attempt to alter an organizational rule. Reviewers should examine actual query parameters as well as spoken acknowledgment.
Write instructions explaining these differences directly. Avoid placing entire manuals in the prompt when the task uses only a portion. Available text volume does not guarantee correct use of every detail. Select resources appropriate to the task, then evaluate whether the agent applies their roles correctly in representative conversations.
Keep the current request clear after corrections
In a fictional case, someone requests Tuesday maintenance, asks about documents and then switches to Thursday. Testing does not finish when the agent says “understood.” The query must use Thursday while preserving the selected service unless that also changed. Acknowledgment and effective state are different pieces of evidence.
Confirm action-relevant fields at the appropriate moment. A brief check of service, branch and time can resolve ambiguity before booking. Do not repeat the entire history; highlight values relevant to the next step. Confirming every incidental detail can make calls tedious without improving decision accuracy.
When conversations include two requests, associate each value with its corresponding request. A contact number for one request should not be copied to another for convenience. Ask a specific question when the reference is unclear. Include test cases with similar names or dates because these expose association mistakes that straightforward examples may miss.
The connected system also needs to check the request and access rights. Do not rely only on the agent remembering a correction. Agree with the integration owner on how the correct request will be identified and an unauthorized change rejected.
Do not let an old result confirm a new request
History is not the only source of confusion. Queries take time, and callers may correct requests while waiting. Later results can correspond to earlier values even when the response looks valid. Evaluation should include this timing relationship instead of assuming every tool return answers the most recent utterance.
In a fictional case, the agent checks availability at the north branch. Before the response arrives, the caller selects the south branch. A time found in the north does not establish availability in the south. The conversation must clarify the association and perform the appropriate query where integration supports it. It should not silently reinterpret the old result.
During review, record event order: request, submitted parameters, correction and returned result. This sequence distinguishes reference problems from service unavailability. Technically successful responses may not answer current intent. Compare the branch, service and date in the result with current confirmed values rather than judging success from response status alone.
Define behavior for operations whose outcomes remain uncertain. Avoiding confirmation differs from claiming nothing happened. Following a write attempt, integration owners should provide a way to inspect state before retrying. This requirement concerns the external operation; conversational fluency cannot establish whether a booking or update was committed.
Plan long conversations without promising unlimited memory
Models have context limits varying with configuration. Do not publish one number as a universal rule without verifying the model in use. Even within a supported limit, long histories can make references less clear and introduce irrelevant information. More capacity does not automatically establish accurate interpretation or reliable action.
The practical response is to reduce task complexity. Separate requests when a caller mixes unrelated topics. Use a short confirmation when changing subjects and explain what remains unfinished. This organization improves caller understanding as well as the next response. It also gives reviewers clearer decision points when diagnosing mistakes.
Do not promise that agents remember previous contacts. Cross-call memory depends on implemented capabilities, available data and authorization. A configured CRM query can retrieve relevant information; that differs from automatically remembering every conversation. Retrieved records still require correct customer identification and appropriate access before disclosure.
Compare short and long conversations containing the same important decisions. If errors appear after many questions and answers, check which information was lost. The solution may be clearer confirmation, a narrower task or asking the connection owner to adjust how information is exchanged.
Build tests requiring reference tracking
Begin with a simple reference case: one request, one correction and one action. Add an intervening question, then introduce two requests with similar information in each. This progression helps locate where association becomes unreliable. It also avoids creating a complicated test whose failure has several indistinguishable causes.
Define expected outcomes for every scenario. Agents may need to ask again, and that is not necessarily failure. Short clarification is better than acting on the wrong record. Evaluate whether the question removes ambiguity or merely repeats generic wording. When ambiguity cannot be resolved, the expected outcome should preserve the limit rather than reward guessing.
Inspect parameters and final states for configured operations. For informational answers, compare applied rules and qualifying context. Do not reduce evaluation to naturalness ratings: requests can sound polished while remaining incorrect. Record the evidence showing whether the right request, person and time were used.
Include changed intent, negation and late correction. A caller may say, “I do not want to cancel; I want to know the deadline.” Responses should use current intent and respect operational limits. Preserve failing examples for retesting after changes. Also include a successful control case, ensuring that added caution does not make straightforward requests unnecessarily difficult to complete.
Use review to improve task design
Describe investigation findings concretely. “Used Tuesday after the Thursday correction” conveys more than “forgot context.” This record allows staff to determine whether a change fixes the observed behavior. It also helps distinguish incorrect state selection from recognition errors or external system responses.
Compare instructions, sources and tools in the tested revision. Integration changes can alter result identification; new instructions can introduce competing decisions. Without configuration records, staff cannot reliably attribute improvement or deterioration. Preserve enough detail to reproduce critical cases without retaining unrelated personal information.
Collect examples requiring callers to repeat information. Some expose reference failures; others show unclear initial questions. Review service experience as well, because confusing tasks force more corrections even when data is maintained accurately. A better question may reduce both caller effort and the number of competing details in history.
Expand scope only after current cases pass evaluation. Being able to discuss many topics does not establish precise execution across them. A clear verifiable scope provides a useful foundation for growth. Keep representative long-call failures in subsequent reviews so fixes remain effective when other resources or instructions change. Use actual observed outcomes to decide whether more capacity is needed rather than treating context size as the default explanation.
