How to plan customer service capacity for voice agents
Plan peaks, integrations and human continuation without equating credits with concurrent capacity.
- Author
- Tigy AI team
- Published
- Updated
To plan voice support, estimate peak call arrivals, call duration and how many conversations happen at once. Include business-system lookups and work left for staff. In Tigy AI, verify platform, telephony and external-connection limits with their owners. Available credits or known monthly volume do not guarantee capacity to handle every call during a peak.
Voice capacity: volume, peaks and concurrency
Separate volume from concurrency. Volume counts calls over a period; concurrency counts conversations active at the same time. Arrival peaks combined with long calls can need much more capacity than the same volume spread throughout the day. Use representative intervals, observed duration and the share of tool-dependent tasks for estimates.
Use matching agent, workspace and report period. Concurrency estimates need peak conditions rather than daily averages.
Check each dependency's limits
Confirm applicable capacity and limits for telephony providers, your instance and external APIs. Slow or limited lookups can impair service even when calls connect.
Do not infer published capacity from an example or plan. Verify operational conditions with service owners.
Include usage and human teams
Check billing state and available usage before tests. Authorization for new runs and available duration differ from concurrent call capacity.
Plan human queues for handoffs and pending records. More intake requires enough capacity to complete team-dependent work.
Expand with observation and criteria
Test under appropriate conditions and monitor failures, response time, outcomes and recovery. Agree on criteria to stop expansion if experience worsens.
This planning guide does not guarantee Tigy scale, availability or numeric limits. Record verified conditions for reassessment after changes.
Test a limited representative peak
Use authorized environments and volumes with representative proportions of simple questions, lookups and routing. Greeting-only loads do not measure tools.
Track delay, refusal, errors and downstream work. Rate-limited integrations should not retry continuously.
Record dependencies and scale gradually once bottlenecks and failure behavior are understood. Limited tests do not establish unlimited capacity.
Plan for short periods rather than monthly totals alone
Monthly call totals do not reveal how many people try to obtain service simultaneously. Moderate volume may be distributed unevenly, with peaks after campaigns or opening time. Observe arrivals over operationally useful intervals and measure task duration. Identify periods concentrating demand and the intentions common in each period.
Use observed information where available. Record contact attempts, answered calls, abandonment, and repeat contact. Reports limited to completed conversations can conceal people who never obtained access. Without history, make assumptions explicit and use a bounded pilot to measure. Initial estimates should not become capacity commitments. Show uncertainty alongside the conditions selected for publication.
In a fictional example, one business receives many status inquiries early in the afternoon. These take less time than profile changes but arrive together. Another receives fewer calls, with long conversations requiring external retrieval. Similar monthly totals can therefore create different requirements. Planning from totals mixes arrivals, duration, and external dependencies. Separate those factors to identify which component needs attention.
Consider the journey after the first call. Failures causing new attempts increase demand without representing new customers. Clearer next steps may reduce repetition, but the effect must be measured. Record campaigns, policy changes, and events affecting the audience when comparing periods. Capacity should follow actual demand rather than projected growth from a comfortable average. Preserve peak intervals separately from quieter ones so a monthly improvement cannot conceal worsening service during the exact hours that matter most. When estimating new campaigns, also include the expected duration mix instead of assuming every additional contact resembles the shortest existing inquiry.
Use duration to understand overlapping conversations
Duration connects arrivals with occupancy. In an illustrative calculation, sixty calls uniformly distributed over one hour, each lasting five minutes, create three hundred conversation minutes within that period. Dividing by sixty minutes yields an average overlap of five conversations. That is arithmetic under stated assumptions, not a guarantee that five positions serve every moment. Arrivals can cluster and durations can vary.
Inspect distributions rather than average duration alone. Some tasks finish quickly while exceptions spend much longer in retrieval or routing. Those cases may occupy capacity during peaks. Separate tasks and examine long conversations. Determine whether time was necessary for safe completion or resulted from repetition, silence, and attempts lacking alternatives. Removing important confirmation may improve occupancy while worsening outcomes.
Do not infer platform limits from this calculation. Verify actual telephony, execution, and external-service conditions with environment owners. Planning must respect observed and authorized limits; it assumes neither unlimited concurrency nor a fixed call allowance per account. Document excess-call handling where restrictions apply. The overload experience is part of service and requires testing.
Compare assumptions through a team spreadsheet or report. Change arrivals, duration, and exception proportion separately to reveal sensitivity. Small duration increases may matter more under concentrated arrivals. The objective is not a definitive number from a few inputs but identifying which assumption needs measurement and which condition needs protection before expansion. Keep the calculation labeled as a planning estimate rather than presenting it as an observed production measurement. Once real intervals are available, compare estimate and observation to improve the next planning cycle.
Plan the dependencies serving tool requests
Conversational capacity does not remove limits in consulted systems. Agents may depend on calendars, CRM, authentication, and order lookup. Many simultaneous conversations can slow or trigger rejection from shared services. Inventory dependencies by task and identify external calls required by normal and exceptional cases. Retries also consume those resources.
Discuss limits and test environments with each service owner. Do not place load on production systems without authorization and appropriate conditions. Small tests can assess controlled delay or rejection. Observe speech while waiting, closure when retrieval does not finish, and writes with lost responses. Agents should communicate actual state rather than repeatedly promising near completion without evidence.
Use reduced outputs and bounded tasks when they improve the contract. Retrieving only needed facts can reduce unnecessary processing and exposure. Optimizations must remain compatible with responsible systems and validated outcomes. Do not remove essential fields merely to reduce latency. Be careful with reusing variable data: general policies and individual slot availability have different freshness requirements.
Define alternatives when dependencies compromise tasks. Informational inquiries may remain available while booking requires routing. This preserves part of service without announcing unavailable capabilities. Record decision ownership and resumption conditions. Operational capacity includes understandable degradation, boundaries, and recovery rather than maximum demonstration volume alone. Check how one unavailable dependency affects unrelated intentions: a calendar failure should not necessarily prevent the agent from answering approved address questions, but the configuration must make that distinction explicit and testable. Document any shared failure that genuinely requires broader containment.
Include human handling in demand forecasts
When conversations end in transfer or later review, human capacity remains important. Agents can receive more contacts than staff can complete. Routing that ignores opening hours and availability merely moves queues. Plan case volume, delivered context, and time needed to recover uncertain operations. That effort is absent from metrics limited to automated duration.
In a fictional situation, the agent handles simple questions and routes complex changes. A campaign increasing those changes raises the routing proportion. Staff needs that information before exposure grows. Do not use historical transfer rates from periods dominated by simple inquiries as universal forecasts. Segment intentions and define pending-work priorities when demand exceeds availability.
Verify receipt of pending cases. Accurate summaries must reach real destinations with confirmed information and operation state. If callers repeat everything, human time may remain high. If writes have unknown outcomes, recipients need to check systems before executing again. Integrations and shift handover must support that verification.
Define honest experiences when destinations are unavailable. The agent may identify another channel or receive a request under existing procedure. Do not promise immediate connection or precise callbacks without approved conditions. Test peaks with absent staff and cases requiring later follow-up. Capacity includes those outcomes because answered calls without continuity can produce new contacts and enlarge the next period's queue. Compare forecast routed volume against observed receipts after the pilot, checking whether cases were lost, duplicated, or placed in an unexpected queue. This verifies that capacity planning covers the receiving operation rather than stopping at the agent's closing message.
Increase exposure with defined stopping criteria
An expansion plan should identify observed indicators and conditions stopping growth. Track access, completion by task, external failures, repeat contact, and human pending work. Define criteria with service owners rather than imposing one arbitrary duration threshold across intentions. Serious failures can require containment despite good averages.
Start with exposure the team can review. Compare equivalent periods and record events changing demand. When assumptions fail, revise planning before expansion. Explain which dependencies supported testing, which remained constrained, and which recovery was verified. Preserve fictional examples for repetition after changes.
Useful planning produces an explainable operating condition: supported tasks, circumstances, dependencies, and alternatives. That brief supports evidence-based growth without turning approximate arithmetic into unlimited-service promises. Assign ownership of the next capacity review and identify triggers beyond the calendar, including new campaigns, changed tool behavior, or a rising share of complex requests. These changes can invalidate earlier estimates while total call volume still looks ordinary. Review peak-period evidence separately, since service deterioration concentrated in a short window deserves action even when the full-day average appears stable. Finally keep the overload alternative tested as configuration evolves; a plan to redirect callers is useful only while the destination and wording remain accurate.
