Abstract
Contact centers increasingly pair Large Language Models (LLMs) with human agents, yet real world deployments struggle to balance quality, safety, latency, and cost. We present a hybrid orchestration framework that coordinates LLM skills, retrieval components, and human expertise across the conversation lifecycle-triage, assist, escalate, and resolve. The orchestrator uses policy signals (risk, user intent, account tier), live system metrics (latency, token cost), and outcome feedback to decide among autonomous generation, tool augmented responses, or agent collaboration. We instantiate the framework in three configurations: (i) Agent Assist co-pilot for drafting and summarization, (ii) Human in the Loop RAG for grounded answers with citation coverage monitoring, and (iii) Collaborative Escalation where LLMs propose plans and humans approve high risk actions. In offline and live shadow evaluations on multi domain customer service datasets and simulated queues, the framework improves first contact resolution and average handle time while maintaining safety and brand tone. We contribute: (1) a general orchestration policy with guardrails, (2) a cost quality-latency trade off analysis, (3) counterfactual ablations showing when human collaboration adds value, and (4) an evaluation protocol combining operational KPIs with human factors measures (e.g., workload and trust). We release prompts, configs, and anonymized evaluation scripts to facilitate reproducible studies of human-AI collaboration in service settings.
Showing the abstract — retrieve the full paper via the Exa API.