Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale
Edesio Alcoba, Kevin Rossell, Aman Gupta, Shao Tang, Jiwoo Hong, Pabel Carrillo-Mendoza, Wanderson Conceição Ferreira, Alvaro Tedeschi, Zayd Simjee, Shreya Rajpal, Bruno Finardi Hime, Christian Sousa, Luis Moneda, Herbert Fei, Daniel Silva, Rohan Ramanath
Abstract
Customer experience (CX) agents use tools and large language models to address customer requests and guide conversational interactions with an organization's products. Improving these agents, especially in regulated industries, is difficult: they must detect intent, follow complex operational policies and use tools reliably. Manual end-to-end testing offers limited coverage, while live experiments expose customers to failures that can erode trust. We present a hypothesis-driven simulation workflow for screening candidate CX agents before deployment. Synthetic customers react to agent responses and simulated tool outputs enable multi-step agentic workflows without invoking production backends. We use the Snowglobe simulator on Nubank's Card Delivery agent and its expanded successor, Card Management - Nubank's highest-volume chat-support agent in Brazil. Across 4 deployed versions, simulated and production version-level binary evaluator scores show high correlation. Simulation-guided iteration increased transactional net promoter score (tNPS) by 36.69 points in a live A/B test. We also screened open-weight configurations in over 16,000 simulated conversations. In a subsequent live A/B test, the selected model increased self-service rate (SSR) by 8.82 percentage points to the highest level observed at Nubank, with no statistically significant change in tNPS. Simulation made broad exploration of models, reasoning settings, and prompts feasible without customer exposure, enabling production improvements that would have been impractical to pursue through live experimentation alone.
Create a lesson
Related papers
GRASP: Generating, Revising, and Assessing for Strategic Planning with Agentic AI
Arunabh Srivastava, Mohammad A., Khojastepour et al.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
HEXIS: Compiling Skills into Extended Finite State Machines
Minghao LI
PrivDrift: Auditing User-Secret Leakage Under Topic Drift in Active LLM Conversations
Luciano Maldonado
Self-Play Pretraining with Zero Data
Aditya Cowsik, Kfir Dolev, Michael Y. Li et al.
SciWalker: Synthesizing Scientific Coding Problems with Operator Graphs and Execution Feedback
Chenxi Li, Wenxuan Zeng, Yun Luo et al.