Skip to content

AdaT2: Adaptive Test Transformations for Black-Box Boundary Testing of Conversational Agents

Liting Lin, Boxi Yu, Qinghua Xu, Yuzhong Zhang, Lionel Briand, Emir Muñoz

cs.SEarXiv:2610.10141

Abstract

Conversational agents based on large language models (LLMs) must comply with policies. Each condition in a policy draws a boundary between user requests, and the agent must behave differently on its two sides. We present AdaT2, which extracts statements from the agent's replies in exploratory conversations with an LLM acting as the user, and uses the statements to guide boundary test generation. Each statement describes one condition and the behavior expected when the condition holds. Besides plain tests guided by single statements, AdaT2 writes transformed tests guided by pairs of a statement and a test transformation instruction, such as "omit one required input". The instruction of a pair can move a test to the other side of the statement's boundary or to another boundary. The statements and instructions form far more pairs than a run can try, and many pairs are not applicable. Adaptive pair selection therefore chooses the statement of each pair by novelty and the instruction with the bandit algorithm Bayes-UCB, which learns from whether earlier pairs yielded a test and whether the agent passed it according to an LLM judge. Our benchmark counts the two sides of each boundary separately and distinguishes boundaries explicitly defined by the agent's prompt, tool code, or knowledge base from boundaries that the agent's LLM infers from domain knowledge. On four domains of τ3-bench, 62.7% to 83.3% of AdaT2's tests are valid boundary tests whose expected behavior is explicitly defined, higher in every domain than AgentEval's (47.7% to 68.1%). Transformed tests add 13 to 46 explicitly defined boundaries that plain tests miss. As regression tests, AdaT2's test suites detect all eight seeded policy faults in the airline domain and four of eight in the retail domain, and AgentEval's test suites, with fewer than a third as many tests, detect five and two, respectively.

Create a lesson