Pith. sign in

REVIEW 2 cited by

Automated test generation to evaluate tool-augmented LLMs as conversational AI agents

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.15934 v2 pith:BQND6NDD submitted 2024-09-24 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords llmsagentsconversationsevaluateprocedurestesttool-augmentedconversational
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Tool-augmented LLMs are a promising approach to create AI agents that can have realistic conversations, follow procedures, and call appropriate functions. However, evaluating them is challenging due to the diversity of possible conversations, and existing datasets focus only on single interactions and function-calling. We present a test generation pipeline to evaluate LLMs as conversational AI agents. Our framework uses LLMs to generate diverse tests grounded on user-defined procedures. For that, we use intermediate graphs to limit the LLM test generator's tendency to hallucinate content that is not grounded on input procedures, and enforces high coverage of the possible conversations. Additionally, we put forward ALMITA, a manually curated dataset for evaluating AI agents in customer support, and use it to evaluate existing LLMs. Our results show that while tool-augmented LLMs perform well in single interactions, they often struggle to handle complete conversations. While our focus is on customer support, our method is general and capable of AI agents for different domains.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems

    cs.CL 2025-01 conditional novelty 6.0 of 10

    IntellAgent uses a policy graph, synthetic event generation, and user simulation to automatically create benchmarks that rank conversational AI agents similarly to tau-bench.

  2. Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A survey that classifies AI agent evaluation benchmarks along environment and capability axes, and proposes five traits that distinguish agents from chatbots.

Pith tools