Pith. sign in

REVIEW 4 cited by

Interpretable User Satisfaction Estimation for Conversational Systems with Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.12388 v2 pith:VTPU5U5R submitted 2024-03-19 cs.IR cs.AI

classification cs.IRcs.AI
keywords satisfactionuserconversationalinterpretablesystemsapproachesestimationlanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Accurate and interpretable user satisfaction estimation (USE) is critical for understanding, evaluating, and continuously improving conversational systems. Users express their satisfaction or dissatisfaction with diverse conversational patterns in both general-purpose (ChatGPT and Bing Copilot) and task-oriented (customer service chatbot) conversational systems. Existing approaches based on featurized ML models or text embeddings fall short in extracting generalizable patterns and are hard to interpret. In this work, we show that LLMs can extract interpretable signals of user satisfaction from their natural language utterances more effectively than embedding-based approaches. Moreover, an LLM can be tailored for USE via an iterative prompting framework using supervision from labeled examples. The resulting method, Supervised Prompting for User satisfaction Rubrics (SPUR), not only has higher accuracy but is more interpretable as it scores user satisfaction via learned rubrics with a detailed breakdown.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Verifiable Self-Evolution for Open-Ended Dialogue Skills via Future-Feedback Prediction

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Future-feedback skill evolution makes dialogue-skill self-improvement verifiable offline by learning to predict the logged resolved/unresolved user signal rather than scoring counterfactual replies.

  2. How can we assess human-agent interactions? Case studies in software agent design

    cs.AI 2025-10 conditional novelty 6.0 of 10

    PULSE combines sparse human ratings with prediction-powered inference to cut confidence intervals by ~40% and shows LLM choice matters more than scaffolding for user satisfaction.

  3. Conversation Progress Guide : UI System for Enhancing Self-Efficacy in Conversational AI

    cs.HC 2025-01 reject novelty 5.0 of 10

    A progress bar with subtask markers added to a conversational AI chat measurably increased task-specific self-efficacy in a 22-person user study.

  4. PACT: A Contract-Theoretic Framework for Pricing Agentic AI Services Powered by Large Language Models

    cs.GT 2025-05 conditional novelty 4.0 of 10

    PACT models agentic AI services as a menu of quality-price contracts and uses contract theory to show incentive-compatible, individually rational pricing, with numerical examples for cybersecurity log analysis.

Pith tools