REVIEW 3 cited by
A Preliminary Evaluation of ChatGPT for Zero-shot Dialogue Understanding
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Zero-shot dialogue understanding aims to enable dialogue to track the user's needs without any training data, which has gained increasing attention. In this work, we investigate the understanding ability of ChatGPT for zero-shot dialogue understanding tasks including spoken language understanding (SLU) and dialogue state tracking (DST). Experimental results on four popular benchmarks reveal the great potential of ChatGPT for zero-shot dialogue understanding. In addition, extensive analysis shows that ChatGPT benefits from the multi-turn interactive prompt in the DST task but struggles to perform slot filling for SLU. Finally, we summarize several unexpected behaviors of ChatGPT in dialogue understanding tasks, hoping to provide some insights for future research on building zero-shot dialogue understanding systems with Large Language Models (LLMs).
Forward citations
Cited by 3 Pith papers
-
LLMCL-GEC: Advancing Grammatical Error Correction with LLM-Driven Curriculum Learning
An LLM-scored easy-to-hard curriculum for training grammatical error correction models yields small but consistent F0.5 gains over one-shot and length-based training.
-
LLM-Assisted Automated Deductive Coding of Dialogue Data: Leveraging Dialogue-Specific Characteristics to Enhance Contextual Understanding
An LLM-assisted pipeline that separates communicative acts from events, uses multi-model voting, and applies a consistency check reaches Cohen's kappa above 0.80 with human coders on a small student-dialogue corpus.
-
RBF++: Quantifying and Optimizing Reasoning Boundaries across Measurable and Unmeasurable Capabilities for Chain-of-Thought Reasoning
RBF++ models reasoning limits as harmonic-mean boundaries and uses them to explain, predict, and improve chain-of-thought performance across 38 models and 13 tasks.
Discussion (0). Continue with ORCID to comment.