Pith. sign in

REVIEW 2 cited by

PUB: A Pragmatics Understanding Benchmark for Assessing LLMs' Pragmatics Capabilities

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.07078 v1 pith:VJRZ7HON submitted 2024-01-13 cs.CL

classification cs.CL
keywords pragmaticscapabilitiesmodelstasksunderstandingbenchmarkperformanceconsisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

LLMs have demonstrated remarkable capability for understanding semantics, but they often struggle with understanding pragmatics. To demonstrate this fact, we release a Pragmatics Understanding Benchmark (PUB) dataset consisting of fourteen tasks in four pragmatics phenomena, namely, Implicature, Presupposition, Reference, and Deixis. We curated high-quality test sets for each task, consisting of Multiple Choice Question Answers (MCQA). PUB includes a total of 28k data points, 6.1k of which have been created by us, and the rest are adapted from existing datasets. We evaluated nine models varying in the number of parameters and type of training. Our study indicates that fine-tuning for instruction-following and chat significantly enhances the pragmatics capabilities of smaller language models. However, for larger models, the base versions perform comparably with their chat-adapted counterparts. Additionally, there is a noticeable performance gap between human capabilities and model capabilities. Furthermore, unlike the consistent performance of humans across various tasks, the models demonstrate variability in their proficiency, with performance levels fluctuating due to different hints and the complexities of tasks within the same dataset. Overall, the benchmark aims to provide a comprehensive evaluation of LLM's ability to handle real-world language tasks that require pragmatic reasoning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels

    cs.CL 2026-08 conditional novelty 6.0 of 10

    ChronoLens uses feature-aligned crosscoders to show that historical language change has comparable magnitude across linguistic levels within a language, but divergent timing and direction across five parliamentary languages.

  2. QUENCH: Measuring the gap between Indic and Non-Indic Contextual General Reasoning in LLMs

    cs.CL 2024-12 conditional novelty 6.0 of 10

    QUENCH introduces a generation-based quiz benchmark with masked entities and rationales, and documents a consistent Indic-versus-non-Indic performance gap across seven LLMs.

Pith tools