Pith. sign in

REVIEW 2 cited by

Recognizing Limits: Investigating Infeasibility in Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.05873 v3 pith:J7SY3POF submitted 2024-08-11 cs.CL

classification cs.CL
keywords llmstasksinfeasiblecapabilitiesmodelslanguagelargeperformance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) have shown remarkable performance in various tasks but often fail to handle queries that exceed their knowledge and capabilities, leading to incorrect or fabricated responses. This paper addresses the need for LLMs to recognize and refuse infeasible tasks due to the requests surpassing their capabilities. We conceptualize four main categories of infeasible tasks for LLMs, which cover a broad spectrum of hallucination-related challenges identified in prior literature. We develop and benchmark a new dataset comprising diverse infeasible and feasible tasks to evaluate multiple LLMs' abilities to decline infeasible tasks. Furthermore, we explore the potential of increasing LLMs' refusal capabilities with fine-tuning. Our experiments validate the effectiveness of the trained models, suggesting promising directions for improving the performance of LLMs in real-world applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Anatomy to Smells: An Empirical Study of SKILL.md in Agent Skills

    cs.SE 2026-07 unverdicted novelty 7.0 of 10

    Over 99% of real-world SKILL.md files contain skill smells (violations of authoring best practices), and those smells almost never disappear as the skills evolve.

  2. A Simple and Effective Method for Uncertainty Quantification and OOD Detection

    cs.LG 2025-08 reject novelty 3.0 of 10

    A single-model OOD detection method that measures feature-space density with a Gaussian kernel (IPF) reports AUROC 93.18 on CIFAR-10 vs SVHN, a marginal gain over DDU's 92.90, with the kernel width selected to maximize AUROC.

Pith tools