Pith. sign in

REVIEW 3 cited by

Large Language Models Struggle to Describe the Haystack without Human Help: Human-in-the-loop Evaluation of Topic Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.14748 v2 pith:D42YRUKW submitted 2025-02-20 cs.CL

classification cs.CL
keywords modelslargehumantraditionaldataexplorationtopicdatasets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A common use of NLP is to facilitate the understanding of large document collections, with a shift from using traditional topic models to Large Language Models. Yet the effectiveness of using LLM for large corpus understanding in real-world applications remains under-explored. This study measures the knowledge users acquire with unsupervised, supervised LLM-based exploratory approaches or traditional topic models on two datasets. While LLM-based methods generate more human-readable topics and show higher average win probabilities than traditional models for data exploration, they produce overly generic topics for domain-specific datasets that do not easily allow users to learn much about the documents. Adding human supervision to the LLM generation process improves data exploration by mitigating hallucination and over-genericity but requires greater human effort. In contrast, traditional. models like Latent Dirichlet Allocation (LDA) remain effective for exploration but are less user-friendly. We show that LLMs struggle to describe the haystack of large corpora without human help, particularly domain-specific data, and face scaling and hallucination limitations due to context length constraints.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ProxAnn: Use-Oriented Evaluations of Topic Models and Document Clustering

    cs.CL 2025-07 conditional novelty 7.0 of 10

    An LLM-based proxy can reproduce human relevance judgments about topic model outputs closely enough to substitute for an average human annotator, and classical LDA remains competitive under this test.

  2. PenTest2.0: Towards Autonomous Privilege Escalation Using GenAI

    cs.CR 2025-07 conditional novelty 5.0 of 10

    PenTest2.0 demonstrates that an LLM-driven agent can autonomously suggest and run privilege escalation commands on a purposely vulnerable Linux VM, reaching root in every tested configuration but achieving automatic r...

  3. TopicImpact: Improving Customer Feedback Analysis with Opinion Units for Topic Modeling and Star-Rating Prediction

    cs.CL 2025-07 conditional novelty 4.0 of 10

    Clustering LLM-extracted opinion units instead of whole reviews yields coherent topics, and splitting those units by sentiment before clustering gives the best star-rating prediction, with an average R2 around 0.73.

Pith tools