REVIEW 3 cited by
Large Language Models Struggle to Describe the Haystack without Human Help: Human-in-the-loop Evaluation of Topic Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
A common use of NLP is to facilitate the understanding of large document collections, with a shift from using traditional topic models to Large Language Models. Yet the effectiveness of using LLM for large corpus understanding in real-world applications remains under-explored. This study measures the knowledge users acquire with unsupervised, supervised LLM-based exploratory approaches or traditional topic models on two datasets. While LLM-based methods generate more human-readable topics and show higher average win probabilities than traditional models for data exploration, they produce overly generic topics for domain-specific datasets that do not easily allow users to learn much about the documents. Adding human supervision to the LLM generation process improves data exploration by mitigating hallucination and over-genericity but requires greater human effort. In contrast, traditional. models like Latent Dirichlet Allocation (LDA) remain effective for exploration but are less user-friendly. We show that LLMs struggle to describe the haystack of large corpora without human help, particularly domain-specific data, and face scaling and hallucination limitations due to context length constraints.
Forward citations
Cited by 3 Pith papers
-
ProxAnn: Use-Oriented Evaluations of Topic Models and Document Clustering
An LLM-based proxy can reproduce human relevance judgments about topic model outputs closely enough to substitute for an average human annotator, and classical LDA remains competitive under this test.
-
PenTest2.0: Towards Autonomous Privilege Escalation Using GenAI
PenTest2.0 demonstrates that an LLM-driven agent can autonomously suggest and run privilege escalation commands on a purposely vulnerable Linux VM, reaching root in every tested configuration but achieving automatic r...
-
TopicImpact: Improving Customer Feedback Analysis with Opinion Units for Topic Modeling and Star-Rating Prediction
Clustering LLM-extracted opinion units instead of whole reviews yields coherent topics, and splitting those units by sentiment before clustering gives the best star-rating prediction, with an average R2 around 0.73.
Discussion (0). Continue with ORCID to comment.