REVIEW 4 cited by
Creating a Fine Grained Entity Type Taxonomy Using LLMs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In this study, we investigate the potential of GPT-4 and its advanced iteration, GPT-4 Turbo, in autonomously developing a detailed entity type taxonomy. Our objective is to construct a comprehensive taxonomy, starting from a broad classification of entity types - including objects, time, locations, organizations, events, actions, and subjects - similar to existing manually curated taxonomies. This classification is then progressively refined through iterative prompting techniques, leveraging GPT-4's internal knowledge base. The result is an extensive taxonomy comprising over 5000 nuanced entity types, which demonstrates remarkable quality upon subjective evaluation. We employed a straightforward yet effective prompting strategy, enabling the taxonomy to be dynamically expanded. The practical applications of this detailed taxonomy are diverse and significant. It facilitates the creation of new, more intricate branches through pattern-based combinations and notably enhances information extraction tasks, such as relation extraction and event argument extraction. Our methodology not only introduces an innovative approach to taxonomy creation but also opens new avenues for applying such taxonomies in various computational linguistics and AI-related fields.
Forward citations
Cited by 4 Pith papers
-
How Well Do LLMs Generate Taxonomies in the SE Domain? A Multi-perspective Evaluation Framework
An empirical study shows a quality-efficiency trade-off in LLM-generated SE taxonomies: TnT-LLM approaches human quality but over-generates, while CLIMB is 15-40x faster yet weaker on latent-concept categories.
-
Automatic Generation of a Cryptography Misuse Taxonomy Using Large Language Models
An LLM-driven pipeline detected and classified cryptographic API misuse across 3,492 programs, producing a 279-category taxonomy with 36 new categories, and encoded 11 of them into detection rules that expand existing tools.
-
TaxoAdapt: Aligning LLM-Based Multidimensional Taxonomy Construction to Evolving Research Corpora
TaxoAdapt aligns LLM-generated taxonomies to a corpus by classifying papers along task, method, dataset, evaluation, and domain dimensions, then expanding the tree based on paper density.
-
An Agentic Hybrid Top-Down and Bottom-Up Approach to Knowledge Graph Generation
A hybrid LLM pipeline grounds recognized skills in Wikidata and uses agentic reflection to synthesize and organize emerging long-tail skills into a multilingual knowledge graph.
Discussion (0). Continue with ORCID to comment.