Pith. sign in

REVIEW 4 cited by

Creating a Fine Grained Entity Type Taxonomy Using LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.12557 v1 pith:55LOVYBE submitted 2024-02-19 cs.CL

classification cs.CL
keywords taxonomyentityextractiongpt-4classificationcreationdetailedprompting
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In this study, we investigate the potential of GPT-4 and its advanced iteration, GPT-4 Turbo, in autonomously developing a detailed entity type taxonomy. Our objective is to construct a comprehensive taxonomy, starting from a broad classification of entity types - including objects, time, locations, organizations, events, actions, and subjects - similar to existing manually curated taxonomies. This classification is then progressively refined through iterative prompting techniques, leveraging GPT-4's internal knowledge base. The result is an extensive taxonomy comprising over 5000 nuanced entity types, which demonstrates remarkable quality upon subjective evaluation. We employed a straightforward yet effective prompting strategy, enabling the taxonomy to be dynamically expanded. The practical applications of this detailed taxonomy are diverse and significant. It facilitates the creation of new, more intricate branches through pattern-based combinations and notably enhances information extraction tasks, such as relation extraction and event argument extraction. Our methodology not only introduces an innovative approach to taxonomy creation but also opens new avenues for applying such taxonomies in various computational linguistics and AI-related fields.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. How Well Do LLMs Generate Taxonomies in the SE Domain? A Multi-perspective Evaluation Framework

    cs.SE 2026-08 conditional novelty 6.0 of 10

    An empirical study shows a quality-efficiency trade-off in LLM-generated SE taxonomies: TnT-LLM approaches human quality but over-generates, while CLIMB is 15-40x faster yet weaker on latent-concept categories.

  2. Automatic Generation of a Cryptography Misuse Taxonomy Using Large Language Models

    cs.CR 2025-09 conditional novelty 6.0 of 10

    An LLM-driven pipeline detected and classified cryptographic API misuse across 3,492 programs, producing a 279-category taxonomy with 36 new categories, and encoded 11 of them into detection rules that expand existing tools.

  3. TaxoAdapt: Aligning LLM-Based Multidimensional Taxonomy Construction to Evolving Research Corpora

    cs.CL 2025-06 conditional novelty 6.0 of 10

    TaxoAdapt aligns LLM-generated taxonomies to a corpus by classifying papers along task, method, dataset, evaluation, and domain dimensions, then expanding the tree based on paper density.

  4. An Agentic Hybrid Top-Down and Bottom-Up Approach to Knowledge Graph Generation

    cs.CL 2026-08 conditional novelty 4.0 of 10

    A hybrid LLM pipeline grounds recognized skills in Wikidata and uses agentic reflection to synthesize and organize emerging long-tail skills into a multilingual knowledge graph.

Pith tools