Pith. sign in

REVIEW 7 cited by

Large Language Models for Generative Information Extraction: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.17617 v3 pith:WTJFCLMZ submitted 2023-12-29 cs.CL

classification cs.CL
keywords generativelanguagellmstasksworksexplorationextractiongithub
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Information extraction (IE) aims to extract structural knowledge from plain natural language texts. Recently, generative Large Language Models (LLMs) have demonstrated remarkable capabilities in text understanding and generation. As a result, numerous works have been proposed to integrate LLMs for IE tasks based on a generative paradigm. To conduct a comprehensive systematic review and exploration of LLM efforts for IE tasks, in this study, we survey the most recent advancements in this field. We first present an extensive overview by categorizing these works in terms of various IE subtasks and techniques, and then we empirically analyze the most advanced methods and discover the emerging trend of IE tasks with LLMs. Based on a thorough review conducted, we identify several insights in technique and promising research directions that deserve further exploration in future studies. We maintain a public repository and consistently update related works and resources on GitHub (\href{https://github.com/quqxui/Awesome-LLM4IE-Papers}{LLM4IE repository})

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 24 citations worldwide. Full citation record

  1. LA-RL: Label-Aware Self-Reflection for Reinforcement Learning in Information Extraction

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Label-aware diagnostic reflection plus two-stage outcome GRPO improves same-backbone IE F1 over SFT, with larger gains under relation-extraction domain shift.

  2. Enhancing Automatic Term Extraction with Large Language Models via Syntactic Retrieval

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Syntactic similarity retrieval of demonstrations improves LLM-based automatic term extraction in cross-domain settings, but gains are modest and in-domain lexical retrieval is often competitive or better.

  3. MPL: Multiple Programming Languages with Large Language Models for Information Extraction

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Using multiple programming languages as code-style prompts during fine-tuning improves LLM information extraction accuracy over single-language prompting.

  4. Visual Text Mining with Progressive Taxonomy Construction for Environmental Studies

    cs.HC 2025-02 conditional novelty 6.0 of 10

    An interactive text-mining system combines a three-step LLM prompting pipeline with a consistency-based uncertainty chart for progressive DPSIR taxonomy refinement.

  5. Media Bias Detector: Designing and Implementing a Tool for Real-Time Selection and Framing Bias Analysis in News Coverage

    cs.HC 2025-02 conditional novelty 5.0 of 10

    An LLM-driven, near-real-time dashboard labels US news articles by topic, lean, and tone, and user studies suggest it helps experts and consumers explore selection and framing bias.

  6. (Towards) Scalable Reliable Automated Evaluation with Large Language Models

    cs.CL 2026-07 conditional novelty 4.0 of 10

    Multi-LLM pairwise Elo ranking with adjustable consensus thresholds produces rankings of competency profiles that average Spearman ρ≈0.83 with expert judgments.

  7. Effective and Efficient Schema-aware Information Extraction Using On-Device Large Language Models

    cs.CL 2025-05 conditional novelty 4.0 of 10

    DLISC, a dual-LoRA two-stage schema-aware extraction method with incremental schema caching, reports better F1 and lower latency than three RAG baselines on two IE datasets, though the comparison lacks error bars and code.

Pith tools