Pith. sign in

REVIEW 1 cited by

Fast Statistical Parsing of Noun Phrases for Document Indexing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv cmp-lg/9702009 v1 pith:YBRNLPIS submitted 1997-02-12 cmp-lg cs.CL

classification cmp-lgcs.CL
keywords documentindexingparsingphrasestechniquesapplicationbeencollection
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Information Retrieval (IR) is an important application area of Natural Language Processing (NLP) where one encounters the genuine challenge of processing large quantities of unrestricted natural language text. While much effort has been made to apply NLP techniques to IR, very few NLP techniques have been evaluated on a document collection larger than several megabytes. Many NLP techniques are simply not efficient enough, and not robust enough, to handle a large amount of text. This paper proposes a new probabilistic model for noun phrase parsing, and reports on the application of such a parsing technique to enhance document indexing. The effectiveness of using syntactic phrases provided by the parser to supplement single words for indexing is evaluated with a 250 megabytes document collection. The experiment's results show that supplementing single words with syntactic phrases for indexing consistently and significantly improves retrieval performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ERU-KG: Efficient Reference-aligned Unsupervised Keyphrase Generation

    cs.CL 2025-05 conditional novelty 6.0 of 10

    ERU-KG uses reference-trained SPLADE term importances plus neighbor-document noun phrases to generate present and absent keyphrases without keyphrase labels, and reports strong benchmark and retrieval results.

Pith tools