Pith. sign in

REVIEW 2 cited by

Metadata-based Data Exploration with Retrieval-Augmented Generation for Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.04231 v1 pith:7XOZVMAL submitted 2024-10-05 cs.IR

classification cs.IR
keywords datadatasetsexplorationmodelstasksdiscoveryenhancegeneration
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Developing the capacity to effectively search for requisite datasets is an urgent requirement to assist data users in identifying relevant datasets considering the very limited available metadata. For this challenge, the utilization of third-party data is emerging as a valuable source for improvement. Our research introduces a new architecture for data exploration which employs a form of Retrieval-Augmented Generation (RAG) to enhance metadata-based data discovery. The system integrates large language models (LLMs) with external vector databases to identify semantic relationships among diverse types of datasets. The proposed framework offers a new method for evaluating semantic similarity among heterogeneous data sources and for improving data exploration. Our study includes experimental results on four critical tasks: 1) recommending similar datasets, 2) suggesting combinable datasets, 3) estimating tags, and 4) predicting variables. Our results demonstrate that RAG can enhance the selection of relevant datasets, particularly from different categories, when compared to conventional metadata approaches. However, performance varied across tasks and models, which confirms the significance of selecting appropriate techniques based on specific use cases. The findings suggest that this approach holds promise for addressing challenges in data exploration and discovery, although further refinement is necessary for estimation tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Making Sense of Data in the Wild: Data Analysis Automation at Scale

    cs.IR 2025-01 conditional novelty 5.0 of 10

    A multi-agent LLM system with retrieval-augmented generation automatically curates datasets from Zenodo and Hugging Face, yielding small retrieval gains and a confounded synthetic-data improvement.

  2. Multiple Abstraction Level Retrieve Augment Generation

    cs.CL 2025-01 conditional novelty 4.0 of 10

    MAL-RAG retrieves document, section, paragraph, and multi-sentence chunks together and claims a 25.7% improvement in AI-judged answer correctness on glycoscience questions over single-level RAG.

Pith tools