REVIEW 2 major objections 6 minor 179 references
This survey argues that open dataset search has outgrown keyword matching and is now defined by a two-way partnership with large language models.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
A structured review of open dataset search across tabular, spatial, JSON, graph, and vector data, plus the two-way relationship with LLMs.
T0 review reviewed 2026-08-05 challenge →
load-bearing objection A useful modality-aware survey of dataset search that mostly delivers, but the 'dataset search for LLM' half of its central pitch is asserted rather than shown. the 2 major comments →
A Survey on Open Dataset Search in the LLM Era: Retrospectives and Perspectives
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The paper's contribution is a structured synthesis: modern open dataset search is no longer a metadata-matching problem but a content-aware, modality-specific retrieval problem, and the most forward-looking direction is the two-way coupling between LLMs and dataset search. On one side, LLMs automate dataset construction, cleaning, and transformation, and enable natural-language queries, semantic planning, and relevance estimation. On the other side, dataset search supports LLMs by selecting datasets for retrieval-augmented generation and for data selection during pretraining, fine-tuning, and in-context learning. The survey develops this thesis across five data modalities, offering formal de
What carries the argument
The organizing instrument is a modality taxonomy: tabular, vector, spatial, JSON document, and graph datasets, each with a formal definition (for example, a vector dataset is a set of embedding vectors, a spatial dataset is a set of georeferenced points, a JSON document is a nested key-value tree, and a graph dataset is a vertex-edge pair). The paper maps each modality to its dominant similarity measure—MaxSim for vector sets, Earth mover's distance and Hausdorff distance for spatial data, tree edit distance for JSON, and graph edit distance or maximum common subgraph for graphs—along with indexing and acceleration techniques. Over this taxonomy it overlays the two-way LLM relationship, trea
Load-bearing premise
The survey assumes that the research landscape is accurately represented by its modality-based partition and by the particular selection of papers, so that the taxonomy and the open-problems list correspond to the field's true shape rather than to the authors' reading of it.
What would settle it
A systematic sweep of published dataset search work that finds a material body of research on a modality the survey does not index (e.g., time-series or audio datasets), or that shows the LLM-for-search systems the survey cites are not adopted or cited by downstream RAG or data-selection work, would weaken both the completeness of the modality taxonomy and the claim that the LLM-dataset search relationship is genuinely two-way.
If this is right
- Example-based and natural-language queries will increasingly replace keyword-only interfaces in dataset search systems.
- LLM-based schema inference, cleaning, and transformation should be treated as part of the dataset search pipeline, not as separate data-engineering tasks.
- Dataset search will become a core stage in LLM development, feeding retrieval-augmented generation and data-selection pipelines with queryable, filterable data.
- Standardized benchmarks are missing for spatial and vector dataset search, and quality control is not yet integrated into search; closing these gaps is a prerequisite for practical progress.
- Privacy-preserving similarity calculation, indexing over encrypted data, and federated search remain open problems that need new index and acceleration designs.
Where Pith is reading between the lines
- If the mutual-benefit thesis holds, dataset search could become a closed-loop controller for LLM training: a model's judged failures could seed queries for corrective datasets, and the retrieved data could then be evaluated by the model, forming a feedback loop; the paper mentions agentic integration but does not formalize this loop.
- The modality taxonomy suggests that cross-modal dataset search is the natural stress test: a shared embedding space for tables, JSON trees, graphs, and spatial points would unify the separate similarity measures, and a cross-modal join and union benchmark would directly test the value of such a unified representation.
- The paper's emphasis on content over metadata implies that search quality may be robust to missing or poor metadata; a direct experiment comparing ranking quality on datasets with full metadata versus deliberately stripped metadata would quantify how much content signals compensate.
- The recurring use of sketch and hash approximations across tabular, vector, and spatial search suggests a common abstraction—set containment at scale—that could be factored into a shared data-discovery index, potentially simplifying future system design.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This survey reviews recent research on open dataset search, moving beyond keyword- and metadata-based retrieval. It organizes the field by data modality (tabular, spatial, JSON, graph, and vector datasets), surveys query-by-example and natural-language search techniques, and discusses the interplay between LLMs and dataset search. The claimed central contribution is a two-way relationship: LLMs improve dataset search through query understanding, semantic planning, and interactive guidance, while dataset search supports LLMs through retrieval-augmented generation (RAG) and data selection. The survey also identifies open problems, including privacy-preserving search, task-oriented integration, cross-modal discovery, federated search, and benchmarks.
Significance. If the reference selection is representative and the reverse-direction claims are adequately evidenced, this survey would fill a useful gap by providing a modality-aware, LLM-centric overview of open dataset search. The paper's strengths are its broad taxonomy, formal definitions, reproduction of key equations (e.g., MaxSim, Hausdorff distance, JSON tree edit distance, GED/MCS), comparative tables that summarize representative methods, and an explicit roadmap of open challenges. It also makes the useful expository choice of distinguishing LLM-for-search from search-for-LLM. However, as a survey, its conclusions rest on the accuracy of its secondary summaries and on the completeness and neutrality of its reference selection, and neither can be fully verified from the manuscript. The most load-bearing weakness is the reverse-direction chapter, whose RAG examples do not actually demonstrate dataset-search-backed RAG.
major comments (2)
- [§5.3, Table 7] The claimed direction 'advances in dataset search can support LLMs by enabling more effective integration into RAG frameworks' is not supported by the evidence in this subsection. After noting that RAG's retrieval component 'naturally aligns' with dataset search, the text explicitly states that the subsection 'instead showcases recent applications of RAG across diverse domains.' Table 7 lists 15 RAG applications (e.g., Self-BioRAG, GraphQA, InstructRAG), but these retrieve scientific papers, knowledge graphs, or domain-specific text corpora—not open datasets as defined in Definition 1.1, and not through any dataset-search system surveyed in Sections 3–4. None of the cited examples shows a dataset-search engine (e.g., Auctus, DBF, Starmie, LOTUS) supplying the retrieval source of a RAG pipeline. The abstract and Section 1.3 stake a core contribution on this mutual-benefit claim, so the re
- [§5.3, Data Selection for LLM] The data-selection discussion similarly overstates the integration of dataset search into LLM data pipelines. The paragraph cites generic data-selection methods (LESS, Dolma, MateS) and then identifies only one method—Wang et al. [138]—that actually integrates query-driven dataset search into a selection pipeline. The claim that 'recent advances have begun to integrate dataset search into data selection pipelines' is supported by a single concrete example. Given that Section 6.2 itself mentions other systems (DeepResearchGym, STARK, GPT-Instructor) that embed dataset retrieval in end-to-end pipelines, the authors could strengthen this subsection substantially by moving those examples here or by adding additional published cases. Without that, the reverse-direction evidence remains one example, which is not enough to sustain the survey's broader 'mutually beneficial relationship' framing.
minor comments (6)
- [Table 1] The 'Vector' row lists '[99]' (BioVSS) as a data source, but [99] is a search method, not an open data repository. The underlying Microsoft Academic Graph [127] appears to be the actual source. Please correct the citation or rephrase the row.
- [Definition 1.2] Definition 1.2 covers keyword and exemplar queries but not natural-language queries, which become a major theme in Section 5.2. Consider extending the definition or adding a remark that NL queries are handled as a separate query type later in the survey.
- [§4.1] The vector-dataset-search section includes several systems originally designed for passage/document retrieval (COLBERT, PLAID, SLIM, COIL, CITADEL, XTR). Since the survey's title and abstract are about dataset search, a brief justification of why token-vector collections over documents are treated as vector datasets would help readers accept this categorization.
- [Figure 2] Figure 2 is dense and is referenced only once. Adding pointers from the pipeline stages to the corresponding sections (e.g., query mechanisms, similarity calculation, indexing) would improve navigability.
- [§4.2] The discussion of spatial dataset search states that the field lacks a standardized benchmark, but no reference is given for this claim. A citation or a brief explanation of which evaluation gaps exist would strengthen the point.
- [§5.3] In text immediately before Table 7, the phrase 'retrieval-augmented generation' and the surrounding discussion could more precisely distinguish between retrieving documents/corpora and retrieving datasets. As written, the section's scope drifts from dataset search to general RAG.
Circularity Check
No significant circularity: the survey's claims are descriptive taxonomies and literature syntheses, not derivations or fitted predictions, and its self-citations are not load-bearing.
full rationale
This is a survey paper, so the standard circularity failure modes (fitting a parameter and renaming it a prediction, defining X in terms of Y, importing a uniqueness theorem from the authors' prior work) do not apply. The paper's central contribution is a modality-aware organization of existing dataset-search techniques and a discussion of the LLM/dataset-search relationship. Definitions 1.1/1.2, 3.1, 4.1–4.4 are formal definitions, not derived results. Sections 3 and 4 survey external, peer-reviewed methods (e.g., LSH Ensemble, JOSIE, Starmie, COLBERT, DBF, JEDI, A*GED) and attribute each method to its original publication; no equation in the survey is shown to reduce to an input by construction. The authors do cite their own prior works (e.g., [97,99,154–156]), but these are included as surveyed contributions alongside many third-party works, and no load-bearing conclusion rests solely on a self-citation. The 'mutually beneficial relationship' claim in §5 is expository; the skeptical concern that §5.3's RAG examples retrieve documents/corpora rather than open datasets is an internal evidence gap about the support for one direction of that relationship, not circular reasoning, since the claim is not derived from itself. The paper even concedes in §5.3 that it 'instead showcases recent applications of RAG across diverse domains,' which is a scope limitation, not a circular step. Overall, the survey's organization and gap analysis are self-contained descriptive claims, so the circularity score is 0.
Axiom & Free-Parameter Ledger
axioms (3)
- domain assumption The modality-based taxonomy (tabular, spatial, JSON, graph, vector) is a valid way to organize the dataset search literature.
- domain assumption The cited descriptions of prior systems and methods are accurate and representative.
- standard math Standard mathematical definitions used in the survey (e.g., Hausdorff distance Eq. 4, EMD Eq. 5, GED Def. 4.5, MCS Def. 4.6) are correct and apply as stated.
Cite this review
Pith. "Pith review of A Survey on Open Dataset Search in the LLM Era: Retrospectives and Perspectives." pith.science (2026). https://pith.science/paper/K4AFYL3X
@misc{pith2026250900728,
author = {Pith},
title = {Pith review of: A Survey on Open Dataset Search in the LLM Era: Retrospectives and Perspectives},
year = {2026},
howpublished = {\url{https://pith.science/paper/K4AFYL3X}},
note = {Machine review of arXiv:2509.00728}
}
read the original abstract
High-quality datasets are typically required for accomplishing data-driven tasks, such as training medical diagnosis models, predicting real-time traffic conditions, or conducting experiments to validate research hypotheses. Consequently, open dataset search, which aims to ensure the efficient and accurate fulfillment of users' dataset requirements, has emerged as a critical research challenge and has attracted widespread interest. Recent studies have made notable progress in enhancing the flexibility and intelligence of open dataset search, and large language models (LLMs) have demonstrated strong potential in addressing long-standing challenges in this area. Therefore, a systematic and comprehensive review of the open dataset search problem is essential, detailing the current state of research and exploring future directions. In this survey, we focus on recent advances in open dataset search beyond traditional approaches that rely on metadata and keywords. From the perspective of dataset modalities, we place particular emphasis on example-based dataset search, advanced similarity measurement techniques based on dataset content, and efficient search acceleration techniques. In addition, we emphasize the mutually beneficial relationship between LLMs and open dataset search. On the one hand, LLMs help address complex challenges in query understanding, semantic modeling, and interactive guidance within open dataset search. In turn, advances in dataset search can support LLMs by enabling more effective integration into retrieval-augmented generation (RAG) frameworks and data selection processes, thereby enhancing downstream task performance. Finally, we summarize open research problems and outline promising directions for future work. This work aims to offer a structured reference for researchers and practitioners in the field of open dataset search.
Figures
Reference graph
Works this paper leans on
-
[1]
Abbas Acar, Hidayet Aksu, A Selcuk Uluagac, and Mauro Conti. 2018. A survey on homomorphic encryption schemes: Theory and implementation. Comput. Surveys 51, 4 (2018), 1–35
2018
-
[2]
Uchenna Akujuobi and Xiangliang Zhang. 2017. Delve: a dataset-driven scholarly search and analysis system.ACM SIGKDD Explorations Newsletter 19, 2 (2017), 36–46
2017
-
[3]
Alawwad, Areej Alhothali, Usman Naseem, Ali Alkhathlan, and Amani Jamal
Hessa A. Alawwad, Areej Alhothali, Usman Naseem, Ali Alkhathlan, and Amani Jamal. 2025. Enhancing textual textbook question answering with large language models and retrieval augmented generation. Pattern Recognition 162 (2025), 111332
2025
-
[4]
Alon Albalak, Yanai Elazar, Sang Michael Xie, Shayne Longpre, Nathan Lambert, Xinyi Wang, Niklas Muennighoff, Bairu Hou, Liangming Pan, Haewon Jeong, et al. 2024. A survey on data selection for language models. arXiv:2402.16827 (2024)
Pith/arXiv arXiv 2024
-
[5]
Eric Anderson, Jonathan Fritz, Austin Lee, Bohou Li, Mark Lindblad, Henry Lindeman, Alex Meyer, Parth Parmar, Tanvi Ranade, Mehul A Shah, et al. 2024. The Design of an LLM-powered Unstructured Analytics System. arXiv:2409.00847 (2024)
Pith/arXiv arXiv 2024
-
[6]
Simran Arora, Brandon Yang, Sabri Eyuboglu, Avanika Narayan, Andrew Hojel, Immanuel Trummer, and Christopher Ré. 2023. Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes. Proceedings of the VLDB Endowment 17, 2 (2023), 92–105
2023
-
[7]
Mohamed-Amine Baazizi, Dario Colazzo, Giorgio Ghelli, and Carlo Sartiani. 2019. Parametric schema inference for massive JSON datasets. The VLDB Journal 28, 4 (2019), 497–521
2019
-
[8]
Mohamed-Amine Baazizi, Dario Colazzo, Giorgio Ghelli, and Carlo Sartiani. 2019. Schemas and types for JSON data: from theory to practice. In Proceedings of the 2019 International Conference on Management of Data (SIGMOD) . 2060–2063
2019
-
[9]
Jiyang Bai and Peixiang Zhao. 2021. TaGSim: Type-aware Graph Similarity Learning and Computation. Proceedings of the VLDB Endowment 15, 2 (2021), 335–347
2021
-
[10]
Tianyi Bai, Ling Yang, Zhen Hao Wong, Fupeng Sun, Xinlin Zhuang, Jiahui Peng, Chi Zhang, Lijun Wu, Jiantao Qiu, Wentao Zhang, Binhang Yuan, and Conghui He. 2025. Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL) . 9465–9491
2025
-
[11]
Omar Benjelloun, Shiyu Chen, and Natasha Noy. 2020. Google Dataset Search by the Numbers. In International Semantic Web Conference (ISWS) . 667–682
2020
-
[12]
Chandra Sekhar Bhagavatula, Thanapon Noraset, and Doug Downey. 2015. Tabel: Entity linking in web tables. In International Semantic Web Conference (ISWS). 425–441
2015
-
[13]
Fabian Biester, Mohamed Abdelaal, and Daniel Del Gaudio. 2024. Llmclean: Context-aware tabular data cleaning via llm-generated ofds. In European Conference on Advances in Databases and Information Systems (ADBIS) , Vol. 2186. 68–78
2024
-
[14]
Tobias Bleifuß, Leon Bornemann, Dmitri V Kalashnikov, Felix Naumann, and Divesh Srivastava. 2021. The Secret Life of Wikipedia Tables. In Proceedings of the 2nd Workshop on Search, Exploration, and Analysis in Heterogeneous Datastores (SEA-Data 2021) co-located with 47th International Conference on Very Large Data Bases , Vol. 2929. 20–26
2021
-
[15]
Alex Bogatu, Alvaro A A Fernandes, Norman W Paton, and Nikolaos Konstantinou. 2020. Dataset Discovery in Data Lakes. In36th IEEE International Conference on Data Engineering (ICDE) . 709–720
2020
-
[16]
Pierre Bourhis, Juan L Reutter, Fernando Suárez, and Domagoj Vrgoč. 2017. JSON: data model, query languages and schema specification. In Proceedings of the 36th ACM SIGMOD-SIGACT-SIGAI Symposium on Principles of Database Systems (PODS) . 123–135
2017
-
[17]
Pierre Bourhis, Juan L Reutter, and Domagoj Vrgoč. 2020. JSON: Data model and query languages. Information Systems 89 (2020), 101478
2020
-
[18]
Dan Brickley, Matthew Burgess, and Natasha Noy. 2019. Google Dataset Search: Building a search engine for datasets in an open Web ecosystem. In The World Wide Web Conference (WWW) . 1365–1375
2019
-
[19]
Horst Bunke and Kim Shearer. 1998. A graph distance metric based on the maximal common subgraph. Pattern recognition letters 19, 3-4 (1998), 255–259
1998
-
[20]
Sonia Castelo, Rémi Rampin, Aécio Santos, Aline Bessa, Fernando Chirigati, and Juliana Freire. 2021. Auctus: a dataset search engine for data discovery and augmentation. Proceedings of the VLDB Endowment 14, 12 (2021), 2791–2794
2021
-
[21]
Lijun Chang, Xing Feng, Xuemin Lin, Lu Qin, Wenjie Zhang, and Dian Ouyang. 2020. Speeding up GED verification for graph similarity search. In 36th IEEE International Conference on Data Engineering (ICDE) . 793–804
2020
-
[22]
Lijun Chang, Xing Feng, Kai Yao, Lu Qin, and Wenjie Zhang. 2022. Accelerating graph similarity search via efficient GED computation. IEEE Transactions on Knowledge and Data Engineering 35, 5 (2022), 4485–4498
2022
-
[23]
Ming-Fang Chang, John Lambert, Patsorn Sangkloy, Jagjeet Singh, Slawomir Bak, Andrew Hartnett, De Wang, Peter Carr, Simon Lucey, Deva Ramanan, et al. 2019. Argoverse: 3d tracking and forecasting with rich maps. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 8748–8757
2019
-
[24]
Adriane Chapman, Elena Simperl, Laura Koesten, George Konstantinidis, Luis-Daniel Ibáñez, Emilia Kacprzak, and Paul Groth. 2020. Dataset search: a survey. The VLDB Journal 29, 1 (2020), 251–272
2020
-
[25]
Minxiao Chen, Haitao Yuan, Nan Jiang, Zhifeng Bao, and Shangguang Wang. 2024. Urban Traffic Accident Risk Prediction Revisited: Regionality, Proximity, Similarity and Sparsity. InProceedings of the 33rd ACM International Conference on Information and Knowledge Management (CIKM) . 281–290
2024
-
[26]
Minxiao Chen, Haitao Yuan, Nan Jiang, Zhihan Zheng, Zhifeng Bao, Ao Zhou, Jiaxin Jiang, and Shangguang Wang. 2025. S-MGHSTN: Towards An Effective Streaming Traffic Accident Risk Prediction Framework. IEEE Transactions on Knowledge and Data Engineering 37, 7 (2025), 4285 – 4298
2025
-
[27]
Qiaosheng Chen, Jiageng Chen, Xiao Zhou, and Gong Cheng. 2024. Enhancing Dataset Search with Compact Data Snippets. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR) . 1093–1103
2024
-
[28]
Qiaosheng Chen, Weiqing Luo, Zixian Huang, Tengteng Lin, Xiaxia Wang, Ahmet Soylu, Basil Ell, Baifan Zhou, Evgeny Kharlamov, and Gong Cheng. 2024. ACORDAR 2.0: A Test Collection for Ad Hoc Dataset Retrieval with Densely Pooled Datasets and Question-Style Queries. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in ...
2024
-
[29]
Wei-Hao Chen, Weixi Tong, Amanda Case, and Tianyi Zhang. 2025. Dango: A Mixed-Initiative Data Wrangling System using Large Language Model. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI) . 1–28
2025
-
[30]
Zui Chen, Zihui Gu, Lei Cao, Ju Fan, Samuel Madden, and Nan Tang. 2023. Symphony: Towards Natural Language Query Answering over Multi-modal Data Lakes. In 13th Conference on Innovative Data Systems Research (CIDR) . 1–7
2023
-
[31]
Zhiyu Chen, Mohamed Trabelsi, Jeff Heflin, Yinan Xu, and Brian D. Davison. 2020. Table Search Using a Deep Contextualized Language Model. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval (SIGIR) , Jimmy X. Huang, Yi Chang, Xueqi Cheng, Jaap Kamps, Vanessa Murdock, Ji-Rong Wen, and Yiqun Liu...
2020
-
[32]
Mingyue Cheng, Yucong Luo, Jie Ouyang, Qi Liu, Huijie Liu, Li Li, Shuo Yu, Bohou Zhang, Jiawei Cao, Jie Ma, et al . 2025. A survey on knowledge-oriented retrieval-augmented generation. arXiv:2503.10677 (2025)
Pith/arXiv arXiv 2025
-
[33]
João Coelho, Jingjie Ning, Jingyuan He, Kangrui Mao, Abhijay Paladugu, Pranav Setlur, Jiahe Jin, Jamie Callan, João Magalhães, Bruno Martins, et al. 2025. Deepresearchgym: A free, transparent, and reproducible evaluation sandbox for deep research. arXiv:2505.19253 (2025)
arXiv 2025
-
[34]
Tianji Cong, Fatemeh Nargesian, and HV Jagadish. 2023. Pylon: Semantic table union search in data lakes. arXiv:2301.04901 (2023)
Pith/arXiv arXiv 2023
-
[35]
Arash Dargahi Nobari and Davood Rafiei. 2024. Dtt: An example-driven tabular transformer for joinability by leveraging large language models. Proceedings of the ACM on Management of Data 2, 1 (2024), 1–24
2024
-
[36]
Anish Das Sarma, Lujun Fang, Nitin Gupta, Alon Halevy, Hongrae Lee, Fei Wu, Reynold Xin, and Cong Yu. 2012. Finding related tables. In Proceedings of the 2012 International Conference on Management of Data (SIGMOD) . 817–828
2012
-
[37]
Auriol Degbelo and Brhane Bahrishum Teka. 2019. Spatial search strategies for open government data: A systematic comparison. In Proceedings of the 13th Workshop on Geographic Information Retrieval (GIR) . 1–10
2019
-
[38]
Yuhao Deng, Chengliang Chai, Lei Cao, Qin Yuan, Siyuan Chen, Yanrui Yu, Zhaoze Sun, Junyi Wang, Jiajun Li, Ziqi Cao, et al. 2024. LakeBench: A Benchmark for Discovering Joinable and Unionable Tables in Data Lakes. Proceedings of the VLDB Endowment 17, 8 (2024), 1925–1938
2024
-
[39]
Yuyang Dong, Kunihiro Takeoka, Chuan Xiao, and Masafumi Oyamada. 2021. Efficient Joinable Table Discovery in Data Lakes: A High-Dimensional Similarity-Based Approach. In 37th IEEE International Conference on Data Engineering (ICDE) . 456–467
2021
-
[40]
Yuyang Dong, Chuan Xiao, Takuma Nozawa, Masafumi Enomoto, and Masafumi Oyamada. 2023. DeepJoin: Joinable Table Discovery with Pre-Trained Language Models. Proceedings of the VLDB Endowment 16, 10 (2023), 2458–2470
2023
-
[41]
Dominik Durner, Viktor Leis, and Thomas Neumann. 2021. JSON tiles: Fast analytics on semi-structured data. InProceedings of the 2021 International Conference on Management of Data (SIGMOD) . 445–458
2021
-
[42]
Mohamed Y Eltabakh, Mayuresh Kunjir, Ahmed K Elmagarmid, and Mohammad Shahmeer Ahmad. 2023. Cross Modal Data Discovery over Structured and Unstructured Data Lakes. Proceedings of the VLDB Endowment 16, 11 (2023), 3377–3390
2023
-
[43]
Joshua Engels, Benjamin Coleman, Vihan Lakshman, and Anshumali Shrivastava. 2023. DESSERT: an efficient algorithm for vector set search with vector set queries. Advances in Neural Information Processing Systems 36 (2023), 67972–67992
2023
-
[44]
Mahdi Esmailoghli, Christoph Schnell, Renée J Miller, and Ziawasch Abedjan. 2025. BLEND: A Unified Data Discovery System. In 41st IEEE International Conference on Data Engineering (ICDE) . 737–750
2025
-
[45]
Grace Fan, Jin Wang, Yuliang Li, and Renée J Miller. 2023. Table discovery in data lakes: State-of-the-art and future directions. In Companion of the 2023 International Conference on Management of Data (SIGMOD/PODS) . 69–75
2023
-
[46]
Grace Fan, Jin Wang, Yuliang Li, Dan Zhang, and Renée J Miller. 2023. Semantics-Aware Dataset Discovery from Data Lakes with Contextualized Column-Based Representation Learning. Proceedings of the VLDB Endowment 16, 7 (2023), 1726–1739
2023
-
[47]
Ju Fan, Zihui Gu, Songyue Zhang, Yuxin Zhang, Zui Chen, Lei Cao, Guoliang Li, Samuel Madden, Xiaoyong Du, and Nan Tang. 2024. Combining small language models and large language models for zero-shot nl2sql. Proceedings of the VLDB Endowment 17, 11 (2024), 2750–2763
2024
-
[48]
Lizhou Fan, Sara Lafia, Lingyao Li, Fangyuan Yang, and Libby Hemphill. 2023. DataChat: Prototyping a conversational agent for dataset search and visualization. Proceedings of the Association for Information Science and Technology 60, 1 (2023), 586–591
2023
-
[49]
Shangbin Feng, Chan Young Park, Yuhan Liu, and Yulia Tsvetkov. 2023. From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL). 11737–11762
2023
-
[50]
Pedro Arthur de Fernandes Vasconcelos, Wensttay de Sousa Alencar, Victor Hugo da Silva Ribeiro, Natarajan Ferreira Rodrigues, and Fabio de Gomes Andrade. 2017. Enabling Spatial Queries in Open Government Data Portals. In International Conference on Electronic Government and the Information Systems Perspective (EGOVIS) . 64–79
2017
-
[51]
Raul Castro Fernandez, Ziawasch Abedjan, Famien Koko, Gina Yuan, Samuel Madden, and Michael Stonebraker. 2018. Aurum: A data discovery system. In 34th IEEE International Conference on Data Engineering (ICDE) . 1001–1012
2018
-
[52]
Raul Castro Fernandez, Jisoo Min, Demitri Nava, and Samuel Madden. 2019. Lazo: A cardinality-based method for coupled estimation of jaccard similarity and containment. In 35th IEEE International Conference on Data Engineering (ICDE) . 1190–1201
2019
-
[53]
Benjamin Feuer, Yurong Liu, Chinmay Hegde, and Juliana Freire. 2024. ArcheType: A Novel Framework for Open-Source Column Type Annotation Using Large Language Models. Proceedings of the VLDB Endowment 17, 9 (2024)
2024
-
[54]
Juliana Freire, Grace Fan, Benjamin Feuer, Christos Koutras, Yurong Liu, Eduardo Peña, Aécio S. R. Santos, Cláudio T. Silva, and Eden Wu. 2025. Large Language Models for Data Discovery and Integration: Challenges and Opportunities. IEEE Data Engineering Bulletin 49, 1 (2025), 3–31
2025
-
[55]
Luyu Gao, Zhuyun Dai, and Jamie Callan. 2021. COIL: Revisit Exact Lexical Match in Information Retrieval with Contextualized Inverted List. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL) . 3030–3042
2021
-
[56]
Adamu Garba and Shangli Wu. 2025. Snippet-based result merging in federated search. Journal of Information Science 51, 3 (2025), 623–637
2025
-
[57]
Adamu Garba, Shengli Wu, and Shah Khalid. 2023. Federated search techniques: an overview of the trends and state of the art. Knowledge and Information Systems 65, 12 (2023), 5065–5095
2023
-
[58]
Saheli Ghosh, Ahmed Eldawy, and Shipra Jais. 2019. Aid: An adaptive image data index for interactive multilevel visualization. In 35th IEEE International Conference on Data Engineering (ICDE) . 1594–1597
2019
-
[59]
Saheli Ghosh, Tin Vu, Mehrad Amin Eskandari, and Ahmed Eldawy. 2019. UCR-STAR: The UCR spatio-temporal active repository. SIGSPATIAL Special 11, 2 (2019), 34–40
2019
-
[60]
Victor Giannakouris and Immanuel Trummer. 2025. SwellDB: Dynamic Query-Driven Table Generation with Large Language Models. InCompanion of the 2025 International Conference on Management of Data . 95–98
2025
-
[61]
Youdi Gong, Guangzhen Liu, Yunzhi Xue, Rui Li, and Lingzhong Meng. 2023. A survey on dataset quality in machine learning. Information and Software Technology 162 (2023), 107268
2023
-
[62]
Tobias Grubenmann, Abraham Bernstein, Dmitry Moor, and Sven Seuken. 2018. Financing the web of data with delayed-answer auctions. In The World Wide Web Conference (WWW). 1033–1042
2018
-
[63]
Ruiqi Guo, Philip Sun, Erik Lindgren, Quan Geng, David Simcha, Felix Chern, and Sanjiv Kumar. 2020. Accelerating large-scale inference with anisotropic vector quantization. In International Conference on Machine Learning (ICML) . 3887–3896
2020
-
[64]
Stefan Hahmann and Dirk Burghardt. 2013. How much information is geospatially referenced? Networks and cognition. International Journal of Geographical Information Science 27, 6 (2013), 1171–1189
2013
-
[65]
Rihan Hai, Christos Koutras, Christoph Quix, and Matthias Jarke. 2023. Data lakes: A survey of functions and systems. IEEE Transactions on Knowledge and Data Engineering 35, 12 (2023), 12571–12590
2023
-
[66]
Zifei FeiFei Han, Jionghao Lin, Ashish Gurung, Danielle R Thomas, Eason Chen, Conrad Borchers, Shivang Gupta, and Kenneth R Koedinger. 2024. Improving Assessment of Tutoring Practices using Retrieval-Augmented Generation. Proceedings of Machine Learning Research 257 (2024), 66–76
2024
-
[67]
Xiaoxin He, Yijun Tian, Yifei Sun, Nitesh Chawla, Thomas Laurent, Yann LeCun, Xavier Bresson, and Bryan Hooi. 2024. G-retriever: Retrieval- augmented generation for textual graph understanding and question answering. Advances in Neural Information Processing Systems 37 (2024), 132876–132907
2024
-
[68]
Thomas Hervey, Sara Lafia, and Werner Kuhn. 2021. Search Facets and Ranking in Geospatial Dataset Search. In 11th International Conference on Geographic Information Science (GIScience) , Vol. 177. 5:1–5:15
2021
-
[69]
Sirui Hong, Yizhang Lin, Bang Liu, Bangbang Liu, Binhao Wu, Ceyao Zhang, Danyang Li, Jiaqi Chen, Jiayi Zhang, Jinlin Wang, Li Zhang, Lingyao Zhang, Min Yang, Mingchen Zhuge, Taicheng Guo, Tuo Zhou, Wei Tao, Robert Tang, Xiangtao Lu, Xiawu Zheng, Xinbing Liang, Yaying Fei, Yuheng Cheng, Yongxin Ni, Zhibin Gou, Zongze Xu, Yuyu Luo, and Chenglin Wu. 2025. Da...
2025
-
[70]
Xuming Hu, Shen Wang, Xiao Qin, Chuan Lei, Zhengyuan Shen, Christos Faloutsos, Asterios Katsifodimos, George Karypis, Lijie Wen, and Philip S Yu. 2023. Automatic table union search with tabular representation learning. In Findings of the Association for Computational Linguistics . 3786–3800
2023
-
[71]
Yuntong Hu, Zhihan Lei, Zhongjie Dai, Allen Zhang, Abhinav Angirekula, Zheng Zhang, and Liang Zhao. 2025. Cg-rag: Research question answering by citation graph retrieval-augmented llms. In Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR) . 678–687
2025
-
[72]
Benhao Huang, Yingzhuo Yu, Jin Huang, Xingjian Zhang, and Jiaqi Ma. 2024. DCA-Bench: A Benchmark for Dataset Curation Agents. arXiv:2406.07275 (2024)
Pith/arXiv arXiv 2024
-
[73]
Thomas Hütter, Nikolaus Augsten, Christoph M Kirsch, Michael J Carey, and Chen Li. 2022. JEDI: These aren’t the JSON documents you’re looking for.... In Proceedings of the 2022 International Conference on Management of Data (SIGMOD) . 1584–1597
2022
-
[74]
Thomas Hütter, Mateusz Pawlik, Robert Löschinger, and Nikolaus Augsten. 2019. Effective filters and linear time verification for tree similarity joins. In 35th IEEE International Conference on Data Engineering (ICDE) . 854–865
2019
-
[75]
Andra Ionescu, Kiril Vasilev, Florena Buse, Rihan Hai, and Asterios Katsifodimos. 2024. AutoFeat: Transitive Feature Discovery over Join Paths. In 40th IEEE International Conference on Data Engineering (ICDE) . 1861–1873
2024
-
[76]
Minbyul Jeong, Jiwoong Sohn, Mujeen Sung, and Jaewoo Kang. 2024. Improving medical reasoning through retrieval and self-reflection with retrieval-augmented large language models. Bioinformatics 40, Supplement_1 (2024), i119–i129
2024
-
[77]
Shane Culpepper
Daomin Ji, Hui Luo, Zhifeng Bao, and J. Shane Culpepper. 2024. Navigating Data Repositories: Utilizing Line Charts to Discover Relevant Datasets. Proceedings of the VLDB Endowment 17, 12 (2024), 4289–4292
2024
-
[78]
Daomin Ji, Hui Luo, Zhifeng Bao, and J Shane Culpepper. 2025. Table integration in data lakes unleashed: pairwise integrability judgment, integrable set discovery, and multi-tuple conflict resolution. The VLDB Journal 34, 3 (2025), 1–24
2025
-
[79]
Lin Jiang, Junqiao Qiu, and Zhijia Zhao. 2020. Scalable structural index construction for JSON analytics. Proceedings of the VLDB Endowment 14, 4 (2020), 694–707
2020
-
[80]
Xinhui Kang, Wenjie You, and Ying Luo. 2025. Bio-inspired product design system integrating retrieval-augmented question answering and semantic fusion diffusion model. Advanced Engineering Informatics 67 (2025), 103537
2025
This paper was first reviewed by deepseek-v4-flash on August 5, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.