REVIEW 4 major objections 4 minor 63 references
FathomGPT: A Natural Language Interface for Interactively Exploring Ocean Science Data
T0 review · 4 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read FathomGPT gives ocean scientists a natural language interface that retrieves images, taxonomy, and measurements from the FathomNet database, answering 72.08% of 781 real workshop prompts successfully with a median latency of 6.92 seconds.
desk verdict A solid, honest systems paper for an ocean-science database interface; the headline accuracy number is self-evaluated and should be read as provisional, but the system and ablations are real. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the prompt evaluator: a large language model with function calling that decides which of five processing functions handles a prompt and rewrites the prompt to include only the relevant prior conversation, so downstream models see a self-contained question. Supporting it are species knowledge graphs, structured records extracted from encyclopedia text that map each scientific name to aliases, body parts, colors, predators, diet, and habitat; name resolution aligns a prompt-derived subject-relation-object triple against these graphs. Text-to-SQL uses three fine-tuned models specialized for similarity-search, visualization, and image/text/table outputs, plus few-shot schema examples and an error-feedback loop that sends failed SQL to a stronger model for repair. Visualization and image search complete the pipeline: graphing code is generated for charts, and pattern queries use an interactive segmentation model, color-based pattern extraction, and a feature-vector model to rank similar images.
What would settle it
Run FathomGPT on a fresh, unfiltered set of prompts collected from new ocean-science users who have not seen the interface examples, and check whether the success rate stays near 72.08%; a large drop would show that the filtered workshop prompts were not representative. A second check is to count correct name resolutions for prompts that mention predator-prey relations, where the paper reports many empty results.
Extended reading notes
Core claim
The central claim is that a natural language interface can make a large, complex scientific image database practically accessible without sacrificing speed or accuracy. FathomGPT routes every user prompt through a prompt evaluator that decides among five capabilities—name resolution, SQL query generation, taxonomic lookup, visualization code generation, and general information—and rewrites the prompt to fold in relevant conversation context. Common names and morphological descriptions are resolved to scientific names by aligning a prompt-derived knowledge graph against species knowledge graphs generated from encyclopedia text. Query generation uses three fine-tuned text-to-SQL models specialized by output type, and visualization requests generate graphing code in parallel with the database query. On 781 prompts logged from a workshop with domain users, the system produced a successful response 72.08% of the time with a median latency of 6.92 seconds, and ablation studies attribute part of this gain to fine-tuning and part to context-aware prompt modification.
Load-bearing premise
The 781 prompts that remain after filtering the workshop logs are representative of how real users will actually use FathomGPT, so the measured 72.08% success rate and the ablation gains carry over to everyday use.
Editorial extensions
If this is right
- Marine scientists can ask for images, taxonomic trees, measurements, and charts in plain English, removing the SQL and schema expertise barrier that keeps many researchers away from databases like FathomNet.
- Fine-tuning separate text-to-SQL models for different output types yields substantially more correct queries than a single few-shot model, so systems with heterogeneous outputs should expect to specialize their query generators.
- Rewriting each follow-up prompt to embed relevant conversation context improves multi-turn accuracy by 7.27% and cuts token usage by roughly 15.67%, a concrete design choice for conversational database interfaces.
- Knowledge-graph name resolution resolves common names and descriptive queries to scientific names with about 92% of returned names correct, compared with 56% for a direct general-purpose LLM, while avoiding hallucinations of species not present in the database.
- Pattern-based image search lets users highlight a region of an uploaded image and retrieve similar database images, extending access from text queries to visual queries.
Reading between the lines
- The reported success rate is measured on a filtered subset of 781 prompts, so on raw workshop logs that include duplicate, invalid, and adversarial queries the per-prompt success rate would be lower; a follow-up evaluation on unfiltered logs would give a more conservative estimate.
- Because the knowledge graphs are extracted from encyclopedia text, species with sparse coverage will produce empty or partial name-resolution results, especially for predator-prey relations; expanding the source corpus is a direct testable extension.
- The three-way split of fine-tuned text-to-SQL models by output type is a design choice other database interfaces could adopt, but the paper does not isolate whether the benefit comes from specialization or simply from more fine-tuning data per model.
- The same prompt-evaluator architecture could be ported to other relational scientific databases with a schema and a curated knowledge graph, since the pipeline itself is not ocean-specific; the paper notes this possibility but does not demonstrate it.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. FathomGPT is an open-source natural language interface for querying and visualizing the FathomNet ocean image database. The paper describes a multi-stage LLM pipeline that includes a prompt evaluator, a name-resolution component based on knowledge graphs built from Wikipedia, fine-tuned text-to-SQL models, Plotly-based visualization generation, and a pattern-extraction image search. The system was deployed at a FathomNet workshop, where 2,977 prompts were logged; after filtering, 781 prompts were used to evaluate overall accuracy, reporting a 72.08% success rate and a median latency of 6.92 seconds. Two ablation studies claim improvements from fine-tuning and from a context-modification prompting strategy, and a third evaluation compares the knowledge-graph name-resolution method against GPT-4o and vector embeddings.
Significance. If the reported performance is stable, FathomGPT would be a useful contribution to scientific data access: it is a deployed, open-source system with concrete architectural components (prompt evaluator, species knowledge graphs, specialized text-to-SQL models, pattern-based image retrieval) that target real user needs identified with marine scientists. The paper provides genuine empirical material: logged workshop prompts, latency measurements, ablations, and a head-to-head name-resolution comparison showing the knowledge-graph method outperforms GPT-4o on the tested set. However, the central quantitative claims rest on evaluation choices that are not yet fully transparent or externally validated, so the magnitude of the reported gains should be treated with caution until the methodology is strengthened.
major comments (4)
- [§5.1.2] The headline accuracy of 72.08% (563/781) is measured on a subset of 2,977 workshop prompts after an author-defined filtering step whose criteria are deferred to Supplementary Material, and the responses were labeled by the authors without any reported inter-rater reliability, independent labeler, or confidence interval. Since this figure is the paper's central evidence of system effectiveness, the reader cannot currently distinguish system performance from filtering and labeling judgment. Please provide the full filtering protocol, report agreement statistics if multiple labelers are used, and give a confidence interval for the success rate.
- [§5.2] The fine-tuning ablation reports that the fine-tuned model generated a correct response "14.54% more often" with counts 252 vs. 220. With the total of 335 introduced later in the same section, the absolute success rates are 75.2% and 65.7%, a difference of 9.5 percentage points; 14.54% is the relative improvement over the ablated baseline (32/220). The text does not state whether the reported percentage is absolute or relative, and the 335-prompt subset is not clearly defined as the ablation evaluation set. Please report absolute success rates and error counts for each condition with a consistent denominator, and add a significance test if feasible.
- [§5.3] The context ablation reports "7.27% more queries" with 129 vs. 153 errors out of 483 prompts. The absolute difference is 24/483 = 4.97 percentage points, while 7.27% is the relative reduction over the ablated version's 153 errors. As in §5.2, the reported metric is ambiguous. In addition, the 483-prompt set was obtained by manually filtering 126 conversations with exclusion criteria ("conversations that are not affected by context") that are not operationalized, making the result non-reproducible from the text. Please clarify the filtering rules and report absolute performance in both arms.
- [§5.4] The name-resolution evaluation treats empty results as correct when the authors judge that no matching species exists in FathomNet (e.g., "predators of hexanchus griseus"), but no ground-truth negative set or sampling methodology is described, and the manual correctness checking of 185 prompts is reported without inter-rater reliability. Because the 92%-vs-56% comparison against GPT-4o uses the same subjective correctness criterion, the conclusion that the knowledge-graph method is superior would be more convincing if the ground truth were independently established, especially for negative cases.
minor comments (4)
- [Table 1] The response times in Table 1 appear to be single examples per prompt type rather than summary statistics; the paper states "most often under 5 seconds" but also reports a median of 6.92 seconds, so clarify whether Table 1 is illustrative or aggregate.
- [§3.3] There is a grammatical error: "enabling our model enables to more effectively integrate with the FathomNet database" should be rephrased.
- [§5.1.2] The sentence "If the final output answered the user's prompt, the response was still be placed into these categories" contains a verb-form error; it should read "was still placed".
- [§6] The claim "To the best of our knowledge, FathomGPT is the first system to include a prompting technique that instructs the LLM to generate modified prompts that incorporate specific contextual elements" appears in §5.3; it would be safer to soften this to avoid an unverifiable first-use claim.
Circularity Check
No significant circularity: the central claims are empirical measurements against FathomNet, workshop logs, and ablated baselines, with no prediction reducing to its own input.
full rationale
FathomGPT is a systems paper whose evaluation is grounded in the FathomNet target database, in 2,977 logged workshop prompts, and in comparisons against ablated variants and GPT-4o/vector-embedding baselines. I found no step in which a claimed result is defined in terms of its own outcome, no parameter fitted to data and then renamed a prediction, and no load-bearing self-citation chain. The filtering of workshop logs to 781 "reasonable" prompts (Section 5) and the manual labeling of responses in Section 5.1.2 are author-controlled evaluation choices and are genuine threats to the stability of the 72.08% accuracy estimate and the ablation deltas, but they are not circular reductions: the rate is an observed measurement, not a quantity constructed to equal its own input. The self-citation of FathomNet [24] is the system's intended data source rather than evidence used to justify the paper's conclusions. The paper's own noted limitations (prompt-evaluator sensitivity and pattern-extraction limits for subtly patterned species) are also empirical caveats, not circular reasoning. Under the standard that circularity must be exhibited as an explicit reduction or self-citation chain, no significant circularity is present.
Assumptions & free parameters
free parameters (2)
- Few-shot example count in text-to-SQL prompts =
2
- Number of specialized SQL models =
3
assumptions (3)
- domain assumption The FathomNet database schema is a stable relational model that can be queried with SQL.
- domain assumption Wikipedia text contains accurate and sufficiently complete species characteristics to build useful knowledge graphs.
- domain assumption OpenAI API function calling and fine-tuning behave as documented during the study period.
invented entities (2)
-
FathomGPT system
independent evidence
-
Species Knowledge Graph
Cite this review
Pith. "Pith review of FathomGPT: A Natural Language Interface for Interactively Exploring Ocean Science Data." pith.science (2026). https://pith.science/paper/ERCDNJY2
@misc{pith2026241202784,
author = {Pith},
title = {Pith review of: FathomGPT: A Natural Language Interface for Interactively Exploring Ocean Science Data},
year = {2026},
howpublished = {\url{https://pith.science/paper/ERCDNJY2}},
note = {Machine review of arXiv:2412.02784}
}
read the original abstract
We introduce FathomGPT, an open source system for the interactive investigation of ocean science data via a natural language interface. FathomGPT was developed in close collaboration with marine scientists to enable researchers to explore and analyze the FathomNet image database. FathomGPT provides a custom information retrieval pipeline that leverages OpenAI's large language models to enable: the creation of complex queries to retrieve images, taxonomic information, and scientific measurements; mapping common names and morphological features to scientific names; generating interactive charts on demand; and searching by image or specified patterns within an image. In designing FathomGPT, particular emphasis was placed on enhancing the user's experience by facilitating free-form exploration and optimizing response times. We present an architectural overview and implementation details of FathomGPT, along with a series of ablation studies that demonstrate the effectiveness of our approach to name resolution, fine tuning, and prompt modification. We also present usage scenarios of interactive data exploration sessions and document feedback from ocean scientists and machine learning experts.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[1]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Floren- cia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. GPT-4 Technical Report. arXiv preprint arXiv:2303.08774 (2023)
arXiv 2023
-
[2]
Carole C. Baldwin. 2013. The phylogenetic significance of colour patterns in marine teleost larvae. Zoological Journal of the Linnean Society 168, 3 (2013), 496–563
work page 2013
-
[3]
Katherine LC Bell, Jennifer Szlosek Chow, Alexis Hope, Maud C Quinzin, Kat A Cantner, Diva J Amon, Jessica E Cramp, Randi D Rotjan, Lehua Kamalu, Asha de Vos, et al. 2022. Low-cost, deep-sea imaging and analysis tools for deep-sea exploration: A collaborative design study. Frontiers in Marine Science 9, 873700 (2022)
work page 2022
-
[4]
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. 2021. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258 (2021)
arXiv 2021
-
[5]
Antoine Bordes, Nicolas Usunier, Sumit Chopra, and Jason Weston. 2015. Large- scale simple question answering with memory networks. arXiv preprint arXiv:1506.02075 (2015)
arXiv 2015
-
[6]
Stephen Brade, Bryan Wang, Mauricio Sousa, Sageev Oore, and Tovi Grossman
-
[7]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in Neural Information Processing Systems 33 (2020), 1877–1901
2020
-
[8]
Weihao Chen, Chun Yu, Huadong Wang, Zheng Wang, Lichen Yang, Yukun Wang, Weinan Shi, and Yuanchun Shi. 2023. From Gap to Synergy: Enhancing Contextual Understanding through Human-Machine Collaboration in Personalized Systems. In Proceedings of the ACM Symposium on User Interface Software and Technology . 110:1–15
work page 2023
Show all 63 references
-
[9]
Mark J Costello, Philippe Bouchet, Geoff Boxshall, Kristian Fauchald, Dennis Gordon, Bert W Hoeksema, Gary CB Poore, Rob WM van Soest, Sabine Stöhr, T Chad Walter, et al. 2013. Global coordination and standardisation in marine biodiversity through the World Register of Marine ...
2013
-
[10]
Crall, Charles V
Jonathan P. Crall, Charles V. Stewart, Tanya Y. Berger-Wolf, Daniel I. Ruben- stein, and Siva R. Sundaresan. 2013. Hotspotter — Patterned species instance recognition. IEEE Workshop on Applications of Computer Vision (W ACV) (2013), 230–237
2013
-
[11]
Alison Crosby, Eric Coughlin Orenstein, Susan E Poulton, Katherine LC Bell, Benjamin Woodward, Henry Ruhl, Kakani Katija, and Angus G Forbes. 2023. Designing Ocean Vision AI: An Investigation of Community Needs for Imaging- based Ocean Conservation. InProceedings of the CHI Co...
2023
-
[12]
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. Ima- geNet: A large-scale hierarchical image database. InIEEE Conference on Computer Vision and Pattern Recognition . 248–255
2009
-
[13]
Post, Nash E
Simone Des Roches, David M. Post, Nash E. Turley, Joseph K. Bailey, Andrew P. Hendry, Michael T. Kinnison, Jennifer A. Schweitzer, and Eric P. Palkovacs. 2017. The ecological importance of intraspecific variation. Nature Ecology & Evolution 2, 1 (2017), 57–64
2017
-
[14]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. BERT: Pre-training of deep bidirectional transformers for language understanding.arXiv preprint arXiv:1810.04805 (2018)
2018 arXiv
-
[15]
Victor Dibia. 2023. LIDA: A tool for automatic generation of grammar-agnostic visualizations and infographics using large language models. arXiv preprint arXiv:2303.02927 (2023)
2023 arXiv
-
[16]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv prepri...
2020 arXiv
-
[17]
Dawei Gao, Haibin Wang, Yaliang Li, Xiuyu Sun, Yichen Qian, Bolin Ding, and Jingren Zhou. 2023. Text-to-SQL empowered by large language models: A bench- mark evaluation. arXiv preprint arXiv:2308.15363 (2023)
2023 arXiv
-
[18]
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, and Haofen Wang. 2023. Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997 (2023)
2023 arXiv
-
[19]
Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger
-
[20]
Mina Huh, Yi-Hao Peng, and Amy Pavel. 2023. GenAssist: Making image genera- tion accessible. In Proceedings of the ACM Symposium on User Interface Software and Technology. 38:1–17
2023
-
[21]
Wonseok Hwang, Jinyeong Yim, Seunghyun Park, and Minjoon Seo. 2019. A comprehensive exploration on WikiSQL with table-aware word contextualization. arXiv preprint arXiv:1902.01069 (2019)
2019 arXiv
-
[22]
Dow, and Haijun Xia
Peiling Jiang, Jude Rayan, Steven P. Dow, and Haijun Xia. 2023. Graphologue: Exploring Large Language Model Responses with Interactive Diagrams. In Pro- ceedings of the ACM Symposium on User Interface Software and Technology. 3:1–20
2023
-
[23]
Yongyao Jiang and Chaowei Yang. 2024. Is ChatGPT a Good Geospatial Data An- alyst? Exploring the Integration of Natural Language into Structured Query Lan- guage within a Spatial Database. ISPRS International Journal of Geo-Information 13, 1 (2024), 26
2024
-
[24]
Kakani Katija, Eric Orenstein, Brian Schlining, Lonny Lundsten, Kevin Barnard, Giovanna Sainz, Oceane Boulais, Megan Cromwell, Erin Butler, Benjamin Wood- ward, et al. 2022. FathomNet: A global image database for enabling artificial intelligence in the ocean. Scientific Report...
2022
-
[25]
YuHe Ke, Liyuan Jin, Kabilan Elangovan, Hairil Rizal Abdullah, Nan Liu, Alex Tiong Heng Sia, Chai Rick Soh, Joshua Yi Min Tung, Jasmine Chiat Ling Ong, and Daniel Shu Wei Ting. 2024. Development and Testing of Retrieval Augmented Generation in Large Language Models — A Case St...
2024 arXiv
-
[26]
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al
-
[27]
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al
-
[28]
Haoyang Li, Jing Zhang, Hanbing Liu, Ju Fan, Xiaokang Zhang, Jun Zhu, Renjie Wei, Hongyan Pan, Cuiping Li, and Hong Chen. 2024. CodeS: Towards Building Open-source Language Models for Text-to-SQL. arXiv preprint arXiv:2402.16347 (2024)
2024 arXiv
-
[29]
In Proceedings of the IEEE/CVF International Conference on Computer Vision
Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 4015–4026
-
[30]
Yuanyuan Liang, Keren Tan, Tingyu Xie, Wenbiao Tao, Siyuan Wang, Yunshi Lan, and Weining Qian. 2024. Aligning Large Language Models to a Domain-specific Graph Database. arXiv preprint arXiv:2402.16567 (2024)
2024 arXiv
-
[31]
Nick McKenna and Priyanka Sen. 2023. KGQA without retraining. In Proceedings of The Fourth Workshop on Simple and Efficient Natural Language Processing (SustaiNLP). 212–218
2023
-
[32]
Miguel Murguía-Romero, Bernardo Serrano-Estrada, Enrique Ortiz, and José Luis Villaseñor. 2021. Taxonomic identification keys on the web: Tools for better knowledge of biodiversity. Revista Mexicana de Biodiversidad 92 (2021)
2021
-
[33]
Jinyang Li, Binyuan Hui, Ge Qu, Jiaxi Yang, Binhua Li, Bowen Li, Bailin Wang, Bowen Qin, Ruiying Geng, Nan Huo, et al . 2024. Can LLM already serve as a database interface? A big bench for large-scale database grounded text-to-SQLs. FathomGPT: A Natural Language Interface for ...
2024
-
[34]
OpenAI. 2024. https://platform.openai.com/docs/guides/fine-tuning. Accessed: 2024-04-01
2024
-
[35]
OpenAI. 2024. https://platform.openai.com/docs/models/gpt-3-5-turbo. Accessed: 2024-04-01
2024
-
[36]
OpenAI. 2024. https://platform.openai.com/docs/models/gpt-4-and-gpt-4-turbo. Accessed: 2024-04-01
2024
-
[37]
OpenAI. 2024. https://platform.openai.com/docs/guides/function-calling. Ac- cessed: 2024-04-01
2024
-
[38]
Junwoo Park, Youngwoo Cho, Haneol Lee, Jaegul Choo, and Edward Choi. 2021. Knowledge graph-based question answering with electronic health records. In Machine Learning for Healthcare Conference . 36–53
2021
-
[39]
Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H Miller, and Sebastian Riedel. 2019. Language models as knowledge bases? arXiv preprint arXiv:1909.01066 (2019)
2019 arXiv
-
[40]
Plotly.com. 2024. https://plotly.com/graphing-libraries/. Accessed April 3rd, 2024
2024
-
[41]
Jason Parham, Charles Stewart, Jonathan Crall, Daniel Rubenstein, Jason Holm- berg, and Tanya Berger-Wolf. 2018. An animal detection pipeline for identifica- tion. IEEE Winter Conference on Applications of Computer Vision (W ACV) (2018), 1075–1083
2018
-
[42]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research 21, 140 (2020), 1–67
2020
-
[43]
Laria Reynolds and Kyle McDonell. 2021. Prompt programming for large lan- guage models: Beyond the few-shot paradigm. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems . 314:1–7
2021
-
[44]
Charles V Stewart, Jason R Parham, Jason Holmberg, and Tanya Y Berger-Wolf
-
[45]
Mohammadreza Pourreza and Davood Rafiei. 2024. DIN-SQL: Decomposed in-context learning of text-to-SQL with self-correction. Advances in Neural Information Processing Systems 36 (2024), 36339–36348
2024
-
[46]
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014. Sequence to sequence learning with neural networks. Advances in Neural Information Processing systems 27 (2014)
2014
-
[47]
Mingxing Tan and Quoc Le. 2019. Efficientnet: Rethinking model scaling for convolutional neural networks. In International conference on machine learning . PMLR, 6105–6114
2019
-
[48]
Mingxing Tan and Quoc Le. 2021. Efficientnetv2: Smaller models and faster training. In International Conference on Machine Learning . 10096–10106
2021
-
[49]
Yuan Tian, Weiwei Cui, Dazhen Deng, Xinjing Yi, Yurun Yang, Haidong Zhang, and Yingcai Wu. 2024. ChartGPT: Leveraging LLMs to Generate Charts from Abstract Natural Language. IEEE Transactions on Visualization and Computer Graphics (2024), 1–15
2024
-
[50]
Yuze Sun, Yifei Cai, Yunkai Shen, Qianchi Zhang, Xiaolong Feng, Mengmeng Yin, and Dongmei Li. 2021. The COVID-19 question answering system based on knowledge graph. In 2021 IEEE/ACIS 20th International Fall Conference on Computer and Information Science . 215–220
2021
-
[51]
Pere-Pau Vázquez. 2024. Are LLMs ready for Visualization? arXiv preprint arXiv:2403.06158 (2024)
2024 arXiv
-
[52]
Wilco Verberk. 2012. Explaining General Patterns in Species Abundance and Distributions. Nature Education Knowledge 3 (2012), 38
2012
-
[53]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems 35 (2022), 24824–24837
2022
-
[54]
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2022. ReAct: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629 (2022)
2022 arXiv
-
[55]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in Neural Information Processing Systems 30 (2017)
2017
-
[56]
Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, et al. 2018. Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-SQL task. arXiv preprint arXiv:1809.08887 (2018)
2018 arXiv
-
[57]
Victor Zhong, Caiming Xiong, and Richard Socher. 2017. Seq2SQL: Generating structured queries from natural language using reinforcement learning. arXiv preprint arXiv:1709.00103 (2017)
2017 arXiv
-
[58]
Vélez-Rubio, Alejandro Fallabrino, Patricia Zárate, Maike Heidemeyer, Daniel A
Rocío Álvarez Varas, David Véliz, Gabriela M. Vélez-Rubio, Alejandro Fallabrino, Patricia Zárate, Maike Heidemeyer, Daniel A. Godoy, and Hugo A. Benítez. 2019. Identifying genetic lineages through shape: An example in a cosmopolitan marine turtle species using geometric morpho...
2019
-
[60]
Xuchen Yao and Benjamin Van Durme. 2014. Information extraction over struc- tured data: Question answering with Freebase. In Proceedings of the Association for Computational Linguistics. 956–966
2014
-
[2017]
In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
Densely connected convolutional networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 4700–4708
-
[2020]
In Proceedings of the 34th International Conference on Neural Information Processing Systems
Retrieval-augmented generation for knowledge-intensive NLP tasks. In Proceedings of the 34th International Conference on Neural Information Processing Systems. 793:9459–9474
-
[2021]
arXiv preprint arXiv:2106.10377 (2021)
The animal ID problem: Continual curation. arXiv preprint arXiv:2106.10377 (2021)
2021 arXiv
-
[2023]
In Proceedings of the ACM Symposium on User Interface Software and Technology
Promptify: Text-to-Image Generation through Interactive Prompt Explo- ration with Large Language Models. In Proceedings of the ACM Symposium on User Interface Software and Technology . 96:1–14
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.