{"id":"58d0a632-0952-4cd1-8057-23bbafd72b67","arxiv_id":"2412.02784","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"FathomGPT is an open-source natural language interface for the FathomNet ocean image database, with ablations showing improved text-to-SQL accuracy through fine-tuning and prompt modification.","lead":"FathomGPT is a natural-language chatbot that lets researchers query the FathomNet ocean image database, resolve species names, generate charts, and search images by pattern. The paper reports latency and ablation results showing higher accuracy from fine-tuned SQL generation and prompt rewriting.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 72.08% success rate rests on the authors' own filtering of 2,977 workshop prompts and their own labeling of responses; with no inter-rater reliability or held-out prompt set reported, the headline number is not yet stable.","rationale":"The reader identified representativeness of the 781 filtered prompts as the weakest assumption. I agree with that, and I extend it: representativeness is not the only unverified link. The success labels are also self-assigned, and the filtering rules are not fully specified in the main text. The strongest claim, that FathomGPT generates a successful response for 72.08% of realistic prompts, requires both the prompt set and the outcome labels to be inter-subjective. The paper gives no inter-rater reliability, no held-out evaluation, and no error bars for the headline accuracy. This is a load-bearing concern because the central contribution is an empirical system claim, not a mathematical derivation; the correctness risk is medium, as the reader judged. I do not think the paper should be rejected: the workshop deployment, the open-source code, the latency measurements, and the ablation comparisons are genuine evidence, and the system very likely works as described for many prompts. But the specific figure of 72.08% should be treated as an upper-bound estimate pending independent labeling and a more clearly specified filtering pipeline. This reinforces rather than changes the reader's CONDITIONAL verdict, so I mark the verdict as unchanged.","tokens_in":19880,"tokens_out":5938,"duration_ms":64728,"concrete_test":"Obtain from the authors the full unfiltered 2,977-prompt log and the exact filter rules. Have two independent annotators who did not build FathomGPT apply the paper's success/error taxonomy to all 781 filtered responses, or to a pre-specified random sample of at least 200, without seeing the system's own labels. Report per-annotator success rates, inter-annotator agreement (Cohen's kappa), and a bootstrap confidence interval for the difference from 72.08%. If the independent success rate differs by more than 5 percentage points from 72.08%, or if kappa is below 0.6, the headline metric should be re-estimated on a fresh held-out prompt set before being taken at face value.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim depends on two author-controlled steps. Section 5 says that 2,977 workshop prompts were filtered to 781 'reasonable' prompts by removing duplicates, invalid inputs, and 'hack' attempts, with the filtering methodology deferred to Supplementary Material. Section 5.1.2 then states that 'we labeled each response' as successful or erroneous, yielding 72.08%. Both the inclusion criterion and the outcome labels are therefore set by the people who built the system. If the filter preferentially removes prompts that stress the system, and if the labels are applied leniently, the reported success rate can overstate performance on future real usage. The ablations inherit the same weakness: the fine-tuning gain of 14.54% and the context-modification gain of 7.27% are measured against the same author-labeled set, so rater bias could affect both arms in the same direction. This is not an internal contradiction, and the paper does provide real evidence: a deployed system, 2,977 logged prompts, and measured latencies. But the evidence as written does not separate system performance from rater judgment and prompt-selection judgment, so the headline 72.08% is not yet a stable estimate of real-world success.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"FathomGPT is an open-source natural language interface for querying and visualizing the FathomNet ocean image database. The paper describes a multi-stage LLM pipeline that includes a prompt evaluator, a name-resolution component based on knowledge graphs built from Wikipedia, fine-tuned text-to-SQL models, Plotly-based visualization generation, and a pattern-extraction image search. The system was deployed at a FathomNet workshop, where 2,977 prompts were logged; after filtering, 781 prompts were used to evaluate overall accuracy, reporting a 72.08% success rate and a median latency of 6.92 seconds. Two ablation studies claim improvements from fine-tuning and from a context-modification prompting strategy, and a third evaluation compares the knowledge-graph name-resolution method against GPT-4o and vector embeddings.","tokens_in":20122,"tokens_out":3238,"duration_ms":32635,"significance":"If the reported performance is stable, FathomGPT would be a useful contribution to scientific data access: it is a deployed, open-source system with concrete architectural components (prompt evaluator, species knowledge graphs, specialized text-to-SQL models, pattern-based image retrieval) that target real user needs identified with marine scientists. The paper provides genuine empirical material: logged workshop prompts, latency measurements, ablations, and a head-to-head name-resolution comparison showing the knowledge-graph method outperforms GPT-4o on the tested set. However, the central quantitative claims rest on evaluation choices that are not yet fully transparent or externally validated, so the magnitude of the reported gains should be treated with caution until the methodology is strengthened.","major_comments":[{"comment":"The headline accuracy of 72.08% (563/781) is measured on a subset of 2,977 workshop prompts after an author-defined filtering step whose criteria are deferred to Supplementary Material, and the responses were labeled by the authors without any reported inter-rater reliability, independent labeler, or confidence interval. Since this figure is the paper's central evidence of system effectiveness, the reader cannot currently distinguish system performance from filtering and labeling judgment. Please provide the full filtering protocol, report agreement statistics if multiple labelers are used, and give a confidence interval for the success rate.","section":"§5.1.2"},{"comment":"The fine-tuning ablation reports that the fine-tuned model generated a correct response \"14.54% more often\" with counts 252 vs. 220. With the total of 335 introduced later in the same section, the absolute success rates are 75.2% and 65.7%, a difference of 9.5 percentage points; 14.54% is the relative improvement over the ablated baseline (32/220). The text does not state whether the reported percentage is absolute or relative, and the 335-prompt subset is not clearly defined as the ablation evaluation set. Please report absolute success rates and error counts for each condition with a consistent denominator, and add a significance test if feasible.","section":"§5.2"},{"comment":"The context ablation reports \"7.27% more queries\" with 129 vs. 153 errors out of 483 prompts. The absolute difference is 24/483 = 4.97 percentage points, while 7.27% is the relative reduction over the ablated version's 153 errors. As in §5.2, the reported metric is ambiguous. In addition, the 483-prompt set was obtained by manually filtering 126 conversations with exclusion criteria (\"conversations that are not affected by context\") that are not operationalized, making the result non-reproducible from the text. Please clarify the filtering rules and report absolute performance in both arms.","section":"§5.3"},{"comment":"The name-resolution evaluation treats empty results as correct when the authors judge that no matching species exists in FathomNet (e.g., \"predators of hexanchus griseus\"), but no ground-truth negative set or sampling methodology is described, and the manual correctness checking of 185 prompts is reported without inter-rater reliability. Because the 92%-vs-56% comparison against GPT-4o uses the same subjective correctness criterion, the conclusion that the knowledge-graph method is superior would be more convincing if the ground truth were independently established, especially for negative cases.","section":"§5.4"}],"minor_comments":[{"comment":"The response times in Table 1 appear to be single examples per prompt type rather than summary statistics; the paper states \"most often under 5 seconds\" but also reports a median of 6.92 seconds, so clarify whether Table 1 is illustrative or aggregate.","section":"Table 1"},{"comment":"There is a grammatical error: \"enabling our model enables to more effectively integrate with the FathomNet database\" should be rephrased.","section":"§3.3"},{"comment":"The sentence \"If the final output answered the user's prompt, the response was still be placed into these categories\" contains a verb-form error; it should read \"was still placed\".","section":"§5.1.2"},{"comment":"The claim \"To the best of our knowledge, FathomGPT is the first system to include a prompting technique that instructs the LLM to generate modified prompts that incorporate specific contextual elements\" appears in §5.3; it would be safer to soften this to avoid an unverifiable first-use claim.","section":"§6"}],"recommendation":"major_revision","confidential_remarks":"The paper describes a genuinely useful and open system, and the workshop deployment is a real strength. However, the evaluation methodology as written does not yet support the precision of the headline numbers. I would encourage the editor to require the authors to report confidence intervals, inter-rater reliability or independent labeling, and unambiguous absolute effect sizes for the ablations. These are fixable within the scope of a revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know about FathomGPT is that it is a real, deployed system with a real evaluation, not a demo. The authors built a natural-language interface over the FathomNet image database, logged nearly 3,000 prompts from 174 workshop attendees, and report latency and ablation data. The architecture is sensible: a prompt evaluator routes to specialized text-to-SQL models, a knowledge graph for name resolution, and a pattern/image search component. The paper is honest about what it did not do.\n\nWhat's new is modest but genuine. The specific integration of LLM-based prompt-to-SQL, Wikipedia-derived species knowledge graphs with graph alignment, and interactive pattern search on a scientific image database is not in the prior literature. The prompt modification technique—having the LLM rewrite a follow-up prompt to include relevant context before downstream function calls—is a small but useful trick, and the ablation shows a 7.27% accuracy gain with 15.67% fewer tokens. The name resolution comparison against GPT-4o is also a concrete result: KG-based resolution got 92% of names correct versus 56% for GPT-4o, with the latter hallucinating often.\n\nThe soft spots are concentrated in the headline accuracy number. The 72.08% success rate comes from a self-filtered set of 781 prompts (out of 2,977) and author-assigned labels. The filtering methodology is deferred to supplementary material, and there is no inter-rater reliability or independent labeling. That makes 72.08% a provisional estimate, not a stable one. The stress-test note is right that the filter and labels are controlled by the same people who built the system, and the ablations inherit that bias. But this is not a hidden flaw—the paper says the filtering happened and points to the supplement. It is a limitation of a single-workshop evaluation, not a sign of bad faith.\n\nThe more substantive gap is that there is no comparison to a non-LLM baseline, like a search UI or a human writing SQL. That would have made the claim \"this helps ocean scientists\" much stronger. As it stands, the paper shows the system works under its own scoring, and that components help, but not that it beats existing tools.\n\nWho is this for? Anyone building LLM interfaces over scientific databases. It is a good example of what works and what breaks in practice. I'd bring it to a reading group as a case study and would cite it if I were working on text-to-SQL or scientific data interfaces. It deserves a serious referee. Send it out; the evaluation weaknesses are fixable with a held-out set and second labeler, and the system itself is worth the community's attention.","headline":"A solid, honest systems paper for an ocean-science database interface; the headline accuracy number is self-evaluated and should be read as provisional, but the system and ablations are real.","tokens_in":20690,"tokens_out":1734,"would_cite":true,"duration_ms":17139,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"FathomGPT gives ocean scientists a natural language interface that retrieves images, taxonomy, and measurements from the FathomNet database, answering 72.08% of 781 real workshop prompts successfully with a median latency of 6.92 seconds.","keywords":["natural language interfaces","ocean science data","text-to-SQL","knowledge graph name resolution","image retrieval","data visualization","large language models","FathomNet"],"falsifier":"Run FathomGPT on a fresh, unfiltered set of prompts collected from new ocean-science users who have not seen the interface examples, and check whether the success rate stays near 72.08%; a large drop would show that the filtered workshop prompts were not representative. A second check is to count correct name resolutions for prompts that mention predator-prey relations, where the paper reports many empty results.","tokens_in":1620,"feed_emoji":"🌊","tokens_out":2016,"duration_ms":77770,"temperature":0.7,"pith_summary":"This paper introduces FathomGPT, an open-source natural language interface that lets researchers query a large ocean image database by typing ordinary English instead of writing SQL. The system promises to remove a barrier that keeps many marine scientists and enthusiasts from using rich scientific databases. The authors report that the system produced a successful response for 72.08% of 781 prompts collected from a real user workshop, with a median response time of 6.92 seconds. They argue that a combination of context-aware prompt rewriting, fine-tuned text-to-SQL models, and knowledge-graph name resolution is what makes this performance possible. If the approach holds, it points to a template for opening other scientific databases to non-specialists.","feed_headline":"Natural-language ocean database answers 72% of real prompts","feed_subtitle":"Plain English turns into SQL, charts, and image searches over 109,871 marine images in about 7 seconds.","key_machinery":"The load-bearing machinery is the prompt evaluator: a large language model with function calling that decides which of five processing functions handles a prompt and rewrites the prompt to include only the relevant prior conversation, so downstream models see a self-contained question. Supporting it are species knowledge graphs, structured records extracted from encyclopedia text that map each scientific name to aliases, body parts, colors, predators, diet, and habitat; name resolution aligns a prompt-derived subject-relation-object triple against these graphs. Text-to-SQL uses three fine-tuned models specialized for similarity-search, visualization, and image/text/table outputs, plus few-shot schema examples and an error-feedback loop that sends failed SQL to a stronger model for repair. Visualization and image search complete the pipeline: graphing code is generated for charts, and pattern queries use an interactive segmentation model, color-based pattern extraction, and a feature-vector model to rank similar images.","core_discovery":"The central claim is that a natural language interface can make a large, complex scientific image database practically accessible without sacrificing speed or accuracy. FathomGPT routes every user prompt through a prompt evaluator that decides among five capabilities—name resolution, SQL query generation, taxonomic lookup, visualization code generation, and general information—and rewrites the prompt to fold in relevant conversation context. Common names and morphological descriptions are resolved to scientific names by aligning a prompt-derived knowledge graph against species knowledge graphs generated from encyclopedia text. Query generation uses three fine-tuned text-to-SQL models specialized by output type, and visualization requests generate graphing code in parallel with the database query. On 781 prompts logged from a workshop with domain users, the system produced a successful response 72.08% of the time with a median latency of 6.92 seconds, and ablation studies attribute part of this gain to fine-tuning and part to context-aware prompt modification.","pith_inferences":["The reported success rate is measured on a filtered subset of 781 prompts, so on raw workshop logs that include duplicate, invalid, and adversarial queries the per-prompt success rate would be lower; a follow-up evaluation on unfiltered logs would give a more conservative estimate.","Because the knowledge graphs are extracted from encyclopedia text, species with sparse coverage will produce empty or partial name-resolution results, especially for predator-prey relations; expanding the source corpus is a direct testable extension.","The three-way split of fine-tuned text-to-SQL models by output type is a design choice other database interfaces could adopt, but the paper does not isolate whether the benefit comes from specialization or simply from more fine-tuning data per model.","The same prompt-evaluator architecture could be ported to other relational scientific databases with a schema and a curated knowledge graph, since the pipeline itself is not ocean-specific; the paper notes this possibility but does not demonstrate it."],"forward_implications":["Marine scientists can ask for images, taxonomic trees, measurements, and charts in plain English, removing the SQL and schema expertise barrier that keeps many researchers away from databases like FathomNet.","Fine-tuning separate text-to-SQL models for different output types yields substantially more correct queries than a single few-shot model, so systems with heterogeneous outputs should expect to specialize their query generators.","Rewriting each follow-up prompt to embed relevant conversation context improves multi-turn accuracy by 7.27% and cuts token usage by roughly 15.67%, a concrete design choice for conversational database interfaces.","Knowledge-graph name resolution resolves common names and descriptive queries to scientific names with about 92% of returned names correct, compared with 56% for a direct general-purpose LLM, while avoiding hallucinations of species not present in the database.","Pattern-based image search lets users highlight a region of an uploaded image and retrieve similar database images, extending access from text queries to visual queries."],"supporting_citations":[{"why":"Supplies the FathomNet image database with images, bounding boxes, species concepts, and scientific measurements that FathomGPT queries.","marker":"[24]"},{"why":"Documents the community needs and access barriers that motivate building a natural language interface for ocean science data.","marker":"[11]"},{"why":"Provides the function-calling mechanism the prompt evaluator uses to route prompts to the appropriate processing function.","marker":"[33]"},{"why":"Provides the fine-tuning capability used to build the specialized text-to-SQL models.","marker":"[34]"},{"why":"Supplies the interactive segmentation model used to isolate a user-selected pattern before extracting and searching by it.","marker":"[26]"},{"why":"Provides the graphing library that renders the user-requested interactive charts.","marker":"[40]"},{"why":"Supplies the feature-vector model used for ranking pattern-based image similarity searches.","marker":"[48]"},{"why":"Supplies the vision transformer used to precompute image feature vectors for whole-image similarity search.","marker":"[16]"}],"fun_headline_variants":["FathomGPT: plain English meets ocean data with 72% success","Natural language queries to 109,871 ocean images succeed 72%","FathomGPT answers 72% of real ocean science prompts","AI interface unlocks ocean image database for scientists","Ocean database talks back: FathomGPT hits 72% on real queries"],"cache_read_input_tokens":22784,"weakest_assumption_plain":"The 781 prompts that remain after filtering the workshop logs are representative of how real users will actually use FathomGPT, so the measured 72.08% success rate and the ablation gains carry over to everyday use.","fun_headline_variants_meta":{"raw":{"variants":["FathomGPT: plain English meets ocean data with 72% success","Natural language queries to 109,871 ocean images succeed 72%","FathomGPT answers 72% of real ocean science prompts","AI interface unlocks ocean image database for scientists","Ocean database talks back: FathomGPT hits 72% on real queries"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00073,"raw_usage":{"total_tokens":3248,"prompt_tokens":904,"completion_tokens":2344,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":520,"completion_tokens_details":{"reasoning_tokens":2252}},"tokens_in":520,"tokens_out":2344,"duration_ms":16568,"temperature":1.0,"reasoning_tokens":2252,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T23:06:01.755079+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run FathomGPT on a fresh, unfiltered set of prompts collected from new ocean-science users who have not seen the interface examples, and check whether the success rate stays near 72.08%; a large drop would show that the filtered workshop prompts were not representative. A second check is to count correct name resolutions for prompts that mention predator-prey relations, where the paper reports many empty results.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the FathomNet image database with images, bounding boxes, species concepts, and scientific measurements that FathomGPT queries."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the community needs and access barriers that motivate building a natural language interface for ocean science data."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the function-calling mechanism the prompt evaluator uses to route prompts to the appropriate processing function."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the fine-tuning capability used to build the specialized text-to-SQL models."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the graphing library that renders the user-requested interactive charts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the feature-vector model used for ranking pattern-based image similarity searches."}],"review_version":1}