{"id":"5cdeb26c-f6bb-4262-9326-0964c43df4e8","arxiv_id":"2412.15574","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"J-EDI QA is a new 100-image Japanese multiple-choice benchmark for deep-sea organism identification; OpenAI o1 scored 50%, GPT-4o 39%, and non-expert humans about 40%.","lead":"This paper introduces J-EDI QA, a 100-question Japanese benchmark that tests multimodal AI models on identifying deep-sea organisms from JAMSTEC archive images. The best model tested, OpenAI o1, answered 50% correctly, showing that current AI is far from expert-level deep-sea species recognition.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claim that o1 is below expert level is unsupported because no expert baseline was measured; 'expert level' is asserted, not demonstrated.","rationale":"The paper's strongest claim is not merely that o1 scored 50%, but that this score is below expert level. The most load-bearing assumption is that expert-level performance on this benchmark is known to be higher. The reported human-subject result of approximately 40% for non-experts actually complicates the interpretation: o1 outperformed those non-expert humans, so the 'below expert' conclusion depends entirely on an unmeasured expert ceiling. Without an expert baseline, the central quantitative claim cannot be distinguished from a statement about item difficulty. I agree with the reader that data release and representativeness are important limitations, but the missing expert calibration is more directly tied to the abstract's headline interpretation. The reader did mention the absence of an expert baseline in the rationale, so there is partial agreement, though the reader's stated weakest assumption focused on image sampling and label correctness. The recommended verdict remains CONDITIONAL: the benchmark is a useful contribution but requires the expert baseline (and ideally public release of the QA set) before the central claim is fully supported. The proposed concrete test—running the same items through a JAMSTEC expert panel—would settle whether the 'not expert level' claim is valid or whether the benchmark items themselves are the limiting factor.","tokens_in":15817,"tokens_out":2363,"duration_ms":25319,"concrete_test":"Have a panel of JAMSTEC marine biology researchers independently answer the same 100 Japanese multiple-choice questions from the same images, under the same condition of receiving only image and options (no metadata), and record their accuracy and inter-annotator agreement. If expert accuracy is substantially above o1's 50% (e.g., above 75%), the 'not expert level' claim is supported; if expert accuracy is close to 50-60%, the benchmark's difficulty or ambiguity, rather than model deficiency, would explain the result and the central claim would need to be weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that OpenAI o1 achieves 50% on J-EDI QA, indicating that state-of-the-art multimodal models as of December 2024 are 'not yet at an expert level' for deep-sea species comprehension. For that claim to hold, the benchmark must be calibrated against expert performance on the same 100 items. The paper does not provide such a baseline. Section 3.2 reports only two human subjects with 'some knowledge' scoring approximately 40%, and no JAMSTEC expert accuracy is reported. This matters because the abstract's 'expert level' comparison is the interpretive load-bearing step: without an expert ceiling, the 50% figure could mean the items are too difficult or ambiguous for experts too, or that o1 is already at or above non-expert human level. The paper itself acknowledges that some images were difficult to identify 'from the images given alone' (Section 3.2), which makes expert calibration even more necessary. A related but secondary issue is that the benchmark claims to assess Japanese-language comprehension, yet no English-translated version was run, so the Japanese-specific component is not isolated. The most decisive missing control, however, is the expert baseline.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces J-EDI QA, a Japanese-language multiple-choice benchmark for multimodal LLMs, built from 100 deep-sea organism images selected from the JAMSTEC J-EDI archive, with questions and answers authored by JAMSTEC researchers. The authors evaluate OpenAI o1 and GPT-4o, reporting 50% and 39% accuracy respectively, and interpret the o1 result as indicating that state-of-the-art multimodal models are not yet at an expert level for deep-sea species comprehension. The appendix lists URLs, questions, options, and answer keys for all 100 items. The paper also reports that two non-expert human subjects scored approximately 40%, and discusses model reasoning behaviors and limitations.","tokens_in":15998,"tokens_out":4648,"duration_ms":38966,"significance":"If the benchmark is valid and made publicly available, it would fill a specific niche: a domain-specific, Japanese-language multimodal benchmark for deep-sea imagery, which could support development and evaluation of deep-sea organism identification systems. The raw accuracy counts are simple and presumably correct, and the authors are candid about the small size of the dataset and the need for an English-translated parallel assessment. However, the headline claim that current models are 'not yet at an expert level' rests on an unmeasured expert baseline, and the benchmark itself is not yet released, so the paper's current contribution is closer to a proof-of-concept than a fully validated benchmark.","major_comments":[{"comment":"The central claim that o1 'is not yet at an expert level' is not supported by the evidence presented, because no expert human baseline was measured on the same 100 items. Section 3.2 reports only two non-expert human subjects at approximately 40%, and the paper itself states that some images were difficult to identify from the images alone. Without a JAMSTEC expert accuracy on these exact questions, a 50% model score cannot be interpreted as below expert level; it could in principle be at or above the expert ceiling if the image set is unusually hard. The abstract's interpretive sentence should be reworded to state what was actually measured (model accuracy on this benchmark), and the expert-level comparison should either be supported by an explicit expert baseline or explicitly deferred.","section":"Abstract; §3.2"},{"comment":"The benchmark's representativeness is not established. The 100 images are described only as 'selected by JAMSTEC researchers' with no sampling criteria, no annotation protocol, and no inter-annotator agreement, and the compiled 100-image benchmark is not publicly released ('will be made available at a later date'). This makes it impossible for readers to assess whether the 50% figure reflects general deep-sea organism comprehension or a particular, possibly idiosyncratic, selection. The paper should document the selection process, release the benchmark, and ideally report agreement statistics on the expert-authored answers before claiming that the benchmark measures the target competency.","section":"§2.2; Data availability"},{"comment":"All accuracy results are reported as point estimates without confidence intervals or significance tests. With n=100, the standard error of a proportion is approximately 5 percentage points, so the observed o1 versus GPT-4o gap (50% vs 39%) is not clearly distinguishable from sampling variability, and the per-category breakdowns (e.g., 14/20 vs 9/20 for crustaceans) have even wider intervals. The claims that o1 'had a higher percentage of correct answers' and that crustacean performance is 'particularly high' should be qualified with confidence intervals or significance tests.","section":"§3.1; §3.2"},{"comment":"The abstract and introduction describe the benchmark as assessing 'deep-sea species and Japanese terminology comprehension,' but no English-translated version was evaluated, so the Japanese-language component is confounded with visual identification ability. The conclusion acknowledges that an English parallel assessment is needed; this should be presented as an explicit limitation of the current results rather than only as future work.","section":"§1; Conclusion"}],"minor_comments":[{"comment":"Reference [20] is cited for JA-VLM-Bench-In-the-Wild but the listed title is 'Evolutionary Optimization of Model Merging Recipes' by Takuya Akiba et al.; this reference appears incorrect or mismatched.","section":"References"},{"comment":"Typo: 'OenAI o1' should be 'OpenAI o1'.","section":"§3.1"},{"comment":"The table is difficult to use because URLs are broken across lines and some rows appear to have missing or merged option cells (e.g., row 29); providing a machine-readable supplementary file with image IDs, options, and answers would greatly improve usability.","section":"Appendix Table 1"},{"comment":"The phrase 'a selection question was included to facilitate the potential for erroneous responses' is unclear; please clarify what a selection question is and how it relates to the distractor design.","section":"§2.2"},{"comment":"The statement that the benchmark 'will be made available at a later date, but for now they will be available on individual request' is internally inconsistent; specify a concrete release plan or repository.","section":"Data availability"}],"recommendation":"major_revision","confidential_remarks":"The paper is better characterized as a short dataset/benchmark report than a full journal article. Its central quantitative observation (50% on this 100-item set) is easy to verify, but the interpretive claim about expert level is premature without an expert baseline. The mismatched reference [20] and unreleased data also suggest the manuscript is at an early stage. I would consider it for publication after the expert baseline and data release are added; alternatively, it may be more appropriate as a workshop paper or dataset announcement."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"J-EDI QA is a genuine new resource: a 100-question Japanese benchmark for multimodal LLM identification of deep-sea organisms, built from JAMSTEC's J-EDI archive with QA pairs authored by domain researchers. That fills a niche that general and Japanese benchmarks like MMMU, JMMMU, and Heron-Bench don't cover. The raw results — o1 50%, GPT-4o 39% — are simple counts and I have no reason to doubt them. The paper is also refreshingly restrained about its own limitations: it says the dataset is small, notes that some images are hard to identify from the image alone, and reports that two non-expert humans scored about 40%.\n\nThe soft spots are mostly around interpretation and reproducibility. The headline claim that o1 is 'not yet at an expert level' is asserted, not demonstrated. No expert baseline was run on these 100 items, so we have no ceiling to compare against. The two human subjects were 'with some knowledge' and scored ~40%, which is below o1 — so the models may already be above non-expert humans, and possibly close to what experts achieve on deliberately tricky images. That missing control is load-bearing: the abstract's 'expert level' language needs it. Second, 100 questions is a thin sample; there are no confidence intervals or item-level statistics, so the 39 vs 50 gap between models is not clearly meaningful. Third, the selection criteria for images are not documented beyond 'chosen by JAMSTEC researchers,' and the QA pairs are only available on individual request rather than released with the paper, limiting independent verification. Fourth, the claim that this measures Japanese-language comprehension is plausible but not isolated — there is no English-parallel run.\n\nThis is a paper for people building domain-specific or Japanese-language multimodal benchmarks, and for marine biology AI. It is not a major empirical result, but the resource is a legitimate starting point. I would send it to peer review with a clear ask for revision: add an expert baseline (even on a subset), release the QA pairs and image list, report uncertainty, and recalibrate the 'expert level' wording. With those changes, the benchmark could be a useful reference point.","headline":"A genuinely new but small benchmark for deep-sea organism VQA; the 50% result is plausible, but the 'below expert level' claim outruns the evidence because no expert baseline was measured.","tokens_in":16515,"tokens_out":3174,"would_cite":false,"duration_ms":25940,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"J-EDI QA proposes that current multimodal LLMs, including OpenAI o1, cannot yet identify deep-sea organisms at expert level, scoring only 50% on a new Japanese benchmark of 100 images.","keywords":["deep-sea organisms","multimodal LLM","benchmark","J-EDI","Japanese QA","species identification"],"falsifier":"Run the identical 100 questions on the same models using images stripped of the ©JAMSTEC watermark and cropped to remove background habitat cues; if accuracy drops markedly below 50%, the original result was at least partly an artifact of contextual leakage rather than organism identification.","tokens_in":15636,"feed_emoji":"🦑","tokens_out":3852,"duration_ms":32825,"temperature":0.7,"pith_summary":"This paper introduces J-EDI QA, a benchmark of 100 deep-sea images from the JAMSTEC archive, each paired with a four-choice identification question written in Japanese by JAMSTEC researchers. The authors evaluate two multimodal large language models, OpenAI o1 and GPT-4o, and report that o1 answers 50% of the questions correctly while GPT-4o answers 39%. They interpret this as evidence that state-of-the-art general-purpose models, as of December 2024, have not reached expert-level comprehension of deep-sea organisms. The benchmark is intended to support the development of deep-sea-specific LLMs and to serve as a test of public understanding. The central claim is that the 50% result indicates a real gap between current models and expert marine biologists.","feed_headline":"AI models score only 50% on deep-sea species test","feed_subtitle":"A 100-image Japanese benchmark from JAMSTEC shows OpenAI o1 and GPT-4o fall short of expert-level identification.","key_machinery":"The central object is the J-EDI QA benchmark itself: 100 images selected from the J-EDI archive, each with a four-option multiple-choice question in Japanese written by JAMSTEC researchers, plus an expert commentary explaining the correct identification. The evaluation protocol uploads only the image, asks the model to choose an answer and justify it, and scores the percentage of correct choices. This machinery allows the authors to compare model performance against expert answers and against each other, and to separate identification accuracy from the ability to provide a rationale.","core_discovery":"The paper's central claim is that J-EDI QA measures something existing benchmarks do not: the ability to identify deep-sea organisms from survey images, and to do so in Japanese. On this benchmark, the best model tested, OpenAI o1, achieves 50% correct, with GPT-4o at 39%, while two human subjects with some knowledge of deep-sea organisms score about 40%. The authors conclude that current multimodal LLMs have only rudimentary understanding of deep-sea species and remain below the level of JAMSTEC researchers. They further observe that models sometimes rely on contextual cues such as hydrothermal deposits in the background or the ©JAMSTEC watermark, indicating that their reasoning is not purely based on the organism's features.","pith_inferences":["Because all images come from one archive with surveys concentrated around Japan, the 50% figure may not generalize to deep-sea biota from other regions; a geographically diverse sample would test this.","The paper's observation that models use background cues like hydrothermal deposits and the ©JAMSTEC watermark suggests that accuracy could be inflated by contextual leakage; removing those cues would likely lower scores.","The 25% random-guess baseline means 50% is clearly above chance, but the benchmark would be more informative with confidence scores or a measure of how often the model's chosen rationale matches the expert commentary.","A translated English version of the same 100 questions, which the authors mention as future work, would separate language proficiency in Japanese from biological knowledge."],"forward_implications":["If the claim holds, general-purpose multimodal models cannot be relied on for automated identification of deep-sea organisms from images alone.","The benchmark provides a reusable Japanese-language test for tracking progress of deep-sea-specific multimodal LLMs.","The low accuracy motivates training and retrieval-augmented generation with non-digital expert resources such as illustrated field guides.","The J-EDI archive's video data could be used to construct a video-version benchmark, testing temporal and behavioral identification.","The benchmark can also function as a public education and outreach tool for deep-sea biology."],"supporting_citations":[{"why":"Serves as the college-level multimodal benchmark that J-EDI QA builds on and contrasts with expert-level content.","marker":"[2]"},{"why":"MMMU-Pro, with its increased number of choices, is cited as a model for raising benchmark difficulty.","marker":"[15]"},{"why":"Science QA is the scientific image-question benchmark that motivates domain-specific scientific LLM evaluation.","marker":"[18]"},{"why":"Heron-Bench is a Japanese vision-language benchmark whose approach to Japanese terminology J-EDI QA follows.","marker":"[19]"},{"why":"JMMMU is another Japanese multimodal benchmark cited to situate the Japanese-language evaluation in the paper.","marker":"[21]"}],"fun_headline_variants":["Deep-sea AI benchmark stumps top models at 50%","J-EDI QA: Multimodal LLMs fall short on deep-sea species","AI only half-right on deep-sea organism quiz","Deep-sea species test: AI scores 50%, not expert level","New benchmark reveals AI's shallow grasp of deep-sea life"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"A load-bearing premise is that the 100 selected images and their expert-written QA pairs are representative and correctly labeled, so that model accuracy on this set measures general deep-sea organism comprehension.","fun_headline_variants_meta":{"raw":{"variants":["Deep-sea AI benchmark stumps top models at 50%","J-EDI QA: Multimodal LLMs fall short on deep-sea species","AI only half-right on deep-sea organism quiz","Deep-sea species test: AI scores 50%, not expert level","New benchmark reveals AI's shallow grasp of deep-sea life"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000791,"raw_usage":{"total_tokens":3469,"prompt_tokens":912,"completion_tokens":2557,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":528,"completion_tokens_details":{"reasoning_tokens":2470}},"tokens_in":528,"tokens_out":2557,"duration_ms":15483,"temperature":1.0,"reasoning_tokens":2470,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T11:17:23.121715+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the identical 100 questions on the same models using images stripped of the ©JAMSTEC watermark and cropped to remove background habitat cues; if accuracy drops markedly below 50%, the original result was at least partly an artifact of contextual leakage rather than organism identification.","supporting_citations":[],"review_version":1}