Pith. sign in

REVIEW 6 cited by

OceanGPT: A Large Language Model for Ocean Science Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.02031 v8 pith:K76BB2SR submitted 2023-10-03 cs.CL cs.AIcs.CEcs.LGcs.RO

classification cs.CLcs.AIcs.CEcs.LGcs.RO
keywords oceansciencedomainlargellmsoceangptlanguageoceans
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Ocean science, which delves into the oceans that are reservoirs of life and biodiversity, is of great significance given that oceans cover over 70% of our planet's surface. Recently, advances in Large Language Models (LLMs) have transformed the paradigm in science. Despite the success in other domains, current LLMs often fall short in catering to the needs of domain experts like oceanographers, and the potential of LLMs for ocean science is under-explored. The intrinsic reasons are the immense and intricate nature of ocean data as well as the necessity for higher granularity and richness in knowledge. To alleviate these issues, we introduce OceanGPT, the first-ever large language model in the ocean domain, which is expert in various ocean science tasks. We also propose OceanGPT, a novel framework to automatically obtain a large volume of ocean domain instruction data, which generates instructions based on multi-agent collaboration. Additionally, we construct the first oceanography benchmark, OceanBench, to evaluate the capabilities of LLMs in the ocean domain. Though comprehensive experiments, OceanGPT not only shows a higher level of knowledge expertise for oceans science tasks but also gains preliminary embodied intelligence capabilities in ocean technology.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs

    cs.CL 2025-05 conditional novelty 6.0 of 10

    EarthSE provides a two-level QA benchmark and an open-ended dialogue benchmark for Earth science and shows current LLMs perform poorly on both.

  2. LightRouter: Towards Efficient LLM Collaboration with Minimal Overhead

    cs.AI 2025-05 conditional novelty 5.0 of 10

    LightRouter uses short preview outputs to filter a pool of LLMs down to two, then aggregates their full responses, beating ensemble baselines and matching costlier models.

  3. A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications

    cs.AI 2025-06 conditional novelty 4.0 of 10

    A survey of 80+ Deep Research systems that proposes a four-layer taxonomy (foundation models, tool use, planning, synthesis) and compares commercial and open-source implementations.

  4. AquaChat: An LLM-Guided ROV Framework for Adaptive Inspection of Aquaculture Net Pens

    cs.RO 2025-07 conditional novelty 3.0 of 10

    AquaChat translates natural-language commands into symbolic ROV plans executed by a PID controller, with experiments in Gazebo and a pool; the framework runs, but several headline claims are not directly measured.

  5. A Review of Generative AI in Aquaculture: Foundations, Applications, and Future Directions for Smart and Sustainable Farming

    cs.RO 2025-07 conditional novelty 3.0 of 10

    A review that maps generative AI to aquaculture tasks, with a marine robotics case study, but the synthesis is weakened by overstated claims and weak citation support.

  6. BioPars: A Pretrained Biomedical Large Language Model for Persian Biomedical Text Mining

    cs.CL 2025-06 reject novelty 3.0 of 10

    A proposed Persian biomedical LLM, BioPars, is evaluated on medical QA datasets and reported to beat GPT-4 on a self-built Persian QA benchmark, but the training setup is not described.

Pith tools