Pith. sign in

Paper Citation Record · LEDGER

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs

As of 10 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 1 inbound Pith citation observation for arXiv:2505.17139.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17139 v3

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:10:21.111443Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:22:23.855247Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T00:22:26.516274Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact3
  • verified fuzzy10
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 12e6ab02-ab96-4435-9f59-007f7f320169 · outbound

This paper cites GPT-4 Technical Report.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.566305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.566305Z digest=sha256:892fc685fe26223c58d7d88b7b5fd9171f563d6bc606a111b73c74e02d1123cd

Observation 6f7f9f8a-0dc7-4e7c-b5ab-106c5b6bd07a · outbound

This paper cites OceanGPT: A Large Language Model for Ocean Science Tasks.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs OceanGPT: A Large Language Model for Ocean Science Tasks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.576960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.576960Z digest=sha256:2aad57793e5fa3d8f41ad7bb46e60a7973d903612eaf202c7228fc4ad8ca2eda

Observation 3b9a4a9b-84d1-49dc-8890-d6f090abcec9 · outbound

This paper cites Matplotlib and seaborn.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Matplotlib and seaborn

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:22.734997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:10:20.586447Z digest=sha256:52fbd23a60c75458bd050cdba29b7ba41d99f2bb3669d202b8c152bbca82402b

Observation ee12df61-3140-427d-89bd-3baa7c2563f7 · outbound

This paper cites This reference does not exist: an exploration of llm citation accuracy and relevance.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs This reference does not exist: an exploration of llm citation accuracy and relevance

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:22.704132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:10:20.597692Z digest=sha256:d5e0719d00e802fcc18357e2801018dfb236dc12ff185d49523f9ab5d0dca372

Observation 03694014-52e2-4d23-a59a-ca3afde7f50d · outbound

This paper cites SciAssess: Benchmarking LLM Proficiency in Scientific Literature Analysis.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs SciAssess: Benchmarking LLM Proficiency in Scientific Literature Analysis

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.605939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.605939Z digest=sha256:937c87537c2d73948258b0eff980a3063d9d40214c2f36baf38e3869dbd87053

Observation 5c6626de-6fc7-4472-9abd-8611b34178a7 · outbound

This paper cites On the design and analysis of llm-based algorithms.arXiv preprint arXiv:2407.14788, 2024.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs On the design and analysis of llm-based algorithms.arXiv preprint arXiv:2407.14788, 2024

Reference 6

Resolution
verified exact
raw_fallback, observed 2026-08-07T15:10:22.199410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:10:20.613696Z digest=sha256:9f3bf75a9430e9914525d6f3be17ad73edfb076e6732028e76a845ac47e4eee8

Observation 5c3a3ed5-4562-4bb9-a41d-72eecfe0a441 · outbound

This paper cites Grok, gemini, chatgpt and deepseek: Com- parison and applications in conversational artificial intelligence.INTELIGENCIA ARTIFICIAL, 2(1), 2025.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Grok, gemini, chatgpt and deepseek: Com- parison and applications in conversational artificial intelligence.INTELIGENCIA ARTIFICIAL, 2(1), 2025

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:22.679570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:10:20.624596Z digest=sha256:c39be5426bc9119d54b3c6d3f71c774ef77a9787354ae79eb9401193be0ec861

Observation ddd7c2a3-4a07-4ef7-8503-14a8219b0570 · outbound

This paper cites K2: A foundation language model for geoscience knowledge understanding and utilization.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs K2: A foundation language model for geoscience knowledge understanding and utilization

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:22.653051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:10:20.633444Z digest=sha256:50c843c94762f2582c334281ca004bad0684af8f6d7fba23990de917a8d23ed1

Observation 3a77f33f-629b-42ae-aa1d-af2ec3f14e6e · outbound

This paper cites A deep learning model based on bert and sentence transformer for semantic keyphrase extraction on big social data.IEEE Access, 9:165252–165261, 2021.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs A deep learning model based on bert and sentence transformer for semantic keyphrase extraction on big social data.IEEE Access, 9:165252–165261, 2021

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:22.629090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:10:20.642177Z digest=sha256:c3a090d4585ab4520c83c046a5792bda292facbb4b9a096fab601055c323c19c

Observation 8e52f789-699d-415e-8981-809fe1d85431 · outbound

This paper cites SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.649206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.649206Z digest=sha256:9c8039bce56c9d022c8b1d53ebd4d0f9bf2ee2cea5e9812f046659cede805317

Observation 901f78dc-51e1-4aa5-89ae-83a1174aeb2f · outbound

This paper cites The impact factor.Current contents, 25(20):3–7, 1994.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs The impact factor.Current contents, 25(20):3–7, 1994

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:22.599611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:10:20.656055Z digest=sha256:52a4f99a93d1b6894aa2c1d4664ebafbc5c2286b26dec30dfb48ee151e58fcf6

Observation 558f8956-d940-4cbb-ac35-562400adbf7a · outbound

This paper cites The Llama 3 Herd of Models.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.662840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.662840Z digest=sha256:8bdc0c60135c52f08998dd363bda2c8311cc2f8397a82196fe2cc99f5d11b2c2

Observation a51b57c8-8b18-4140-ba82-5509949b35fe · outbound

This paper cites Llm-based code generation method for golang compiler testing.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Llm-based code generation method for golang compiler testing

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:22.566077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:10:20.674468Z digest=sha256:57ef4a3255e795641eeaa7a2ba1c11d81e0a0188e5ec8c38c4fa7925c0a9d026

Observation 95be0cb6-704e-402e-886b-af5bf384528f · outbound

This paper cites OpenDataLab: Empowering General Artificial Intelligence with Open Datasets.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs OpenDataLab: Empowering General Artificial Intelligence with Open Datasets

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.683625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.683625Z digest=sha256:902f4eb010c89a44b216c7122468dd1a864fec57bf2b236a86383bcc8495418c

Observation 91fc2325-9511-40b3-a819-e546cb97576a · outbound

This paper cites The Accuracy, Robustness, and Readability of LLM-Generated Sustainability-Related Word Definitions.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs The Accuracy, Robustness, and Readability of LLM-Generated Sustainability-Related Word Definitions

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:10:22.028408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:10:20.695723Z digest=sha256:d42950ebc8698b80b7d44e5be906b5d3c375079fc233885763a0747299e9480b

Observation b1fc3332-5753-4732-b197-51241d8d9f52 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Measuring Massive Multitask Language Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.704332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.704332Z digest=sha256:37e392cab6a79bef3d190752b6a00ae56e4cf5ae981901ff7b2e105484de48b2

Observation 78b29462-ae4a-485d-b659-4a848770e410 · outbound

This paper cites The era5 global reanalysis.Quarterly journal of the royal meteorological society, 146(730):1999–2049, 2020.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs The era5 global reanalysis.Quarterly journal of the royal meteorological society, 146(730):1999–2049, 2020

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.713964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.713964Z digest=sha256:fcb7407e02e23e14870864bb413aeed6f266a4f2920fbf71001131ff4d211348

Observation 281212e1-4636-41b7-acc7-a9c31a696dea · outbound

This paper cites Gpt-4o: The cutting-edge advancement in multimodal llm.Authorea Preprints, 2024.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Gpt-4o: The cutting-edge advancement in multimodal llm.Authorea Preprints, 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.723245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.723245Z digest=sha256:a33c5487a83901dace7047ce24609d3a88eaa014570457468ac889e9ffc4c63d

Observation ebe8e811-81a1-49bf-b046-a2957d4f76d5 · outbound

This paper cites Enhancing Large Language Models with Climate Resources.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Enhancing Large Language Models with Climate Resources

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.734284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.734284Z digest=sha256:91318404160ead3b232b6a6e13b799d8725123c3fccb2453793e914896021c9c

Observation afffe1b5-8a9f-49a9-9371-5b9914af7231 · outbound

This paper cites an unresolved cited work.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:10:22.494299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:10:20.757242Z digest=sha256:292e9b0a171a2e8e0fa8d3577b4c61ff7eea5ffdeca82987bf98ddd1f7e23f43

Observation 74aae74f-2757-4bb5-b553-06560836c397 · outbound

This paper cites Iterative large language models evolution through self-critique.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Iterative large language models evolution through self-critique

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:22.471709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:10:20.766466Z digest=sha256:e6b899a853c3d59ac609da48044d439103bd81a1b15ec1023f6e503393c42fed

Observation 6bc77e47-f0a3-4788-b6d7-8306b957569e · outbound

This paper cites LLM with Relation Classifier for Document-Level Relation Extraction.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs LLM with Relation Classifier for Document-Level Relation Extraction

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:10:21.956033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:10:20.777262Z digest=sha256:9152b10ce23c162c31200595f37515fe78663b6e5f87fe0b2098f8f06a8c505b

Observation fe3cb26d-1bf5-4cb4-8269-3a815cbc21c2 · outbound

This paper cites Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.798959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.798959Z digest=sha256:860c14f353928e9bd940f772d9dbd78496464de4abbd9db4d28ead1121d1dacf

Observation 872ab3e7-b821-49cc-87ec-146b1e0ff023 · outbound

This paper cites an unresolved cited work.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:10:22.445208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:10:20.819684Z digest=sha256:01ce821e84587e18a6a2182a3fefdb33f184c8c2dd7596faccd36af6ba9c233b

Observation 056aff60-4819-4244-93b8-6950fc322455 · outbound

This paper cites DeepSeek-V3 Technical Report.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs DeepSeek-V3 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.829089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.829089Z digest=sha256:7582049517a6afbeb02ececae73324d16f2cadf1e36adfc1e2304dd608595114

Observation 62bd0b92-6d7e-4daa-b372-7e9b269ae79d · outbound

This paper cites The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.840356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.840356Z digest=sha256:7b512455ff6af92da1cca4e18c9364bf3a09fbe341a457751e2963b4b1c31546

Observation 6531b314-a0a7-4985-8138-d996dd5fdde2 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521, 2022.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521, 2022

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.852395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.852395Z digest=sha256:46d307bb7b0b732f0b0edca779a7a74e88465078d4d9fcf4c489332f2c27c7ce

Observation 8f689b85-f820-48e3-a3e8-f959eef2b0ac · outbound

This paper cites The five environmental spheres.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs The five environmental spheres

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:22.399166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:10:20.859905Z digest=sha256:85cf2b8506c6a59506cd44d71f1c6360c60584c153ffc699c7672deb201714e9

Observation bd80d183-00d0-4c2f-86be-20dacafbcab6 · outbound

This paper cites ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.868023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.868023Z digest=sha256:91c8bdac6decacca14cad4d9e0545f901c5934ed85c656d90a0721053f5843c1

Observation 45cf0f75-ab81-4122-ba23-88b942d26cd5 · outbound

This paper cites Seafloorai: A large-scale vision- language dataset for seafloor geological survey.Advances in Neural Information Processing Systems, 37:22107–22123, 2024.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Seafloorai: A large-scale vision- language dataset for seafloor geological survey.Advances in Neural Information Processing Systems, 37:22107–22123, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:22.378751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:10:20.876565Z digest=sha256:d61b95d9dce1580d07715adcf2063c1fae1f1ee8ba2e0b25753577a815a94161

Observation 4c1b245e-0f1d-4ba5-bac3-493010190ebf · outbound

This paper cites Is Temperature the Creativity Parameter of Large Language Models?.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Is Temperature the Creativity Parameter of Large Language Models?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.885364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.885364Z digest=sha256:617a7be7da40250a3f18819a62814d6ce1e5fda98e308c2870661c98404006c7

Observation c50de15f-b858-4496-af4f-72d51508e97a · outbound

This paper cites Humanity's Last Exam.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Humanity's Last Exam

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.895287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.895287Z digest=sha256:8bdb3e9254ca814301e64433b254b1126ce7470226e3cd808f2c8796aabf747e

Observation 13166c0e-916b-40bc-b5bc-c9685bceabd0 · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Gpqa: A graduate-level google-proof q&a benchmark

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.903533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.903533Z digest=sha256:b08e38c56d37258d655eb1fea392da38167e2e242bff57962cb1810805450d00

Observation 7d5f817d-5d5f-484e-9271-46d7e81fbf16 · outbound

This paper cites From Calculation to Adjudication: Examining LLM judges on Mathematical Reasoning Tasks.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs From Calculation to Adjudication: Examining LLM judges on Mathematical Reasoning Tasks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.913666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.913666Z digest=sha256:d25dee2e8b0db0fc503995a9b96d59a0575d8cc6cf5bf1a130708e05ae6aa6c2

Observation 2ea61cff-348e-4ac5-8f3c-92b75c277f79 · outbound

This paper cites Galactica: A Large Language Model for Science.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Galactica: A Large Language Model for Science

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.923953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.923953Z digest=sha256:4bebe744c8f8d551cafa6defa5edb999e05727a5a06aa537b427eb7462d6143a

Observation 673f422b-1eae-4dc7-b4f9-06cf2fd9c73e · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Gemini: A Family of Highly Capable Multimodal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.951647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.951647Z digest=sha256:117bab2ef7705f38961cacdb6d2c3f7a5293e998d6c0dd706827deffd725870a

Observation bf4b7597-e101-4f0e-a4d5-bdccec3f0e9d · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.957403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.957403Z digest=sha256:45cbc9218f48fc0a910a30782ad25b3d638e877a6d9c7c922a70c966d28bcb9b

Observation acd440be-6d2c-4734-9ae5-b41998e399ba · outbound

This paper cites Large language models in medicine.Nature medicine, 29(8):1930–1940, 2023.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Large language models in medicine.Nature medicine, 29(8):1930–1940, 2023

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.967014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.967014Z digest=sha256:07750cd8a78ecce05c1a99659fd2eeb4b01ee2db9dbd6b1a26d2455e1448cb99

Observation 8756a6bd-781f-4f4d-9f0a-649e2e411800 · outbound

This paper cites ClimaText: A Dataset for Climate Change Topic Detection.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs ClimaText: A Dataset for Climate Change Topic Detection

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.976823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.976823Z digest=sha256:c989abe3a93b61ed5f078c278670a0f973ae0368b72ef975f610b1f03bc4443e

Observation 357c6f56-b065-4774-afe6-059089e7894a · outbound

This paper cites MinerU: An Open-Source Solution for Precise Document Content Extraction.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs MinerU: An Open-Source Solution for Precise Document Content Extraction

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.984124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.984124Z digest=sha256:41bdb7ba02f9650c16469d2f142a7eb557463ddd3d8e30fd8c3f41aec240be2d

Observation 12c2eb19-6730-442c-bc1a-4f20aabc21d0 · outbound

This paper cites SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.991495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.991495Z digest=sha256:daff88b13adaed04a051596f97bb081843dd59bf67dad58a306b8557071f79a6

Observation 28c22c42-0dcb-4e72-a4ef-d3c17236d7a5 · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:21.000707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:21.000707Z digest=sha256:78d06e79501966808d9d04dfe41c01865eedebb697ce2383d40008f2b8642b94

Observation aa62ad69-0af6-4ff4-ba72-3673f29240f7 · outbound

This paper cites ClimateBert: A Pretrained Language Model for Climate-Related Text.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs ClimateBert: A Pretrained Language Model for Climate-Related Text

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:21.012346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:21.012346Z digest=sha256:9ef6db973bb8e3a87e015c57c70bbb52a0c7f096b6a740f9abacc3f6f5b477b5

Observation ee202403-8be4-410d-aa55-a0fe69b11eba · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Chain-of-thought prompting elicits reasoning in large language models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:21.021049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:21.021049Z digest=sha256:5531de2c274cd544d478c438c7275daec685452cfbd0a6d815880d931f597b92

Observation 882da276-877c-4230-8329-583e317e64c8 · outbound

This paper cites Measuring and Reducing LLM Hallucination without Gold-Standard Answers.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Measuring and Reducing LLM Hallucination without Gold-Standard Answers

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:21.033271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:21.033271Z digest=sha256:b762da9da81af330a72d85a0ee790a97937a8dd9b2d471f4dc2de24c4f7e39d4

Observation a3fa1203-58fb-4250-8f7a-2351f2c302e5 · outbound

This paper cites Generate-on-Graph: Treat LLM as both Agent and KG in Incomplete Knowledge Graph Question Answering.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Generate-on-Graph: Treat LLM as both Agent and KG in Incomplete Knowledge Graph Question Answering

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:21.043894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:21.043894Z digest=sha256:b2566b62688e73b0c97e866df54cd285b6ab8953b07fa983c1edd497add8dae4

Observation 0a166818-bfdd-4da6-9091-d13638846a20 · outbound

This paper cites The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:21.055953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:21.055953Z digest=sha256:019b5e9b53682b81fddcab810e3e4a19d2b2aa9ad2d867155104d7f27dd803ad

Observation b950559e-db08-45cb-9739-4e1e6c3a1833 · outbound

This paper cites Qwen2.5 Technical Report.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Qwen2.5 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:21.069062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:21.069062Z digest=sha256:45337527318e362244f2ebcc5ffbc3e26d6b32d7f004521fd197f380674cf96d

Observation 066a04ea-38ca-43f4-b377-dab3c81ac489 · outbound

This paper cites Moose-chem: Large language models for rediscovering unseen chemistry scientific hypotheses.arXiv preprint arXiv:2410.07076, 2024.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Moose-chem: Large language models for rediscovering unseen chemistry scientific hypotheses.arXiv preprint arXiv:2410.07076, 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:21.082662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:21.082662Z digest=sha256:4174fce130c636fb731436d8c11b1c00b9d56db785a71a3a46ef32755e103dfc

Observation f19a0ba4-1cee-4e09-b7f4-adc0898306ea · outbound

This paper cites EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:21.090064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:21.090064Z digest=sha256:d6c441dc06c599152df4a25a4ba19b941e6fb9efd7f9598c53a1e3cca491d2d4

Observation e5745806-da5d-4446-a65c-17f1f2f9bb1a · outbound

This paper cites ChemLLM: A Chemical Large Language Model.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs ChemLLM: A Chemical Large Language Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:21.095513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:21.095513Z digest=sha256:e0d66ceade4acf1b3b0b33595979131605d76671cd5f07b0b6b2cf02d9df295c

Observation 315d0c97-3918-4b71-8097-c69c1a985ad6 · outbound

This paper cites Towards LLM-based Fact Verification on News Claims with a Hierarchical Step-by-Step Prompting Method.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Towards LLM-based Fact Verification on News Claims with a Hierarchical Step-by-Step Prompting Method

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:21.101708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:21.101708Z digest=sha256:e276eb3cd4d7b809c787cfb629f7509e4af336d86b8affb64663d109ad153393

Observation de14b375-f8d0-4057-97a3-1c5fa38b8394 · outbound

This paper cites GeoGPT: Understanding and Processing Geospatial Tasks through An Autonomous GPT.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs GeoGPT: Understanding and Processing Geospatial Tasks through An Autonomous GPT

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:21.111443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:21.111443Z digest=sha256:6fa4b2d7ec48f3396a227a3dd90b8ce4b6b994aecb197a9c9484753349374205

Pith citing papers

Observation e1b52b5b-524f-4269-898e-da894a8fef87 · inbound

A Vision for Geo-Temporal Deep Research Systems: Towards Comprehensive, Transparent, and Reproducible Geo-Temporal Information Synthesis cites this paper.

A Vision for Geo-Temporal Deep Research Systems: Towards Comprehensive, Transparent, and Reproducible Geo-Temporal Information Synthesis EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:22:26.646516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:22:23.855247Z digest=sha256:bf30be2b9736b1d124e932104752ee592ac02367c01aec6abe60b112a79bcd46