Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T23:41:46.847782Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 2 inbound Pith citation observations for arXiv:2602.12984.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T23:41:46.847782Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-27T09:34:09.347912Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-03T11:28:04.326064Z
49 of 49 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 74473242-8dee-46df-bf85-2a44e35342cf · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents On Evaluation of Embodied Navigation Agents
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 950e1056-e214-47dd-8521-8f0845cb7eff · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents System card: Claude opus 4 & claude sonnet 4
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc4b5753-9acd-4630-802b-d7c3fcf6e0a1 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents Claude sonnet 4.5 system card.https://www.anthropic.com/claude-sonne t-4-5-system-card, 2025
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e40eec4-fa03-4bb0-bf45-28bb5653ea85 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents SciMaster: Towards General-Purpose Scientific AI Agents, Part I. X-Master as Foundation: Can We Lead on Humanity's Last Exam?
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e29a8cc-7b09-435b-94b2-a796665c3b17 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dc69148-fa01-465f-bef5-f6d5a7a2240b · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ca25276-d045-401e-a83c-c106fa18b3ee · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0663a81-39ac-4ee8-b7e8-d3f351d37055 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4976358a-5bc9-445a-bfb7-55cb644d8af7 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents Discoveryworld: A virtual environ- ment for developing and evaluating automated scientific discovery agents.Advances in Neural Information Processing Systems, 37:10088–10116, 2024
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a84e9cb5-7ffb-4bf3-afd3-a6f9e5cf63c3 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2188affb-e4be-4da2-934d-62c4c193b34b · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents AgentBench: Evaluating LLMs as Agents
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03a0ca64-dbeb-45e7-b0fd-21cb65174b20 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents Search-o1: Agentic Search-Enhanced Large Reasoning Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2adc9e63-4b74-4682-9bcc-435936b1d7d0 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents SciAgent: Tool-augmented Language Models for Scientific Reasoning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2644418-bd9b-4aeb-916e-a46e51f3f435 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents Learn to explain: Multimodal reasoning via thought chains for sciencequestionanswering
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9df05970-bd16-4978-bedd-dafad7ef3e9d · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng-Jie Ji, Vishnu Suresh, Ion Stoica, and Joseph E
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c320aa6-47e9-46bd-84c8-a3e4cfb63e72 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents GPT-4 Technical Report
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d12ad95-afbf-4ad3-8b9d-dc75de34e558 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13636209-cf9b-4ac8-aa01-4d39739d02c0 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents OpenAI GPT-5 System Card
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b829228-cf1d-4d46-8ee0-fac425ef68e2 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51aa645c-fae6-4fa5-a76c-97b7b2b589f8 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents A Survey of AI for Materials Science: Foundation Models, LLM Agents, Datasets, and Tools
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 822f1f30-8980-4930-92de-533dc8e382ca · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents SciPy 1.0--Fundamental Algorithms for Scientific Computing in Python
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06f3a84d-5ff8-4037-8741-988b642024a5 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents Gemini: A Family of Highly Capable Multimodal Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5da8c59e-ac14-4cad-88ec-70bf662a7fb7 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents From AI for science to agentic science: A survey on autonomous scientific discovery.CoRR, abs/2508.14111, 2025
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2328cb1-f3cd-4bd8-88e7-2a7dd8f8ff5f · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents AgentGym: Evaluating and training large language model-based agents across diverse environments
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18793b37-a88c-4d30-aeee-1d82db5fb1e6 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 043f1a4b-b753-4a0b-9e06-afe4d804ee16 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd321c8a-80c4-48f9-8e85-7e8113c85983 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents On the Tool Manipulation Capability of Open-source Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 521c8bd2-31d2-4ddf-bc5f-0bfa840077e2 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ac982ff-d9a9-42a2-b187-1956894abcec · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents Qwen3 Technical Report
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f315465b-7099-4a1d-8722-cbb6ef739bb4 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents ReAct: Synergizing Reasoning and Acting in Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ad184a2-420a-4a78-be22-8aa42cfb9da2 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents Wang, Peifeng Ruan, Donghan Yang, Tao Wang, Guanghua Xiao, Xin Liu, Carl Yang, Yang Xie, and Wenqi Shi
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a366daec-3d43-43b2-a6f9-54bebbd1f34a · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents debug-gym: A Text-Based Environment for Interactive Debugging
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3022e7e4-6289-4b75-b337-7e9140606382 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a2a993a-c2e0-4d3b-825b-f272da28d61e · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents URLhttps://arxiv.org/abs/24 06.12045
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20337007-c0ef-4811-a432-d37b4836ab93 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents Sciinstruct: a self-reflective instruction annotated dataset for training scientific language models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3a7ae48-876a-4096-9503-ade0391bc9f1 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5312309a-044c-49d6-bfc6-5dc8604930c6 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4e71bd1-acc7-45b5-8874-68bee738df54 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents Scientists’ first exam: Probing cognitive abilities of mllm via perception, understanding, and reasoning.arXiv preprint arXiv:2506.10521, 2025
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e2daef1-1a3a-45d7-9cdb-109302624488 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents WebArena: A Realistic Web Environment for Building Autonomous Agents
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df752ca1-612f-4c9f-8f17-37053c441d98 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents Unresolved cited work
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1a4de21-0e95-4915-a80c-9eeca4213bd3 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents •Standard Return:{’result’: main_value, ’metadata’: {...}}(e.g., units, status flags, data sources)
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5ff4846-6c47-45d5-a582-582625dc238e · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents •Type Hints:All tools must provide complete Python type hints for parameters and return values
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0a8b7f6-be47-432e-b27b-768bd2e50fac · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents •Query:Retrieve hierarchical facts/records from external resources or local indices and return normalized fields for downstream steps
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 128b0d44-45b7-4905-b6de-2b5d57aed577 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents must include X
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62b6cd3d-2f28-46bd-8a08-6aed599e8cc9 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents Unresolved cited work
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3848077d-eac6-431d-85a0-b01af5be0ee1 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents Unresolved cited work
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 039c8a1a-3706-4739-84ad-c9d4779a0ae3 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents Tool-round limit exceeded; stopped automatically
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ac35669-a6d2-4d1a-adbe-3287f9bab8c2 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents GPT-4 Technical Report
Reference 774
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23f30786-a418-42b2-961e-5a0b4ebefc36 · outbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents Unresolved cited work
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34e2cafd-a6ce-4171-a250-d665cac90c96 · inbound
AI scientists produce results without reasoning scientifically SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4a5d332b-b3ee-410f-93d2-9b9884877e8d · inbound
Benchmarking AI Agents for Addressing Scientific Challenges Across Scales SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.