Pith. sign in

Paper Citation Record · LEDGER

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing

As of 12 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2608.07437.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07437 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:39:03.692877Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved39
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 02ee2bb2-7fe8-4c06-b2ee-ce4641c13f46 · outbound

This paper cites Stability analysis of fluid flows using Lagrangian Perturbation Theory (LPT): application to the plane Couette flow.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Stability analysis of fluid flows using Lagrangian Perturbation Theory (LPT): application to the plane Couette flow

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.400189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.400189Z digest=sha256:3a0c35ef135a4e10cc1cce94ef0979ea7c56c08d7399bcbf5ed43153e7cce13f

Observation f153cbbe-5435-4fe6-bc70-a6848ac07d3b · outbound

This paper cites Data-Copilot: Bridging Billions of Data and Humans with Autonomous Workflow.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Data-Copilot: Bridging Billions of Data and Humans with Autonomous Workflow

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.405648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.405648Z digest=sha256:222be2eda4131c356b7f4f147b96fee18c6cd512b39c8746e46d21f94c1bf467

Observation 0d367f4e-3feb-4514-9909-6a82962ea663 · outbound

This paper cites Deepanalyze: Agentic large language models for autonomous data science.arXiv preprint arXiv:2510.16872, 2025.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Deepanalyze: Agentic large language models for autonomous data science.arXiv preprint arXiv:2510.16872, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.410767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.410767Z digest=sha256:8201f298a9929a8c0296274279d6afbe3219680fce96908177f0c70bb5dddc26

Observation bb3088c0-b3fa-4f21-b3ac-6010955f896f · outbound

This paper cites Data interpreter: An llm agent for data science.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Data interpreter: An llm agent for data science

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.752551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T04:39:03.415655Z digest=sha256:ec7b9808cbffe5266f7c1bbb78bd2f88724e52eba7d7aef1eab6ba53cd3420da

Observation 3688e365-d990-471c-8239-cf1a285070ff · outbound

This paper cites Infiagent-dabench: Evaluating agents on data analysis tasks,.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Infiagent-dabench: Evaluating agents on data analysis tasks,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.736770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T04:39:03.420634Z digest=sha256:bd52a57865afc8d8e52a0f3ece78ed3ebd23063f3e45e7f524dcbbdb4ff3a12d

Observation 644faab7-ad14-4c37-b046-38a7b75ab743 · outbound

This paper cites DABstep: Data Agent Benchmark for Multi-step Reasoning.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing DABstep: Data Agent Benchmark for Multi-step Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.430946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.430946Z digest=sha256:7ca63c3c7b4b89dc2e5e52d6fa76afca87c1ab944cf551b9d4c634ff78d471bc

Observation d847d5b0-9d18-44da-930c-7c95e024f0b1 · outbound

This paper cites DA-code: Agent data science code gen- eration benchmark for large language models.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing DA-code: Agent data science code gen- eration benchmark for large language models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.436223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.436223Z digest=sha256:7248dd643b22a9ecc83dac5184b5f4db279a8aa63315fde0192092c1c683c495

Observation 0a10aadc-bd13-43fb-ab13-3168d2574294 · outbound

This paper cites Are large language models good statisticians?Advances in Neural Information Processing Systems, 37:62697–62731, 2024.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Are large language models good statisticians?Advances in Neural Information Processing Systems, 37:62697–62731, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.720970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T04:39:03.441812Z digest=sha256:363e13117e803a695befe446abdbf8e43a0f9ae2aee4b073105b5d09c78c9a22

Observation 9ff204ad-749c-48ef-a97b-da696b13364a · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing ReAct: Synergizing Reasoning and Acting in Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.446928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.446928Z digest=sha256:e9ee818cadc08d14713ab35785b05295c3b0d96e5c2397a95c0591d15c9a183a

Observation 4ebded3c-7ce2-497e-a8ad-eeff6bf9b079 · outbound

This paper cites Ds-1000: A natural and reliable benchmark for 12 data science code generation.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Ds-1000: A natural and reliable benchmark for 12 data science code generation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.704281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T04:39:03.452234Z digest=sha256:a375670936bfa9418582e87a92b873a585d8e3d12ed2b08aa1216b45a11108e1

Observation c9f45c4b-e365-44ce-94b6-d9d21dcce73d · outbound

This paper cites InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.457120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.457120Z digest=sha256:4534596b58fc44225a3bbc24ffa53e4497920bd3f677f35f79e1308ed7df1321

Observation b7b4fe64-9111-4aac-8e71-5a507dd9ac22 · outbound

This paper cites DataSciBench: An LLM Agent Benchmark for Data Science.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing DataSciBench: An LLM Agent Benchmark for Data Science

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.461634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.461634Z digest=sha256:a5358d8249ebe3023da27568659caefe6c6e20bba17d427171400382ec5b102a

Observation c1557f8b-cbf6-4f32-a05c-31576141a4ab · outbound

This paper cites MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.466801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.466801Z digest=sha256:7f797671a4dad4ecba8e97ae97a855cfdf95ab048f58dea216b5ec3851b21a60

Observation 9a176489-9d48-4993-bd21-04ab319deee1 · outbound

This paper cites Tapilot-Crossing: Benchmarking and Evolving LLMs Towards Interactive Data Analysis Agents.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Tapilot-Crossing: Benchmarking and Evolving LLMs Towards Interactive Data Analysis Agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.471639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.471639Z digest=sha256:5cdc0640f01a7ae890052bf9782ea999b4748731bd1bdcf70e808b6395f5097e

Observation 8837324d-564e-4c35-8c3a-42f6175dc1af · outbound

This paper cites IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.476630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.476630Z digest=sha256:83c8052b1f9e011ed021285c5554b3228e25d42d3bbbba6e339e1dce1591a713

Observation 30637093-fa4e-4b87-9771-5515245c4045 · outbound

This paper cites Fact or fiction: Verifying scientific claims.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Fact or fiction: Verifying scientific claims

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.481558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.481558Z digest=sha256:45d4ad8cda9eb0e110b4b925a799911d22427b6564bf41201c4e0ce987b1a95b

Observation d8348684-9385-4de9-a54d-6444071f1631 · outbound

This paper cites Sciclaimhunt: A large dataset for evidence-based scientific claim verification.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Sciclaimhunt: A large dataset for evidence-based scientific claim verification

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.678309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T04:39:03.486856Z digest=sha256:30204638e5c19e9f91b4bd060bc5b85c68e5e28819fea589d29c9f126af5f642

Observation 967c7018-a6b1-49aa-9c25-6282fef5ebaf · outbound

This paper cites Musciclaims: Multimodal scientific claim verifi- cation.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Musciclaims: Multimodal scientific claim verifi- cation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.663704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T04:39:03.491749Z digest=sha256:c89dbffa8f1f48ab30f86d5ae2f4a2040ab2c6a7d1e3a1771bbcfff31bdc1d37

Observation 75554491-30c5-42cb-bed5-39805595e3d7 · outbound

This paper cites Investi- gating the reproducibility of the social and behavioural sciences.Nature, 652(8108):126–134, 2026.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Investi- gating the reproducibility of the social and behavioural sciences.Nature, 652(8108):126–134, 2026

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.648515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T04:39:03.496445Z digest=sha256:f51ad09a73b8fcb563c8a48322d63c77e0c13d441ac0dfeb6dc4084fb69e4a79

Observation 2e9418d9-cd96-49f7-b8f7-bbdd9163863e · outbound

This paper cites Towards end-to-end automation of ai research.Nature, 651(8107):914–919, 2026.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Towards end-to-end automation of ai research.Nature, 651(8107):914–919, 2026

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.500807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.500807Z digest=sha256:bb1713b0b79121fa155e61140e68e293e1b2bd4c29c4008303f1f5f114f0b250

Observation bb557469-8e70-4498-822e-5276a4c060ee · outbound

This paper cites AI-Researcher: Autonomous Scientific Innovation.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing AI-Researcher: Autonomous Scientific Innovation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.505323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.505323Z digest=sha256:b68c9ecdc27955aa73700476e52fc3fdb0b77500015677b922311e6396697298

Observation 2e96c553-6fcf-4e65-aa12-39d281943fde · outbound

This paper cites The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.509944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.509944Z digest=sha256:527b7aaa6783c7a79c91eec53e37bab420ea82377196a56d0a3e5091c35f66ff

Observation 7097c805-7d63-451c-8122-a5b98befa13a · outbound

This paper cites Development economics field experiments (dfeep).

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Development economics field experiments (dfeep)

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.620868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T04:39:03.514880Z digest=sha256:612c1e1b92bbd3881eaba0d4a192aca31d6a7917d31f210b67eb7c59ed0d25b6

Observation 7af7ec38-d793-45e9-ade6-72a95d998e44 · outbound

This paper cites The cbio cancer genomics portal: an open platform for exploring multidimensional cancer genomics data.Cancer discovery, 2(5):401–404, 2012.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing The cbio cancer genomics portal: an open platform for exploring multidimensional cancer genomics data.Cancer discovery, 2(5):401–404, 2012

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.605005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T04:39:03.519472Z digest=sha256:78811a1a5a741f58083b3b41eb38900e22befdc0500add8caf289f7d533ac403

Observation ba341f71-2d39-4448-8fd2-9d69c49fc373 · outbound

This paper cites Integrative analysis of complex cancer genomics and clinical profiles using the cbioportal.Science signal- ing, 6(269):pl1–pl1, 2013.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Integrative analysis of complex cancer genomics and clinical profiles using the cbioportal.Science signal- ing, 6(269):pl1–pl1, 2013

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.588856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T04:39:03.524397Z digest=sha256:510f47b72421195482eee735dcc15675978c669b2c0da9127f661fbb1f014ecd

Observation 5b5ffc14-a414-42a7-b78d-11cf301e6476 · outbound

This paper cites BioDSA-1K: Benchmarking Data Science Agents for Biomedical Research.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing BioDSA-1K: Benchmarking Data Science Agents for Biomedical Research

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.528813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.528813Z digest=sha256:9164e8736e513acd120b6673a650c903ba1304ac79c9f09b091b6e106d0e9ef7

Observation 06e28d36-989a-43de-885b-022acebdf441 · outbound

This paper cites an unresolved cited work.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-10T04:39:04.573011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T04:39:03.533331Z digest=sha256:5e5dbfd9c6c7aa6d8e1ed5820bfcc60886fe977f82aaae94df0383b64d829351

Observation 3509a304-b336-450a-a2e1-85062991a107 · outbound

This paper cites Introducing claude sonnet 4.6.https://www.anthropic.com/news/ claude-sonnet-4-6, February 2026.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Introducing claude sonnet 4.6.https://www.anthropic.com/news/ claude-sonnet-4-6, February 2026

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.556765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T04:39:03.537708Z digest=sha256:c4f50a36d2fc24e786771cbb87846eddb20a6332e9cc74fdb098f7b44560dbe8

Observation d71be0a8-203a-4c60-a02d-d6a67a1c3ce9 · outbound

This paper cites Executable code actions elicit better llm agents.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Executable code actions elicit better llm agents

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.542160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.542160Z digest=sha256:a98b7d6a6eccf671caf69436b92e21eb39948a23c91c9221058edf64960ccd89

Observation 50d4aaaa-401c-49ae-be49-6e26115c50f2 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.546532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.546532Z digest=sha256:87c69d2110d4e783a9a04c763a07e8ef0d8cc0819d68e5a175945e2cf60db468

Observation da19215c-905d-4a00-b296-589b5f48acdf · outbound

This paper cites Qwen2.5-Coder Technical Report.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Qwen2.5-Coder Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.550965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.550965Z digest=sha256:823af2fb9b771a010da497cb8265ca5cf10a7577cfeb8e6d4b61db56b428b2b6

Observation 26ec706c-2bf8-4b48-8181-2fbbf363f3cb · outbound

This paper cites Introducing gpt-5.4.https://openai.com/index/introducing-gpt-5-4/, March.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Introducing gpt-5.4.https://openai.com/index/introducing-gpt-5-4/, March

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.555768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.555768Z digest=sha256:34e7b90786e32519336e4dca9e4d73c58ecb70086b1ce9e3d793dc61029260b2

Observation 71e73ab9-e116-4aed-86a6-0400ae6b77c6 · outbound

This paper cites Deepseek-v4: Towards highly efficient million-token context intelligence.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Deepseek-v4: Towards highly efficient million-token context intelligence

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.512901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T04:39:03.564690Z digest=sha256:10f9a65220063df02f5a02da67759ecedd280a331c3fefd418f00b9e6cc64c65

Observation 8c37b627-0242-4fa6-a6dc-129a9b438cbd · outbound

This paper cites Introducing gpt-oss.https://openai.com/index/introducing-gpt-oss/, 2025.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Introducing gpt-oss.https://openai.com/index/introducing-gpt-oss/, 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.496607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T04:39:03.569100Z digest=sha256:304b71b140a7e4106c7665e69509578f9506fd2b2cda6d63fd6d53796b68549d

Observation 6402fc39-cc60-48fe-9bb9-27fbc343b7b8 · outbound

This paper cites Qwen3-coder-30b-a3b-instruct.https://huggingface.co/Qwen/ Qwen3-Coder-30B-A3B-Instruct, 2025.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Qwen3-coder-30b-a3b-instruct.https://huggingface.co/Qwen/ Qwen3-Coder-30B-A3B-Instruct, 2025

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.480535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T04:39:03.573317Z digest=sha256:03c8e5d507f8445947afb6ec4780bc3d690910cd184b65410728c781d8e8f517

Observation df0caddc-0ea6-49f8-8c9c-80b35a7704a1 · outbound

This paper cites Qwen3 Technical Report.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Qwen3 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.577579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.577579Z digest=sha256:2112c424dc74cc16793bc7436e0247fa1d4f384083201e79097d2fa4d4961ccf

Observation 38312006-f686-462e-ae24-0e77b00ab9de · outbound

This paper cites Scaling generalist data- analytic agents, 2026.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Scaling generalist data- analytic agents, 2026

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.582193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.582193Z digest=sha256:21166646652f998ed74e57ace679a1d46f2769b97b55c2e1bdb4d5ba8d02d817

Observation c8b03bf1-d1ee-4461-a21f-d5d7b661b266 · outbound

This paper cites Accessed: 2026-05-06.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Accessed: 2026-05-06

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.465596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T04:39:03.586686Z digest=sha256:03ea881e51f69c8a6a674eb7e3ef02da6f64ffb6d56bb0f24eb8db9d936b1d81

Observation e7de4fcd-e169-406b-91ea-3254483e4b57 · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Self-Refine: Iterative Refinement with Self-Feedback

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.591340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.591340Z digest=sha256:2f9640c68b6066251873a7390c7e97d1a950fa228fd87c629f3abc2f88d7370b

Observation c44e2879-0050-436a-96b8-0081c40c9409 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.Advances in neural informa- tion processing systems, 36:8634–8652, 2023.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Reflexion: Language agents with verbal reinforcement learning.Advances in neural informa- tion processing systems, 36:8634–8652, 2023

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.450651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T04:39:03.596337Z digest=sha256:07bc654857e06599d87f31b88e55496c3d75815b31d4e15fe939e4c11706b325

Observation d7ab10fe-dee9-4e4b-be02-276bb43b150b · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.600566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.600566Z digest=sha256:fabcf3c31ca73da56c764c9acaa6f19aa85935bc5f0b29923a80a3635ea0261f

Observation c65ef131-db0a-42f3-ae75-2e46db8a5d69 · outbound

This paper cites Swe-agent: Agent-computer interfaces enable automated soft- ware engineering.Advances in Neural Information Processing Systems, 37:50528–50652, 2024.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Swe-agent: Agent-computer interfaces enable automated soft- ware engineering.Advances in Neural Information Processing Systems, 37:50528–50652, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.434898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T04:39:03.605328Z digest=sha256:88b64407351e394b2eafc11c3a3811a66c5d647c85f862f86a31c78a062643d3

Observation 000e14c4-7303-4d23-bcd9-1596e22aecd7 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Proximal Policy Optimization Algorithms

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.609725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.609725Z digest=sha256:350576a96e98128a04f1d315bf2fea21ed4c575f124ae40784f912f5c269c207

Observation c488ede0-f512-4ae6-9050-0a36aa648941 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.614383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.614383Z digest=sha256:afcbabbc3b1ac3235dfbc0fa16d43fb08647aeab77f53fde1adc113c0f05bf4c

Observation aaa26ca6-40be-4162-acf0-908c80fce4d3 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.Advances in neural information processing systems, 36:53728–53741, 2023.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Direct preference optimization: Your language model is secretly a reward model.Advances in neural information processing systems, 36:53728–53741, 2023

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.618872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.618872Z digest=sha256:1b65b9198ff92f7e0294d9b1ca54831ec736aa508b748d9d55303f8673dce04c

Observation ba795ca9-c6bf-41fb-af74-c838cfb26782 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.623408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.623408Z digest=sha256:9e55e53b7e3fad6028e1213e64f8ae6fa6fecbedb36a113349361fae064ef713

Observation 53f52013-4298-4d13-9832-06e4463ba865 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.628209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.628209Z digest=sha256:b865f46de14835dd175355844ab1250229b2d057627b7df8a93285baff45a97c

Observation 3c61b79c-030c-42a0-88ff-cbac19d0b4fb · outbound

This paper cites ToRL: Scaling Tool-Integrated RL.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing ToRL: Scaling Tool-Integrated RL

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.633077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.633077Z digest=sha256:af50d59fe216e080dcbc03829038ed7abb11a27ed825ade710146d2bdd3dd850

Observation e6ce7ab3-ca53-44a0-960d-951f42109cd4 · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.637815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.637815Z digest=sha256:502f8dd307070dde5e397024f7581d9e14e51e48ff1d5d9abfe925ba8ebf45f8

Observation 3a6b4ba9-845b-4056-be59-28eb4a5da243 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing ToolRL: Reward is All Tool Learning Needs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.642514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.642514Z digest=sha256:85ae4ba832095bfcf73c92db381a5af6c1967b33b0b78807df09c2a7650b56c3

Observation ca4ee6f6-5d92-472b-95e6-2c5810b604f1 · outbound

This paper cites Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.647525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.647525Z digest=sha256:23318d2283094a68085777426bead3b61d5162805c6407c3bfb00f8791574245

Observation 65e2fbaa-1e39-42b7-8721-101cae37630c · outbound

This paper cites Effects of cognitive behavioral therapy and cash transfers on older persons living alone in india: a randomized trial.Annals of internal medicine, 176(5):632–641, 2023.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Effects of cognitive behavioral therapy and cash transfers on older persons living alone in india: a randomized trial.Annals of internal medicine, 176(5):632–641, 2023

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.399794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T04:39:03.652822Z digest=sha256:9908a416b17795475486fd13e93753b8399fd2a9faf522dad52c245a2b4cf006

Observation 5b30023b-5725-4082-80ba-ac74832d7b4f · outbound

This paper cites Genomic characterization of metastatic patterns from prospective clinical sequenc- ing of 25,000 patients.Cell, 185(3):563–575, 2022.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Genomic characterization of metastatic patterns from prospective clinical sequenc- ing of 25,000 patients.Cell, 185(3):563–575, 2022

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.384842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T04:39:03.657534Z digest=sha256:23ba880da2f59374e240620fc311a3f8d53918d8818eff1b96e049e1ea063458

Observation 5254a881-e8a1-488e-9cd5-2826d602f635 · outbound

This paper cites The support prognostic model: Objective estimates of survival for seriously ill hospitalized adults.Annals of internal medicine, 122(3):191–203, 1995.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing The support prognostic model: Objective estimates of survival for seriously ill hospitalized adults.Annals of internal medicine, 122(3):191–203, 1995

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.369442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T04:39:03.662098Z digest=sha256:a227698792cb8b064234f6ad2559917d48ab30fac24e6cf5bb1455d539601e9a

Observation f15cf1cd-b341-43da-b393-5e0ccc2186a6 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.667845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.667845Z digest=sha256:e9f4a302c556210212ca1705e7ee46fe0bb6a6a73788f2e1552fcdd94d97bdba

Observation 8c18f6e8-df20-41d5-b0d1-6c30635501c6 · outbound

This paper cites claim”: “drug improves patient outcome.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing claim”: “drug improves patient outcome

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.353710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T04:39:03.672908Z digest=sha256:3c5c138cad9d3d1816ba2582c02f6be541f3b0cd0b1db4b6e3d87f11f0ad2fb4

Observation adbe0b92-7e51-427f-8525-b05ec970620f · outbound

This paper cites an unresolved cited work.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-10T04:39:04.336736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T04:39:03.678589Z digest=sha256:368a9210fd31e0fdf42439232f550acbc876e1adf318d7180407d0e6dbc467ca

Observation ddf3d3cd-3bf0-42c1-b65d-4b12483816bc · outbound

This paper cites an unresolved cited work.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-10T04:39:04.322049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T04:39:03.683802Z digest=sha256:ae85a25540cb597d7f568c710453b1a2ffbebe2a829dcbc44ee2694841b03f35

Observation 93bd0286-2b0f-47e0-8e6b-13789a89d806 · outbound

This paper cites an unresolved cited work.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-10T04:39:04.306832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T04:39:03.688424Z digest=sha256:497a6c87d899ba62400e55be60b829bc895f69d29205f90430441291d3318b7e

Observation 507ce3fc-5327-40f6-b266-2abf02d2f7e1 · outbound

This paper cites an unresolved cited work.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-10T04:39:04.290784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T04:39:03.692877Z digest=sha256:edc8ecf36f49a98914c9381cdb4e68f3b6f1c2ba7ebbc8210fa4dd3133c24df5

Observation 163b0ec8-d481-4081-8989-a5e184270d91 · outbound

This paper cites an unresolved cited work.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Unresolved cited work

Reference 2026

Resolution
parse uncertain
no resolver link, observed 2026-08-10T04:39:03.560376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.560376Z digest=sha256:5e6940ef98d2fcae1ba75436c66f2e8ab4c55b73d2e1406ef1ae1dfe028ed0f6

Pith citing papers

No inbound Pith citation observations are available.