Pith. sign in

Paper Citation Record · LEDGER

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing

As of 11 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2608.07437.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07437 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:39:03.692877Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved39
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 02ee2bb2-7fe8-4c06-b2ee-ce4641c13f46 · outbound

This paper cites Stability analysis of fluid flows using Lagrangian Perturbation Theory (LPT): application to the plane Couette flow.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Stability analysis of fluid flows using Lagrangian Perturbation Theory (LPT): application to the plane Couette flow

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.400189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.400189Z digest=sha256:a7cf3b564cdb7714f2745193e6af7b5ae00dba821f0aa8bfeda0158a445a9c9f

Observation f153cbbe-5435-4fe6-bc70-a6848ac07d3b · outbound

This paper cites Data-Copilot: Bridging Billions of Data and Humans with Autonomous Workflow.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Data-Copilot: Bridging Billions of Data and Humans with Autonomous Workflow

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.405648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.405648Z digest=sha256:97ea4f6333c5b84f7a9f7ccd758f0c7e0e279a9aa54b845060ed9e547c948c38

Observation 0d367f4e-3feb-4514-9909-6a82962ea663 · outbound

This paper cites Deepanalyze: Agentic large language models for autonomous data science.arXiv preprint arXiv:2510.16872, 2025.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Deepanalyze: Agentic large language models for autonomous data science.arXiv preprint arXiv:2510.16872, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.410767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.410767Z digest=sha256:a143f3658106e5c07889408c135e80dd4983968f85198725ccaa6dcd94c599a9

Observation bb3088c0-b3fa-4f21-b3ac-6010955f896f · outbound

This paper cites Data interpreter: An llm agent for data science.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Data interpreter: An llm agent for data science

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.752551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T04:39:03.415655Z digest=sha256:a2a75be73de93f3fab3ddfa1c95887a0e53c0b5bf7c10212984ed2e6164d6dc3

Observation 3688e365-d990-471c-8239-cf1a285070ff · outbound

This paper cites Infiagent-dabench: Evaluating agents on data analysis tasks,.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Infiagent-dabench: Evaluating agents on data analysis tasks,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.736770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T04:39:03.420634Z digest=sha256:8527afd39b19e0968ba442bb6063dcc6e9e0fa214f815992ff57f16f0745f6af

Observation 644faab7-ad14-4c37-b046-38a7b75ab743 · outbound

This paper cites DABstep: Data Agent Benchmark for Multi-step Reasoning.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing DABstep: Data Agent Benchmark for Multi-step Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.430946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.430946Z digest=sha256:082a3b34d592b465e575d70e64b42fe0c7420e2350947c1de2ccd59eca9c9989

Observation d847d5b0-9d18-44da-930c-7c95e024f0b1 · outbound

This paper cites DA-code: Agent data science code gen- eration benchmark for large language models.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing DA-code: Agent data science code gen- eration benchmark for large language models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.436223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.436223Z digest=sha256:51e58300ac0732b999f45470eb9ec7755b1d173137fad4b9cd06688c9ed8a664

Observation 0a10aadc-bd13-43fb-ab13-3168d2574294 · outbound

This paper cites Are large language models good statisticians?Advances in Neural Information Processing Systems, 37:62697–62731, 2024.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Are large language models good statisticians?Advances in Neural Information Processing Systems, 37:62697–62731, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.720970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T04:39:03.441812Z digest=sha256:6ad6bb45abbf790e0e00531d2bd2d280b068f80c72e416c33b7af1ffa1d97f93

Observation 9ff204ad-749c-48ef-a97b-da696b13364a · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing ReAct: Synergizing Reasoning and Acting in Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.446928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.446928Z digest=sha256:c22436d7159bf6085cc3246e9c0e4a3277a611da390717a32aeade555461de1e

Observation 4ebded3c-7ce2-497e-a8ad-eeff6bf9b079 · outbound

This paper cites Ds-1000: A natural and reliable benchmark for 12 data science code generation.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Ds-1000: A natural and reliable benchmark for 12 data science code generation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.704281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T04:39:03.452234Z digest=sha256:0c23c920c4ad96ee592c42a786fccb49419d9d487df2724e14b59422797970c4

Observation c9f45c4b-e365-44ce-94b6-d9d21dcce73d · outbound

This paper cites InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.457120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.457120Z digest=sha256:8545992cd8ab35392689b38dee10466c35b2f00969d7881a63e560c14f80d731

Observation b7b4fe64-9111-4aac-8e71-5a507dd9ac22 · outbound

This paper cites DataSciBench: An LLM Agent Benchmark for Data Science.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing DataSciBench: An LLM Agent Benchmark for Data Science

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.461634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.461634Z digest=sha256:bd09821f4acacac58bc58b563d81a4dc6f52507580e3f7bb271d19cf1a0f284d

Observation c1557f8b-cbf6-4f32-a05c-31576141a4ab · outbound

This paper cites MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.466801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.466801Z digest=sha256:a4ef3cbe51c2da7559568c14334eee2eadff3ace15aa257fbaea603c0212300e

Observation 9a176489-9d48-4993-bd21-04ab319deee1 · outbound

This paper cites Tapilot-Crossing: Benchmarking and Evolving LLMs Towards Interactive Data Analysis Agents.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Tapilot-Crossing: Benchmarking and Evolving LLMs Towards Interactive Data Analysis Agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.471639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.471639Z digest=sha256:2a410b5a536ee7bdf2ea810c3b7dd032ddc9b1860ab204c6eb7453ab05b84690

Observation 8837324d-564e-4c35-8c3a-42f6175dc1af · outbound

This paper cites IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.476630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.476630Z digest=sha256:1d5812782a357815e46136aac826b2cea64ccfa264590d905ed13dccf1d74b9e

Observation 30637093-fa4e-4b87-9771-5515245c4045 · outbound

This paper cites Fact or fiction: Verifying scientific claims.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Fact or fiction: Verifying scientific claims

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.481558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.481558Z digest=sha256:1aa18e6982b6d9607f2901386dc81fb4c5a28d1126c5b97861b9f3c932c55e17

Observation d8348684-9385-4de9-a54d-6444071f1631 · outbound

This paper cites Sciclaimhunt: A large dataset for evidence-based scientific claim verification.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Sciclaimhunt: A large dataset for evidence-based scientific claim verification

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.678309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T04:39:03.486856Z digest=sha256:c3c088e7b78b84c388a19da27a2641bd0d4495e9ddc146a79845f46fe7320f9c

Observation 967c7018-a6b1-49aa-9c25-6282fef5ebaf · outbound

This paper cites Musciclaims: Multimodal scientific claim verifi- cation.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Musciclaims: Multimodal scientific claim verifi- cation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.663704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T04:39:03.491749Z digest=sha256:e7e7c7c4f965f8f55bb6232ae10c3f8b352cc550e6da7371a80f9ccee1a5ba94

Observation 75554491-30c5-42cb-bed5-39805595e3d7 · outbound

This paper cites Investi- gating the reproducibility of the social and behavioural sciences.Nature, 652(8108):126–134, 2026.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Investi- gating the reproducibility of the social and behavioural sciences.Nature, 652(8108):126–134, 2026

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.648515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T04:39:03.496445Z digest=sha256:2b6d458c25ecf029115a75b2294a0a2443b01359ff210e4cdfda7a6a21da088a

Observation 2e9418d9-cd96-49f7-b8f7-bbdd9163863e · outbound

This paper cites Towards end-to-end automation of ai research.Nature, 651(8107):914–919, 2026.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Towards end-to-end automation of ai research.Nature, 651(8107):914–919, 2026

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.500807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.500807Z digest=sha256:96734e0c9b32a61f83ce4b62248fc7aef3228294607f28abea721fc2cceb9282

Observation bb557469-8e70-4498-822e-5276a4c060ee · outbound

This paper cites AI-Researcher: Autonomous Scientific Innovation.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing AI-Researcher: Autonomous Scientific Innovation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.505323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.505323Z digest=sha256:59a78317102f7bb8295f33644f5a1232ce708f030e789d2a5f9c6ea39885769d

Observation 2e96c553-6fcf-4e65-aa12-39d281943fde · outbound

This paper cites The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.509944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.509944Z digest=sha256:afae0f125cc273cebcaa39f1450cfd5b3fbf6514793489304d3878ac7fae47f6

Observation 7097c805-7d63-451c-8122-a5b98befa13a · outbound

This paper cites Development economics field experiments (dfeep).

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Development economics field experiments (dfeep)

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.620868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T04:39:03.514880Z digest=sha256:9fd8d39756d6fe76832bdac77e5cc40a008abd0658e0fa22f144e93ecd996a58

Observation 7af7ec38-d793-45e9-ade6-72a95d998e44 · outbound

This paper cites The cbio cancer genomics portal: an open platform for exploring multidimensional cancer genomics data.Cancer discovery, 2(5):401–404, 2012.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing The cbio cancer genomics portal: an open platform for exploring multidimensional cancer genomics data.Cancer discovery, 2(5):401–404, 2012

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.605005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T04:39:03.519472Z digest=sha256:f073e9d7e1f429b46c888025003033af1df0ea6d7e432903bd179082ec4a18f2

Observation ba341f71-2d39-4448-8fd2-9d69c49fc373 · outbound

This paper cites Integrative analysis of complex cancer genomics and clinical profiles using the cbioportal.Science signal- ing, 6(269):pl1–pl1, 2013.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Integrative analysis of complex cancer genomics and clinical profiles using the cbioportal.Science signal- ing, 6(269):pl1–pl1, 2013

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.588856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T04:39:03.524397Z digest=sha256:69d02a93d6de4374bfc9ef90d060011aac3e71f865423664dfe66d27f91b8f72

Observation 5b5ffc14-a414-42a7-b78d-11cf301e6476 · outbound

This paper cites BioDSA-1K: Benchmarking Data Science Agents for Biomedical Research.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing BioDSA-1K: Benchmarking Data Science Agents for Biomedical Research

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.528813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.528813Z digest=sha256:dd412a38fc1de51db47950fa4e375f11536298230bf92944a1c357adbcb0f206

Observation 06e28d36-989a-43de-885b-022acebdf441 · outbound

This paper cites an unresolved cited work.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-10T04:39:04.573011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T04:39:03.533331Z digest=sha256:ce60b7bd19613f36c3a51440cb42a89dae9b5009c831b8fc86788e16fef513b4

Observation 3509a304-b336-450a-a2e1-85062991a107 · outbound

This paper cites Introducing claude sonnet 4.6.https://www.anthropic.com/news/ claude-sonnet-4-6, February 2026.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Introducing claude sonnet 4.6.https://www.anthropic.com/news/ claude-sonnet-4-6, February 2026

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.556765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T04:39:03.537708Z digest=sha256:01e9fc05cb338ead559931b9e7ac6829d10476e57a73fa236a665fc974dbdfeb

Observation d71be0a8-203a-4c60-a02d-d6a67a1c3ce9 · outbound

This paper cites Executable code actions elicit better llm agents.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Executable code actions elicit better llm agents

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.542160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.542160Z digest=sha256:37a51b9a2e99a0cf675c70efc3852955b9b4c476ab977680511dd49a9e0bb0b8

Observation 50d4aaaa-401c-49ae-be49-6e26115c50f2 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.546532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.546532Z digest=sha256:247cc6b7d4341a825b750174e83cd8ce6301a447a0407cfbae138b8ff9a825e6

Observation da19215c-905d-4a00-b296-589b5f48acdf · outbound

This paper cites Qwen2.5-Coder Technical Report.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Qwen2.5-Coder Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.550965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.550965Z digest=sha256:329ecb8a42f44aa987eeb373cab77466d7477ab8072d5f20901bfae4bb45f86a

Observation 26ec706c-2bf8-4b48-8181-2fbbf363f3cb · outbound

This paper cites Introducing gpt-5.4.https://openai.com/index/introducing-gpt-5-4/, March.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Introducing gpt-5.4.https://openai.com/index/introducing-gpt-5-4/, March

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.555768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.555768Z digest=sha256:e49c1705fe7b899e52cf05c20672f92e91206be34a6ccee4935f90cff44189c6

Observation 71e73ab9-e116-4aed-86a6-0400ae6b77c6 · outbound

This paper cites Deepseek-v4: Towards highly efficient million-token context intelligence.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Deepseek-v4: Towards highly efficient million-token context intelligence

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.512901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T04:39:03.564690Z digest=sha256:1cc281078576a3d33b3d006f2c589427ee96eeaee511445cf5848d0ec9e09098

Observation 8c37b627-0242-4fa6-a6dc-129a9b438cbd · outbound

This paper cites Introducing gpt-oss.https://openai.com/index/introducing-gpt-oss/, 2025.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Introducing gpt-oss.https://openai.com/index/introducing-gpt-oss/, 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.496607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T04:39:03.569100Z digest=sha256:f7af65783a3253fb3b8c5f3de64c1cbe277e5725a2f90b4b223ceb445322522a

Observation 6402fc39-cc60-48fe-9bb9-27fbc343b7b8 · outbound

This paper cites Qwen3-coder-30b-a3b-instruct.https://huggingface.co/Qwen/ Qwen3-Coder-30B-A3B-Instruct, 2025.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Qwen3-coder-30b-a3b-instruct.https://huggingface.co/Qwen/ Qwen3-Coder-30B-A3B-Instruct, 2025

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.480535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T04:39:03.573317Z digest=sha256:e7b8348e729e8ba547b3dad0b1cb212e02b3b7b22dc822cb639900e8e3d2da55

Observation df0caddc-0ea6-49f8-8c9c-80b35a7704a1 · outbound

This paper cites Qwen3 Technical Report.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Qwen3 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.577579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.577579Z digest=sha256:4c232b4747ebc05acc94f96786ede1a0dcdb8c069407f46a93dd4507deb29b1e

Observation 38312006-f686-462e-ae24-0e77b00ab9de · outbound

This paper cites Scaling generalist data- analytic agents, 2026.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Scaling generalist data- analytic agents, 2026

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.582193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.582193Z digest=sha256:93ee1369c9bf28e955aa33d164d93258c927d645574276d4a0badf2f940889fb

Observation c8b03bf1-d1ee-4461-a21f-d5d7b661b266 · outbound

This paper cites Accessed: 2026-05-06.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Accessed: 2026-05-06

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.465596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T04:39:03.586686Z digest=sha256:78393cc9746904c4f7f935da2b54e70aa7ad570fc3bdb7032e650928bec6e742

Observation e7de4fcd-e169-406b-91ea-3254483e4b57 · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Self-Refine: Iterative Refinement with Self-Feedback

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.591340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.591340Z digest=sha256:1bbdf22bee209fbc99ee080ed9bde6d87b76388c866693d7e76bd2a2918e1ffe

Observation c44e2879-0050-436a-96b8-0081c40c9409 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.Advances in neural informa- tion processing systems, 36:8634–8652, 2023.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Reflexion: Language agents with verbal reinforcement learning.Advances in neural informa- tion processing systems, 36:8634–8652, 2023

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.450651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T04:39:03.596337Z digest=sha256:a67a7e75b2ce4011da84cd56eb04471690bd045ee8ee7e243c9979284dc6a253

Observation d7ab10fe-dee9-4e4b-be02-276bb43b150b · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.600566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.600566Z digest=sha256:40149495af92bacef633b44ecc8b9ae6b7877253d553db19be75c9706ed8f21e

Observation c65ef131-db0a-42f3-ae75-2e46db8a5d69 · outbound

This paper cites Swe-agent: Agent-computer interfaces enable automated soft- ware engineering.Advances in Neural Information Processing Systems, 37:50528–50652, 2024.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Swe-agent: Agent-computer interfaces enable automated soft- ware engineering.Advances in Neural Information Processing Systems, 37:50528–50652, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.434898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T04:39:03.605328Z digest=sha256:109bd854e0362319b642bd99b8f0d970efd58fd85bf1f10b5a8470eae200e27b

Observation 000e14c4-7303-4d23-bcd9-1596e22aecd7 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Proximal Policy Optimization Algorithms

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.609725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.609725Z digest=sha256:992b08f2282934bd76289a5ad0b9f6518fb45995abfdb78bd81ff2a0356a596f

Observation c488ede0-f512-4ae6-9050-0a36aa648941 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.614383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.614383Z digest=sha256:df492a246459a14e5e7182d6f954ad6e61084a924786ddf813b3ab7b3b712acd

Observation aaa26ca6-40be-4162-acf0-908c80fce4d3 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.Advances in neural information processing systems, 36:53728–53741, 2023.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Direct preference optimization: Your language model is secretly a reward model.Advances in neural information processing systems, 36:53728–53741, 2023

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.618872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.618872Z digest=sha256:fc1742db36819d03f7f30793deacc4c7f726197e829da83953207af76646d897

Observation ba795ca9-c6bf-41fb-af74-c838cfb26782 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.623408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.623408Z digest=sha256:48e62c89180dd797ade45bba01d41fc0df5e6e334b1ace528a0b9e983cadad37

Observation 53f52013-4298-4d13-9832-06e4463ba865 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.628209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.628209Z digest=sha256:27cd16a57d00e12a7374045df0ec44e6843096c4338aa1200bd44d9dcf2e56dd

Observation 3c61b79c-030c-42a0-88ff-cbac19d0b4fb · outbound

This paper cites ToRL: Scaling Tool-Integrated RL.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing ToRL: Scaling Tool-Integrated RL

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.633077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.633077Z digest=sha256:cb09be613e90597bea989cd6a41fb05669ef440dd90428754a0ef9df643bc992

Observation e6ce7ab3-ca53-44a0-960d-951f42109cd4 · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.637815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.637815Z digest=sha256:dc9f7bf427c507c8b92c4115772f8b1bb21f6bbc03e24da1d19ee3fbfe07be64

Observation 3a6b4ba9-845b-4056-be59-28eb4a5da243 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing ToolRL: Reward is All Tool Learning Needs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.642514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.642514Z digest=sha256:930e8098760ea4e15ea8e614097348741dd60e21c66382ac8be9d26391cb6e71

Observation ca4ee6f6-5d92-472b-95e6-2c5810b604f1 · outbound

This paper cites Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.647525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.647525Z digest=sha256:3cc23342fb8372eefe9bb4b248f70ccc6bd981a1e8077f2e2c3ce4a7135bf5fd

Observation 65e2fbaa-1e39-42b7-8721-101cae37630c · outbound

This paper cites Effects of cognitive behavioral therapy and cash transfers on older persons living alone in india: a randomized trial.Annals of internal medicine, 176(5):632–641, 2023.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Effects of cognitive behavioral therapy and cash transfers on older persons living alone in india: a randomized trial.Annals of internal medicine, 176(5):632–641, 2023

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.399794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T04:39:03.652822Z digest=sha256:59964182be804eef26d6f9e3c533acff4e7dca6897cb7697af712cac36a1d49a

Observation 5b30023b-5725-4082-80ba-ac74832d7b4f · outbound

This paper cites Genomic characterization of metastatic patterns from prospective clinical sequenc- ing of 25,000 patients.Cell, 185(3):563–575, 2022.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Genomic characterization of metastatic patterns from prospective clinical sequenc- ing of 25,000 patients.Cell, 185(3):563–575, 2022

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.384842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T04:39:03.657534Z digest=sha256:1bb3d8e5fb2fd65a04cadd713b6357391f4d0775872e42996073fa8341e79df2

Observation 5254a881-e8a1-488e-9cd5-2826d602f635 · outbound

This paper cites The support prognostic model: Objective estimates of survival for seriously ill hospitalized adults.Annals of internal medicine, 122(3):191–203, 1995.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing The support prognostic model: Objective estimates of survival for seriously ill hospitalized adults.Annals of internal medicine, 122(3):191–203, 1995

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.369442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T04:39:03.662098Z digest=sha256:83fcee349d22a195b813d1a463dcac40fd98aa456a2b20ebd6cc4f2dfc09f884

Observation f15cf1cd-b341-43da-b393-5e0ccc2186a6 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.667845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.667845Z digest=sha256:2d4db6fc27f0a34e29a679a4d0b168b86827c3013888002966471190640ca1af

Observation 8c18f6e8-df20-41d5-b0d1-6c30635501c6 · outbound

This paper cites claim”: “drug improves patient outcome.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing claim”: “drug improves patient outcome

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.353710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T04:39:03.672908Z digest=sha256:d4bb03a8926db9f6558b4af8f94ad58c94eaca25d0c424ed60bca1ff585310de

Observation adbe0b92-7e51-427f-8525-b05ec970620f · outbound

This paper cites an unresolved cited work.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-10T04:39:04.336736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T04:39:03.678589Z digest=sha256:88fa619c06448eac6a1bb315c708515d44b54649e3f682601d24411ee7bdb216

Observation ddf3d3cd-3bf0-42c1-b65d-4b12483816bc · outbound

This paper cites an unresolved cited work.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-10T04:39:04.322049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T04:39:03.683802Z digest=sha256:9c65614b143a888210c5f82cdedbd5955e91fb8cf599d5d9261198b89ac3efdb

Observation 93bd0286-2b0f-47e0-8e6b-13789a89d806 · outbound

This paper cites an unresolved cited work.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-10T04:39:04.306832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T04:39:03.688424Z digest=sha256:f1cf7fc62f08116c713a841fec8be5cf3fa708d6a457b1acaa74967da7635549

Observation 507ce3fc-5327-40f6-b266-2abf02d2f7e1 · outbound

This paper cites an unresolved cited work.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-10T04:39:04.290784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T04:39:03.692877Z digest=sha256:abd9db987fa06b5cddec1e9694bddbb979fed5d2c5da81048106a352a13fad8e

Observation 163b0ec8-d481-4081-8989-a5e184270d91 · outbound

This paper cites an unresolved cited work.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Unresolved cited work

Reference 2026

Resolution
parse uncertain
no resolver link, observed 2026-08-10T04:39:03.560376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.560376Z digest=sha256:87401fd337f68cc2f477a390ff0dbf7247cdc36847ea504ba4948f3770d16995

Pith citing papers

No inbound Pith citation observations are available.