Pith. sign in

Paper Citation Record · LEDGER

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science

As of 23 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 5 inbound Pith citation observations for arXiv:2505.07889.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.07889 v4

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:33:24.554504Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T07:22:25.349398Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T20:47:23.051268Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4ef48869-c7f9-4272-a554-53979c5e0b89 · outbound

This paper cites GPT-4 Technical Report.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T22:33:24.030229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:33:24.030229Z digest=sha256:05057d7f82646eb4f2e858111cc6ddbafc327fe913e5428412663dd6f2c61318

Observation cc59d3f1-5c72-42ab-af7b-a5ebdd53defe · outbound

This paper cites PaLM 2 Technical Report.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science PaLM 2 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T22:33:24.035130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:33:24.035130Z digest=sha256:ea484d78379da9c227ca47b24bb027f41925a5e7e934ee02c46975a54a782a1f

Observation bfc01981-4d2b-4918-b3ba-f16dfcee5a7b · outbound

This paper cites Claude 3.7 sonnet system card, Feb 2025.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science Claude 3.7 sonnet system card, Feb 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:33:25.802465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:33:24.039516Z digest=sha256:9a340bdaf1334c1f1c48c9edd8797052afbe6006199c42b9677e4b72167468f9

Observation 482e647f-e966-4a15-939d-5d24d7e421f4 · outbound

This paper cites Cardbiomedbench: A benchmark for evaluating large language model performance in biomedical research.bioRxiv, pages 2025–01, 2025.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science Cardbiomedbench: A benchmark for evaluating large language model performance in biomedical research.bioRxiv, pages 2025–01, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:33:25.791612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:33:24.043316Z digest=sha256:05a9f6f8640c4c1552cdd85318f367a18e263dc23744a820c95fe5e3e756e08c

Observation c72c7b79-7a63-4aaa-8664-167a84d8c4a8 · outbound

This paper cites Autonomous chemical research with large language models.Nature, 624(7992):570–578, 2023.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science Autonomous chemical research with large language models.Nature, 624(7992):570–578, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T22:33:24.047620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:33:24.047620Z digest=sha256:b0dd314cf7553a06e9164196290a3ba04fbd8c3deb97caef4a772f6bff0920a9

Observation 56e6fa41-a771-483f-aaf6-b4d00359f140 · outbound

This paper cites Language models are few-shot learners.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science Language models are few-shot learners

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:33:25.681398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:33:24.051899Z digest=sha256:348b472ab466e769fed3dc1d788867ff05459ca74d23d19d3af8ae34260991c4

Observation e2775839-3772-49b5-9fdf-e59b6b01065f · outbound

This paper cites Probio: A protocol-guided multimodal dataset for molecular biology lab.Advances in Neural Information Processing Systems, 36:41543–41571, 2023.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science Probio: A protocol-guided multimodal dataset for molecular biology lab.Advances in Neural Information Processing Systems, 36:41543–41571, 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:33:25.606712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:33:24.056027Z digest=sha256:10c3dbee816e1735a8df2451fe8d74b0011a0bcbbb0cee559a54bd3a632d8afe

Observation a16224dd-1988-49ef-b422-450ba6ef2f69 · outbound

This paper cites Introducing gemini 2.0: Our new ai model for the agentic era, Dec 2024.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science Introducing gemini 2.0: Our new ai model for the agentic era, Dec 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:33:25.522604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:33:24.059497Z digest=sha256:2d74791b454779c6081b53181fdec8e25ca3a2912529f18db60b80437bbba9ff

Observation 27bfc0ea-2f04-4e5b-a1d7-de2e19d0ff58 · outbound

This paper cites Introducing gemini 2.5 pro experimental (model id gemini-2.5-pro-exp-03-25).

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science Introducing gemini 2.5 pro experimental (model id gemini-2.5-pro-exp-03-25)

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:33:25.511214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:33:24.063042Z digest=sha256:2f835c26ddbcb5d9955db4d7faaea4c5159996a17b7dfffaa215da1783f1b432

Observation dd86a53a-1e3d-401b-a59d-c5c222ab10b4 · outbound

This paper cites Keybert: Minimal keyword extraction with bert., 2020.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science Keybert: Minimal keyword extraction with bert., 2020

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:33:25.500250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:33:24.067096Z digest=sha256:04851e93691b820a9672f5e0f84ee5d990bf0bcfbd4a75c364365c34d027bbc7

Observation 26573b03-3ecc-4e7d-a820-5a83d1ce8a46 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T22:33:24.071249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:33:24.071249Z digest=sha256:63ad3bb1e7d71b343ca46540bde6686f8c59eb2f28762c51e740932427a67999

Observation 9ccf2c3c-635c-4251-8ca1-5807ef8d2619 · outbound

This paper cites Biolp-bench: Measuring understanding of biological lab protocols by large language models.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science Biolp-bench: Measuring understanding of biological lab protocols by large language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:33:25.489654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:33:24.075045Z digest=sha256:6a60c46d6d5023d2b1675b64d066d9045699075f0a06b65c26e0990cda890a88

Observation 4386f693-a154-4904-a871-ebaf76bcd542 · outbound

This paper cites A comprehensive evaluation of large language models on benchmark biomedical text processing tasks.Computers in biology and medicine, 171:108189, 2024.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science A comprehensive evaluation of large language models on benchmark biomedical text processing tasks.Computers in biology and medicine, 171:108189, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:33:25.478126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:33:24.078626Z digest=sha256:03e5d5e838a60dcc72cfe9ea625748ff6aa1a9db2fcec1556147c1b52990c35e

Observation 2632edba-e4b9-4627-b43c-2634e3f632f9 · outbound

This paper cites Pubmedqa: A dataset for biomedical research question answering.EMNLP 2019, 2019.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science Pubmedqa: A dataset for biomedical research question answering.EMNLP 2019, 2019

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:33:25.404547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:33:24.082176Z digest=sha256:1b577f6bd7697055d9cb267ad189f4fcdb53932508889edee9920c007e13228e

Observation 9ef99e5b-04b3-49da-b4f2-aa4e84c5aaa5 · outbound

This paper cites Large language models are zero-shot reasoners.Advances in neural information processing systems, 35:22199–22213, 2022.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science Large language models are zero-shot reasoners.Advances in neural information processing systems, 35:22199–22213, 2022

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T22:33:24.118123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:33:24.118123Z digest=sha256:fbe80bf871e76a669a1e8a40a74e6fc0eecc15f6b7b31771e1f2f8b8689aff34

Observation fe76b110-c84f-4464-8320-63463f1e24c7 · outbound

This paper cites LAB-Bench: Measuring Capabilities of Language Models for Biology Research.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science LAB-Bench: Measuring Capabilities of Language Models for Biology Research

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T22:33:24.166885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:33:24.166885Z digest=sha256:20c5f415f14ccc5d768f7afad0f1c0fcee5f9f783685b8be5832f90be68e5e6a

Observation 2b18c5d4-a79f-4cc8-9f86-874115103139 · outbound

This paper cites Biobert: a pre-trained biomedical language representation model for biomedical text mining.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science Biobert: a pre-trained biomedical language representation model for biomedical text mining

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T22:33:24.222169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:33:24.222169Z digest=sha256:502bb5a407fd33ae0108fde702b23a2e7ce5e6b8b999f0b1f15817235c0ae446

Observation b612ac02-2311-4d73-94f1-62a8e16b5137 · outbound

This paper cites Can large language models reason about medical questions?Patterns, 5(3), 2024.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science Can large language models reason about medical questions?Patterns, 5(3), 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:33:25.269643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:33:24.277179Z digest=sha256:9c950c7ed6af8153e6e4cc1ff4ef57329d8523a56c3804ec169810aab6c00bad

Observation 0ab93968-88a9-479b-9c01-077aeb514815 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T22:33:24.330357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:33:24.330357Z digest=sha256:c496f8768ec6fbdd31f09d7ced318b3c32b318405dc3443b2dbd63900b5ed2fb

Observation 3fed51ad-583f-48e6-b1fd-43ba99d26226 · outbound

This paper cites Biogpt: generative pre-trained transformer for biomedical text generation and mining.Briefings in Bioinformatics, 23(6):bbac409, 2022.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science Biogpt: generative pre-trained transformer for biomedical text generation and mining.Briefings in Bioinformatics, 23(6):bbac409, 2022

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:33:25.182310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:33:24.388337Z digest=sha256:5ad56f8a7b1d0dcd52eba6177367a8739bbd2175be293a3d16b438764b562494

Observation ab4bce7d-b830-4ab8-9844-f5395712c16c · outbound

This paper cites BioMedGPT: Open Multimodal Generative Pre-trained Transformer for BioMedicine.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science BioMedGPT: Open Multimodal Generative Pre-trained Transformer for BioMedicine

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T22:33:24.399333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:33:24.399333Z digest=sha256:52885141a15d95f349035942d2a1f9fc8695f0e6c9cb4f573419484bbc07575c

Observation 37d42c6d-bd1a-4df5-85f0-793c9d74234a · outbound

This paper cites Bixbench: a comprehensive benchmark for llm-based agents in computational biology.arXiv preprint arXiv:2503.00096, 2025.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science Bixbench: a comprehensive benchmark for llm-based agents in computational biology.arXiv preprint arXiv:2503.00096, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T22:33:24.419662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:33:24.419662Z digest=sha256:ce54fcc66e9c8fbd6816140adab33df0bee93940a32d40ea5c69f6bb7ffe7d0b

Observation f95fd62a-c51a-4737-bc82-783b07465ad8 · outbound

This paper cites Laboratory automation and high-throughput biology, 2024.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science Laboratory automation and high-throughput biology, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:33:25.172610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:33:24.425395Z digest=sha256:b3f073aaa46214b534dce516f1ea96a08ca0881ee744f3a3cae2f9bdc937ae12

Observation 646a9bc7-982a-4d76-89fb-95d5d86769d1 · outbound

This paper cites Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T22:33:24.437379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:33:24.437379Z digest=sha256:688bcd7f7ea66626acaeef6647f8be7fd3d1f36c6dfdf62410f7baeb1f465308

Observation d53d2074-19c5-4fa2-bd36-f874e87b6e5e · outbound

This paper cites Openai announces gpt-4 turbo, 2023.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science Openai announces gpt-4 turbo, 2023

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:33:25.159787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:33:24.445630Z digest=sha256:3386c3a4c1b4cf8052eae8e133d482048e88a763938ac36e7b9285c7cf7ca06e

Observation 327a174e-6835-4788-b762-5dda3502423c · outbound

This paper cites Introducing openai o3 and o4-mini.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science Introducing openai o3 and o4-mini

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:33:25.144263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:33:24.450152Z digest=sha256:5910713ebf2df3a43b90dbcb605dc7f53a776c57d7434734e80454f4564751c9

Observation 8c18fcf0-9521-4352-ab23-d3b361c5d92c · outbound

This paper cites Bioplanner: automatic evaluation of llms on protocol planning in biology.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science Bioplanner: automatic evaluation of llms on protocol planning in biology

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:33:25.125544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:33:24.456245Z digest=sha256:eba19f1bc92a5115ef07e137232031d42258cf41bcff6a100a20d41e77fc0524

Observation 23ad8e47-517a-4e07-a782-3b76fe763067 · outbound

This paper cites Qwen2.5 -72b-instruct, 2024.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science Qwen2.5 -72b-instruct, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:33:25.100255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:33:24.459348Z digest=sha256:7778d6c0e96f84b79971285b252b787dd853be834a008d913ea92f4044f06b70

Observation 8141858d-35af-4c05-ac07-c1c8296b72cd · outbound

This paper cites Qwq-32b, 2025.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science Qwq-32b, 2025

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:33:25.088365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:33:24.462848Z digest=sha256:8c65be248d03a6b523d5ffb43f6c169cd76d0d98351a3af4264d07e3bc9d3f24

Observation 28bbac65-ea09-4246-9588-ff164292dad3 · outbound

This paper cites Towards scientific intelligence: A survey of llm-based scientific agents.arXiv preprint arXiv:2503.24047, 2025.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science Towards scientific intelligence: A survey of llm-based scientific agents.arXiv preprint arXiv:2503.24047, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T22:33:24.466494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:33:24.466494Z digest=sha256:13975228f80b3b2a7fe4a447094e4517d9ad6c16541500d412ca69a1cd310c2e

Observation 82146d74-9232-46be-a82a-7a24a3cb7648 · outbound

This paper cites Toward expert-level medical question answering with large language models.Nature Medicine, pages 1–8, 2025.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science Toward expert-level medical question answering with large language models.Nature Medicine, pages 1–8, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T22:33:24.469288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:33:24.469288Z digest=sha256:7f93cf93af643baab7d21e7cdec033e688b9de2a73ecf90271bca6970656d4f1

Observation d418fada-59b7-4993-82a9-8326cddcb5ec · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science Gemini: A Family of Highly Capable Multimodal Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T22:33:24.472820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:33:24.472820Z digest=sha256:aa0a451e24c5248e7ae933ecef22d505ba1af992e7fc0df61805f625e09292a5

Observation d7ce34c2-ff60-41f3-9daf-27129c411f16 · outbound

This paper cites Large language models in medicine.Nature medicine, 29(8):1930–1940, 2023.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science Large language models in medicine.Nature medicine, 29(8):1930–1940, 2023

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T22:33:24.476436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:33:24.476436Z digest=sha256:3bdb1d852cac1ab60089cbe53bc0a5a2d33e6a146a22d69dec675fc66c8e4cd5

Observation 85bb4287-4c4e-4b41-9dc3-49f380de8a00 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T22:33:24.479666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:33:24.479666Z digest=sha256:c6239185f00be0feb32740e8e8499865ca5cf7519991aef013317ca33a867867

Observation 5caf1e77-09dc-4e02-9204-7aa3370c991f · outbound

This paper cites An overview of the bioasq large-scale biomedical semantic indexing and question answering competition.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science An overview of the bioasq large-scale biomedical semantic indexing and question answering competition

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:33:25.058938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:33:24.483551Z digest=sha256:158cfecef63fe8069200b9c76536f3d2d8e3eb509c5f1ef7985f9f6bd1722405

Observation 4d5c6e2f-1ebb-4486-b229-fc13c499554c · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T22:33:24.486899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:33:24.486899Z digest=sha256:7376b839c978032f2e43fc0a062a9fb98ff16961800b4a061c859a2e1c16f25e

Observation 047d94bc-00c6-4307-841a-53dd228c3ccf · outbound

This paper cites Advancing Multimodal Medical Capabilities of Gemini.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science Advancing Multimodal Medical Capabilities of Gemini

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T22:33:24.490545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:33:24.490545Z digest=sha256:7cdb4ddcc00476a5ae25115fb29538dc2f1ce7592b05e926ed3e9428def28d69

Observation 37528c27-9505-4911-8ad2-b5deb8bf9387 · outbound

This paper cites BioBART: Pretraining and Evaluation of A Biomedical Generative Language Model.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science BioBART: Pretraining and Evaluation of A Biomedical Generative Language Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T22:33:24.493818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:33:24.493818Z digest=sha256:8633807ef45e098683931cb659514910ea5d5cab5cc0002cc451654005df9ce6

Observation 0f158872-d2a1-48d3-9766-ddbafec2954f · outbound

This paper cites Benchmarking large language models on safety risks in scientific laboratories.Nature Machine Intelligence, 2026.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science Benchmarking large language models on safety risks in scientific laboratories.Nature Machine Intelligence, 2026

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:33:25.037402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:33:24.509725Z digest=sha256:0959ad4cb2fcdfa3b4c90f8f96164ecad50ef9f581ea295e191e30ae5dae22fc

Observation 00c3fdb2-6f00-44be-bef9-dc76724c0ff6 · outbound

This paper cites MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T22:33:24.554504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:33:24.554504Z digest=sha256:74c59d2bb05f73a5d3ef92de31d49d1726cbfafe9e74e602de7b6631248459ce

Pith citing papers

Observation 8a6c22ae-54da-485f-aa4d-b980f06fa848 · inbound

BioXArena: Benchmarking LLM Agents on Multi-Modal Biomedical Machine Learning Tasks cites this paper.

BioXArena: Benchmarking LLM Agents on Multi-Modal Biomedical Machine Learning Tasks BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-28T02:22:20.781065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T19:31:32.334837Z digest=sha256:9d356eb4a55103e93101653e0f92c0c28ce55ec7fbb43c49b7b91f6dc7153ac5

Observation 09b484fe-03b4-46c2-b55f-666496e43373 · inbound

Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches cites this paper.

Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science

Reference 169

Resolution
verified exact
arxiv_id, observed 2026-07-28T02:22:20.781065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T17:35:01.285534Z digest=sha256:9c6a9a21ff85d309f3aaf98f2d229006c7c61cdf9db55d58a3ba3195afaa1615

Observation f291752a-eb15-4872-a448-dad8c0f8df51 · inbound

Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches cites this paper.

Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science

Reference 182

Resolution
verified exact
arxiv_id, observed 2026-07-28T02:22:20.781065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T07:22:25.349398Z digest=sha256:29196c43595afea8111cb65b90d8423d3cc004a94272cd13ba9ccb610494fe9b

Observation 241fbc7c-ba67-4fdf-9979-5cf8b7839721 · inbound

SCI-PRM: A Tool Aware Process Reward Model for Scientific Reasoning Verification cites this paper.

SCI-PRM: A Tool Aware Process Reward Model for Scientific Reasoning Verification BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-28T02:22:20.781065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T06:38:37.969788Z digest=sha256:ef4ab640d0cc43bff3a558c7650bfd83afbe1ac3a7424704948b90a85929981c

Observation cd8caa39-9c92-487b-b70b-1771377efa11 · inbound

DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoning cites this paper.

DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoning BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-07-28T02:22:20.781065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T20:08:29.208550Z digest=sha256:170df2ff21443b1dda23ab27fc970d2efc71505a9701499c4a7abdb30757f1d2