Pith. sign in

Paper Citation Record · LEDGER

RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents

As of 19 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 6 inbound Pith citation observations for arXiv:2504.18565.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.18565 v2

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:39:05.891159Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:11:40.348373Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T01:43:56.971386Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3b3bcf62-8be0-4edd-890e-7d60530b4a9f · outbound

This paper cites an unresolved cited work.

RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:39:06.063069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:39:05.863782Z digest=sha256:29b6bcfd7aa257a1a7a29a28acefa27a5112e3660b94187ef103bb54e0d41006

Observation 786d1325-9141-4436-8772-8a4d247ce192 · outbound

This paper cites MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering.

RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:39:05.835106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:39:05.835106Z digest=sha256:cdc58348728e337d48075081ed8238dffbb706e55b12b757d0ef03b68ab1398d

Observation 79a0daf4-d199-4639-8685-6a90fb08eb91 · outbound

This paper cites an unresolved cited work.

RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:39:06.038684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:39:05.870962Z digest=sha256:50dd94d9ddfbdc27f988b9fc5631b60c8e8a9a3d9e071276d6a76bd937bed513

Observation 41d88582-92b8-489f-a944-8fe389e71dac · outbound

This paper cites DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents.

RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:39:05.843538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:39:05.843538Z digest=sha256:8b526b00ab7ab37e505bf37e1f15fb5cef26e880bbb814a522dc57653defa18b

Observation 0dee1437-3a61-4dd9-b2b0-39985c3dd6e4 · outbound

This paper cites Frontier AI systems have surpassed the self-replicating red line.

RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents Frontier AI systems have surpassed the self-replicating red line

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:39:05.851538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:39:05.851538Z digest=sha256:18f1acca38c7a5be3808fb5767de12c3210807e5305bfe72813eb12f2e145de1

Observation 10a4323b-0f4c-47d5-8400-e75c766a3c6d · outbound

This paper cites CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models.

RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:39:05.855707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:39:05.855707Z digest=sha256:ede31b3806158f688b2d6747b51bd7d975d83008e2c53e21f0c858ba4d4ee8e3

Observation e550fcf3-8a9b-474e-889a-092e6f0bc35b · outbound

This paper cites code": "sf login.

RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents code": "sf login

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:39:06.074847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:39:05.859921Z digest=sha256:e396aaaabaaecc535fdf03789231bacac8555e537c33206fc46b7950a471c8d8

Observation 51438e66-9efe-4133-ba4e-d3f62eb29f97 · outbound

This paper cites We can do this by listing and capturing the file's exact size in bytes.

RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents We can do this by listing and capturing the file's exact size in bytes

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:39:06.051643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:39:05.867110Z digest=sha256:09836bd0d7fa9a1773f9d8f7ad7afce41b335594ff63cb27e474ff7ef4304318

Observation b810d2b1-b732-47b5-a3c6-c135bda83907 · outbound

This paper cites answer":.

RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents answer":

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:39:06.025696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:39:05.874580Z digest=sha256:cbdb6969ed14a96bc87c85bf5c5f9a56cb09ae8290d7b81219902d7752340d5a

Observation 15fa266c-bfa2-4b14-978f-356aea19c366 · outbound

This paper cites an unresolved cited work.

RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:39:06.013309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:39:05.880474Z digest=sha256:a520727995e787c914c79ce6fe06bcb757394a22c8d6d6b69abffab6c0d455cf

Observation 6267ff5c-5976-4e80-8d7f-3c63d9460c64 · outbound

This paper cites an unresolved cited work.

RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:39:06.001417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:39:05.883789Z digest=sha256:1f2b3325787da8221bdb20842258727b8e9b1aed52a0e6f986388aa7028e4294

Observation b53b4283-60b0-4932-90b8-d05b5e79ca06 · outbound

This paper cites an unresolved cited work.

RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:39:05.990034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:39:05.887264Z digest=sha256:0a37ca44aec318e43f164a5420e3fbf6a7a0f9bde031233ea56109cca7cb9680

Observation 2e02dfce-789f-4dcc-9921-aa1666af9d85 · outbound

This paper cites an unresolved cited work.

RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:39:05.979077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:39:05.891159Z digest=sha256:d5ef9d65fdc148f03a82d4445aade5be7dd58d2cab2a35bc7b645a28295976c3

Observation 09790895-3e1a-445d-aa34-e1004c166f77 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents Evaluating Large Language Models Trained on Code

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-16T11:39:05.839443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:39:05.839443Z digest=sha256:1e75b2effbfcfb12852b916cf056472b23c637866c13b73ea471592777cb2641

Observation 6d3af73f-0d79-472b-a3c9-721740e20f2f · outbound

This paper cites an unresolved cited work.

RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:39:06.087194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:39:05.847608Z digest=sha256:dc2ce43721148f6a400217005522a47d82dbd1149d55a1b116ad745238d652c1

Observation 7b20b0db-2c63-4c19-bc8c-bda2a4517338 · outbound

This paper cites Jun Shern Chan, Neil Chowdhury, Oliver Jaffe, James Aung, Dane Sherburn, Evan Mays, Giulio Starace, Kevin Liu, Leon Maksin, Tejal Patwardhan, Lilian Weng, and Aleksander M ˛ adry.

RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents Jun Shern Chan, Neil Chowdhury, Oliver Jaffe, James Aung, Dane Sherburn, Evan Mays, Giulio Starace, Kevin Liu, Leon Maksin, Tejal Patwardhan, Lilian Weng, and Aleksander M ˛ adry

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:39:06.098315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:39:05.830494Z digest=sha256:f5b6abbc94e88250a52ce3259e66e77a1ac0b4c3cc453f242d54623cd740b6b6

Pith citing papers

Observation fa19d45b-42b9-4813-a6bc-984737b7cbcb · inbound

Adversarial Attacks on Robotic Vision Language Action Models cites this paper.

Adversarial Attacks on Robotic Vision Language Action Models RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:40.348373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:40.348373Z digest=sha256:da35eeb1b39da6bc24fdf0064a3995474c2f4757676f9a9841066f8c8fd8b012

Observation 5daba2b2-56c1-428a-9a14-c27588fa87d8 · inbound

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework cites this paper.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:33.918076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:33.918076Z digest=sha256:b12a497b3cbf4974c0681140a6729f7cc60603dfa3baaa978410021ef26d9f5b

Observation 1caa89ce-d703-4800-bc51-1c1fc2ae7ed8 · inbound

Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report cites this paper.

Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:14:21.791055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:14:21.791055Z digest=sha256:aabc87610398879d0509b535d5255fd73143ea0f868064991375e60e89a8d378

Observation 9ccabc5e-e7f6-4027-916a-ad7104dc2836 · inbound

Ensemble Monitoring for AI Control: Diverse Signals Outweigh More Compute cites this paper.

Ensemble Monitoring for AI Control: Diverse Signals Outweigh More Compute RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:13:43.355158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-20T20:12:03.715605Z digest=sha256:aa1033e02518f10ed1ca06d1f0838ff06f41e18e00cabfbd03b7724afe927c73

Observation 6457398a-7e11-4d66-841d-cf2780bdb754 · inbound

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents cites this paper.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.973932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:e6e914ba4c44d66b2a908e47dd67959f9d43a90e2e46747855a7909a1cbc2916

Observation fe5c5c71-a3a5-4da5-82cd-0910ce99f5d6 · inbound

Risky Business: Measuring The Faithfulness-Safety Tension cites this paper.

Risky Business: Measuring The Faithfulness-Safety Tension RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T13:40:54.013590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:40:54.013590Z digest=sha256:b35c8dc890da925b6f207c681e9511dde5d005c78cff6a2705b994a017702d14