Pith. sign in

Paper Citation Record · LEDGER

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts

As of 20 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2608.09567.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09567 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:59:21.169998Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact3
  • verified fuzzy11
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a1282627-8509-4760-946d-533861cd7381 · outbound

This paper cites Repeatability in computer systems research.Com- munications of the ACM, 59(3):62–69, 2016.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Repeatability in computer systems research.Com- munications of the ACM, 59(3):62–69, 2016

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.741603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:59:21.105922Z digest=sha256:83e96ad810a09e4206dbc7743f85aa86508a431b9b3b1d9534a93f73dd781520

Observation d1dfbc5e-d755-4bcb-9f3a-b38169a24809 · outbound

This paper cites PentestGPT: Evaluating and harnessing large lan- guage models for automated penetration testing.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts PentestGPT: Evaluating and harnessing large lan- guage models for automated penetration testing

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.731735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:59:21.109554Z digest=sha256:db552ed7f935adc30a40b1d8b6be2a68c3d82216963a003d990dac4025842a42

Observation 60a5c369-9de9-4bbd-a35c-2fda88132fa2 · outbound

This paper cites PoC-Adapt: Semantic-Aware Automated Vulnerability Reproduction with LLM Multi-Agents and Reinforcement Learning-Driven Adaptive Policy.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts PoC-Adapt: Semantic-Aware Automated Vulnerability Reproduction with LLM Multi-Agents and Reinforcement Learning-Driven Adaptive Policy

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-11T14:59:21.625299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:59:21.112758Z digest=sha256:f4f94b92dbc45973c61ded5934eb18c986cca916d9fbb5a65e3b2bc5b28124f6

Observation c9316492-b658-4e01-b4c2-1351facc0ad4 · outbound

This paper cites Good News for Script Kiddies? Evaluating Large Language Models for Automated Exploit Generation.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Good News for Script Kiddies? Evaluating Large Language Models for Automated Exploit Generation

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-11T14:59:21.609935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:59:21.115774Z digest=sha256:38ce46c5ea4810bf5a1b3784e7e66643ffec341f659b2136803643135135f0e8

Observation 06240eb3-caa5-4ae1-8982-09a13c19c859 · outbound

This paper cites SEC-bench: Automated bench- marking of LLM agents on real-world software security tasks, 2025.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts SEC-bench: Automated bench- marking of LLM agents on real-world software security tasks, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:21.119045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:21.119045Z digest=sha256:ad7640f45c323cce66cb10ca5990cdae0f88f8dec219ae52c81feed15594713b

Observation 6f542bdc-2dd2-4900-9858-1aab64c12444 · outbound

This paper cites Automated vulnerability validation and verification: A large language model approach, 2025.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Automated vulnerability validation and verification: A large language model approach, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:21.121999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:21.121999Z digest=sha256:f434905d4a1706273009fb3f0e12660ab887161825722254a1dfa1a0cbdc660c

Observation 5803089d-ea40-4083-a907-0e16f02c80f0 · outbound

This paper cites Shell or nothing: Real-world benchmarks and memory-activated agents for automated penetration testing, 2025.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Shell or nothing: Real-world benchmarks and memory-activated agents for automated penetration testing, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:21.124600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:21.124600Z digest=sha256:cd64edfe5223636da2957371a4993cc86729d79a7cd59670b6d7c1c273961024

Observation e7107df3-6e64-43ba-adec-137281b4aaa5 · outbound

This paper cites Cve-2020-1967.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Cve-2020-1967

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.721813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:59:21.127846Z digest=sha256:0d06754cfe0d5c3dbd70d63821d632790842f88c6d060e19d658f60480394280

Observation 7a1bbd3a-9b00-42dc-9851-03062c6b97e7 · outbound

This paper cites Cve-2021-31162.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Cve-2021-31162

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.709414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:59:21.130988Z digest=sha256:d78ed4f1edd570850a7d948bb33f815c2c503764803762801ee4debd84b73b98

Observation 95fb2193-4302-4002-919a-d4d807e5b642 · outbound

This paper cites Cve-2021-44228.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Cve-2021-44228

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.698634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:59:21.133376Z digest=sha256:0ec973a7ab8a0bd46089f0df399bfa3b9bc0c7460bcd6bf52f5f38d65fc3fe46

Observation 00495009-4a96-4cb2-93a4-dbefd281da6b · outbound

This paper cites Cve-2022-22816.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Cve-2022-22816

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.689690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:59:21.135474Z digest=sha256:0d0f24d5d62e2b3b9af50f4fa8e3e333f4de82b93dd81cc67b8c6a0b7545151c

Observation a96a511a-d63a-49de-bbc3-012a39c3ad09 · outbound

This paper cites Cve-2023-0217.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Cve-2023-0217

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.679183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:59:21.137405Z digest=sha256:67301c933f696954a2a2c294f3152972806c20b5873526928b208f36c24f5e78

Observation 75fdb339-dfdf-40ce-88dc-bcc9626e46a1 · outbound

This paper cites Cve-2023-25676.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Cve-2023-25676

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.667701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:59:21.139624Z digest=sha256:2abdcd18fd31245194b0bf11d791cfaaba56a34e344739e369657491e0aad11c

Observation 051932f2-7da9-4c96-a5b9-452c4a7a54de · outbound

This paper cites Cve-2025-30223.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Cve-2025-30223

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.656506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:59:21.142159Z digest=sha256:5e2fc5d08e8358c875818661a68e05731c26fb983d7e6c4dda870b795939ba81

Observation 121e8746-34d5-4006-b594-5d76ec9184cd · outbound

This paper cites FaultLine: Automated Proof-of-Vulnerability Generation Using LLM Agents.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts FaultLine: Automated Proof-of-Vulnerability Generation Using LLM Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:21.144083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:21.144083Z digest=sha256:74207725439b64c3f64485be1e6f6f27c36406417122612bc424f819b234c42c

Observation cc4577a1-a9f2-45d0-83c6-dd8b8b9e68ce · outbound

This paper cites ”Get in Researchers; We’re Measuring Reproducibility”: A reproducibility study of machine learning papers in tier 1 se- curity conferences.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts ”Get in Researchers; We’re Measuring Reproducibility”: A reproducibility study of machine learning papers in tier 1 se- curity conferences

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:21.146899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:21.146899Z digest=sha256:4d9369af45ab64df1355b45ce0213532f4bb16b020057cc81405ce3e8f813581

Observation 0e329e66-dde9-4e57-8ad4-3a7fb2764e8c · outbound

This paper cites Artificial Intelligence as the New Hacker: Developing Agents for Offensive Security.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Artificial Intelligence as the New Hacker: Developing Agents for Offensive Security

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-11T14:59:21.381995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:59:21.150018Z digest=sha256:d209739e1d7d0ce453aa8555879697e16cf9a1ed686cca7750a7714abcb508a5

Observation d484c919-ee1f-4bdb-8d38-5aea5222ff80 · outbound

This paper cites Contract- Tinker: LLM-empowered vulnerability repair for real-world smart contracts.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Contract- Tinker: LLM-empowered vulnerability repair for real-world smart contracts

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.646448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:59:21.153315Z digest=sha256:dd1396c7e475d15cb27790d600c4ba43ace3906936b32819bc46d15363441395

Observation ebc02d0a-a208-4ec6-8639-f29b534e7d7c · outbound

This paper cites PATCHEV AL: A new benchmark for evaluating LLMs on patching real-world vulnerabilities, 2025.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts PATCHEV AL: A new benchmark for evaluating LLMs on patching real-world vulnerabilities, 2025

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:21.156187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:21.156187Z digest=sha256:cec3e92dada6a44812596881d5bcaf4a4336b9941b198ac1080a5558ec198655

Observation 9ece5547-9224-41f4-aff0-b4fa6c490bc6 · outbound

This paper cites AutoPT: How Far Are We from the End2End Automated Web Penetration Testing?.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts AutoPT: How Far Are We from the End2End Automated Web Penetration Testing?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:21.159703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:21.159703Z digest=sha256:ea10199d678c01bfb989fbd587f528c5a7c7779377c8a95778e01b062c640e52

Observation 3c4d597c-d324-4f3c-a450-2a8e05b64f1a · outbound

This paper cites Patch validation in automated vulnerability repair, 2026.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Patch validation in automated vulnerability repair, 2026

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:21.164056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:21.164056Z digest=sha256:aa41fb42ac5b607b235a308ac967ea1589d7397444ffda54c6e7a9668ae845f3

Observation 43af46b6-aeb0-426a-878b-a0fd40de25c3 · outbound

This paper cites Zhang, Neil Perry, Riya Dulepet, Joey Ji, Celeste Menders, Justin W.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Zhang, Neil Perry, Riya Dulepet, Joey Ji, Celeste Menders, Justin W

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.636028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:59:21.166818Z digest=sha256:85601a70679d12b50399b11278c2ff2fd11edbdd46ddfe1133446e2e8c639775

Observation 93a5464a-3e34-4fb2-80e2-97f2de3efcc9 · outbound

This paper cites CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:21.169998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:21.169998Z digest=sha256:dacd4ed5422305c6e59dc354c82290c0e188f8f5b0a0e4a52d1acc3930744fd0

Pith citing papers

No inbound Pith citation observations are available.