Pith. sign in

Paper Citation Record · LEDGER

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark

As of 20 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2608.11469.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11469 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:19:09.555981Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact6
  • verified fuzzy6
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b1eb8ad7-90b9-4218-b1a5-c45a73908c59 · outbound

This paper cites Vulnerability Detection with Code Language Models: How Far Are We?.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark Vulnerability Detection with Code Language Models: How Far Are We?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T14:19:09.463443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:19:09.463443Z digest=sha256:da57f657e83248882f29a4763e4aff4e95274c58a968c0a1ad2c3b4b8283671e

Observation b91c5c13-c24f-4c0f-b0f9-44e36b867e0b · outbound

This paper cites Livecodebench: Holistic and contamination free evalua- tion of large language models for code.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark Livecodebench: Holistic and contamination free evalua- tion of large language models for code

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:19:10.117159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:19:09.482962Z digest=sha256:f089f14dd0d496a86742ab145551c348e22a31b380761e2f6e5459da7fc1e65b

Observation add1ed68-6e4b-4d14-a6cd-5e7bcb1d6cca · outbound

This paper cites REStack: A Large-Scale Dataset of Reverse Engineering Discussions from Stack Exchange.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark REStack: A Large-Scale Dataset of Reverse Engineering Discussions from Stack Exchange

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:19:09.995528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:19:09.488937Z digest=sha256:f47eb20b6a7f2d8e8b6c43c432b4b9cf612f17c7aac6fa1ed958172739625151

Observation 89e9e2b8-9263-4ed2-ba8a-f2fbeccc797a · outbound

This paper cites REFORGE: A Method for Benchmarking LLMs' Reverse Engineering Capabilities in Decompiled Binary Function Naming.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark REFORGE: A Method for Benchmarking LLMs' Reverse Engineering Capabilities in Decompiled Binary Function Naming

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:19:09.966622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:19:09.494438Z digest=sha256:a94f8b6ab3ac2f7f645b17644011573d91ba2c7a64e860d84bc31ad9d795e575

Observation fd51eebc-b6e7-4fdf-b5c8-a32a1b2c1bfc · outbound

This paper cites SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks?.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks?

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:19:09.932406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:19:09.500035Z digest=sha256:53a44d6f590a0cf61d2ba50e0c99ec9eaef1448eea186385c6728f6e2231e63c

Observation c552c635-48c1-4632-b541-da7b525265cf · outbound

This paper cites ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T14:19:09.506467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:19:09.506467Z digest=sha256:acc2e882b0433392b4f697c04d5de60463f303b0b61dc47e86f8f2621e547f2e

Observation 3dafe5fe-d4a7-48a8-a715-060c674c64e5 · outbound

This paper cites VulDetectBench: Evaluating the Deep Capability of Vulnerability Detection with Large Language Models.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark VulDetectBench: Evaluating the Deep Capability of Vulnerability Detection with Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T14:19:09.511903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:19:09.511903Z digest=sha256:384c4709c0f767adc0eb1666c651eee1d90a94ed70eab096fb9c9c633d901a81

Observation 2385ced4-2f48-4612-9641-74581e10eab3 · outbound

This paper cites Patch-to-poc: A systematic study of agentic llm systems for linux kernel n-day reproduction.arXiv preprint arXiv:2602.07287,.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark Patch-to-poc: A systematic study of agentic llm systems for linux kernel n-day reproduction.arXiv preprint arXiv:2602.07287,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T14:19:09.517108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:19:09.517108Z digest=sha256:00d360f4ba8c8dc512cf8278560d0f8902a11e01cc526c538264107a38ca4883

Observation 33f44ef8-e248-460e-82d0-050a37b5e4c2 · outbound

This paper cites ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T14:19:09.533769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:19:09.533769Z digest=sha256:47880a663eb2ff2934220024b17c33d7187dbfb57e46951bfe78199124d16ccc

Observation 59ea09b2-ab52-4a31-920b-e7bd3b2cc4d7 · outbound

This paper cites REBENCH: A Procedural, Fair-by-Construction Benchmark for LLMs on Stripped-Binary Types and Names (Extended Version).

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark REBENCH: A Procedural, Fair-by-Construction Benchmark for LLMs on Stripped-Binary Types and Names (Extended Version)

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:19:09.718401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:19:09.538836Z digest=sha256:d4db653d9cca8c2850a91d4c5725ae75395c033c230a94e00d2642d996cc735e

Observation fbfa9057-47ef-4d9a-9aec-a9d5f361fe27 · outbound

This paper cites Cybench: A framework for evaluating cyber- security capabilities and risks of language models.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark Cybench: A framework for evaluating cyber- security capabilities and risks of language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:19:10.084545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:19:09.544423Z digest=sha256:53ea253c6d01fcb9412bdb536c4859327aa420162e5cb36aabfb12685c3e4a17

Observation 868f2da6-28b3-4df1-bd39-6bb51e0ea382 · outbound

This paper cites CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T14:19:09.549877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:19:09.549877Z digest=sha256:9e0904d51ab552a7ec02b87177eb6804620c86a14a161e32d95e92dea957ce19

Observation 735dac0f-0924-4810-95f8-6c070051ae83 · outbound

This paper cites Training language model agents to find vulnerabilities with ctf-dojo.arXiv preprint arXiv:2508.18370,.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark Training language model agents to find vulnerabilities with ctf-dojo.arXiv preprint arXiv:2508.18370,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T14:19:09.555981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:19:09.555981Z digest=sha256:7417a3351e578c4334ec38da3b242fe77442982b1b932fc8147faea64b54b82d

Observation 1aa92109-06f1-4f88-88dd-7ba8b870c5d6 · outbound

This paper cites CREBench: Evaluating Large Language Models in Cryptographic Binary Reverse Engineering.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark CREBench: Evaluating Large Language Models in Cryptographic Binary Reverse Engineering

Reference 1994

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:19:10.063822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:19:09.451651Z digest=sha256:cedc57fafca28badd7dbf7b980b9c1ede75e60116ac354407d5ac23eb55faf75

Observation 1efe0ece-f99c-458b-bf16-0dbf85826279 · outbound

This paper cites The concept assignment problem in program understanding.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark The concept assignment problem in program understanding

Reference 2005

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:19:10.175322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:19:09.441883Z digest=sha256:f2371a69190b8c77af8579c204e68690951d6f7af285b7e9a6d1211477d2394e

Observation c75fbb3c-ea5c-49f6-8c25-ae639e94bf46 · outbound

This paper cites Benchmarking binary type inference techniques in decompilers.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark Benchmarking binary type inference techniques in decompilers

Reference 2008

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:19:10.101141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:19:09.521963Z digest=sha256:0ae584c59681ad0ef35462710b654baaf323002766db6b3ace76e6bf90254008

Observation 2292012d-bf99-4722-addb-1cc959017dc0 · outbound

This paper cites De- compilebench: A comprehensive benchmark for evaluating decompilers in real-world scenarios.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark De- compilebench: A comprehensive benchmark for evaluating decompilers in real-world scenarios

Reference 2011

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:19:10.154010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:19:09.470783Z digest=sha256:5c6de44fe186be3ecb27f178efbf4e63bc18973df9e2c796fc3e63cfb90fa071

Observation 34d9a0fe-8c3f-4e3e-9226-5a02dc55cb1d · outbound

This paper cites Cy- bergym: Evaluating ai agents’ real-world cybersecurity capabilities at scale.arXiv preprint arXiv:2506.02548,.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark Cy- bergym: Evaluating ai agents’ real-world cybersecurity capabilities at scale.arXiv preprint arXiv:2506.02548,

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-15T14:19:09.528096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:19:09.528096Z digest=sha256:968f66d9020362c80debdadb3929e6ca8bb2e42b5b02397490e5b478793d5431

Observation 8574ff04-1778-4338-95ee-1823f150f486 · outbound

This paper cites Look what you made us patch: 2025 zero-days in review.https://cloud.google.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark Look what you made us patch: 2025 zero-days in review.https://cloud.google

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:19:10.134531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:19:09.476129Z digest=sha256:e1cf99dd2b7f3b13b088405668e93abf9034fde7789258f30a588ce1e27803c5

Observation 919065b6-03ce-4b70-968e-669227f06a61 · outbound

This paper cites CrackMeBench: Binary Reverse Engineering for Agents.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark CrackMeBench: Binary Reverse Engineering for Agents

Reference 2026

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:19:10.035977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:19:09.458239Z digest=sha256:2ee38112272d3895515122b7bb6455324d27d1089575490a8a8140443abe7f60

Pith citing papers

No inbound Pith citation observations are available.