Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T06:02:32.399075Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 0 inbound Pith citation observations for arXiv:2604.17159.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T06:02:32.399075Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
12 of 12 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8927117d-cc4c-4fb4-88ca-d6e6cba95f30 · outbound
Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks Generative AI in cybersecurity: A com- prehensive review of LLM applications and vulnerabilities
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 479d3493-16c0-438e-8fbb-2b1b43ddccdd · outbound
Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks LLM Agents can Autonomously Hack Websites
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 77506b14-1d7b-41c9-ab74-0a552213cd1e · outbound
Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks PentestGPT: Evaluating and harnessing large language models for automated penetration testing
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f1e731c9-9870-46d5-a69c-b3b9c6ce0b86 · outbound
Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks OCCULT: Evaluating Large Language Models for Offensive Cyber Operation Capabilities
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 465b323a-806e-4e04-8059-d42cec2a3e14 · outbound
Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7613c64e-2c0d-4576-8aea-5fcbcae9637c · outbound
Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks EnIGMA: Interactive Tools Substantially Assist LM Agents in Finding Security Vulnerabilities
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ec0a22bb-13f5-4249-81da-1e0d1341dad6 · outbound
Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks D-CIPHER: Dynamic Collaborative Intelligent Multi-Agent System with Planner and Heterogeneous Executors for Offensive Security
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a1693e07-42de-44bf-889c-eaf38d71ce4b · outbound
Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks CRAKEN: Cybersecurity LLM Agent with Knowledge-Based Execution
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4bb081e1-779b-4042-a3d7-129c3a192a6f · outbound
Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks An Empirical Evaluation of LLMs for Solving Offensive Security Challenges
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 670e3a77-f49b-4e6a-b904-6acda6aaf242 · outbound
Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f925fd48-69c5-4323-a651-8aae7f8b7b15 · outbound
Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks CTFusion: A CTF-based benchmark for LLM agent evaluation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4342aafa-f030-4bd0-9318-cb02d203a8f3 · outbound
Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks Zhang, Joey Ji, Celeste Menders, Riya Dulepet, T
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
No inbound Pith citation observations are available.