Pith. sign in

Paper Citation Record · LEDGER

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)

As of 14 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2608.04317.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04317 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:58:17.376083Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact0
  • verified fuzzy37
  • unresolved29
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e9e9ffb8-750e-4b83-8c2f-bd45f20644c3 · outbound

This paper cites https://www.atomicredteam.io/.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) https://www.atomicredteam.io/

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:18.297825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.162099Z digest=sha256:5fd886538f4065f02074bea21986a8f3dad5681c19978b53a4617250dc63a266

Observation 55d89643-12c0-4fe2-b587-e12a11e90d67 · outbound

This paper cites https://aicyberchallenge.com/ overview/.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) https://aicyberchallenge.com/ overview/

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:18.288443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.166233Z digest=sha256:d00c62dfa118949920674e61d6c18295f88352d3260e176e9903392779ac79b8

Observation 01e142a8-f26c-45f9-9bbe-4eb79a5a53e1 · outbound

This paper cites https://attack.mitre.org/.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) https://attack.mitre.org/

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:18.279178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.169561Z digest=sha256:6006438aea3a4522305bc13e6527bc21e896faef3e922c91023a49001b7a17d8

Observation 79f9f41b-f184-4565-abc2-5c6a8e94bdcd · outbound

This paper cites EnIGMA: Interactive tools substantially assist LM agents in finding security vulnerabilities.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) EnIGMA: Interactive tools substantially assist LM agents in finding security vulnerabilities

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:18.270066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.173238Z digest=sha256:756a0be44fe51dfc4b21cb4689ddfa2edb52ea83cf0e57b1cbecd22ce0f5eebf

Observation 067df76b-02b2-4127-a1bc-ba247e54a6f4 · outbound

This paper cites Back to basics: Revisiting REINFORCE-style optimization for learning from human feedback in LLMs.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Back to basics: Revisiting REINFORCE-style optimization for learning from human feedback in LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.176790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.176790Z digest=sha256:9c13d62f0dd9227118aaff0b2851102ee9844fa61e079a9733b8e24ad47c6db6

Observation 141545ae-3ce0-4a1b-9c95-bb69a16af683 · outbound

This paper cites Ctibench: a benchmark for evaluating llms in cyber threat intelligence.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Ctibench: a benchmark for evaluating llms in cyber threat intelligence

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:18.255302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.180734Z digest=sha256:59a81efe2c08f8a345a5640a94eed21c0dd0f2a4eced84e20ae01737ac8fb43a

Observation 489af948-f565-4c76-9a3a-e2ed8ea9bf0b · outbound

This paper cites Claude Mythos Preview red.anthropic.com — red.anthropic.com.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Claude Mythos Preview red.anthropic.com — red.anthropic.com

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:18.245978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.185307Z digest=sha256:0e775ec629df01287873c89626170631ae1cfff5bff2173f84ecb9854365caa0

Observation a94a278f-1b97-4c27-a052-f8d65eea29d9 · outbound

This paper cites CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.189295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.189295Z digest=sha256:968377de4bf2e89e1679f5a3669599de265d1e46b78e29b4aa453fa204b73ce4

Observation 7cdc2bd7-d8d6-4bb1-8996-da6ea68649a7 · outbound

This paper cites Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.193944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.193944Z digest=sha256:842959cf0b37adb8f4cc6d3779d629de0f69103412b32b7510a2932de17b9ffe

Observation 66b4479a-f346-4677-b5e0-5a1b4346697f · outbound

This paper cites Large language models are autonomous cyber defenders.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Large language models are autonomous cyber defenders

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:18.236818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.197620Z digest=sha256:3f9e14cb56a41dfd04f3ebd7731d8daab929870df9abff748a269d31235f5d8b

Observation eafcbef2-bf1a-42ce-8839-962fd2b7ef32 · outbound

This paper cites Agentverse: Facilitating multi-agent collaboration and exploring emergent behaviors.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Agentverse: Facilitating multi-agent collaboration and exploring emergent behaviors

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.200744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.200744Z digest=sha256:7b9b4ef93d47449d0f8d40fd348fc0457170f816c055d1eea10fcec782ceef03

Observation f56c22bb-c346-46b6-a51d-8999397414a7 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Training Verifiers to Solve Math Word Problems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.203952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.203952Z digest=sha256:0a9572083cf1d873c91adc607f4b316346e2fc4ad09ea09ae4810588504d36e5

Observation c44250ce-64e6-41b6-9e34-c9837de6b25f · outbound

This paper cites PentestGPT: Evaluating and harnessing large language models for automated penetration testing.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) PentestGPT: Evaluating and harnessing large language models for automated penetration testing

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:18.222351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.207606Z digest=sha256:e6d90ba50b203130b0624a30d82bec1877030db404c3fcc7dddb4446f1cf9c45

Observation 20e8ebcb-43a6-4c2f-9524-793b567730e6 · outbound

This paper cites LLM Agents can Autonomously Exploit One-day Vulnerabilities.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) LLM Agents can Autonomously Exploit One-day Vulnerabilities

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.210586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.210586Z digest=sha256:93f555c7bc84c658d0c8718c05912c127929a5ef861db06e761e76f426da2ffd

Observation 8d7de844-f4f5-43d7-aa20-28fb6659842a · outbound

This paper cites LLM Agents can Autonomously Hack Websites.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) LLM Agents can Autonomously Hack Websites

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.213871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.213871Z digest=sha256:0ac72e3875d2646ded7205d9afd4092cbfc8e4fb2afed1c890ef57f8ecf41322

Observation 537a98f3-3e25-4ac9-9f28-20ec16beef8b · outbound

This paper cites Graphplanner: Graph- based agentic routing for LLMs.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Graphplanner: Graph- based agentic routing for LLMs

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:18.213523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.217256Z digest=sha256:ef0283ca861c7121f5d2f3e4dc4f5b367f540176a7b5287a4be366a08461b5b1

Observation 8d073036-155f-4786-b95f-11dcadf22c59 · outbound

This paper cites Redcode: Risky code execution and generation benchmark for code agents.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Redcode: Risky code execution and generation benchmark for code agents

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:18.203990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.220501Z digest=sha256:ad1f751334c3c01e7198e98d40366dc780c7cbe8d49ec9974497c3443cd443fd

Observation 39a0136a-f926-47b7-a4fa-5a049b9ef4b5 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.223565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.223565Z digest=sha256:0abb56d7bbfcfe210c9f3e5b84f81e69501125d67f0446712cdc0667dfa5dfd7

Observation 392eb2c5-267a-48d7-8cfd-b33f5b60f0b6 · outbound

This paper cites Getting pwn’d by ai: Penetration testing with large language models.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Getting pwn’d by ai: Penetration testing with large language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:18.193871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.227059Z digest=sha256:3763975c6e378d8670b7124bb232b4d7d2e8535851921ad6d48847dca65f2b0b

Observation c775929a-b989-4410-9ac8-1cc62c818250 · outbound

This paper cites Llms as hackers: Autonomous linux privilege escalation attacks.Empirical Software Engineering, 31(3):70, 2026.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Llms as hackers: Autonomous linux privilege escalation attacks.Empirical Software Engineering, 31(3):70, 2026

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:18.183206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.230065Z digest=sha256:4d97ebe21d411003fa42406d96569d9d11c075e9380e5eda933ebaa228e75056

Observation a9a08bcc-f95d-4e36-9e3e-b8c4f572151a · outbound

This paper cites Deepmath-103k: A large-scale, challenging, decontaminated, and verifiable mathematical dataset for advancing reasoning.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Deepmath-103k: A large-scale, challenging, decontaminated, and verifiable mathematical dataset for advancing reasoning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:18.173065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.233285Z digest=sha256:45438380909cc5a92784d0702856e173c187db1bcee7efb2b5a720623bec1c11

Observation c10c72e4-b6db-44d3-b056-381995d46494 · outbound

This paper cites Qwen2.5-Coder Technical Report.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Qwen2.5-Coder Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.236382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.236382Z digest=sha256:c9d825601a93669efd14eb78b39f21f974d160abe7bd18089163e5082c471bf0

Observation 84a5e156-aac6-4e3a-88ed-2c2efa7cef9b · outbound

This paper cites GPT-4o System Card.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) GPT-4o System Card

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.239797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.239797Z digest=sha256:7bc0c57016449afc8c623f2213b2709d4cc85ef5a3f1b1961a5d497fdcb84440

Observation 28054fda-b27f-4aaf-8f4b-81dd9ae96b37 · outbound

This paper cites Agentic ai for cyber defense: Llm-guided hierarchical multi-agent reinforcement learning.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Agentic ai for cyber defense: Llm-guided hierarchical multi-agent reinforcement learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:18.160033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.243121Z digest=sha256:3734c5feab1be38affeba397a938b25889e56ac1e60d5ae1465537a4fe0dd1e3

Observation 828afaeb-c042-4d3c-8910-8c58d6057b1c · outbound

This paper cites Search-r1: Training LLMs to reason and leverage search engines with reinforcement learning.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Search-r1: Training LLMs to reason and leverage search engines with reinforcement learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:18.148367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.246095Z digest=sha256:35a46408c3038ca55acb9278a313e816db4d87207989719b088cfd12d111be85

Observation 5157025d-7461-46e7-b441-5d66539fec3a · outbound

This paper cites Exploring the efficacy of multi-agent reinforcement learning for autonomous cyber defence: A cage challenge 4 perspective.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Exploring the efficacy of multi-agent reinforcement learning for autonomous cyber defence: A cage challenge 4 perspective

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.249096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.249096Z digest=sha256:24cf20e2b7528766044787d814af6d0ea35a37dab3e47b80c94cea797ddc6c66

Observation e6e1b05a-f842-4dbb-b1e7-d144c3b9786f · outbound

This paper cites Automated cyber defense with generalizable graph-based reinforcement learning agents.arXiv preprint arXiv:2509.16151, 2025.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Automated cyber defense with generalizable graph-based reinforcement learning agents.arXiv preprint arXiv:2509.16151, 2025

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.252347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.252347Z digest=sha256:96505727c730fd0968ef7715fc220f5c4ebeca61e609017b91d44ab20f665dbf

Observation eea105c7-6006-490e-be8b-80e29d897860 · outbound

This paper cites Large language models are zero-shot reasoners.Advances in neural information processing systems, 35:22199–22213, 2022.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Large language models are zero-shot reasoners.Advances in neural information processing systems, 35:22199–22213, 2022

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.255777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.255777Z digest=sha256:067647ada7b17ad625056c6725301c4fffce432fdacf50175008bc081bb6ce85

Observation 002fa239-7cad-40a3-9abc-bd5b8d3cfe68 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Gonzalez, Hao Zhang, and Ion Stoica

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.258881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.258881Z digest=sha256:5ecc5741d7b69fe6f0b8359975588a22e2d519ea998d55df2bb0fc41ce44ba0f

Observation e9181f4b-d3ac-47de-9116-2de7e18ab6a0 · outbound

This paper cites In-the-flow agentic system optimization for effective planning and tool use.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) In-the-flow agentic system optimization for effective planning and tool use

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.910451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.261962Z digest=sha256:9124b21958a4af970ca97b0d2b85a744783b476db206f027a4f9fcfbe079cbef

Observation 122e0f90-d9be-4245-9c27-eee78120d9bd · outbound

This paper cites Code as policies: Language model programs for embodied control.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Code as policies: Language model programs for embodied control

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.264984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.264984Z digest=sha256:1d1a98fc4b70f733cb9197a50db1400ecf1f4f538fb4b70723c1086c6e9e5b55

Observation 74634654-66b4-4099-b9a5-9ac2f6dd52b0 · outbound

This paper cites Et-bert: A contextualized datagram representation with pre-training transformers for encrypted traffic classification.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Et-bert: A contextualized datagram representation with pre-training transformers for encrypted traffic classification

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.896149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.267969Z digest=sha256:6b151fe229c567c8df286bf665e9a0c7340df60d12866ac4e413d2d43f838588

Observation 602324bb-d5ad-459d-9e75-1357f935d447 · outbound

This paper cites Cyberbench: A multi-task benchmark for evaluating large language models in cybersecurity.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Cyberbench: A multi-task benchmark for evaluating large language models in cybersecurity

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.886939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.271027Z digest=sha256:3c04165ec792389382d5d60b50ed903cd1e1e05f111df7c555ce2a48fae60e4e

Observation 7addd55e-ce6c-400f-837c-9d12548168d0 · outbound

This paper cites Visual-rft: Visual reinforcement fine-tuning.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Visual-rft: Visual reinforcement fine-tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.273922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.273922Z digest=sha256:6be9733b4981e827fe306507f8fd36a5cf9304dda495c84a96b82e0023dd14b2

Observation fe77814f-b49e-4ca1-a9d0-27a9fc5cc743 · outbound

This paper cites Contrasting centralized and decentralized critics in multi-agent reinforcement learning.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Contrasting centralized and decentralized critics in multi-agent reinforcement learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.872569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.276936Z digest=sha256:af17b814d402cf3aff4ad47021782d11ea945cb1a123a198dc6bde0337321806

Observation 3113a622-7603-43f8-a95c-d5faabd251b3 · outbound

This paper cites Eureka: Human-level reward design via coding large language models.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Eureka: Human-level reward design via coding large language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.279999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.279999Z digest=sha256:32b26b943ab0c7b5f16a359e6d2689578a513dcf3d267fadcacdeae67b9704a0

Observation 19bfc18e-0300-400e-9f37-06b0db33acbd · outbound

This paper cites Ray: A distributed framework for emerging {AI} applications.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Ray: A distributed framework for emerging {AI} applications

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.283080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.283080Z digest=sha256:770f29425766a08ff931b70c89d49bc5287128674794b55531ec1ecf2d0ba7a1

Observation 3457e4e5-f661-4b23-8580-db69c17dbeaf · outbound

This paper cites Experience with emerald to date.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Experience with emerald to date

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.853501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.285979Z digest=sha256:b5c90b155ac3a88f9d04c98d6ce9052b02d2e2c9021e87dd5f317855b1ee7c39

Observation bbccf353-cf5e-42ab-b437-e87fa106e2e5 · outbound

This paper cites Towards a high fidelity training environment for autonomous cyber defense agents.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Towards a high fidelity training environment for autonomous cyber defense agents

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.843871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.289324Z digest=sha256:27693fc292ca61b925a52c9aba78ee7e2d833945408b205c824e7bdc0887cbd1

Observation 6de97952-a83d-4464-a26e-8cf90eb500df · outbound

This paper cites Proximal Policy Optimization Algorithms.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Proximal Policy Optimization Algorithms

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.292586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.292586Z digest=sha256:cca604bffea7054e02039e2a6c3a4de8af47f39dae91ce8bf397a86dbcdbf26c

Observation a004ecfb-09e6-44c3-986f-43b6cbf9ea81 · outbound

This paper cites Nyu ctf bench: A scalable open-source benchmark dataset for evaluating llms in offensive security.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Nyu ctf bench: A scalable open-source benchmark dataset for evaluating llms in offensive security

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.834168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.295843Z digest=sha256:24198ee73fa6debc74690c4c860de7c3020847bda33634f9bdb70f62030813e6

Observation 3a10cc6c-cbfd-46c2-a99b-629db9076462 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.298765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.298765Z digest=sha256:ec26d36c78ec74f8c8b96c306cefd1b8a059c3b6f473b0a60310940e32467cd1

Observation 97c601a8-7810-4d7f-93d1-c46ca867f66b · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Hybridflow: A flexible and efficient rlhf framework

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.301997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.301997Z digest=sha256:ef5611195496c49c9f2b762f02cf47bf1338ee932d7ee1ae08eb9c3a378685dd

Observation b6fbeb53-181b-4e61-9d1c-dbe37dd0d1ce · outbound

This paper cites Hierarchical multi-agent reinforcement learning for cyber network defense.Reinforcement Learning Journal, 6:790–810, 2025.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Hierarchical multi-agent reinforcement learning for cyber network defense.Reinforcement Learning Journal, 6:790–810, 2025

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.819801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.305358Z digest=sha256:dd4088329f5937adc32431d4ab1d97da758d695b5f7ab47cc38026230af91219

Observation 0d1bded1-97d5-40eb-8495-315f36c22d2b · outbound

This paper cites A taxonomy of intrusion response systems.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) A taxonomy of intrusion response systems

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.810253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.308716Z digest=sha256:82b90d39b509a8f152d5fed6acf93bd6e5a837fd88b5088f53cd0ea94f727fc9

Observation 2f99b4cb-176e-44b6-8d9e-9af9daa5eac8 · outbound

This paper cites Redsage: A cybersecurity generalist LLM.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Redsage: A cybersecurity generalist LLM

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.800806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.311692Z digest=sha256:9ff2dacf3fb2968d577e1cc81134aa24855acbceae4ac909db9cf911a83f47cf

Observation 294053bb-943e-470b-9ada-fd18ddb5bad4 · outbound

This paper cites Cyberbattlesim, 2021.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Cyberbattlesim, 2021

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.791420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.314759Z digest=sha256:6860f446b3596fa862ea772f88c8c49b4d0b3b5e0327b470a70130782aa6d257

Observation 1782e2f7-4dd3-4e2a-9cab-129cb7519f89 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Qwen2.5: A party of foundation models, September 2024

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.317755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.317755Z digest=sha256:3125504223be5d0982315ac030ac92a307cc6e17584961f54b4649df2f79d3cf

Observation 52f80ab4-0d91-46e9-be5e-44d05215a9f5 · outbound

This paper cites CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.320654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.320654Z digest=sha256:d3d6999befa536b203d473b70626732cdf0363c0c55c352566b4aba28f43f811

Observation c66d7f38-f65f-4667-bcd0-2e9622ff5b4e · outbound

This paper cites SPPO: Sequence-Level PPO for Long-Horizon Reasoning Tasks.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) SPPO: Sequence-Level PPO for Long-Horizon Reasoning Tasks

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.324049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.324049Z digest=sha256:c27da8e4ef0b71d2c98cc1ee9219aaa9fc56043e57e8d9e42d2187740deebd32

Observation 2c23d796-9356-45ec-a517-60c7e6bfa970 · outbound

This paper cites SymRTLO: Enhancing RTL code optimization with LLMs and neuron-inspired symbolic reasoning.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) SymRTLO: Enhancing RTL code optimization with LLMs and neuron-inspired symbolic reasoning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.776191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.327331Z digest=sha256:08dc63d6e2ba871cba60f201ddf6fb34267922175da66803ba59b5ba5661aa03

Observation 5dfeadbe-4bab-4196-a7e0-b38991215f8c · outbound

This paper cites Cyber- gym: Evaluating AI agents’ real-world cybersecurity capabilities at scale.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Cyber- gym: Evaluating AI agents’ real-world cybersecurity capabilities at scale

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.330596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.330596Z digest=sha256:fded9ca68674387e6ccbc1cba1f9205011633262866500951635d75b947a4a42

Observation 7d765deb-4598-47d6-9abe-0b9c9fe38835 · outbound

This paper cites SWE-RL: Advancing LLM reasoning via reinforcement learning on open software evolution.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) SWE-RL: Advancing LLM reasoning via reinforcement learning on open software evolution

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.762294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.333659Z digest=sha256:78cd4a7979e548cde347d7a707afb39e0df92295a7a3188ad4207e6d7bd0c17b

Observation 7967ed37-304e-4d30-8fef-a0d5f8595db7 · outbound

This paper cites Autogen: Enabling next-gen LLM applications via multi-agent conversations.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Autogen: Enabling next-gen LLM applications via multi-agent conversations

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.752947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.336740Z digest=sha256:22b7bb32e87933574aa971985e3b5ff39c40bd55c9420d52269d61eca7598335

Observation eee7a847-1c5c-464e-8a27-af119d650e66 · outbound

This paper cites Qwen2 Technical Report.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Qwen2 Technical Report

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.339747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.339747Z digest=sha256:99010ecbd9e064562249f59c20201e67e87ea0f919e736ac601549d8a5431c83

Observation c3698d0f-5488-44a5-81a1-c9642f22a97c · outbound

This paper cites Intercode: Standardizing and benchmarking interactive coding with execution feedback.Advances in Neural Information Processing Systems, 36:23826–23854, 2023.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Intercode: Standardizing and benchmarking interactive coding with execution feedback.Advances in Neural Information Processing Systems, 36:23826–23854, 2023

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.343099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.343099Z digest=sha256:77c875b34d8aa258e4cb5ac360fe1da82aeb4ce8ac9050e6ce86edb83835331f

Observation f1dc883e-3319-46a9-bdee-54080f81ded9 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) React: Synergizing reasoning and acting in language models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.346037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.346037Z digest=sha256:84e880c695b1afd95bfc8987e8b28387470df85330da541e75a18491276d3a86

Observation 1e522d46-fdde-43c2-bcf0-37c6e7d88933 · outbound

This paper cites Primus: A pioneering collection of open-source datasets for cybersecurity LLM training.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Primus: A pioneering collection of open-source datasets for cybersecurity LLM training

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.734052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.348903Z digest=sha256:93770afe2699f47c7be356c8a4d8184356d0e7bc48e24bee1d2587097bfa7dd0

Observation 1caec6f9-8983-4b60-8d5d-737a854fc042 · outbound

This paper cites ACECODER: Acing coder RL via automated test-case synthesis.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) ACECODER: Acing coder RL via automated test-case synthesis

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.724919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.351778Z digest=sha256:84ae8071554231089ce180a21d5fa1cc468c25ff6f276580e6d4bc9ba7d6736a

Observation 367b9b5e-445a-46ef-b1b6-e52272cd4fa2 · outbound

This paper cites Ho, and Percy Liang.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Ho, and Percy Liang

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.354780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.354780Z digest=sha256:19dc164e68c1e17e72d0d0fbaa6471cbb4e558ab703252789f2ba7a1931c9bd6

Observation 13fc7108-c6b0-42d2-9814-7d4992243321 · outbound

This paper cites Abdi, William Blum, and Muhammad Abdul-Mageed.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Abdi, William Blum, and Muhammad Abdul-Mageed

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.710547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.357824Z digest=sha256:843d6fc94811f396d574451fce8af262d925d26427f1181c98e019116be708c4

Observation e13e38df-282a-43f7-a6e7-425f58fe3047 · outbound

This paper cites Yet another traffic classifier: A masked autoencoder based traffic transformer with multi- level flow representation.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Yet another traffic classifier: A masked autoencoder based traffic transformer with multi- level flow representation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.700836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.360906Z digest=sha256:e5e1f0b96c78299f91688fef0b1ec7c21a97c9143223cdf0399f74817af52ef0

Observation 0ad2c876-7da4-447a-9c50-02f548644f98 · outbound

This paper cites Curran Associates Inc., Red Hook, NY , USA, 2019.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Curran Associates Inc., Red Hook, NY , USA, 2019

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.690889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.363877Z digest=sha256:8d626fde5c93d717739a9c018106ab16756144f9749a065feaf2d15948903298

Observation 7c3dd15d-4a6b-4b36-974e-70e9254d75d6 · outbound

This paper cites More than just functional: LLM-as-a-critique for efficient code generation.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) More than just functional: LLM-as-a-critique for efficient code generation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.681374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.366857Z digest=sha256:d30c2a6e49d89974411bab5790f8015a0cc25a6dc3c0bc6c2072406f6ee8db15

Observation 9750a2ff-fd21-4e2c-8c61-455322623ec2 · outbound

This paper cites CVE-bench: A benchmark for AI agents’ ability to exploit real-world web application vulnerabilities.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) CVE-bench: A benchmark for AI agents’ ability to exploit real-world web application vulnerabilities

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.671444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.369963Z digest=sha256:db33c6f67dc29a2e5dcf0c462eb718ad9a735df4844efc2d4f8592f98b763ae0

Observation bf259fbf-ae1a-4a4e-8ced-017136dbe910 · outbound

This paper cites Teams of LLM agents can exploit zero-day vulnerabilities.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Teams of LLM agents can exploit zero-day vulnerabilities

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:58:17.661621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.372961Z digest=sha256:5c7c5a0c97e276c581a70fe7d4b8135b8f002dd5984abd797dcacfc20713be20

Observation c46cf609-583b-4ba5-b8a1-54ebd9e9fbf7 · outbound

This paper cites Cyber-zero: Training cybersecurity agents without runtime.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Cyber-zero: Training cybersecurity agents without runtime

Reference 67

Resolution
malformed identifier
raw_fallback, observed 2026-08-08T19:58:17.476910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:58:17.376083Z digest=sha256:54037182f7cbe11a6580356c2f43a2292ab61559cdd8633b16c2a763308d76a8

Pith citing papers

No inbound Pith citation observations are available.