Pith. sign in

Paper Citation Record · LEDGER

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models

As of 18 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 3 inbound Pith citation observations for arXiv:2506.14682.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14682 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:53:33.486755Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:25:15.771731Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T23:51:45.015967Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved25
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5786a459-44e7-455b-ae16-95ad0238f8f1 · outbound

This paper cites IRIS: LLM-Assisted Static Analysis for Detecting Security Vulnerabilities.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models IRIS: LLM-Assisted Static Analysis for Detecting Security Vulnerabilities

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.591918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.591918Z digest=sha256:f83da92942cbf92cd582d276435964c6759b227b8bee8f0c940da210d68dc5da

Observation d6d8958e-1ec8-4462-9ae4-d15fa45b3c4b · outbound

This paper cites LLMs in Software Security: A Survey of Vulnerability Detection Techniques and Insights.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models LLMs in Software Security: A Survey of Vulnerability Detection Techniques and Insights

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.622526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.622526Z digest=sha256:6bdc753c055b2330592ec00f49eab68dcb6d5fadf9200cf29fdbbd3018d8ed89

Observation a2bdf34d-1c59-44e8-ba29-1d6dd7929034 · outbound

This paper cites An Empirical Evaluation of LLMs for Solving Offensive Security Challenges.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.627164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.627164Z digest=sha256:c3dc3ec9099799962b4c8a5befd0a02629e8c287df61e9cefd5108384178f023

Observation 8cdf4899-13d1-4341-9b92-d3897291a8c3 · outbound

This paper cites LLM Agents can Autonomously Hack Websites.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models LLM Agents can Autonomously Hack Websites

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.631740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.631740Z digest=sha256:b7adc7253f706309b09f123d93dcf627bcd5c021a1ac14e902390e25d0d8375a

Observation b8956961-49fd-4354-b263-91394b2e0055 · outbound

This paper cites Llm4decompile: Decompilingbinarycodewithlargelanguagemodels,.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Llm4decompile: Decompilingbinarycodewithlargelanguagemodels,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:34.298021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:53:32.637031Z digest=sha256:d64fbce00b24a6cf278942fc1855ae43405d3cd9780e424685a4087f85eca351

Observation f6cc0e37-7723-497a-a17c-3d861bd3a300 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Evaluating Large Language Models Trained on Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.644504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.644504Z digest=sha256:09eeeda714e06307640376bac6e2d080c58116df84b91b72df9b3f7c44231e64

Observation b17f629e-780e-4ca8-b4e4-0430f937be60 · outbound

This paper cites ATLAS - Adversarial Threat Landscape for Artificial-Intelligence Systems.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models ATLAS - Adversarial Threat Landscape for Artificial-Intelligence Systems

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:34.287871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:53:32.674375Z digest=sha256:3c8e845f82018c601a00bf924e0f80010580944762b61e9347ed7ccd2fdbd4cc

Observation a10781cf-654c-441a-ac0e-f37410e03917 · outbound

This paper cites OWASP Top Ten for Large Language Model Applications.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models OWASP Top Ten for Large Language Model Applications

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:34.276513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:53:32.679927Z digest=sha256:48262e53b9df80330dfd75867b0cc9e2a1a126f5f00f590fe1191d95f15ad741

Observation cd5fef5e-240c-472a-953b-0c62d239a656 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Measuring Massive Multitask Language Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.684251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.684251Z digest=sha256:2cad08dd081ee1b07ffeff02cf7ff8b4ca34ab7d40c2fd191433b43f687a7e04

Observation ee12daaa-f90a-435c-a0a1-70b9f18a4511 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Training Verifiers to Solve Math Word Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.689085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.689085Z digest=sha256:eee43d6e330db1e11912b209c02ce97b03297e386bd99fa41d17a978ebf308e2

Observation 786fffbf-ad22-4c29-9c6b-fc7f941845ed · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.693437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.693437Z digest=sha256:b99e63899fc83b3fce49378e74d48ef2ce9d7ab0e1edb8bbcea629cc7817931e

Observation 1aec31e9-929e-4d0b-91d4-37a68ad00980 · outbound

This paper cites OpenAI o1 System Card.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models OpenAI o1 System Card

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:34.262382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:53:32.697337Z digest=sha256:9064c748ca53061a4897a6bc920976c9e5c5cfdcfea6cb35959fee4bf8873f03

Observation 648e0d2c-4834-4711-b100-5d48e8b50c1e · outbound

This paper cites Achieving >97% on GSM8K: Deeply Understanding the Problems Makes LLMs Better Solvers for Math Word Problems.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Achieving >97% on GSM8K: Deeply Understanding the Problems Makes LLMs Better Solvers for Math Word Problems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.701223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.701223Z digest=sha256:1038339ec596c65788eb0dd501cbd35ef37c939420ba8a6af117e58df2984ad6

Observation 37b18dbc-7422-4518-9dc3-0f954283512b · outbound

This paper cites Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive Principles.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive Principles

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.705945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.705945Z digest=sha256:0091fbf6523cb039bc8497973cc2fa17eb0df7ee6c132e4689313c57c8aba31b

Observation 415e0e21-ec74-46e4-b7a1-af7ae258fcce · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.725506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.725506Z digest=sha256:f035ef7489f27aa1192b3c8462fa72ed16b5ea1e7806c28351ae71935a22da9a

Observation 08056cae-2d1e-4ed9-962f-f7d42c40eb71 · outbound

This paper cites Jimenez, John Yang, Kai Liu, and Aleksander Madry.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Jimenez, John Yang, Kai Liu, and Aleksander Madry

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:34.249705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:53:32.756599Z digest=sha256:95407c04040553f19063e38a74d5ed0b2460f0ebbcc0813fdb0a8e10e97db62f

Observation 113acd9c-9a00-4443-9b82-41d4088c02d2 · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.783357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.783357Z digest=sha256:15521d89e519388148b3397a07922a3c149b984cade5b0733fefeb9ed0e1ad01

Observation a39166fe-2e49-45b5-96c4-593488160ad8 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models AgentBench: Evaluating LLMs as Agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.791730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.791730Z digest=sha256:c300517caa79f0b6c4fccf05fd4f0232629ce75f2720534e2fd23e49841b8cf3

Observation 1ac4ff36-5be3-4839-8875-091815b144b3 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.795634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.795634Z digest=sha256:5fc2933ed066ad89984440e09c044e12801f3dcf3368f565f98283080476572f

Observation b8383111-e426-45e0-b10a-14028c21132d · outbound

This paper cites Mind2Web: Towards a Generalist Agent for the Web.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Mind2Web: Towards a Generalist Agent for the Web

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.803200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.803200Z digest=sha256:99e9cd8ba7d5ebab196e6ef614e62ca37bbc048b281e477c80d44d68749f2ab5

Observation 1608e2a3-8a37-42b2-9663-684db5f7e669 · outbound

This paper cites SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.807624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.807624Z digest=sha256:076482a9360c550174b5f12b6aa2f7e41911f2ef3c8853af57bd8a5e8bb127b5

Observation dd6e4679-7a17-4487-8954-34b71f13326a · outbound

This paper cites an unresolved cited work.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:53:34.149152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:53:32.814010Z digest=sha256:b860e133db2a64561aa1a2f7c5a522bb8978797e2286a0c50eed5f8e52eae31b

Observation 0a2b9505-3ecd-448a-985a-de94353d94c7 · outbound

This paper cites InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.840618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.840618Z digest=sha256:a2cafead977f4966ea49535bd37b1bd7a915b7a08e0878749968b0e1e4ae8e29

Observation ef0eb5b6-ae4b-416e-a836-0b3f445ab6ca · outbound

This paper cites NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.863240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.863240Z digest=sha256:01b772f55e9205e871ddb42b52cb7dbf21e696f5c37d2ec4502fda57dfe428a8

Observation 8cbe8746-024a-480d-8048-9548adcda81f · outbound

This paper cites EnIGMA: Interactive Tools Substantially Assist LM Agents in Finding Security Vulnerabilities.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models EnIGMA: Interactive Tools Substantially Assist LM Agents in Finding Security Vulnerabilities

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.885270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.885270Z digest=sha256:f0b1ca7405af8c7f9fd32190aee636eac49f4aeb8bee39b6e1eb3c1ac5d81329

Observation 24625837-5955-4a37-a26f-c90c28372f39 · outbound

This paper cites AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.995201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.995201Z digest=sha256:e63853c8b6efcd13272738c794080f27b23bcafee6da50435f6b10c4727a2567

Observation 38c7678a-0a84-41de-9196-0ece445f7681 · outbound

This paper cites Picoctf learning.https://www.picoctf.org/, 2024.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Picoctf learning.https://www.picoctf.org/, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:34.077377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:53:33.103438Z digest=sha256:8c689afd8cc2cecdba28183b92d0ee236f0dd74ac274d7dcac1ab5877444cd15

Observation 54ec6960-1b49-4018-b144-1c0b1af26992 · outbound

This paper cites Ai village capture the flag @ defcon31.https://kaggle.com/ competitions/ai-village-capture-the-flag-defcon31, 2023.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Ai village capture the flag @ defcon31.https://kaggle.com/ competitions/ai-village-capture-the-flag-defcon31, 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:34.026552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:53:33.122420Z digest=sha256:273095f09faa0e26a995e3387af722d77e2723a1ed59f760355949116e5c2078

Observation d877e2e1-e64c-4c89-8521-0132a06329a0 · outbound

This paper cites Jupyter datascience notebook docker image, 2025.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Jupyter datascience notebook docker image, 2025

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:34.016063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:53:33.131197Z digest=sha256:a8aaceedcd0ccc708bbee00be9d4a4fa4920743844d3edc53ffa42735b5e388f

Observation 72825d09-5f60-42c1-8844-985a27434988 · outbound

This paper cites Optimizing Large Language Model Hyperparameters for Code Generation.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Optimizing Large Language Model Hyperparameters for Code Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:33.154793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:33.154793Z digest=sha256:d4f0012fa04b66952879618ef5845549f9b25b0264c968fb987010f34cb2bef7

Observation dbdd6fc6-cfb3-4bf4-9e40-81635bd17218 · outbound

This paper cites The Automation Advantage in AI Red Teaming.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models The Automation Advantage in AI Red Teaming

Reference 31

Resolution
malformed identifier
no resolver link, observed 2026-08-15T19:53:33.168800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:33.168800Z digest=sha256:b36b65bfca81673a3b842fc707a10d5272fa043ecd8099681199d19dc962e717

Observation 4d833146-c9cd-4247-aa55-ccfec3df252c · outbound

This paper cites For example : ‘t = turtle.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models For example : ‘t = turtle

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:33.981683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:53:33.281100Z digest=sha256:8feef4bfbc792c9a287eb9e86ef32722b39b0807f749b8052c5dd4859a8d8021

Observation 69d706f8-82b6-4039-bad0-b471a24bdfc3 · outbound

This paper cites S Y S T E M _ C O M M A N D _ E X E C U T E D _ V I A _ E X E C.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models S Y S T E M _ C O M M A N D _ E X E C U T E D _ V I A _ E X E C

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:33.970997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:53:33.295834Z digest=sha256:4e0b263c1edd269d8f5cac2db09132bbf27556a230f77033cad72bbf259fa7a2

Observation ec659ead-6b3c-40ff-8c8f-0b0fbabd57e1 · outbound

This paper cites an unresolved cited work.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:53:33.959915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:53:33.306754Z digest=sha256:f348b1975b25c5d27fcf05e36e4433adb3a5ea43c6dc158bc8adaca9c9490827

Observation d19cd96b-6aff-4075-9e4c-415cf88b77b7 · outbound

This paper cites For example : ‘t.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models For example : ‘t

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:33.949128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:53:33.310734Z digest=sha256:3dda5de80cb849f7edb5fe3b5e70a1a440597da143501e0a1d8c272b0e7b5806

Observation 1953971f-4bdf-445d-a1dc-cdb47053335a · outbound

This paper cites " " response = query ( prompt ) print ( response ) </ execute - code > <result idx=0 success=True> ’output’:.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models " " response = query ( prompt ) print ( response ) </ execute - code > <result idx=0 success=True> ’output’:

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:33.911454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:53:33.314088Z digest=sha256:f1ec76f476bbb5eceda59afdf09f0a894efe4c4a279318f9e9b5a8d2332a200b

Observation 54452e34-cca5-47bf-84db-6e10008c0787 · outbound

This paper cites an unresolved cited work.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:53:34.004403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:53:33.325454Z digest=sha256:f123807ec423b5a2a606456d8d094116a355eb7a6bbbb3022b26d1e468e4a102

Observation 26208294-c958-4518-9989-3308f74bc1ef · outbound

This paper cites an unresolved cited work.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Unresolved cited work

Reference 41

Resolution
parse uncertain
raw_fallback, observed 2026-08-15T19:53:33.992721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:53:33.353697Z digest=sha256:1ca129e3dd9966d01fa19f5ed3c25dea43975c24f4fdd4450e6e9f0e353ab56f

Observation a264a8e2-edc5-4510-bf98-42b4c350f91e · outbound

This paper cites Example : ‘t = turtle.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Example : ‘t = turtle

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:33.840388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:53:33.388384Z digest=sha256:361da3bd8c65b1633e4f003a932a6b07fc3044c46cad016d00b309c66e55224c

Observation f768c676-7a52-4781-8769-fc5bad636043 · outbound

This paper cites echo J A G A _ L L M _ G E N E R A T E D _ O S _ C O M M A N D.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models echo J A G A _ L L M _ G E N E R A T E D _ O S _ C O M M A N D

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:33.794439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:53:33.430804Z digest=sha256:9170ee9fa689669d7a9c2a6225db09322d50f977aee5d9db162e98726794c87f

Observation 46ed088a-d71e-42bc-a856-3b11fd256ca3 · outbound

This paper cites system () ‘ call , make the turtle ‘t ‘ draw a square of side 10.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models system () ‘ call , make the turtle ‘t ‘ draw a square of side 10

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:33.783066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:53:33.465619Z digest=sha256:f3f5a539261ea8a3544d4449a42cac0bf5a02d8152a549820a708a94699be603

Observation ef0f6acf-4d21-45a9-ba3c-2e6328803dec · outbound

This paper cites " " response = requests . post ( f.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models " " response = requests . post ( f

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:33.770156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:53:33.486755Z digest=sha256:e4f0a31fb91f98f8da3b86d9c5d699f31f8cecf784878c867d00ce483ab06326

Observation c505c07d-fa71-48f9-a6f4-360ef26a4f7d · outbound

This paper cites LLM4Decompile: Decompiling Binary Code with Large Language Models.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models LLM4Decompile: Decompiling Binary Code with Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.640743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.640743Z digest=sha256:785f33b5de964efe749cd492f33fe5b8f9bda9ce46867c1abb50cef56a4250a0

Pith citing papers

Observation d2a4ce5b-b37f-44f8-b429-c30af3ddcabc · inbound

Running in CIRCLE? A Simple Benchmark for LLM Code Interpreter Security cites this paper.

Running in CIRCLE? A Simple Benchmark for LLM Code Interpreter Security AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:25:15.771731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:25:15.771731Z digest=sha256:91c328788ab4796a55f83ba48e0cc1f9461429345bbff864a9015a41f9977e15

Observation 3dff4f9f-4458-46a4-955f-7c6083e29f80 · inbound

PoCo: Agentic Proof-of-Concept Exploit Generation for Smart Contracts cites this paper.

PoCo: Agentic Proof-of-Concept Exploit Generation for Smart Contracts AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T00:09:33.025834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:09:33.025834Z digest=sha256:8cbdbe5a8106ffa8ea2075656fe8c2ad7f1e6df906a1d31a37fbeb915ed68df5

Observation 898cf871-8391-4070-931f-3fb691354ef7 · inbound

Redefining AI Red Teaming in the Agentic Era: From Weeks to Hours cites this paper.

Redefining AI Red Teaming in the Agentic Era: From Weeks to Hours AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:51:45.056581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-07T16:06:18.057868Z digest=sha256:e31946932676ac522a12e3a70ca2bf6ba97020fd79ed19e756508353e5a9152d