Pith. sign in

Paper Citation Record · LEDGER

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation

As of 10 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2608.03166.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03166 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:29:58.290673Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact2
  • verified fuzzy3
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch8

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b9fac2ea-4937-4cc2-8d6d-b81eafb9675c · outbound

This paper cites The Oscars of AI Theater: A Survey on Role-Playing with Language Models.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation The Oscars of AI Theater: A Survey on Role-Playing with Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:55.891371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:55.891371Z digest=sha256:c23d5eb7867aba9b61fa832f7fdee3f14b88479c2cda0857442729bdd0b5369f

Observation 00c93cdc-df59-47ba-bbb0-569dfd3da8a1 · outbound

This paper cites ASSERT: Automated Safety Scenario Red Teaming for Evaluating the Robustness of Large Language Models.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation ASSERT: Automated Safety Scenario Red Teaming for Evaluating the Robustness of Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:55.943379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:55.943379Z digest=sha256:05640d398c9a9edfbe3952017a5c0216f8940d570512ff5db9266a6b9acdbd8c

Observation 103f45a2-66bb-4496-98f7-7f48abc91fd2 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:56.028755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:56.028755Z digest=sha256:63b7bec35f63ce5d40a61926942b70e93322adfb331973ad53b8b0ff33482d9a

Observation 59846c01-4ae4-47c3-9ac5-60a8dff937ac · outbound

This paper cites RoleMRC: A Fine-Grained Composite Benchmark for Role-Playing Language Agents,.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation RoleMRC: A Fine-Grained Composite Benchmark for Role-Playing Language Agents,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:30:00.431231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:29:56.129225Z digest=sha256:d4cde4eada93b9562e1aa6d561e03798849400fc01b224ce5e4e1d38352240ce

Observation 153fcb74-d969-4170-8d67-9e08ce1803b1 · outbound

This paper cites AART: AI-Assisted Red-Teaming with Diverse Data Generation for New LLM-powered Applications.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation AART: AI-Assisted Red-Teaming with Diverse Data Generation for New LLM-powered Applications

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:56.208408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:56.208408Z digest=sha256:9f0e24ad63c6c2251b7833ec002089f714e73329e3191491cef157c1510ebb8c

Observation 8416447f-0537-483c-91d7-93dbd9f4b2fc · outbound

This paper cites Adversarial Testing in LLMs: Insights into Decision-Making Vulnerabilities.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Adversarial Testing in LLMs: Insights into Decision-Making Vulnerabilities

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-06T00:29:59.822893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:29:56.341529Z digest=sha256:a50eba623b7c748d10fc7d65a80576f3a89447872b11bfb6cd6cb40715f2ed71

Observation 4dded492-67ed-4285-83f1-5d72175ab62e · outbound

This paper cites Red Teaming Large Language Models: A Comprehensive Review and Critical Analysis,.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Red Teaming Large Language Models: A Comprehensive Review and Critical Analysis,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:30:00.272974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:29:56.408327Z digest=sha256:68cf70286f708f661f8282c7ba1c9b633ac5ad5c9594ed31db6e64fcd272ffb5

Observation aa86cdb8-632d-499e-98c9-a29c2ba83f91 · outbound

This paper cites Security of LLM-based Agents: Attacks, Defenses, and Applications,.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Security of LLM-based Agents: Attacks, Defenses, and Applications,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:30:00.131587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:29:56.485119Z digest=sha256:f5c7611548c7a805b987263b3bf480f3372b9fc8e86341cc6f49ca161714bae4

Observation 4dc2bad5-597b-4a54-b856-a8b2916e6f24 · outbound

This paper cites Markov-Enhanced Clustering for Long Document Summarization: Tackling the 'Lost in the Middle' Challenge with Large Language Models.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Markov-Enhanced Clustering for Long Document Summarization: Tackling the 'Lost in the Middle' Challenge with Large Language Models

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T00:29:59.970446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:29:56.532477Z digest=sha256:b32c573b1b91e048a398915c1423e8df991a023788791f9713ddb778028a9f36

Observation 51ee87da-f0b5-42b3-9169-f8cfd2021453 · outbound

This paper cites Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:56.632621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:56.632621Z digest=sha256:bf6b8c555feb580f494efb48991af27accc93585e45b3d4d102d95567a778b09

Observation 593966c0-7a4e-4ad2-93af-ea9b594737a8 · outbound

This paper cites RedDebate: Safer Responses Through Multi-Agent Red Teaming Debates.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation RedDebate: Safer Responses Through Multi-Agent Red Teaming Debates

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:56.700795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:56.700795Z digest=sha256:130b44a23bae045d331426a8f4d69871d92e0e72e99824e6dcba6d1f719b6b12

Observation edb960e9-8d02-478f-ab19-7cf4cadd7b88 · outbound

This paper cites MART: Improving LLM Safety with Multi-round Automatic Red-Teaming.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation MART: Improving LLM Safety with Multi-round Automatic Red-Teaming

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:56.778717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:56.778717Z digest=sha256:b69ea68787b61776616919770196ced877e7f97a93084c31486d0af530e4fb7f

Observation fcd9b88f-f8d7-48cb-8a23-07b0bcc1fd3d · outbound

This paper cites Evil Geniuses: Delving into the Safety of LLM-based Agents.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Evil Geniuses: Delving into the Safety of LLM-based Agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:56.858992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:56.858992Z digest=sha256:890e1c32a3ae0065594d41c7ba029df42c090539aed3b82df83f3b3fb403430d

Observation 2a562869-30fa-40f6-a310-093383e3a0f5 · outbound

This paper cites RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:56.930215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:56.930215Z digest=sha256:3de39d0b013def63d7ed494da32b2655975d07f2b4d375fe7d7ce93bcaa268bd

Observation d28070b9-00f4-4e73-b1d8-8654f92726e6 · outbound

This paper cites Encounter-based model of a run-and-tumble particle with stochastic resetting.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Encounter-based model of a run-and-tumble particle with stochastic resetting

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T00:29:59.614995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:29:56.998828Z digest=sha256:77a98dcc64891fab0ff44ba92a5c1f931e45a717b12ace1bbcf0cefda8d2ec69

Observation be359958-702e-4cb7-a28d-851c99af979f · outbound

This paper cites Exposing Weak Links in Multi-Agent Systems under Adver- sarial Prompting,.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Exposing Weak Links in Multi-Agent Systems under Adver- sarial Prompting,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:57.065596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:57.065596Z digest=sha256:62e0f9f8cb09a0d925bbc67d40c896026c54c8d6c8b3836eb5b6db1868835d08

Observation 6913ec36-c166-4117-bbd4-745fc6aace07 · outbound

This paper cites Iwasawa module of the cyclotomic $\mathbb{Z}_{2}$-extension of certain real quadratic fields.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Iwasawa module of the cyclotomic $\mathbb{Z}_{2}$-extension of certain real quadratic fields

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T00:29:59.270567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:29:57.300906Z digest=sha256:01610308b9c04552d364bdf58664bce6baa860e63cd93411dce20d194a9a29e8

Observation c6ca3504-9d98-4ad8-8114-0bcf1183b914 · outbound

This paper cites SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:57.375647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:57.375647Z digest=sha256:276a362240fd9529ab60794ac560b1ea51edef7b59510128acb022c3f70f5782

Observation a3dc3093-1dfc-4dff-8728-9dbef0276e2e · outbound

This paper cites From Prompt Injections to Protocol Exploits: Threats in LLM-Powered AI Agent Workflows,.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation From Prompt Injections to Protocol Exploits: Threats in LLM-Powered AI Agent Workflows,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:57.468919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:57.468919Z digest=sha256:8e7ac3daf848d505521ce2a2384ab876bea1c192737f58a5dc6737f3c272e932

Observation 891d8d25-88bf-48f7-bdc6-ff1dad7283dc · outbound

This paper cites Stable Diffusion For Aerial Object Detection.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Stable Diffusion For Aerial Object Detection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:57.573900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:57.573900Z digest=sha256:59f18d9b802689fd0a17ffa0b97305118c2ee63b53dcf5a52adf2fa88259a138

Observation dc053066-a483-4ae5-860e-62584c6b65a2 · outbound

This paper cites Online Adaptive Traversability Estimation through Interaction for Unstructured, Densely Vegetated Environments.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Online Adaptive Traversability Estimation through Interaction for Unstructured, Densely Vegetated Environments

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T00:29:59.083036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:29:57.693775Z digest=sha256:0e75543148d3ee251957994c4daa1639795f64dc9ac7d59264e462b39a0ad2e9

Observation 3974ba24-355a-435e-8cb6-3c8a78a87e94 · outbound

This paper cites Learning to Adapt to Position Bias in Vision Transformer Classifiers.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Learning to Adapt to Position Bias in Vision Transformer Classifiers

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T00:29:58.912857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:29:57.773801Z digest=sha256:0f785ec032b6b97f032710253396f44b4f4668ee1d1641005b0f12768aa36603

Observation ca0e29d1-e053-4708-b07e-129e648b92ac · outbound

This paper cites Compound Expression Recognition via Large Vision-Language Models.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Compound Expression Recognition via Large Vision-Language Models

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T00:29:58.788800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:29:58.020373Z digest=sha256:bfcc87e49040aa6fbd1fe18975d0c7ce02e4417954daf9ba5cc8030311063dce

Observation b455f210-342a-402c-a79c-e6bb5e790e4b · outbound

This paper cites Context-Aware Two-Step Training Scheme for Domain Invariant Speech Separation.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Context-Aware Two-Step Training Scheme for Domain Invariant Speech Separation

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-06T00:29:58.622840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:29:58.108401Z digest=sha256:61bd9275226036536d45837227bbc45a190b4a51c20e211218a9e92370e3c28f

Observation 5a58fb64-3080-4b1b-b706-afef034d9e26 · outbound

This paper cites The Dark Side of Human Feedback: Poisoning Large Language Models via User Inputs.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation The Dark Side of Human Feedback: Poisoning Large Language Models via User Inputs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:58.203660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:58.203660Z digest=sha256:3471dba70dae3fa08e8a1d53380271d307000c1509740430262ff751a53291c4

Observation d7788856-23e0-4090-9d8f-4f6cd1ed82a9 · outbound

This paper cites Lower bounds on the $\ell$-rank of ideal class groups.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Lower bounds on the $\ell$-rank of ideal class groups

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T00:29:58.420619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:29:58.290673Z digest=sha256:f109133c35139aacc14aa13021b2746f7270c15d1ebbe6899b2af5a5b747b52b

Observation 58b5660d-6260-46e1-908c-065a195af10c · outbound

This paper cites Mitigating Label Noise on Graph via Topological Sample Selection.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Mitigating Label Noise on Graph via Topological Sample Selection

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:57.946969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:57.946969Z digest=sha256:b176ba809642d95382b763e186c2e7a1593a61397b20f450fab386775bfdcc78

Observation d839405d-1902-4356-b701-0d250ea816f7 · outbound

This paper cites Evaluating the Impact of Verbal Multiword Expressions on Machine Translation.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Evaluating the Impact of Verbal Multiword Expressions on Machine Translation

Reference 2025

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T00:29:59.419281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:29:57.241797Z digest=sha256:04f9b2670748d3d6c6803453a696947cdd8996db9b14a771821dc71de7889b7e

Pith citing papers

No inbound Pith citation observations are available.