Pith. sign in

Paper Citation Record · LEDGER

PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 44 inbound Pith citation observations for arXiv:2306.04528.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.04528 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 44 of 44 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T20:08:13.602909Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

50
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2b2e6116-13c7-4016-8e76-f738ee64014d · inbound

Universal and Transferable Adversarial Attacks on Aligned Language Models cites this paper.

Universal and Transferable Adversarial Attacks on Aligned Language Models PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-24T07:44:08.483210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:02490e0f7a549b2763317cbe1dcbba7250c2ddcafb2ccb203223bf315c552a70

Observation d4a95b09-3832-4a2f-bc31-9fac787648ee · inbound

Baseline Defenses for Adversarial Attacks Against Aligned Language Models cites this paper.

Baseline Defenses for Adversarial Attacks Against Aligned Language Models PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:24:40.082946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T23:24:39.835347Z digest=sha256:971b22ec3d03c53a2a6957ee038195599f9f20a16db0569bcf626d9953c7924f

Observation c781a3b8-e332-4632-8ca8-f7ec654ec293 · inbound

Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models cites this paper.

Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-12T14:21:16.477948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T14:21:16.453610Z digest=sha256:3d058e7d96381d5f995e3201c0b8cadd61fc5ea3694e1f312c0e2c95eb45aa45

Observation f2ed5e7c-03b1-4f84-a0fd-c4e85a3b4c55 · inbound

EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers cites this paper.

EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 137

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:11:49.585490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-16T06:11:49.475825Z digest=sha256:a6ab9f972792bc8c1062ce5849b455abc50e3cdf9aeb787744d3456f08234622

Observation 6503384a-e3c4-4163-9717-5493a50b0b00 · inbound

TrustLLM: Trustworthiness in Large Language Models cites this paper.

TrustLLM: Trustworthiness in Large Language Models PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 172

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:17:08.543894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T11:17:08.108565Z digest=sha256:45535631cb28e1a3e5e7a97be78b7eecf99d956a5f073e072d735bcb2cc1763a

Observation f5db751b-c80a-4b12-a3b9-929b9718e4c5 · inbound

Whispers in the Machine: Confidentiality in Agentic Systems cites this paper.

Whispers in the Machine: Confidentiality in Agentic Systems PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:03:53.865582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T03:59:03.972043Z digest=sha256:aa7ac97b46fe5494f2705d336bfa15f7d9f0f0d05fa912eccbe4f613b7e7aced

Observation 6f1e8eaf-774d-4692-a572-11f1d10e44ce · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 191

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:41.169831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:ca6babb3a7cb146dd32e1c127b0eb351d8cc8e48ebbc153d82c8196bf3376369

Observation b7e0043e-1fed-4f3a-b51b-9971c85280c1 · inbound

Trustworthiness in Retrieval-Augmented Generation Systems: A Survey cites this paper.

Trustworthiness in Retrieval-Augmented Generation Systems: A Survey PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:08:25.916992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T21:08:11.787013Z digest=sha256:2c3ac5035d4a9d51ab4ddb0c9e994959953bf6b98a373151d40564e9d97da013

Observation cfbb0258-1cdc-4033-91d5-b4190fddb289 · inbound

CodeSCM: Causal Analysis for Multi-Modal Code Generation cites this paper.

CodeSCM: Causal Analysis for Multi-Modal Code Generation PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T20:08:13.602909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:08:13.602909Z digest=sha256:774e3a2a3679f4b75803bfe944823647083d3a672cd8a9f34491ebc1729ee86d

Observation b7267df2-c433-47f8-99f8-1efbc8a7c390 · inbound

Benchmarking Prompt Sensitivity in Large Language Models cites this paper.

Benchmarking Prompt Sensitivity in Large Language Models PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.553837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.553837Z digest=sha256:2fdc7f24b33ae4c502c9cb0b78c1c5195a93aa21bd28fc962b82a3c181e16ae1

Observation 5ae89962-a314-4ef0-ba48-b8b0beaff677 · inbound

Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences cites this paper.

Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-08T10:23:06.309263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:23:06.309263Z digest=sha256:235479e037625add07ecbe74a62607a3469e2e1a9650ce27b724415a9b571ba7

Observation defb6ea1-6aed-4bd9-b8e7-c2bba422d85b · inbound

The Science of Evaluating Foundation Models cites this paper.

The Science of Evaluating Foundation Models PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.927960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.927960Z digest=sha256:6aaf7b1930174847521cd243f52129ee025d5d8681bb1bc319324f64324dee9e

Observation 3abcd471-5878-46b6-b2a7-526041342659 · inbound

A Representation Level Analysis of NMT Model Robustness to Grammatical Errors cites this paper.

A Representation Level Analysis of NMT Model Robustness to Grammatical Errors PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:04.449448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:04.449448Z digest=sha256:a4e15694e0f18025e5e1577047770c3895c9e2837d12cebf20512c85509826db

Observation 02365128-d5a3-4427-853b-65671b5f5a03 · inbound

Adversarial Attacks against Closed-Source MLLMs via Feature Optimal Alignment cites this paper.

Adversarial Attacks against Closed-Source MLLMs via Feature Optimal Alignment PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 63

Resolution
malformed identifier
no resolver link, observed 2026-08-07T13:35:56.355578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:56.355578Z digest=sha256:bb50adfc97be9aac0d17a3c2a9c64816d4573b440013c175a2068b5fcd2306d5

Observation 1f3d2782-3fef-4ef9-acbe-fb7db3bb5404 · inbound

Look Within or Look Beyond? A Theoretical Comparison Between Parameter-Efficient and Full Fine-Tuning cites this paper.

Look Within or Look Beyond? A Theoretical Comparison Between Parameter-Efficient and Full Fine-Tuning PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:55.842329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:15:55.842329Z digest=sha256:e692e6d9905cd0e9097f1c5495446042071d18e71c9617eee2d266f16dd9497a

Observation a07116a8-6252-46a8-9d49-de7c94e931c5 · inbound

Evaluating LLMs Robustness in Less Resourced Languages with Proxy Models cites this paper.

Evaluating LLMs Robustness in Less Resourced Languages with Proxy Models PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:52.719510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:52.719510Z digest=sha256:0939e8064e79397476f76f719e0266d17d185e731dda9082036a5b27bb44dca8

Observation aa8fdce4-2789-41cc-a44b-cb4d9432a2dc · inbound

Evaluating and Improving Robustness in Large Language Models: A Survey and Future Directions cites this paper.

Evaluating and Improving Robustness in Large Language Models: A Survey and Future Directions PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 245

Resolution
unresolved
no resolver link, observed 2026-08-07T05:42:31.720393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:42:31.720393Z digest=sha256:c6838f3561312137545cf7f84b34e8555cfd73f6f7333d03611da4e2ab583d0c

Observation 92ea6742-49b2-4302-9880-1e77754532cb · inbound

Investigating the Robustness of Retrieval-Augmented Generation at the Query Level cites this paper.

Investigating the Robustness of Retrieval-Augmented Generation at the Query Level PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:55.041443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:55.041443Z digest=sha256:6e1468756f98820ba685467f9f72ec2fd37dcbf1947ad8b153d669a6408692d9

Observation 47108bc9-e6b0-48df-a5cc-ffe314623e52 · inbound

Explicit Vulnerability Generation with LLMs: An Investigation Beyond Adversarial Attacks cites this paper.

Explicit Vulnerability Generation with LLMs: An Investigation Beyond Adversarial Attacks PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:45:58.463121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:45:58.463121Z digest=sha256:33b9d40722854d9f2608e4d4c48c7e48ab2eedf51a6528bc0c21f7a685e4a014

Observation 15fb529e-3caf-4a08-bb0b-0efca3b5e99f · inbound

Agent Identity Evals: Measuring Agentic Identity cites this paper.

Agent Identity Evals: Measuring Agentic Identity PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.882060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.882060Z digest=sha256:ec77e9465ef42bfa413e3cc7886c7c1c545162dfaddef6cc8ab55545d36cefc7

Observation 7b63df69-81fe-40e7-a062-bae3a724ff7f · inbound

Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs cites this paper.

Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T05:59:28.243698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:59:28.243698Z digest=sha256:c6df740c0cdd695acee82566dfbdad217feb528a5c97f591af8b903f98fd9f9b

Observation 0495afe7-d37d-4a4b-8fe3-b7e7c29526c7 · inbound

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications cites this paper.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.373904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.373904Z digest=sha256:90856f506f6abb04497b45efc596d37e58650a6bbf7cef15574c609ec51121a8

Observation 3bf2096c-140c-4017-84eb-0dd8394f2a8b · inbound

Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations cites this paper.

Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:47:15.338790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T03:43:18.987241Z digest=sha256:e80861463ecaa6d92e9772c9aec6eed8d0885a994a45a2efc52e1401c9a35dc7

Observation abd144a7-b5e5-4d2b-8c87-04c8cdef2ffd · inbound

PEEM: Prompt Engineering Evaluation Metrics for Interpretable Joint Evaluation of Prompts and Responses cites this paper.

PEEM: Prompt Engineering Evaluation Metrics for Interpretable Joint Evaluation of Prompts and Responses PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:00:03.073494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T13:57:41.428695Z digest=sha256:aec806221ccaf333baa1718fc09ea751ae81da6c3c7e17a72f457711ffe388ea

Observation 282e714c-1f2b-4b3a-ac98-aedba4179bcb · inbound

Automated Framework to Evaluate and Harden LLM System Instructions against Encoding Attacks cites this paper.

Automated Framework to Evaluate and Harden LLM System Instructions against Encoding Attacks PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-13T14:38:58.879673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:38:58.879673Z digest=sha256:b0226d7a45c50643f0474466b943243aecbc4901268a4a9ac12791dedbc50ad3

Observation ed3491bb-86d1-4f3e-aa00-134f2634b233 · inbound

Measuring Representation Robustness in Large Language Models for Geometry cites this paper.

Measuring Representation Robustness in Large Language Models for Geometry PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:38:10.484726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T19:35:32.531660Z digest=sha256:596f2c80e2c350a4035d36b97a7989960d52e1a72285d2cfb741b67384a6a0c4

Observation 11f307d6-7263-46c5-8027-d796552a41b7 · inbound

Characterizing Paraphrase-Induced Failures in Lean 4 Autoformalization cites this paper.

Characterizing Paraphrase-Induced Failures in Lean 4 Autoformalization PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:36:08.464012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T08:33:34.423179Z digest=sha256:53f7791487b04b57148f8f65ddfa39695f172e0069a9f43c8477f13a1787f7ae

Observation be275719-33b2-4a79-a3bc-77875de9734d · inbound

Characterizing Paraphrase-Induced Failures in Lean 4 Autoformalization cites this paper.

Characterizing Paraphrase-Induced Failures in Lean 4 Autoformalization PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T00:53:53.714254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T00:49:58.959331Z digest=sha256:4224f23d39f657d855d38e8b2309f31aaa4407734ac2b43c8db2c70df92f7062

Observation 615c03fa-9e15-4524-9310-144f87348904 · inbound

A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework cites this paper.

A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:51:09.098195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T07:53:13.746141Z digest=sha256:82b182117343043feaf69cc4f83b963399ede9b974b13070bc8176502eca51b7

Observation 91f507b1-67ba-48d8-bccc-2c084d245b36 · inbound

When Prompt Under-Specification Improves Code Correctness: An Exploratory Study of Prompt Wording and Structure Effects on LLM-Based Code Generation cites this paper.

When Prompt Under-Specification Improves Code Correctness: An Exploratory Study of Prompt Wording and Structure Effects on LLM-Based Code Generation PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:17:07.253922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T03:00:26.137401Z digest=sha256:dd13889c876448c239a1d224f9c3a443c63d797566baea2f35b239574c40debc

Observation b9a6b961-0c06-4f58-8018-8c947129bdfd · inbound

Paraphrase-Induced Output-Mode Collapse: When LLMs Break Character Under Semantically Equivalent Inputs cites this paper.

Paraphrase-Induced Output-Mode Collapse: When LLMs Break Character Under Semantically Equivalent Inputs PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:16:07.328226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T16:26:54.068997Z digest=sha256:47d6b164c6f8cd9950dcae71f8523a09f8d78b69205f95e9e4328a8d81fe8fac

Observation 0810c81a-95cc-4e59-8a2b-0761e6fcc17e · inbound

Paraphrase-Induced Output-Mode Collapse: When LLMs Break Character Under Semantically Equivalent Inputs cites this paper.

Paraphrase-Induced Output-Mode Collapse: When LLMs Break Character Under Semantically Equivalent Inputs PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:16:16.124945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T02:13:48.672711Z digest=sha256:ad54a891d91c59cc5f86dffcf40189796ef635988f7249f38994da7414c8158f

Observation 8f2c0e70-523f-4127-bf25-7274bc3202c0 · inbound

BiAxisAudit: A Novel Framework to Evaluate LLM Bias Across Prompt Sensitivity and Response-Layer Divergence cites this paper.

BiAxisAudit: A Novel Framework to Evaluate LLM Bias Across Prompt Sensitivity and Response-Layer Divergence PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:21:26.819643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T03:26:18.375974Z digest=sha256:ba84c1010cb014564849aa087f2300b068c127e66b8b17d78e7202be13c46691

Observation cb4f7e11-7e6d-44ef-a76b-25a5eccf2e66 · inbound

Consistency as a Testable Property: Statistical Methods to Evaluate AI Agent Reliability cites this paper.

Consistency as a Testable Property: Statistical Methods to Evaluate AI Agent Reliability PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:41:21.819932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:41:15.286881Z digest=sha256:2c41064fa2b401c25513407c8b16d7c04a4bc5c4a13d82a5723cb9d898ce55f1

Observation ee06c8c2-73dc-445b-95b2-7422e7f31eb1 · inbound

Can we trust LLM Self-Explanations for Entity Resolution? cites this paper.

Can we trust LLM Self-Explanations for Entity Resolution? PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:56:15.179442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T16:13:14.653064Z digest=sha256:9deeba2d9b69c1e215e50f2e2959b434862ab2f8df133b53222b96bd4929438e

Observation f4b554b0-5b1b-4c6d-bf9d-daca796d5c9e · inbound

Dive into Ambiguity: A*-Inspired Multi-Agents Commonsense Obfuscation Attack on LLM Prompts cites this paper.

Dive into Ambiguity: A*-Inspired Multi-Agents Commonsense Obfuscation Attack on LLM Prompts PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:26:16.387346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T16:54:12.354178Z digest=sha256:caa9116864f0b1e4506f0674cee83768c67bdc765de2146f55513519d966435c

Observation 5e033c3d-c72d-4347-a826-3d07a3b969b3 · inbound

From Sampled Outcomes to Capability Distributions: Rethinking Supervision for LLM Routing cites this paper.

From Sampled Outcomes to Capability Distributions: Rethinking Supervision for LLM Routing PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 158

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:17:08.772986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T22:54:28.452796Z digest=sha256:051f4fe787f48a8fab021e7100c271244e0c7b8d890198fbe3d6749fcbf27853

Observation 2ec8c2be-f86e-4035-be37-84907a53bfb9 · inbound

Trajectory-Level Redirection Attacks on Vision-Language-Action Models cites this paper.

Trajectory-Level Redirection Attacks on Vision-Language-Action Models PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:18:33.763432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T06:33:53.013076Z digest=sha256:320c7160e072d150ad43a44db8899fb110d287ad0b51a7da0af33ece64c97918

Observation 25f70ffd-2c35-4b17-b7af-c765deb85791 · inbound

Legal Reasoning Is Not Lawyering: Rethinking Legal Benchmarks for Pro Se Access to Justice cites this paper.

Legal Reasoning Is Not Lawyering: Rethinking Legal Benchmarks for Pro Se Access to Justice PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:19:03.950183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T22:26:54.239505Z digest=sha256:767ed540b49df49b95b405ad8c91f6fbd71fe7c7a139d49b931fb296b0ec2bdc

Observation 4448a36d-8194-49fc-9258-dbf5724d335e · inbound

Format Sensitivity Index: Token-Controlled Prompt Wrapper Robustness and Schema Compliance in LLM Benchmarking cites this paper.

Format Sensitivity Index: Token-Controlled Prompt Wrapper Robustness and Schema Compliance in LLM Benchmarking PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 16

Resolution
malformed identifier
no resolver link, observed 2026-07-14T19:17:45.652076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:17:45.652076Z digest=sha256:dc005f08f2d19da5c3d36e9c0615d3fe0c9b62bf6c6447aeb073313eb5a7b56f

Observation 3e67b377-fb5f-4839-9e62-3a2f3fa3b155 · inbound

Format Sensitivity Index: Token-Controlled Prompt Wrapper Robustness and Schema Compliance in LLM Benchmarking cites this paper.

Format Sensitivity Index: Token-Controlled Prompt Wrapper Robustness and Schema Compliance in LLM Benchmarking PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T19:17:45.652076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:17:45.652076Z digest=sha256:fe64c90e25c6567291a6ba554e3424e6ed4ceeb6b3ce7c3a1de4514525bb542d

Observation 0fb91a91-a8a6-4c45-b562-b96172b8ce3a · inbound

Find Before You Fine-Tune: A Diagnostic Study of Small LLMs for Cybersecurity QA cites this paper.

Find Before You Fine-Tune: A Diagnostic Study of Small LLMs for Cybersecurity QA PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T14:35:01.351523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:35:01.351523Z digest=sha256:d211c0f5dc533431c4b3c3389f9bc7a2f664eceeec07b6cb7fe7ecf7b315d31a

Observation d1c5e8b1-a155-4fd8-817c-65fea8b70be4 · inbound

Imprompt: A Language Framework for Prompt Programming cites this paper.

Imprompt: A Language Framework for Prompt Programming PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T07:07:45.246508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:07:45.246508Z digest=sha256:6e712c8773d0beb017a4ae05ff2b4cfb7202a96c24a38adb82b6fb0d1edbcea2

Observation d7353651-c3b0-43ac-b1a7-9db0d51c80dc · inbound

Visual Grounding in Zero-Shot Vision-Language Control cites this paper.

Visual Grounding in Zero-Shot Vision-Language Control PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:11.549507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:57:11.549507Z digest=sha256:cf73c82e8e0dc31933d4b7e883d77b6819021d3520da19c6aa5d830812e51eb9