Pith. sign in

Paper Citation Record · LEDGER

Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2111.02840.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2111.02840 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:55:06.860412Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T01:37:30.523597Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 299597fc-8011-4ba3-900f-bd62627e7964 · inbound

Universal and Transferable Adversarial Attacks on Aligned Language Models cites this paper.

Universal and Transferable Adversarial Attacks on Aligned Language Models Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-24T07:44:08.495253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:8ef96f2af97c573e79d40437623fe094c7650137eebe0e362b9812592f7bab4f

Observation 5eda0ea9-d41e-4ad7-83a9-452544fbe937 · inbound

TrustLLM: Trustworthiness in Large Language Models cites this paper.

TrustLLM: Trustworthiness in Large Language Models Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 267

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:17:08.774235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T11:17:08.108565Z digest=sha256:5a4e2d1a561ac6be16e4c0bcc2430c856a428393b92b919eb389795e3974496a

Observation e0c194f5-dff8-434a-98bb-03e0a6e15350 · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 149

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:41.098865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:8070c8c5d2396653a2a7bd004ee77368703455bc2139299332a474d5aedcfb68

Observation bef0e7fc-336f-4b0a-962b-42dce9246ae8 · inbound

On Adversarial Robustness and Out-of-Distribution Robustness of Large Language Models cites this paper.

On Adversarial Robustness and Out-of-Distribution Robustness of Large Language Models Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:06.860412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:55:06.860412Z digest=sha256:ccf9e10a61cdd2420c2595ea9a24b373876921d29c056297579afe09c39acfcd

Observation 81301fb8-5a9e-45f0-b363-1aa1c5868a03 · inbound

Serving Long-Context LLMs at the Mobile Edge: Test-Time Reinforcement Learning-based Model Caching and Inference Offloading cites this paper.

Serving Long-Context LLMs at the Mobile Edge: Test-Time Reinforcement Learning-based Model Caching and Inference Offloading Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T15:19:49.656199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:19:49.656199Z digest=sha256:95185864008d13fb6454c29e71918da6ca24859e61fccc12a421b77da26cd311

Observation 02a5b6f4-3520-42f2-b258-e7a0263ce9b5 · inbound

SecPE: Secure Prompt Ensembling for Private and Robust Large Language Models cites this paper.

SecPE: Secure Prompt Ensembling for Private and Robust Large Language Models Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:28.830087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:28.830087Z digest=sha256:15837dd378c8c8883481a8aed9d5d9ea27d4c7f487ffadf8ec108d5d9bd67dd6

Observation 51a6cad7-7eab-4dbe-9466-5731b9e54fb6 · inbound

The Science of Evaluating Foundation Models cites this paper.

The Science of Evaluating Foundation Models Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.872968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.872968Z digest=sha256:c59ffbf0e03ea488ecbad6475294796bbdf7164c3b60706f1f74290d5b41558d

Observation a2356488-2d08-4cfa-83ef-c2faa043053c · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 135

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:27.007444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:27.007444Z digest=sha256:505544b368996a84c6398fa2f478a2ee8a89f94fd6a58e51e05b85cbb128dd4c

Observation 6f6c6654-9fed-443f-8b78-24f80ffffc80 · inbound

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models cites this paper.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:03.461424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:03.461424Z digest=sha256:6590cd31b276532ff5e4c939c46b18e4d8948f7eb25037a8f1b294d917e4509b

Observation fc9f0b3b-194d-4e02-932d-47bbb20ff156 · inbound

Evaluating Robustness of Monocular Depth Estimation with Procedural Scene Perturbations cites this paper.

Evaluating Robustness of Monocular Depth Estimation with Procedural Scene Perturbations Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:15.210109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:15.210109Z digest=sha256:76d78f52055a9dcd59ddb508ea2ea0f11f8a3a00d69de244306977982b06e952

Observation 0e1e92a1-6ee0-4d9a-bb6d-3d55a5a0d33b · inbound

Optimus: A Robust Defense Framework for Mitigating Toxicity while Fine-Tuning Conversational AI cites this paper.

Optimus: A Robust Defense Framework for Mitigating Toxicity while Fine-Tuning Conversational AI Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:21:31.072680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-22T12:17:59.458633Z digest=sha256:a272ec96430d0ebe03ce5d45fe54f06b2d859b7f89a3894664b1cf50c41ccaf2

Observation 1581c38e-115f-4014-b61b-b00e5933eef6 · inbound

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training cites this paper.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.496124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.496124Z digest=sha256:f0d97408ac033d37258efa75a5fd361678d8dedebddd7cf8c6ddf066d82db1ec

Observation 770cff7c-c045-40a4-a587-8bee484c7ae1 · inbound

SALMAN: Stability Analysis of Language Models Through the Maps Between Graph-based Manifolds cites this paper.

SALMAN: Stability Analysis of Language Models Through the Maps Between Graph-based Manifolds Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T17:14:34.612856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:14:34.612856Z digest=sha256:9d7380931bb2397ec14833a1590dde71186aaf6ad5840d921f9c5297ba0c1c29

Observation 906feaee-1255-47a8-909b-097d40f3347d · inbound

UniComp: A Unified Evaluation of Large Language Model Compression via Pruning, Quantization and Distillation cites this paper.

UniComp: A Unified Evaluation of Large Language Model Compression via Pruning, Quantization and Distillation Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T03:07:44.414501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:07:44.414501Z digest=sha256:15597e0e73008bd91eb772afdf069e7d008dab96fd3508d26ade64f230962b95

Observation 057b9cff-7b10-40c0-bb0b-d2e1fde96b51 · inbound

Understanding the Prompt Sensitivity cites this paper.

Understanding the Prompt Sensitivity Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:05:23.299039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-10T04:41:03.416429Z digest=sha256:1294f622e32fafdb4881d859f2b6f1d031d49c2e9c3f50b8909f977f011ecca9

Observation 12759f72-b577-4243-8e03-70be1f0978ec · inbound

SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades cites this paper.

SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:33:32.367930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-15T02:31:18.183715Z digest=sha256:3fbbcdc86f091fecf9d91604cedf7768b2217a88f2666a610c28caf52a520ed7

Observation a0a3fae3-2bc8-407c-be7c-f78a3993fab6 · inbound

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak cites this paper.

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T06:19:41.929065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-21T06:16:01.040236Z digest=sha256:8ea051572d40e92dd9432ebd3f332368907d987c5774f87c96bf771c3f424cb7

Observation d6314ba4-b8cf-4cc9-b9b0-f08fb46c7267 · inbound

Testing LLM Arithmetic Reasoning Generalization with Automatic Numeric-Remapping Attacks cites this paper.

Testing LLM Arithmetic Reasoning Generalization with Automatic Numeric-Remapping Attacks Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-02T04:06:35.135367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T09:27:30.923556Z digest=sha256:4e32f2fd01d881eaada6c668219040578320e784390d95f8798dea2d6811ade2

Observation 4fbfb9ed-1b90-4717-92e7-8a60808cc76d · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 91

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.524920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:4493bc649a0f500e309b07737dc5837ffe7661d6d061297134d1323a85bcb32c

Observation ab15359e-50ef-4264-af87-7108b2756630 · inbound

PRA-RAG: Provably Robust Aggregation in Retrieval-Augmented Generation against Retrieval Corruption cites this paper.

PRA-RAG: Provably Robust Aggregation in Retrieval-Augmented Generation against Retrieval Corruption Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:47:27.205626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-07-02T23:42:23.695930Z digest=sha256:43ef6b1467f9829adfff44eecda9ab5a335820b8e89c58239b4dc590cee3137e