Pith. sign in

Paper Citation Record · LEDGER

Adversarial Training for High-Stakes Reliability

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2205.01663.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2205.01663 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:21:58.812008Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

9
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 91962894-06b8-4dcf-b698-faf34f77066a · inbound

Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned cites this paper.

Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned Adversarial Training for High-Stakes Reliability

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:38:08.643501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T01:38:08.362920Z digest=sha256:856b2c98b3657fa29b9cbce9e4eee5db67889406d13b7507ae3448f493d59367

Observation 8ba6c1c8-f416-430f-a4e4-1450569e1f9a · inbound

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints cites this paper.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Adversarial Training for High-Stakes Reliability

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.812008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.812008Z digest=sha256:62673305ee5a84968fcf7fd7742068e90350ec3e7ade30a1bb407a83a9192563

Observation 350a898d-d298-48a8-bd4c-d7a3a13c8b39 · inbound

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training cites this paper.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Adversarial Training for High-Stakes Reliability

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.617525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.617525Z digest=sha256:fac6ddf4bbdd20cf9ca504f8b26e751a8edaf92accd2bee970c56c8dbd9d4144

Observation 9e946724-1d6a-434f-9614-a098a67d34ce · inbound

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces cites this paper.

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces Adversarial Training for High-Stakes Reliability

Reference 278

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:17:54.706634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-14T20:17:01.224864Z digest=sha256:4cdfaf47cd63f2ee7b25d84c5ae3d1002eee7e54eb305e5987888bdf010b218a

Observation 30b11f2d-32db-41f4-a69a-f0936bd6d3e7 · inbound

Safe Inference-Time Alignment via Lagrangian Reward Augmentation cites this paper.

Safe Inference-Time Alignment via Lagrangian Reward Augmentation Adversarial Training for High-Stakes Reliability

Reference 103

Resolution
unresolved
no resolver link, observed 2026-07-12T07:05:47.150308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T07:05:47.150308Z digest=sha256:e42708ae17577d1941e60827e76fa4bc071bc7c96f8aa7e4c78aaffad97483a7