Pith. sign in

Paper Citation Record · LEDGER

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training

As of 8 August 2026, this Paper Citation Record lists 94 of 94 outbound references and 0 inbound Pith citation observations for arXiv:2507.14202.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.14202 v1

Coverage vector

measured 94 of 94 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:35:48.626472Z

measured 94 of 94 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

94 of 94 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved92
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5b6a9238-00de-4972-a6c5-c256e98d095c · outbound

This paper cites online" 'onlinestring :=.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.238850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.238850Z digest=sha256:0c797d1dc6390dc1329de0a10c6530be360821c84509c0a69a4365a3fb517b5f

Observation 25cf31f4-6c6d-409f-b5cb-f7b395d031a2 · outbound

This paper cites write newline.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.390160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.390160Z digest=sha256:e7ece569c1c7ceb889a417fc99fb617a6559a8eacd4b79a279fec312e44cd60e

Observation 5a2af6a0-2e86-4504-8cbc-6ffeed6a1cde · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.424574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.424574Z digest=sha256:445b7f9d1e63a4e9591f487d0d573157f2a62b4d597bf115a42366bbc923e317

Observation bb252c79-5804-47f0-aa0b-2b0227ac2c84 · outbound

This paper cites Generating Natural Language Adversarial Examples.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Generating Natural Language Adversarial Examples

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.437181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.437181Z digest=sha256:90239044ea1e84f626f60f6bc8592bfeed5a753544bf602bac331fa0931270d7

Observation ffcb86b7-123f-42ce-8130-0269d61955ec · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training A General Language Assistant as a Laboratory for Alignment

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.468119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.468119Z digest=sha256:49945b1ed659bf26672966ea618159bffb30f978bb8631526d77ab92cbd7caa1

Observation adcb3d4e-4952-4233-a7a6-2fef21cd219c · outbound

This paper cites Spinning Language Models: Risks of Propaganda-As-A-Service and Countermeasures.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Spinning Language Models: Risks of Propaganda-As-A-Service and Countermeasures

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:35:50.534543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T17:35:47.522560Z digest=sha256:970b0fac021a073f77a62acc395d41416661e3ddb5055c8a2ceadd2f4c1f4f89

Observation 1f8b8625-ea52-45ea-9324-d027089e80ba · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.608128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.608128Z digest=sha256:443b0a081dddf5848e512bb773a6252b926ec41880231b9b6621ba7107b7a8dd

Observation b29ff979-9448-4e35-8bf3-47d149d4744a · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Constitutional AI: Harmlessness from AI Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.723342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.723342Z digest=sha256:f643e9e549f9929f1dfa13c131376d49500d522f4f419c5a3012bb2feab3ad2a

Observation 517e308b-7e38-4685-85ce-e9f649f6d4c3 · outbound

This paper cites Image Hijacks: Adversarial Images can Control Generative Models at Runtime.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Image Hijacks: Adversarial Images can Control Generative Models at Runtime

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.822419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.822419Z digest=sha256:d1eebf3e624dd64777d9bbc88c8bd9314481701c22d9cd5b8ed48655e5a2ff30

Observation d51cbcab-b7d2-41ed-9e8c-b1dc9bbbc282 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.897090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.897090Z digest=sha256:ff97825dc77a1e679ac42c997bf142b809e34251ec40982ceca8ea0e423c98a6

Observation 8202b0ae-a407-4800-8fcc-ad004b22a9d0 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.960711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.960711Z digest=sha256:f488dae6273a25733413ebd5a32eb39dc4e650db6063b73e72740589d7bc40bc

Observation 21677c54-689c-4f87-803a-53f3d739d648 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.970414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.970414Z digest=sha256:170186efe361b686758b99fee3be71249b8aa07fe3ce0af358772b8d168b6451

Observation ce3d1067-2d5b-4af6-bd46-b95bde12f133 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training On the Opportunities and Risks of Foundation Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.975977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.975977Z digest=sha256:8be8af650796d473215adbab5a11a02c0f53318ef73c4d6b6004765536b2c1a0

Observation bafa7bb3-a8c0-48c3-884f-7ca91f77b74e · outbound

This paper cites Evaluating the Susceptibility of Pre-Trained Language Models via Handcrafted Adversarial Examples.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Evaluating the Susceptibility of Pre-Trained Language Models via Handcrafted Adversarial Examples

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.982288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.982288Z digest=sha256:79d8746bb1e819b96e1833855dd801c60624a56e4236ca06eecb4788af01323b

Observation b26876fb-8ead-4a31-9765-0039acb31008 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.989197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.989197Z digest=sha256:19bbbd60254cde3255ded08e88e39ba6bcac48c7363e35dfa9e16895414f4395

Observation 4bb4e62a-f278-4460-86c3-e2119a30d3e8 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.996109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.996109Z digest=sha256:8a3c7c9727d33c8a37ec0bafdc9e7ae6891246b224370f12c566b3297336b9fa

Observation 8a5a0256-0ed2-4a11-a324-0a6ae5dfe732 · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.002781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.002781Z digest=sha256:d95386986a1c480f96c92bed50599319149f95d571b71617b94365b1bb898687

Observation ee47c9ff-dcb0-43a9-a097-9fad4e0e75d8 · outbound

This paper cites Explore, Establish, Exploit: Red Teaming Language Models from Scratch.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.010807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.010807Z digest=sha256:82cee3171dd72d50fddadbee031902a012f65a251c45499c92c8355ab1f7d255

Observation 4fd144ea-56ba-462c-b157-c2c2fccc610b · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.018055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.018055Z digest=sha256:b1fb7b77f5f2d17383c411c6b578da7639c2120b2fc754b052912ec544b596a1

Observation 2f3b1b5c-6c78-4d68-a448-028cd001b60f · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.025371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.025371Z digest=sha256:0c4105b1f1ce682f42dd92a8271e6c4551e4aa10d35ea40860c15cab8f02a927

Observation 670b6015-0d51-4533-b7e0-7978b003fe56 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Evaluating Large Language Models Trained on Code

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.035926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.035926Z digest=sha256:dbf620ab54799d8541c9674267fd6fe6669313ac1e5f75a199a93636a811e072

Observation 3aa09c3e-12c8-4063-9d99-bb9690b6accf · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training PaLM: Scaling Language Modeling with Pathways

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.042843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.042843Z digest=sha256:27d5f90d4666002a74fd22574da1e30a72616c07e411ce84c2370ddbeab2e4da

Observation 3d70f730-9dc9-4af2-bf75-7a0924bd6251 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.053465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.053465Z digest=sha256:e7a422496846c926cb6e70ee28bb70500ea0077168e71c4e9552824dd8fe8397

Observation a11d8663-aa43-4de8-bd1e-a858946f0824 · outbound

This paper cites Supervising strong learners by amplifying weak experts.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Supervising strong learners by amplifying weak experts

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.063524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.063524Z digest=sha256:1f0c3c26100e4635f359bf67548556d339a638312c1e32629a470ff1f0696d06

Observation 9f0d440a-ef33-4ceb-9f85-df8526eb4c94 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.070951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.070951Z digest=sha256:51e756a7a81d14ea4adfc9989b49669cdff5b0c15ad047163442feb9b03a2219

Observation c904d3f8-3630-4b9d-8b71-f52597e821f0 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Training Verifiers to Solve Math Word Problems

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.078087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.078087Z digest=sha256:fe9a8514830e841e9748ccd9a8efe986095383ff4464bd6ebb87696d598923f8

Observation d6639d9f-321d-44b7-ab7f-ad5bad9c5d6b · outbound

This paper cites Multilingual Jailbreak Challenges in Large Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Multilingual Jailbreak Challenges in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.084454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.084454Z digest=sha256:0466689d24a846978d74b8bb93f4df60b3b9afea035e2e72e716d7ba038df36d

Observation 68c61a06-cfae-48e7-abb9-2bb1571c8feb · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.092363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.092363Z digest=sha256:3f70f5ed2e35c24c03868f080ddca2320ddf517445c75a8ceeb27c9607a0b153

Observation 9acddfb3-9cda-4059-94d5-cd8bd0434b18 · outbound

This paper cites Understanding parameter differences between analyses employing nested data subsets.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Understanding parameter differences between analyses employing nested data subsets

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.100821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.100821Z digest=sha256:5c1416580faf5402f9eb9e3b4706e4d9e7487b8a5f4e904e7f69b399c7109e46

Observation 7f34939f-ae70-482c-8392-47ddccbdb6f8 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.109664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.109664Z digest=sha256:9ed46ffdda097d01a900f9073d7e2ff2fbb5e8a74bea77dc5426244562836450

Observation 26c25ec9-a7aa-48cb-96d2-11d2f2c0cce6 · outbound

This paper cites HotFlip: White-Box Adversarial Examples for Text Classification.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training HotFlip: White-Box Adversarial Examples for Text Classification

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.119501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.119501Z digest=sha256:3a9eb999f537a79d45114070a98c98fbe0b8f2788593ebec0839736001c0df8d

Observation 3f696111-d4ca-4c96-85f0-7d428d2af8ab · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.128310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.128310Z digest=sha256:ea355f20dc87d7c1d9e2f871cac0e8327cc7083c0657f651a938c60c61d09f84

Observation c42d8818-3e1d-410f-87ba-4fa7f898a85c · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.134541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.134541Z digest=sha256:667e2409d24cd1bf8e82ffff9b3960f7fbbd7a163d06672c51cb08dae9bfa5fc

Observation 2ce2fcbd-d0d9-4e9e-9491-0c12d0b58c02 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Scaling Laws for Reward Model Overoptimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.142756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.142756Z digest=sha256:ed45bd851acdb6b01c70452868492593a0e5eb334324eb78cbc001e57610a6ec

Observation a90ddecf-99c1-405b-be46-b442ad14d8ed · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.151517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.151517Z digest=sha256:e656a0d56cedc391c1d99616d0db017aac873e20404dab6c6d5a3af49cff92d5

Observation ebcc7130-ae12-485b-8d12-beb2658a661d · outbound

This paper cites Explaining and Harnessing Adversarial Examples.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Explaining and Harnessing Adversarial Examples

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.158428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.158428Z digest=sha256:6e4ff734a8b27781eef8c76d63d89de7f724268cf55fa28ba8d9004345cc25de

Observation b94e92c5-e06a-46e3-8ddb-3c4b945c6be3 · outbound

This paper cites Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.166666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.166666Z digest=sha256:6f3c8c4c07e9b0d2525507fe3b7b5beb5b24486ec346bc283137add8d60669e8

Observation 590e6d0d-c3b9-4aa1-91cb-f62b13467cb5 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.981468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.174906Z digest=sha256:091fde7eb5f18a2ed8eff30cee38d2fbcc4788325504e1f2e8c5935cb0451f8d

Observation 0358867b-52e6-45ba-b59c-97ede6e52fe3 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Measuring Massive Multitask Language Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.184013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.184013Z digest=sha256:5306015f0fc9c3bab8e66fed466e6fb402a322d02931ddb300a4dc2007a41c7e

Observation 28085756-3d06-4400-9799-143a92401037 · outbound

This paper cites Training Compute-Optimal Large Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Training Compute-Optimal Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.190750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.190750Z digest=sha256:4a3c6236ec6ecb2dcab412d1d3ba1dbd42ab98fbb1d508e59d1586eecb018ce8

Observation 6df30fc5-7a42-4e0a-883e-420551bc9683 · outbound

This paper cites AI safety via debate.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training AI safety via debate

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.198213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.198213Z digest=sha256:683fb7650dce58f53b489788faaf4384d334011b2c2796ee43943bc47bcb4cee

Observation c3918a90-c15b-44a3-8567-3f9d1bed0a53 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.954180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.205695Z digest=sha256:b36cc50976f4ee06a7b04e2030c82066a76490c533bf72b40f75917e26aa9f59

Observation 8e4b1b4f-1cda-4008-9ac4-a6828bb2fbb5 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.925417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.215145Z digest=sha256:eb82622e73a4070c5d7367e0a949d2ab77c1719d901a83fbdabf8a2dc592fddd

Observation 8483a57a-aedc-4071-8651-da0a3aacd10c · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.898671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.223231Z digest=sha256:970b12998740589aa4e29f896f43e4b46bb6640e1459a0c0c7132dcbff247703

Observation 72828130-3626-46ec-aaa7-9b4442706d5a · outbound

This paper cites Automatically Auditing Large Language Models via Discrete Optimization.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Automatically Auditing Large Language Models via Discrete Optimization

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.231633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.231633Z digest=sha256:259336037d233ee4a580f9331194235925151bf4c7a26f4b67038302fbc23616

Observation c5bdc0bf-891b-4c3d-9cf6-b5b29d6825b6 · outbound

This paper cites Single-Source Shortest Paths with Negative Real Weights in $\tilde{O}(mn^{8/9})$ Time.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Single-Source Shortest Paths with Negative Real Weights in $\tilde{O}(mn^{8/9})$ Time

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:35:49.762344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.242664Z digest=sha256:cc06c84b45778631ac7aa84b3677d2c1206e1ac00bbd556e0d6a918858419cdf

Observation 8e39d5a0-df53-4a41-813c-fd6d04b2881a · outbound

This paper cites Open Sesame! Universal Black Box Jailbreaking of Large Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.251121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.251121Z digest=sha256:05457a8f3047abe7b31b4d98c9173629f57272628b3847a82fd9db5a27ec763a

Observation 6893840f-8913-456c-8bfe-03992331d013 · outbound

This paper cites Scalable agent alignment via reward modeling: a research direction.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Scalable agent alignment via reward modeling: a research direction

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.260173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.260173Z digest=sha256:317c2cb955d32c4758cf0a6455c392e815c411d680d9748f6eb855e124204425

Observation c2709b8d-4f65-4d21-ab92-6b9473858fa6 · outbound

This paper cites BERT-ATTACK: Adversarial Attack Against BERT Using BERT.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training BERT-ATTACK: Adversarial Attack Against BERT Using BERT

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.267389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.267389Z digest=sha256:1934cfaea360784dbbadd992e086505a18540c5755c922a415dae5d2428150a8

Observation dd64e832-2a45-4da9-bb99-41539ed0ca24 · outbound

This paper cites Let's Verify Step by Step.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Let's Verify Step by Step

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.273778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.273778Z digest=sha256:a569dc265946b49a444693267e9fd74084403bf4e68961e3e80f388777a597bc

Observation b907c4eb-d916-4cbd-8fa8-60a699c361d2 · outbound

This paper cites Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.280960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.280960Z digest=sha256:a7f4d0a4d01bf9ff24797fc821335b742f01cddb256138c6662aa66ccd84d120

Observation 67ea4d27-d6a9-4d67-a875-6c9bd7c3dada · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.289162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.289162Z digest=sha256:8e5c39a6e655a60d75d68d7a1ccffb5f2fbc3bbaf994310550ac2161c6da1dd1

Observation c8382987-8b29-4efe-b8c6-1a8ea5283f86 · outbound

This paper cites Towards Deep Learning Models Resistant to Adversarial Attacks.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Towards Deep Learning Models Resistant to Adversarial Attacks

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.302663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.302663Z digest=sha256:fdaac420b857e9677971d12a5ae6137c69cbe33dbc49a536ef10b83b4d31064e

Observation 455d04b7-778b-4caa-897a-b06d86de4b1b · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.871578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.312375Z digest=sha256:c58261309f1cb17d9dc252ce705cf69344fa8243eb494eb36f8f600634f9617a

Observation 53fcdf3a-c28f-4d2c-89ca-8008a48a9534 · outbound

This paper cites Tree of Attacks: Jailbreaking Black-Box LLMs Automatically.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Tree of Attacks: Jailbreaking Black-Box LLMs Automatically

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.318958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.318958Z digest=sha256:9b8e2b42cb735430d651577550552a71bc24d38a11655a71ca5a171df4fdfff9

Observation 77b7dbcc-5efa-42cb-92e5-75aeb0e2ccae · outbound

This paper cites Teaching language models to support answers with verified quotes.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Teaching language models to support answers with verified quotes

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.325823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.325823Z digest=sha256:2c087fc35ab35e3d772ee591e705dbabe27ff4c0e2b47835157967c0f4a8a5e7

Observation 55775cee-9169-4500-90d2-4eb9a2b99214 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.842475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.334358Z digest=sha256:efbdab304664807ec2128b308a0e5a10be67c48390a4a4a43a821eed69ef65ab

Observation f5cccaa2-679e-4f61-81c4-4e13dd8d6e5f · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.813928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.340570Z digest=sha256:091128fbd477016315f9c015581836f6293a2d5b17be1a5a45b4ca30355b9539

Observation 103ae790-ec3e-4edf-bf4e-122cf26e8719 · outbound

This paper cites TextAttack: A Framework for Adversarial Attacks, Data Augmentation, and Adversarial Training in NLP.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training TextAttack: A Framework for Adversarial Attacks, Data Augmentation, and Adversarial Training in NLP

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.347629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.347629Z digest=sha256:376ee3d18b88a2bbc1929742937d3132bd9ab14f54d6ba7c76e4459e8ca33141

Observation 482fea23-f239-4910-9841-c2975ba97d87 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training WebGPT: Browser-assisted question-answering with human feedback

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.355471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.355471Z digest=sha256:be74813984239af6639918e17b336e907212f3ea1aaaf508cd6583b1836dfead

Observation a701af0a-b1cc-4a10-8629-972734d4f80b · outbound

This paper cites Scalable Extraction of Training Data from (Production) Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Scalable Extraction of Training Data from (Production) Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.362446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.362446Z digest=sha256:d490cfbf882a63ce1cafabc4c5528c170151cc0c398abdd96cc519d8f6a9b08c

Observation 569713c6-6f4f-406f-9e2d-46aca8eaa3e5 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.370906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.370906Z digest=sha256:3aa7b175c61910492303413d6111469c09ff44e6f8c7861070fcce946d29b5bb

Observation 206bfffb-3393-4296-b8ed-1a4e1c1631e4 · outbound

This paper cites Red Teaming Language Models with Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Red Teaming Language Models with Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.379195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.379195Z digest=sha256:ba007520f11e508c35bb38e3f2d5a36f1673bb571f1ee52bf5b8f60da36e3b97

Observation 1811695b-06f0-464a-b606-483a2f9e0eb1 · outbound

This paper cites Discovering Language Model Behaviors with Model-Written Evaluations.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Discovering Language Model Behaviors with Model-Written Evaluations

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.387140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.387140Z digest=sha256:395cb3fcb34e63a2de078ad2bda4fa4c97deec300ece5c85e6d7c2e7fdbb1f60

Observation 3b8e7847-6a86-4fb4-8f3c-f5a359046bd3 · outbound

This paper cites Ignore Previous Prompt: Attack Techniques For Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Ignore Previous Prompt: Attack Techniques For Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.395357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.395357Z digest=sha256:5d908ec711197d56c052a28bb5805e331a3e9fe5257dc2a785831b8c4168ec78

Observation 83499b1b-c548-4c87-9a07-9c519391f9b0 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.775043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.406503Z digest=sha256:f4b2a8d69651a3a4545149778143528cd90583e4dce48faa4f9afc0d40c559b4

Observation 3fb3cb18-7a56-41ee-aad5-997006c0baf9 · outbound

This paper cites Adversarial Training Can Hurt Generalization.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Adversarial Training Can Hurt Generalization

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.413261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.413261Z digest=sha256:9d4d8686d37173a51fdd0a7b2a27f5c59deb812e1bda47265aeeab71399820ca

Observation 453ce8c0-9e8e-4e86-8651-b26a2615ab74 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.752692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.420411Z digest=sha256:be9a87e2f34107a7bdb436abacb115f2b8f2e182e72d1c5806b2e66528166e5d

Observation f1664621-df1b-4744-89eb-98e8f0ef31ba · outbound

This paper cites Proximal Policy Optimization Algorithms.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Proximal Policy Optimization Algorithms

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.427520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.427520Z digest=sha256:07c7ebd08633767c9daf579e8fc4673bc00d2088dd03c5dc9dc65f2f09013d0a

Observation b853babb-024f-485f-bda8-0d6e9d553ba5 · outbound

This paper cites Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.436360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.436360Z digest=sha256:cc42e84010ceb7be2cde85d154245b7158a363d8b961168ec5afdb46ba7cf71e

Observation 09e47225-07a6-48db-9c59-e42421f41fe9 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.445431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.445431Z digest=sha256:470c6d4e43e1fbbcece7f126808cb1ff0d0ac7ebcf3592b88fd1068aabd8c4a3

Observation 0580d262-b859-4833-9516-6a4889a99700 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.454219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.454219Z digest=sha256:d8dec4bf866a5988fc605714cae937c34e659ad8a0e48a4c152cd97dfea3af4e

Observation 4360779b-94b6-4e1e-9494-feb52167a0a3 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.710417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.461142Z digest=sha256:d0a78d1a4ddee420be98601964e562172a5f6af6cc5568efd1ca16560d22ac13

Observation 5123c869-efcf-4be5-b141-c88b49df9fec · outbound

This paper cites Ensemble Adversarial Training: Attacks and Defenses.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Ensemble Adversarial Training: Attacks and Defenses

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.467325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.467325Z digest=sha256:27043a67b67451e2f7b9c2042e52f7e17ee48f0fb997e5e6baffc5baf5b69626

Observation f7629bdb-6602-453e-9138-b15b6227f1d3 · outbound

This paper cites Robustness May Be at Odds with Accuracy.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Robustness May Be at Odds with Accuracy

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.474839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.474839Z digest=sha256:12ec57bf3da11fa00c1a42fd8e8570024c9058dcb32832eba0a4a256ffed07e2

Observation aefb1499-3f67-469c-b28d-49a7afe5f42f · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Solving math word problems with process- and outcome-based feedback

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.481517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.481517Z digest=sha256:647e18b7065da670c38c3bb0d956f93d81cc9c10fe00de074241d07125063e7b

Observation 923ae7d5-0634-43e6-9a9a-5af5f2a7d9bd · outbound

This paper cites Universal Adversarial Triggers for Attacking and Analyzing NLP.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Universal Adversarial Triggers for Attacking and Analyzing NLP

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.488605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.488605Z digest=sha256:18beb93d6491b34d22f8425c56fa74a276e6232a30f062d159e8474ace0abd37

Observation 1581c38e-115f-4014-b61b-b00e5933eef6 · outbound

This paper cites Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.496124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.496124Z digest=sha256:c476435f98c2acd9ffaa5dd9e71472e8f24f87bb0d9bf58db7b1aa0643df511a

Observation c204d952-f567-4cbe-8f4a-21ee9e541aa8 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.688433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.505595Z digest=sha256:8a003ee8792ab8817d6fb79490b236d92156b89427d76aabab617c9e1c4a425b

Observation 6314ad03-312d-443c-ada9-6fc5fa4f6b2f · outbound

This paper cites Natural Language Adversarial Defense through Synonym Encoding.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Natural Language Adversarial Defense through Synonym Encoding

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.513887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.513887Z digest=sha256:8759477d7667328916493602f65ce039a5357181358ae7b6ce795e8748e3bb3f

Observation 0db3733e-7e4a-4a0e-ad85-2382dfd528e3 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.522431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.522431Z digest=sha256:5272ad6d26986347b669667540ab58ae60f4a0314d21a28deeaa2e8cb9191cc8

Observation 51d9b036-4365-4a9c-b6c8-01475b3ad402 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.666517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.537349Z digest=sha256:9e467a79e5aed19af37261e24d8ec32b779f1db42a141ba89c62fe700ad3ae71

Observation 85d265f5-a6dd-4f4c-bada-8c01f9655ca4 · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Jailbroken: How Does LLM Safety Training Fail?

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.545520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.545520Z digest=sha256:d58110571606a02788b1bff5044529fe1c45dbd787ed9836bafb16a1954eb141

Observation 822f62ee-5f4a-449b-8f91-4f6dc81ff76e · outbound

This paper cites Ethical and social risks of harm from Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Ethical and social risks of harm from Language Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.553979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.553979Z digest=sha256:4863506ff7111316407843566346625a51b3d2a4237e80bf0492655a7047c5d1

Observation c54eaadb-dc35-4b7f-a38a-b6c517553455 · outbound

This paper cites Exploring The Landscape of Distributional Robustness for Question Answering Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Exploring The Landscape of Distributional Robustness for Question Answering Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.560955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.560955Z digest=sha256:a172b9eb011e3c78885b8761c1a67be1720656b4b052bde2f2377e22baa1a83e

Observation a8d55785-7081-4caa-a72b-e802dcad7c97 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.642464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.569962Z digest=sha256:de437c5cbf3f4361fcd7c6b91bf9716c5758cda9f540322680a1ab64f774e9a2

Observation 72c27dd0-14f0-497b-baa6-8cddeb7a5a7a · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training ReAct: Synergizing Reasoning and Acting in Language Models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.576790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.576790Z digest=sha256:a3ebed62a816f29d5e0cca6e18261001262159692729057cc9ddb98b5575fae7

Observation 015af25f-f1e0-4a56-bfe7-fea6ed1beb49 · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Low-Resource Languages Jailbreak GPT-4

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.582594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.582594Z digest=sha256:66ab31c69e43b3bc9cb2f2d8562ee99a3d593f0a4686a677e927977e7e3ff262

Observation 84165dfd-1663-421f-965b-d33159b6d948 · outbound

This paper cites GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.591046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.591046Z digest=sha256:4f4db2dc4f9a34dae7d066304fb0bb923bb5626a23c07f68c8fa7023dc004580

Observation 88de92c1-1199-4988-946c-6ef565782d31 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.613453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.596978Z digest=sha256:c273dc728c190267a6824c2e5656e1ebef3572036f8461df9127cb678cda6f89

Observation 8b547ac2-a353-4344-a0a0-525fd1af06e9 · outbound

This paper cites Adversarial Training for Large Neural Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Adversarial Training for Large Neural Language Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.602842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.602842Z digest=sha256:65e6183b5e8eeefc8184fb0807cb15b084c628688f625ce3f0f47e9a208556fa

Observation 2e6d12c9-8c1c-48d0-a294-1eeebdd0ef52 · outbound

This paper cites FreeLB: Enhanced Adversarial Training for Natural Language Understanding.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training FreeLB: Enhanced Adversarial Training for Natural Language Understanding

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.609506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.609506Z digest=sha256:fc109b449e3eae7620368c6571c6c9c21a3136ef2b4c8859aad7586fbdffc527

Observation 350a898d-d298-48a8-bd4c-d7a3a13c8b39 · outbound

This paper cites Adversarial Training for High-Stakes Reliability.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Adversarial Training for High-Stakes Reliability

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.617525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.617525Z digest=sha256:fac6ddf4bbdd20cf9ca504f8b26e751a8edaf92accd2bee970c56c8dbd9d4144

Observation bba77999-6c87-41b9-bf16-cc2d7c3a81ad · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.626472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.626472Z digest=sha256:c97aa37a79c69900325c372182d29f5c72d49052bc047f87a32cde3b7b0c5462

Pith citing papers

No inbound Pith citation observations are available.