Pith. sign in

Paper Citation Record · LEDGER

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training

As of 17 August 2026, this Paper Citation Record lists 94 of 94 outbound references and 0 inbound Pith citation observations for arXiv:2507.14202.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.14202 v1

Coverage vector

measured 94 of 94 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:35:48.626472Z

measured 94 of 94 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

94 of 94 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved92
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5b6a9238-00de-4972-a6c5-c256e98d095c · outbound

This paper cites online" 'onlinestring :=.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.238850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.238850Z digest=sha256:df96ea22d53ca856b127e7c44b9ba33c96955e98392110ecfe200ca36ac8ea4f

Observation 25cf31f4-6c6d-409f-b5cb-f7b395d031a2 · outbound

This paper cites write newline.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.390160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.390160Z digest=sha256:2cdc32955ae0641291d206287e27a1eec1a1add6e5b8cbfda574735a21a66218

Observation 5a2af6a0-2e86-4504-8cbc-6ffeed6a1cde · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.424574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.424574Z digest=sha256:691becd8a4f2d40fae4074c3038885136ec6e5e5a129e6ad88322a4e34f132be

Observation bb252c79-5804-47f0-aa0b-2b0227ac2c84 · outbound

This paper cites Generating Natural Language Adversarial Examples.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Generating Natural Language Adversarial Examples

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.437181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.437181Z digest=sha256:aa01375268d7c49453df662e63e61c2e55dbf21201d8e36ed42cf10202197ef2

Observation ffcb86b7-123f-42ce-8130-0269d61955ec · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training A General Language Assistant as a Laboratory for Alignment

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.468119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.468119Z digest=sha256:75107b8b0522db65afde54d3496bc589cd637a9f29d545f712630aac9cfbf5ce

Observation adcb3d4e-4952-4233-a7a6-2fef21cd219c · outbound

This paper cites Spinning Language Models: Risks of Propaganda-As-A-Service and Countermeasures.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Spinning Language Models: Risks of Propaganda-As-A-Service and Countermeasures

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:35:50.534543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:35:47.522560Z digest=sha256:ea5137f36951527e23e46a9c26f23f4bcbe7ce2d2f30e57435d0060748990067

Observation 1f8b8625-ea52-45ea-9324-d027089e80ba · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.608128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.608128Z digest=sha256:2953935536bf001dfd35d9ec23c18fd7e81ea9ed296585b400f1bba7ffdb3b57

Observation b29ff979-9448-4e35-8bf3-47d149d4744a · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Constitutional AI: Harmlessness from AI Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.723342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.723342Z digest=sha256:0893b424d1cc271ba86ec339bc3ca7f0ef744d8785b57d84a1f00cce2c590aec

Observation 517e308b-7e38-4685-85ce-e9f649f6d4c3 · outbound

This paper cites Image Hijacks: Adversarial Images can Control Generative Models at Runtime.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Image Hijacks: Adversarial Images can Control Generative Models at Runtime

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.822419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.822419Z digest=sha256:7dd21220ec62105cad10053e4e6cea82e36a962c929606bd63c7f45e745e9efe

Observation d51cbcab-b7d2-41ed-9e8c-b1dc9bbbc282 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.897090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.897090Z digest=sha256:68dbfe1feb24aea8b2bf65d393e6f8f6088d52ebab42011d27e7ad9e3e3f4c00

Observation 8202b0ae-a407-4800-8fcc-ad004b22a9d0 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.960711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.960711Z digest=sha256:5bb66638f01ca7399e6a406b7af25c97fd8edc889b01a4a8960e04459676e7f8

Observation 21677c54-689c-4f87-803a-53f3d739d648 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.970414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.970414Z digest=sha256:9b64a422c7bc88de2460ecb50e4ac651981000877308ad254a5f8c98d06737a2

Observation ce3d1067-2d5b-4af6-bd46-b95bde12f133 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training On the Opportunities and Risks of Foundation Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.975977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.975977Z digest=sha256:dcf75fbf9a77cadf0a1267c6457cd5300ba9a1288530b0d47aa6975fb3f7ba3e

Observation bafa7bb3-a8c0-48c3-884f-7ca91f77b74e · outbound

This paper cites Evaluating the Susceptibility of Pre-Trained Language Models via Handcrafted Adversarial Examples.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Evaluating the Susceptibility of Pre-Trained Language Models via Handcrafted Adversarial Examples

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.982288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.982288Z digest=sha256:e46de4bf13c52db29f5ff19b03e4e8dfa211e8cf996cf31bb6c02e817430403c

Observation b26876fb-8ead-4a31-9765-0039acb31008 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.989197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.989197Z digest=sha256:44d3a656c9392328c2beb1b286e32b4c16b549021158c15cd5acd03e92d8953b

Observation 4bb4e62a-f278-4460-86c3-e2119a30d3e8 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.996109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.996109Z digest=sha256:4c376509c27a61534636e60be95301823669dc2302019f051d500448e534d743

Observation 8a5a0256-0ed2-4a11-a324-0a6ae5dfe732 · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.002781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.002781Z digest=sha256:a386de45d0a0892ca917186b76e25fe6c81d05b9edf0e5c8a83c740489b2539b

Observation ee47c9ff-dcb0-43a9-a097-9fad4e0e75d8 · outbound

This paper cites Explore, Establish, Exploit: Red Teaming Language Models from Scratch.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.010807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.010807Z digest=sha256:c0a2c1fa58923e33811c44ceac614345d4aa2f9a507bf42af2f011f21a8db135

Observation 4fd144ea-56ba-462c-b157-c2c2fccc610b · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.018055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.018055Z digest=sha256:459c67a56f27f18b0329bd26fa25ec463ba59c461c8bff7bf8f26130406b55ad

Observation 2f3b1b5c-6c78-4d68-a448-028cd001b60f · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.025371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.025371Z digest=sha256:9a44431d010b13769d38f234de552e13f314b821a89b76c25c97b7b6eb0b218a

Observation 670b6015-0d51-4533-b7e0-7978b003fe56 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Evaluating Large Language Models Trained on Code

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.035926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.035926Z digest=sha256:f7392c9837c5538b30e64b6b6d5be7997b4984a72175c34c408592826dd3828c

Observation 3aa09c3e-12c8-4063-9d99-bb9690b6accf · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training PaLM: Scaling Language Modeling with Pathways

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.042843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.042843Z digest=sha256:bc75bf572b013864a86e85e0a3c93933d62e3e2764ff79d4253d729209832d3d

Observation 3d70f730-9dc9-4af2-bf75-7a0924bd6251 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.053465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.053465Z digest=sha256:0561200809042369b22fb8580f93688f3743703fc1f94d19ed81b82b1fca9ff2

Observation a11d8663-aa43-4de8-bd1e-a858946f0824 · outbound

This paper cites Supervising strong learners by amplifying weak experts.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Supervising strong learners by amplifying weak experts

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.063524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.063524Z digest=sha256:c404de6621fb950797b682838d3be83b94485da29bfd17f86ebdeb4edf36bfb2

Observation 9f0d440a-ef33-4ceb-9f85-df8526eb4c94 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.070951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.070951Z digest=sha256:d7ebe90cc294f4dfdb0665f2684b9dd878505ad5fd896d321607000114ba532a

Observation c904d3f8-3630-4b9d-8b71-f52597e821f0 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Training Verifiers to Solve Math Word Problems

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.078087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.078087Z digest=sha256:3ebd8fa864c5844e763948f69194a9ae83fae7d20dded9b14cb078e652046be2

Observation d6639d9f-321d-44b7-ab7f-ad5bad9c5d6b · outbound

This paper cites Multilingual Jailbreak Challenges in Large Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Multilingual Jailbreak Challenges in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.084454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.084454Z digest=sha256:1ab637946b30b93e1543a06ab63f3b1df6d9e4d9de471f8c9fc0b8099992a3b2

Observation 68c61a06-cfae-48e7-abb9-2bb1571c8feb · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.092363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.092363Z digest=sha256:3bd6f100d9d40e0f3f8bc47333f744bf5400f8a9b50a31f7e540aad952c177b1

Observation 9acddfb3-9cda-4059-94d5-cd8bd0434b18 · outbound

This paper cites Understanding parameter differences between analyses employing nested data subsets.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Understanding parameter differences between analyses employing nested data subsets

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.100821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.100821Z digest=sha256:0dbd49ff863b6727179113fbc1e8f861f09422b07375565e2ba26aaad4944b40

Observation 7f34939f-ae70-482c-8392-47ddccbdb6f8 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.109664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.109664Z digest=sha256:d322178b5a51ec250f0040d8b7bcb33be07f42b7f1bc2ad50e5a2d6e04370c83

Observation 26c25ec9-a7aa-48cb-96d2-11d2f2c0cce6 · outbound

This paper cites HotFlip: White-Box Adversarial Examples for Text Classification.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training HotFlip: White-Box Adversarial Examples for Text Classification

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.119501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.119501Z digest=sha256:1653041d14c7be062004e5c92f712bca23f6ffcb1284a8cb5cc632962dd56dcc

Observation 3f696111-d4ca-4c96-85f0-7d428d2af8ab · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.128310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.128310Z digest=sha256:f1e8e73dd7817502520c0172278c69fe2bc84405d1e83894031ee6e298ad74c3

Observation c42d8818-3e1d-410f-87ba-4fa7f898a85c · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.134541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.134541Z digest=sha256:eb2abeaf5989b0dc71de8ab19fe4d94f2f14c43f0654d0a676d3ceadb80634c9

Observation 2ce2fcbd-d0d9-4e9e-9491-0c12d0b58c02 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Scaling Laws for Reward Model Overoptimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.142756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.142756Z digest=sha256:f3cc3768455ee5f925693dcd1325ccf1d044797386c799ca6b44555b07e51838

Observation a90ddecf-99c1-405b-be46-b442ad14d8ed · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.151517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.151517Z digest=sha256:b6a4b86f1aa74d3bfcd4c59445d94873ea06a160ab3f5de1273b8eb97cb1d7be

Observation ebcc7130-ae12-485b-8d12-beb2658a661d · outbound

This paper cites Explaining and Harnessing Adversarial Examples.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Explaining and Harnessing Adversarial Examples

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.158428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.158428Z digest=sha256:7f4876e97bd0ea85bc439aa895c74363dfa8076fc2746aabf3fe0cfcdbefea05

Observation b94e92c5-e06a-46e3-8ddb-3c4b945c6be3 · outbound

This paper cites Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.166666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.166666Z digest=sha256:779d2ce98621b842b499a3ea44f3bf128f16de58cb7998c404ce30f5472e668d

Observation 590e6d0d-c3b9-4aa1-91cb-f62b13467cb5 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.981468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.174906Z digest=sha256:455ad0b43cefc4442e30576bae00a77a11c57721b0d63655416025475be67e96

Observation 0358867b-52e6-45ba-b59c-97ede6e52fe3 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Measuring Massive Multitask Language Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.184013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.184013Z digest=sha256:bb39565812466a7f189ecc3be097e6f42d998857243b46d4fa0b5eeaaa5b1a81

Observation 28085756-3d06-4400-9799-143a92401037 · outbound

This paper cites Training Compute-Optimal Large Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Training Compute-Optimal Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.190750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.190750Z digest=sha256:c91ab9591c4cfca8d22b8f24f9dd567128d0a40ce35b536474b5ab42747402d5

Observation 6df30fc5-7a42-4e0a-883e-420551bc9683 · outbound

This paper cites AI safety via debate.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training AI safety via debate

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.198213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.198213Z digest=sha256:0b1a5acc3d91d258a903c505b73eaa89f730db7506a2f6696e61d98f25712e09

Observation c3918a90-c15b-44a3-8567-3f9d1bed0a53 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.954180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.205695Z digest=sha256:1245d77278fc95be0f7159b92e95e246cdc63824129908b2dff19f159be39305

Observation 8e4b1b4f-1cda-4008-9ac4-a6828bb2fbb5 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.925417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.215145Z digest=sha256:ee17a110a6b7ffcad60b990c0821678e1c6f20bf006f1ed29987ab4f7a7c3775

Observation 8483a57a-aedc-4071-8651-da0a3aacd10c · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.898671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.223231Z digest=sha256:05e2f68e1776f8fdc43ab9b541b49bbdb66d66e25801f35a7d08473750d8baa6

Observation 72828130-3626-46ec-aaa7-9b4442706d5a · outbound

This paper cites Automatically Auditing Large Language Models via Discrete Optimization.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Automatically Auditing Large Language Models via Discrete Optimization

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.231633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.231633Z digest=sha256:8804f5c3531835c8de3c56207f747a2becb3946eeb378cbfdebc3503b189443d

Observation c5bdc0bf-891b-4c3d-9cf6-b5b29d6825b6 · outbound

This paper cites Single-Source Shortest Paths with Negative Real Weights in $\tilde{O}(mn^{8/9})$ Time.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Single-Source Shortest Paths with Negative Real Weights in $\tilde{O}(mn^{8/9})$ Time

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:35:49.762344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.242664Z digest=sha256:ab51758c1caf2dc921ceadd3d4dc6dd90ea92705988e1ddcc62b573cc9447099

Observation 8e39d5a0-df53-4a41-813c-fd6d04b2881a · outbound

This paper cites Open Sesame! Universal Black Box Jailbreaking of Large Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.251121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.251121Z digest=sha256:d8d97573378fcae9653846840c234914c24c29566d61439a117ec1ea87051eca

Observation 6893840f-8913-456c-8bfe-03992331d013 · outbound

This paper cites Scalable agent alignment via reward modeling: a research direction.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Scalable agent alignment via reward modeling: a research direction

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.260173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.260173Z digest=sha256:d48807bc8e76b1bd164f5eaacd4096b8a449a96f2db17dcf953e30adafb422c9

Observation c2709b8d-4f65-4d21-ab92-6b9473858fa6 · outbound

This paper cites BERT-ATTACK: Adversarial Attack Against BERT Using BERT.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training BERT-ATTACK: Adversarial Attack Against BERT Using BERT

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.267389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.267389Z digest=sha256:0173a6cc0c305a3d0cc0b89898d9c846c1982ace0d081d69d6744021675cb147

Observation dd64e832-2a45-4da9-bb99-41539ed0ca24 · outbound

This paper cites Let's Verify Step by Step.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Let's Verify Step by Step

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.273778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.273778Z digest=sha256:d17b947faf6ace7e42b7a11d89d9c4ac545d7273383a1417ef5f6a712fd6607e

Observation b907c4eb-d916-4cbd-8fa8-60a699c361d2 · outbound

This paper cites Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.280960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.280960Z digest=sha256:ffc7999c810020e266086bc2f8d4b72677aed2e34067f32cb769ff4e8667fa9d

Observation 67ea4d27-d6a9-4d67-a875-6c9bd7c3dada · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.289162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.289162Z digest=sha256:c27fc8776b7a45895e69ffc3fb1dc4b36371117ed19d1da3324ff2d1cc923d1a

Observation c8382987-8b29-4efe-b8c6-1a8ea5283f86 · outbound

This paper cites Towards Deep Learning Models Resistant to Adversarial Attacks.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Towards Deep Learning Models Resistant to Adversarial Attacks

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.302663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.302663Z digest=sha256:b8be8b3c555f892b7c1124f06f215cb20725ceb998d9236a5bea9db45eee893b

Observation 455d04b7-778b-4caa-897a-b06d86de4b1b · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.871578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.312375Z digest=sha256:a9cfd49cca3d65de8944bc3121a00f1c7cc1adf7459575b0601f7ff1287e0a22

Observation 53fcdf3a-c28f-4d2c-89ca-8008a48a9534 · outbound

This paper cites Tree of Attacks: Jailbreaking Black-Box LLMs Automatically.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Tree of Attacks: Jailbreaking Black-Box LLMs Automatically

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.318958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.318958Z digest=sha256:54b2a33eb103782f5ae022e3f745368a6579411eff33d3a8a22ade600ee4dfd0

Observation 77b7dbcc-5efa-42cb-92e5-75aeb0e2ccae · outbound

This paper cites Teaching language models to support answers with verified quotes.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Teaching language models to support answers with verified quotes

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.325823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.325823Z digest=sha256:7c14eb3033d646182b5c31c33a196427b38270dda3401d2f9913c698070646b2

Observation 55775cee-9169-4500-90d2-4eb9a2b99214 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.842475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.334358Z digest=sha256:0b33b6861a355a125dfd4d59b65e6644e6ce82d5061f2f3a31f13b7641a03acb

Observation f5cccaa2-679e-4f61-81c4-4e13dd8d6e5f · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.813928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.340570Z digest=sha256:7888ec85d096c2cc3ef656ab668bd298f43aa2b34d6edce933bc76430b434452

Observation 103ae790-ec3e-4edf-bf4e-122cf26e8719 · outbound

This paper cites TextAttack: A Framework for Adversarial Attacks, Data Augmentation, and Adversarial Training in NLP.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training TextAttack: A Framework for Adversarial Attacks, Data Augmentation, and Adversarial Training in NLP

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.347629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.347629Z digest=sha256:25fde3b4657f7fdf051b4c600cfa3358d978823a68e363b63c03bd678e70a54b

Observation 482fea23-f239-4910-9841-c2975ba97d87 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training WebGPT: Browser-assisted question-answering with human feedback

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.355471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.355471Z digest=sha256:793767f2b18f6d6721700b62c1d731ef69632275798ad2f81b14a460fc8946ef

Observation a701af0a-b1cc-4a10-8629-972734d4f80b · outbound

This paper cites Scalable Extraction of Training Data from (Production) Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Scalable Extraction of Training Data from (Production) Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.362446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.362446Z digest=sha256:62915ade0d9998f4421318bd8f690f8690bb414cebc23d8291300e81a8d9e53d

Observation 569713c6-6f4f-406f-9e2d-46aca8eaa3e5 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.370906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.370906Z digest=sha256:ab332f771b89ae43e726bbd0fd0d3bb869daf6a5bc2b4390ccfe0a175ebd91c2

Observation 206bfffb-3393-4296-b8ed-1a4e1c1631e4 · outbound

This paper cites Red Teaming Language Models with Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Red Teaming Language Models with Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.379195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.379195Z digest=sha256:4a36039f4f729463c6418377a4638154c06cd7bd17ce403e1414f329095298c0

Observation 1811695b-06f0-464a-b606-483a2f9e0eb1 · outbound

This paper cites Discovering Language Model Behaviors with Model-Written Evaluations.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Discovering Language Model Behaviors with Model-Written Evaluations

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.387140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.387140Z digest=sha256:27461f53a5538d31fa452821a7ad378ba19ddb4f64da25d1157f8b3041da68fe

Observation 3b8e7847-6a86-4fb4-8f3c-f5a359046bd3 · outbound

This paper cites Ignore Previous Prompt: Attack Techniques For Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Ignore Previous Prompt: Attack Techniques For Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.395357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.395357Z digest=sha256:c3857501a8e0d59d74df325bab9b457dadc0c1457447ed4e6ec2481d78105e27

Observation 83499b1b-c548-4c87-9a07-9c519391f9b0 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.775043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.406503Z digest=sha256:82527d6f016c7e4d7cbe6f3ceadd138f7ce2e3f7613ca1541626257a6a7e3084

Observation 3fb3cb18-7a56-41ee-aad5-997006c0baf9 · outbound

This paper cites Adversarial Training Can Hurt Generalization.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Adversarial Training Can Hurt Generalization

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.413261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.413261Z digest=sha256:207233472aebcb2f78f5a000251853c982a6a61e07d387108ecbd8af77b56cc9

Observation 453ce8c0-9e8e-4e86-8651-b26a2615ab74 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.752692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.420411Z digest=sha256:5a0c80a8fa2d65088480ecc1258f64fa7330700c9c3726e1d1416fc8303c3faf

Observation f1664621-df1b-4744-89eb-98e8f0ef31ba · outbound

This paper cites Proximal Policy Optimization Algorithms.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Proximal Policy Optimization Algorithms

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.427520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.427520Z digest=sha256:d5e839a79e43c0ec93f971fa5ee635857745d44e03eb1c5c9aaee6a429fac97e

Observation b853babb-024f-485f-bda8-0d6e9d553ba5 · outbound

This paper cites Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.436360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.436360Z digest=sha256:797d14cd276cbb2d5d0138eea7d691f3618b329bcf9c0e35cff71e5bdbe1e8d0

Observation 09e47225-07a6-48db-9c59-e42421f41fe9 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.445431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.445431Z digest=sha256:014ceee594b796cf96c81549c165590a074c93a07ed947bf720aa388c7685f8e

Observation 0580d262-b859-4833-9516-6a4889a99700 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.454219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.454219Z digest=sha256:2976aefb5a3cb97aa1c2e1e05a5ebd4e03fc48b61ae6a4afaa94a86e1a1ab905

Observation 4360779b-94b6-4e1e-9494-feb52167a0a3 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.710417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.461142Z digest=sha256:63cd122b66a70af6dd98e545293d5edb0159cf15ef39b60a2d249b9cc7aace70

Observation 5123c869-efcf-4be5-b141-c88b49df9fec · outbound

This paper cites Ensemble Adversarial Training: Attacks and Defenses.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Ensemble Adversarial Training: Attacks and Defenses

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.467325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.467325Z digest=sha256:5e5ad38808ea5abcce460dc819bb96fecbda6a7b583646a89fa7f9a2aafd2fc6

Observation f7629bdb-6602-453e-9138-b15b6227f1d3 · outbound

This paper cites Robustness May Be at Odds with Accuracy.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Robustness May Be at Odds with Accuracy

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.474839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.474839Z digest=sha256:9f1a66dbdb09edafdb38b108173c23571219fdef6cbdefeef94e6ba3835d4618

Observation aefb1499-3f67-469c-b28d-49a7afe5f42f · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Solving math word problems with process- and outcome-based feedback

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.481517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.481517Z digest=sha256:5409748f6572564c9d5809504068a593ea721b2cc9c6abf942c9f78f43c16216

Observation 923ae7d5-0634-43e6-9a9a-5af5f2a7d9bd · outbound

This paper cites Universal Adversarial Triggers for Attacking and Analyzing NLP.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Universal Adversarial Triggers for Attacking and Analyzing NLP

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.488605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.488605Z digest=sha256:121e412f61b90824575fee1eb81ebf3f67f6ca059c02f4728fc8738414d4824b

Observation 1581c38e-115f-4014-b61b-b00e5933eef6 · outbound

This paper cites Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.496124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.496124Z digest=sha256:e64b001a9855f4c862e265e14cc6dc9b19651a14a25cd0a5a5fc5db8b5658f7c

Observation c204d952-f567-4cbe-8f4a-21ee9e541aa8 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.688433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.505595Z digest=sha256:affa4d4178dfe6f3bf25d3adabbe6fe9435876f8765b9532c1b3a6537cae1e65

Observation 6314ad03-312d-443c-ada9-6fc5fa4f6b2f · outbound

This paper cites Natural Language Adversarial Defense through Synonym Encoding.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Natural Language Adversarial Defense through Synonym Encoding

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.513887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.513887Z digest=sha256:8fc9481d44730a4cda3a2dfadb7609027fb3f375b932ac80aca6cdadad696665

Observation 0db3733e-7e4a-4a0e-ad85-2382dfd528e3 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.522431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.522431Z digest=sha256:b1a3a26667d764ff84a68356cf27ba2308a9372a15dbf63a990a4f0e976aa509

Observation 51d9b036-4365-4a9c-b6c8-01475b3ad402 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.666517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.537349Z digest=sha256:c0b1e9a6e436b5a85aeb5e26519a603e15c8b0b7ec969c5a4d06f3f36074be08

Observation 85d265f5-a6dd-4f4c-bada-8c01f9655ca4 · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Jailbroken: How Does LLM Safety Training Fail?

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.545520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.545520Z digest=sha256:637a4e7d706ea686222a23b4de4851154293dc58c6528649a511d6b5ee8f6956

Observation 822f62ee-5f4a-449b-8f91-4f6dc81ff76e · outbound

This paper cites Ethical and social risks of harm from Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Ethical and social risks of harm from Language Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.553979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.553979Z digest=sha256:5a3d410a6c37b69c36507c42e736700ffcb8e0f1ce6ad765223fafb29ec39f91

Observation c54eaadb-dc35-4b7f-a38a-b6c517553455 · outbound

This paper cites Exploring The Landscape of Distributional Robustness for Question Answering Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Exploring The Landscape of Distributional Robustness for Question Answering Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.560955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.560955Z digest=sha256:4628ce1890a3c63e88bbd0b10daa9762d0dd5ffe75129bd02c290921213d2e5c

Observation a8d55785-7081-4caa-a72b-e802dcad7c97 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.642464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.569962Z digest=sha256:a1767fde488ab94c0c40481c9adefc73ac065cb0a2a684eb8e8a714b7b0b9cc7

Observation 72c27dd0-14f0-497b-baa6-8cddeb7a5a7a · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training ReAct: Synergizing Reasoning and Acting in Language Models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.576790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.576790Z digest=sha256:07dad718cdcf7e916a44e1fdb908f7807e732db8f45a13355c89d4104a1ab573

Observation 015af25f-f1e0-4a56-bfe7-fea6ed1beb49 · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Low-Resource Languages Jailbreak GPT-4

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.582594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.582594Z digest=sha256:0897e86c14a7020727d381010cae9218b3a709441c39fd40350b12bede33768b

Observation 84165dfd-1663-421f-965b-d33159b6d948 · outbound

This paper cites GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.591046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.591046Z digest=sha256:e765e475a3c2734b86187235138af715835adfa6d86b71a3037ca87ebb18c215

Observation 88de92c1-1199-4988-946c-6ef565782d31 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.613453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.596978Z digest=sha256:dc0b0b8b098803b6196e03815777a90585867ea1baa85a34a6e27d595f1e5e62

Observation 8b547ac2-a353-4344-a0a0-525fd1af06e9 · outbound

This paper cites Adversarial Training for Large Neural Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Adversarial Training for Large Neural Language Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.602842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.602842Z digest=sha256:00d04e5d14324f47d79d1d648025d617bfd1fcdd4115153f9042c15a3372fd86

Observation 2e6d12c9-8c1c-48d0-a294-1eeebdd0ef52 · outbound

This paper cites FreeLB: Enhanced Adversarial Training for Natural Language Understanding.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training FreeLB: Enhanced Adversarial Training for Natural Language Understanding

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.609506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.609506Z digest=sha256:6bdc7025ffffba31718120887238b9ff90dc46b50a984b081acc1605d0eb7754

Observation 350a898d-d298-48a8-bd4c-d7a3a13c8b39 · outbound

This paper cites Adversarial Training for High-Stakes Reliability.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Adversarial Training for High-Stakes Reliability

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.617525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.617525Z digest=sha256:cbd55af4d69ba98fbc87321b035b1a1b95c7b322a0a0bb1706238a3ad8187b40

Observation bba77999-6c87-41b9-bf16-cc2d7c3a81ad · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.626472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.626472Z digest=sha256:6b9c87562a593524580356bd9acf8c08620e43a4570047009a8614160e906936

Pith citing papers

No inbound Pith citation observations are available.