Pith. sign in

Paper Citation Record · LEDGER

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning

As of 15 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 3 inbound Pith citation observations for arXiv:2507.14987.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.14987 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:48:01.488303Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:46:20.547519Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:40:06.671432Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 88ba8c8e-c977-4d46-b786-55e5cba25ab7 · outbound

This paper cites Claude 3.5 sonnet model card addendum.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Claude 3.5 sonnet model card addendum

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.229668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.301870Z digest=sha256:bb6811064bf0dfcdb925810ba8a035e9eff48e2c88ef78e27846575d31626d9f

Observation d80e5d73-1e38-4a03-81a6-2d3663420e85 · outbound

This paper cites Claude 3.7 sonnet system card.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Claude 3.7 sonnet system card

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.219762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.305361Z digest=sha256:e01d9529707f2fb565456ed5f177cb3e34f97ba9d3c671d56cc872b0a3455bec

Observation 0bd55893-82c6-4e7f-b58f-c7a20b20f6f0 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Gemma 2: Improving Open Language Models at a Practical Size

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.309058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.309058Z digest=sha256:79d24f335c0c0626736c44ed37823cb91a09d19e4ed09727c5c4022c8d0e808c

Observation 3fe5fa3a-47f8-40a7-a1fb-c09fc21ac8d2 · outbound

This paper cites Qwen2.5 Technical Report.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Qwen2.5 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.312914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.312914Z digest=sha256:7eb59e53917bb727d5126ed94c3aae6556a4f374fdb8282caa2913a4891f8d4b

Observation b7dbd5c6-5592-49c6-ab0c-908c54cd6b8a · outbound

This paper cites DeepSeek-V3 Technical Report.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning DeepSeek-V3 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.316811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.316811Z digest=sha256:78fb46ea45908608af9f3230db1fc12ead39ea172c426a31fcccf3e4f5576873

Observation 3169a0e9-52a7-4e5f-812e-830775b113a0 · outbound

This paper cites Hadi Amini, and Yanzhao Wu.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Hadi Amini, and Yanzhao Wu

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.209551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.320269Z digest=sha256:3a81aacdec78bee3478cfb25b19f9380e54e2d19b06020634f5551d1e943e1ca

Observation e6826717-384c-4b13-b4d2-b0eedd7eb3d9 · outbound

This paper cites Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.323942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.323942Z digest=sha256:7e25f197a8a0a8b21f86dcd3a112450d3f3048111040416f342c3b5224bff215

Observation c4445efa-845a-47c7-91c3-2a9bb63aef44 · outbound

This paper cites Safe + Safe = Unsafe? Exploring How Safe Images Can Be Exploited to Jailbreak Large Vision-Language Models.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Safe + Safe = Unsafe? Exploring How Safe Images Can Be Exploited to Jailbreak Large Vision-Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.327760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.327760Z digest=sha256:3e944723c215b50fdcc7e9b0a7332f30bb445b7161713f242c85b30757856839

Observation 2ce16309-30d6-4c2e-81ab-dbe6b378be84 · outbound

This paper cites Kummerfeld, and Rada Mihalcea.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Kummerfeld, and Rada Mihalcea

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.198848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.331541Z digest=sha256:59ada76d0d7b6acd94e132c179625aff45acd1968e5935e253179ccf1d9414c5

Observation 9e2ec079-d332-444a-9c5b-7f9aabb0ce09 · outbound

This paper cites Towards understanding jailbreak attacks in llms: A representation space analysis.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Towards understanding jailbreak attacks in llms: A representation space analysis

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.187307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.335199Z digest=sha256:8ae85f00d46a6fcd7faabc76e54cbac4d1c53f10830cc1cf26926ffc4428050e

Observation ae7756c7-859a-47ee-9ec7-d1d2f800487f · outbound

This paper cites Uncovering safety risks of large language models through concept activation vector.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Uncovering safety risks of large language models through concept activation vector

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.176509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.340006Z digest=sha256:c6a97451faf9003f3e488505caa1788a5e3a10bbbfe98a4370a59c693706af14

Observation b86170f1-ddd6-431b-80c1-1690a8b0d795 · outbound

This paper cites On prompt-driven safeguarding for large language models.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning On prompt-driven safeguarding for large language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.164319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.344062Z digest=sha256:f94317316a59342d30f188acc022ffad2f6617bd09f227b5e4edc834ddf39196

Observation f572c8e5-e434-4ba6-bbb2-951b37b78fb2 · outbound

This paper cites Nguyen, Jun Sun, and Tat - Seng Chua.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Nguyen, Jun Sun, and Tat - Seng Chua

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.154293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.347270Z digest=sha256:8ee1469604f512274c29c6787dbc8a262f6e45fb6060d6f5fc736b1cd213f6a8

Observation 337aa62b-04b5-42e8-a50f-5fb9a4304c0a · outbound

This paper cites Jailbroken: How does LLM safety training fail? In NeurIPS, 2023.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Jailbroken: How does LLM safety training fail? In NeurIPS, 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.143935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.350463Z digest=sha256:5dac55d73886c89b01da5117992b576030c27f8f1022846d5b81fa757b4c659a

Observation 0064e74f-d67b-4e40-b717-2c99ddc7614c · outbound

This paper cites Safety Reasoning with Guidelines.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Safety Reasoning with Guidelines

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.356987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.356987Z digest=sha256:8e53728a26abc9c3acf5ca1be2bf5e74d9df76ed6af02e288d043d0b64ccfdcf

Observation 03d9e388-4c7d-45ee-9429-c8443e144e39 · outbound

This paper cites Rule based rewards for language model safety.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Rule based rewards for language model safety

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.132567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.360307Z digest=sha256:7a8e13c960d4caca50a156fa80e83e63843b7654ec4f2089f3c4151707733df6

Observation 1865f620-a909-40b7-a460-e45446f6d8f0 · outbound

This paper cites Safety is not only about refusal: Reasoning-enhanced fine-tuning for interpretable LLM safety.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Safety is not only about refusal: Reasoning-enhanced fine-tuning for interpretable LLM safety

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.363452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.363452Z digest=sha256:240ead9d4d848db76ba96c23067c6a068c3222808122a3008e7628685a3630c3

Observation a6aa999a-65fe-4a16-a1af-f85c3a9b6d6d · outbound

This paper cites Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.121098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.366857Z digest=sha256:1ca4bcc5fc3ad14802bec0854d808c6e10736806b387037ae337d927f4dfde7a

Observation 808083e7-b3a3-4e0c-ae05-a92ca6e7c8ac · outbound

This paper cites an unresolved cited work.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.369990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.369990Z digest=sha256:3b5b58657ecba8684dc9809d629a18826e96fd35d40aa8924a6f05a29998e9d2

Observation 259e10b5-eb2d-4174-be50-c6a1a9969e59 · outbound

This paper cites Manning, Stefano Ermon, and Chelsea Finn.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Manning, Stefano Ermon, and Chelsea Finn

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.373404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.373404Z digest=sha256:40743b73b1439c79f56eb04880bce96a5a852eae74808c40ef8e7225db89806f

Observation b502d8ad-ff9b-4891-9fb4-fa4f7923d7b2 · outbound

This paper cites Safe lora: The silver lining of reducing safety risks when finetuning large language models.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Safe lora: The silver lining of reducing safety risks when finetuning large language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.095789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.376537Z digest=sha256:d3b9a6024fdcda1d6db35abb4258296ad90a14b69efe076787de8107e157cc41

Observation c8dd559e-859e-45f9-8052-95d8bd935554 · outbound

This paper cites Zico Kolter, Matt Fredrikson, and Dan Hendrycks.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Zico Kolter, Matt Fredrikson, and Dan Hendrycks

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.085637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.379757Z digest=sha256:91b6d3eae79f84dc9c566c75f98627d79fda17d442cc61fc5d3ad01194d7a84c

Observation c4b5728f-b8ff-4436-83de-53d40cfe133f · outbound

This paper cites Safety Alignment Should Be Made More Than Just a Few Tokens Deep.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.382809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.382809Z digest=sha256:80e76f737452f0767e816debad2c2969d231c0c40256426fdedb25d8a1b6f4dd

Observation 449e52f3-1076-4415-b5b0-64a3b80315c0 · outbound

This paper cites Deliberative Alignment: Reasoning Enables Safer Language Models.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.386002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.386002Z digest=sha256:f68df907ddd33e11a57486b0f00b30eaa5b4da08e571db69160291eee0d67179

Observation 1600e39f-f572-4f6c-a4ea-866b9f0be95d · outbound

This paper cites Enhancing model defense against jailbreaks with proactive safety reasoning.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Enhancing model defense against jailbreaks with proactive safety reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.388982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.388982Z digest=sha256:4d816040c9135efd7e3b2cd099952da8d59631d5ecf71659177c6d0edd6af530

Observation c04d912f-6ad0-46de-9970-157ed05c112d · outbound

This paper cites Reasoning-to-defend: Safety-aware reasoning can defend large language models from jailbreaking.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Reasoning-to-defend: Safety-aware reasoning can defend large language models from jailbreaking

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.392182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.392182Z digest=sha256:0169541a7107f036df359a2ffc7b0dc5b0069e9e3ea5e3fa036cc66aebfeacf8

Observation 6c62134f-a5b8-4b75-8a62-20223ca065e8 · outbound

This paper cites Bikel, Jason E.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Bikel, Jason E

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.076369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.395614Z digest=sha256:6f8ab78b38c8602c7edfa0ece36b210996136372c7a76f32b1ab776ccb702991

Observation d50ed610-393a-461e-9e5c-49de9b04a66c · outbound

This paper cites Does Refusal Training in LLMs Generalize to the Past Tense?.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Does Refusal Training in LLMs Generalize to the Past Tense?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.398901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.398901Z digest=sha256:812c519d54828bf576437aa0c51d3b17c71b1c0072d0e28a89eb2e02517e20e5

Observation 2b78b9f0-7a0e-4039-8e93-d406d8f2ae0c · outbound

This paper cites Fine-tuning aligned language models compromises safety, even when users do not intend to! In ICLR , 2024 b.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Fine-tuning aligned language models compromises safety, even when users do not intend to! In ICLR , 2024 b

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.066996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.402803Z digest=sha256:4b853773581112cd65d183e2cc75965ffc62021e28efe9d6e76e63df951e63a0

Observation f1929073-70ea-4d28-aba7-49977dfbb656 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Constitutional AI: Harmlessness from AI Feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.405908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.405908Z digest=sha256:6f3aca9acd16fd583b4a76c22b59de59ecee2fbcef0bc7651ef239b62819e510

Observation 899225ce-cc30-48db-baea-f4a0d0aee895 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.409511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.409511Z digest=sha256:a7c87ce98f30f59eec7ae2980020f4a7836e6d0dfebf968b1c862fb807d3f384

Observation 360512d6-0b7c-4d22-9ad5-1df3e4542611 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.413147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.413147Z digest=sha256:3602180bfcf29f9e1d5665f0f8297ec6a4c59d7b68061b420fe57d5f424846d7

Observation 81b675cb-451b-41e5-8199-e56ff022df83 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.416557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.416557Z digest=sha256:78997e8ae4a2ae37e576eea882061101df892dbccfd43de8a3ecb8a5eee809ef

Observation d2109089-99c2-41ee-9788-e561901be489 · outbound

This paper cites Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.420708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.420708Z digest=sha256:273a6e12e7ca868ac80333a7ccfb1daeb67d8955ed9988bc22bb355f2e419de9

Observation 9af60e70-f02b-4ebb-8450-da15badc50b0 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.424241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.424241Z digest=sha256:2691da1cc6e2568adb9da879d831d152e43b443e9c676c518606c804997659af

Observation 5a80a100-a6ef-4799-9694-013717548600 · outbound

This paper cites Proximal Policy Optimization Algorithms.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.428077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.428077Z digest=sha256:77dded35d82945067cbb632dd40748fe2457f985ddb62d6ee3aea8fbd26d8760

Observation a5ed7fa2-05ae-42b8-83bd-5e6dc2241560 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.431316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.431316Z digest=sha256:007438d4d870cd23c09f1e746dfd96acbeafdd030341ed6daa745623dc73a082

Observation b6d57317-baef-4e62-ac3f-bf8c51e592c0 · outbound

This paper cites Jordan, and Pieter Abbeel.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Jordan, and Pieter Abbeel

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.056797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.434748Z digest=sha256:ad43744a3d240a0599c6a440f0c0ff5f8b34c140d063933369f6fbd2c74bf51e

Observation 0f6e7206-c031-46f3-8e11-5fbed43cc04d · outbound

This paper cites A strongreject for empty jailbreaks.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning A strongreject for empty jailbreaks

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.437956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.437956Z digest=sha256:d6a737dc1b3153e5f85e6e25abb35846f2de2efcb9a3215cc93107910f4e766d

Observation 05a11772-9b8d-4347-87ad-2ea278f2011e · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.441397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.441397Z digest=sha256:660c59f93431463bbd72e62cbe6280884bcacd4469e4a2490f377b38b48dcf4c

Observation cda7c6b2-92cc-4789-890d-0239ae4548ac · outbound

This paper cites Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.040032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.444984Z digest=sha256:936171005b6c2a57c6680443d7c7bea1c714e89d4ba5ce29070170c7de5f9f30

Observation 04cb8820-72d2-4536-b12f-2c765052765e · outbound

This paper cites an unresolved cited work.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:48:02.029712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.448493Z digest=sha256:239e66f35621adbc1aa8cf9f4ff89bb50dd15d36983bd0ad2cf05555562d505f

Observation 53376565-e138-4c82-bfd6-a4b69299a99a · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.451746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.451746Z digest=sha256:f460f7ac7134cd98d9b7938ffc11e93e0ecd5a8e23c4455b4f85788804c079a7

Observation 2de32363-6a72-45f3-a94f-7fa96500ef52 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.454863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.454863Z digest=sha256:fb66a869a280147eb30f0bcd77c83acc5007c008a49bcd8753be24baf408ff88

Observation ddb9a189-4774-4efc-af9c-4a844b2ae6c8 · outbound

This paper cites Smith, Yejin Choi, and Hanna Hajishirzi.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Smith, Yejin Choi, and Hanna Hajishirzi

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.019937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.458530Z digest=sha256:51eb145b38ebd217c0cc15c233a9435e01d0d3d81d85123426251a6cee405ebd

Observation 6a0dec33-4c25-4381-903b-8a2bfec30c1e · outbound

This paper cites Measuring massive multitask language understanding.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Measuring massive multitask language understanding

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.009395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.461444Z digest=sha256:0d5cf57b3440e6d9d315bfc2e2471af89afb6c3dae5736a958b57e6c0a0b0dea

Observation 9d460353-3ff9-41c4-b0c3-679149ea8655 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.464564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.464564Z digest=sha256:761c93e69999d3004fc82c22166eafb92c4cc9744658285b480d2534cd3568b0

Observation 95f9e80f-c966-49b4-9e9e-1dc38a4d5fe4 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Training Verifiers to Solve Math Word Problems

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.467912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.467912Z digest=sha256:0384fc2a2f43f361abe878e7e3a114a1386df17ce22887f47243d99551997cc4

Observation ffa41b95-17d9-4168-811e-97642a8be57d · outbound

This paper cites The Llama 3 Herd of Models.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning The Llama 3 Herd of Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.471447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.471447Z digest=sha256:6619c82d230d7f34b4ec410d3e7813168433497f67a4b2c4137cfae86cc7067f

Observation a9b04422-6b54-44f1-9b15-d7d196de2a5f · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.475049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.475049Z digest=sha256:bc2cded4d452f7992457b1e9e2f5f43e12cd97b7c94baf9a0ce2c7f332abfed8

Observation 572296ea-4dd0-4882-9ef0-3338dad7fd00 · outbound

This paper cites Free dolly: Introducing the world's first truly open instruction-tuned llm, 2023.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Free dolly: Introducing the world's first truly open instruction-tuned llm, 2023

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.478311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.478311Z digest=sha256:e5e741715afa33d54f81a6eaa903867505efec1d0692cb33418158e10a38f7f2

Observation 7feb1d53-edbb-42e0-9f61-9bff21590fc5 · outbound

This paper cites Xstest: A test suite for identifying exaggerated safety behaviours in large language models.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Xstest: A test suite for identifying exaggerated safety behaviours in large language models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:01.991310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.481650Z digest=sha256:a97940553772f8d4091b51ea0060a49d12a1e664d7df8b18882d0bc762432b35

Observation 22bfb320-4556-4c0c-8a4a-51cafd65d2cc · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.484886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.484886Z digest=sha256:a001c748debb2a2ecd0279e14f2da5583382a9671ee2c8ae7a3279f253ff2b00

Observation 5c2cfbf5-5502-4743-bb2f-114e7d7fc1a7 · outbound

This paper cites Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint, 2024.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint, 2024

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.488303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.488303Z digest=sha256:921cc8e0f5979adef84c735cf428931ca42519da1a609be9e5a4902e8857229b

Pith citing papers

Observation 0ad39314-b1f0-485a-8b5f-70fe4ef8b5c0 · inbound

Self-ReSET: Learning to Self-Recover from Unsafe Reasoning Trajectories cites this paper.

Self-ReSET: Learning to Self-Recover from Unsafe Reasoning Trajectories AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:11:16.129970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T02:07:04.364417Z digest=sha256:4e5bf90f26ebd8359f511202efc9e639390992aff141875d209f066795f2cea2

Observation a1575c0b-db97-42d4-a424-7ccff2530cb6 · inbound

PolicyAlign: Direct Policy-Based Safety Alignment for Large Language Models cites this paper.

PolicyAlign: Direct Policy-Based Safety Alignment for Large Language Models AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:40:06.672952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-25T21:09:19.727723Z digest=sha256:2a0dcef39d4619dccd0842a953490a5605eaf9d10972142f7eee1d93d39bc8a0

Observation 83c5b1b7-a17d-416c-9347-9ab31cd1660e · inbound

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models cites this paper.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T23:46:20.547519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:46:20.547519Z digest=sha256:6a36ab3ce1ff6ee0fd65a429232381cebc754007999ff32c913c96cb2beb1254