Pith. sign in

Paper Citation Record · LEDGER

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning

As of 9 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 3 inbound Pith citation observations for arXiv:2507.14987.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.14987 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:48:01.488303Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:46:20.547519Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:40:06.671432Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 88ba8c8e-c977-4d46-b786-55e5cba25ab7 · outbound

This paper cites Claude 3.5 sonnet model card addendum.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Claude 3.5 sonnet model card addendum

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.229668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.301870Z digest=sha256:c8c6c63d3eddbb75618b1b917b29237342d386d8bc4dcd580a2119e64e1625ed

Observation d80e5d73-1e38-4a03-81a6-2d3663420e85 · outbound

This paper cites Claude 3.7 sonnet system card.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Claude 3.7 sonnet system card

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.219762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.305361Z digest=sha256:1f45e643abe4f46d7730bc3530e74fde4b51089006fe0812199a6d0bb05377d4

Observation 0bd55893-82c6-4e7f-b58f-c7a20b20f6f0 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Gemma 2: Improving Open Language Models at a Practical Size

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.309058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.309058Z digest=sha256:8e7cd87e3b036606d738d974a5002b1f7cde3ce7f8d12ffa832029a9561319d9

Observation 3fe5fa3a-47f8-40a7-a1fb-c09fc21ac8d2 · outbound

This paper cites Qwen2.5 Technical Report.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Qwen2.5 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.312914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.312914Z digest=sha256:54d030094b46106b229567d3be89ab9cd72ed3c0e9b6c59e1253f82b48ad09b3

Observation b7dbd5c6-5592-49c6-ab0c-908c54cd6b8a · outbound

This paper cites DeepSeek-V3 Technical Report.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning DeepSeek-V3 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.316811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.316811Z digest=sha256:2566da5dc5cd0b499b71cba2fb9c72ae873b4527d909cba47a5bb3c344889f32

Observation 3169a0e9-52a7-4e5f-812e-830775b113a0 · outbound

This paper cites Hadi Amini, and Yanzhao Wu.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Hadi Amini, and Yanzhao Wu

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.209551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.320269Z digest=sha256:728b45a6ddf96fa27ca2aceeb1ebf6a4f9554bd98efe42de422a653985966ea0

Observation e6826717-384c-4b13-b4d2-b0eedd7eb3d9 · outbound

This paper cites Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.323942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.323942Z digest=sha256:ac14103569af1e30d7ee7126894462f123bdd310f72c3d1433e5665d578f1770

Observation c4445efa-845a-47c7-91c3-2a9bb63aef44 · outbound

This paper cites Safe + Safe = Unsafe? Exploring How Safe Images Can Be Exploited to Jailbreak Large Vision-Language Models.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Safe + Safe = Unsafe? Exploring How Safe Images Can Be Exploited to Jailbreak Large Vision-Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.327760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.327760Z digest=sha256:837d6eb2a41cb7bb13ef2af19612419c6d9f418167978cdf7ff011312ac9270b

Observation 2ce16309-30d6-4c2e-81ab-dbe6b378be84 · outbound

This paper cites Kummerfeld, and Rada Mihalcea.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Kummerfeld, and Rada Mihalcea

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.198848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.331541Z digest=sha256:39ced999e7bf5fbdc953df621bd414dbaada703df16f8337cb78d75ad45c51e9

Observation 9e2ec079-d332-444a-9c5b-7f9aabb0ce09 · outbound

This paper cites Towards understanding jailbreak attacks in llms: A representation space analysis.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Towards understanding jailbreak attacks in llms: A representation space analysis

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.187307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.335199Z digest=sha256:d706b4370618f487a9590caefe77049101299dc9c14b68fb0fb2f829ebdc62fb

Observation ae7756c7-859a-47ee-9ec7-d1d2f800487f · outbound

This paper cites Uncovering safety risks of large language models through concept activation vector.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Uncovering safety risks of large language models through concept activation vector

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.176509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.340006Z digest=sha256:268556fb44ff96f58c3d2ce0d088bfc9135ae8cd804ee8945682f061b8b23b8c

Observation b86170f1-ddd6-431b-80c1-1690a8b0d795 · outbound

This paper cites On prompt-driven safeguarding for large language models.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning On prompt-driven safeguarding for large language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.164319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.344062Z digest=sha256:cfbd5d96a81e0b3b27666ef6e0ed9550d37baaa2d439c347527b2b1757a423d5

Observation f572c8e5-e434-4ba6-bbb2-951b37b78fb2 · outbound

This paper cites Nguyen, Jun Sun, and Tat - Seng Chua.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Nguyen, Jun Sun, and Tat - Seng Chua

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.154293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.347270Z digest=sha256:3afe9dd79890c3aa15752dd112829870a32ef43f07b58fe0840c70860a142690

Observation 337aa62b-04b5-42e8-a50f-5fb9a4304c0a · outbound

This paper cites Jailbroken: How does LLM safety training fail? In NeurIPS, 2023.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Jailbroken: How does LLM safety training fail? In NeurIPS, 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.143935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.350463Z digest=sha256:955ad8158947c91047aabb551dffb631ef6cf5bd0b7cefb4a1f5210d281814d5

Observation 0064e74f-d67b-4e40-b717-2c99ddc7614c · outbound

This paper cites Safety Reasoning with Guidelines.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Safety Reasoning with Guidelines

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.356987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.356987Z digest=sha256:ac6d923e6549a5e3ee12a490fbcd6e10c45241b950e4777567ad39e543fa9947

Observation 03d9e388-4c7d-45ee-9429-c8443e144e39 · outbound

This paper cites Rule based rewards for language model safety.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Rule based rewards for language model safety

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.132567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.360307Z digest=sha256:51be151a4bd70760cb6174d4e955dac4497d3b9d9dbdc93e8c741b724d4a5678

Observation 1865f620-a909-40b7-a460-e45446f6d8f0 · outbound

This paper cites Safety is not only about refusal: Reasoning-enhanced fine-tuning for interpretable LLM safety.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Safety is not only about refusal: Reasoning-enhanced fine-tuning for interpretable LLM safety

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.363452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.363452Z digest=sha256:ae5f90b5f9c900f0c6b446f33e1a1fc1fc34658e768555d8f01299559fb51fe2

Observation a6aa999a-65fe-4a16-a1af-f85c3a9b6d6d · outbound

This paper cites Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.121098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.366857Z digest=sha256:9f79f7c09028f7f201667bd44eca745817d9287d1f8ffc86842fbb5bd8f81b7e

Observation 808083e7-b3a3-4e0c-ae05-a92ca6e7c8ac · outbound

This paper cites an unresolved cited work.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.369990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.369990Z digest=sha256:b96c8f0235910cacff6d916a693b0e3d6ba82f1242840d3f11d8e25d4f24aca3

Observation 259e10b5-eb2d-4174-be50-c6a1a9969e59 · outbound

This paper cites Manning, Stefano Ermon, and Chelsea Finn.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Manning, Stefano Ermon, and Chelsea Finn

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.373404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.373404Z digest=sha256:60f534bdd1267086633267d58f9105f92803256728a09cef644f25e7372fd0d8

Observation b502d8ad-ff9b-4891-9fb4-fa4f7923d7b2 · outbound

This paper cites Safe lora: The silver lining of reducing safety risks when finetuning large language models.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Safe lora: The silver lining of reducing safety risks when finetuning large language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.095789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.376537Z digest=sha256:6e08aedf162acc42da4d4d775b596a75326afd9b31dcd17472ce02cb13ed5b5d

Observation c8dd559e-859e-45f9-8052-95d8bd935554 · outbound

This paper cites Zico Kolter, Matt Fredrikson, and Dan Hendrycks.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Zico Kolter, Matt Fredrikson, and Dan Hendrycks

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.085637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.379757Z digest=sha256:15b3e71d0128a51f84d0bc2114a06b2036c8cffac5ca8ec4b2c47d36c677ae29

Observation c4b5728f-b8ff-4436-83de-53d40cfe133f · outbound

This paper cites Safety Alignment Should Be Made More Than Just a Few Tokens Deep.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.382809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.382809Z digest=sha256:2d659cfa21bef934130f269189d5899eda9cef42611f0dc74342e31f99efb3d2

Observation 449e52f3-1076-4415-b5b0-64a3b80315c0 · outbound

This paper cites Deliberative Alignment: Reasoning Enables Safer Language Models.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.386002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.386002Z digest=sha256:f41d99b24bd5c54f02d988d34a11c6e064bea0fca43fbb615d1a1a70c7aa220d

Observation 1600e39f-f572-4f6c-a4ea-866b9f0be95d · outbound

This paper cites Enhancing model defense against jailbreaks with proactive safety reasoning.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Enhancing model defense against jailbreaks with proactive safety reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.388982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.388982Z digest=sha256:1ab95df92a8d574fc2d6401a765a815bbe656d365146c611780a4fca06ad87b9

Observation c04d912f-6ad0-46de-9970-157ed05c112d · outbound

This paper cites Reasoning-to-defend: Safety-aware reasoning can defend large language models from jailbreaking.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Reasoning-to-defend: Safety-aware reasoning can defend large language models from jailbreaking

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.392182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.392182Z digest=sha256:78ea0f4955749f550d403d847f310f5728b6efc46c112870668bc051bf6f2cd8

Observation 6c62134f-a5b8-4b75-8a62-20223ca065e8 · outbound

This paper cites Bikel, Jason E.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Bikel, Jason E

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.076369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.395614Z digest=sha256:19765d8e1db5c3b0794f70b24bcbfc7f35ce6b8529de5d996896bc04addd6729

Observation d50ed610-393a-461e-9e5c-49de9b04a66c · outbound

This paper cites Does Refusal Training in LLMs Generalize to the Past Tense?.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Does Refusal Training in LLMs Generalize to the Past Tense?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.398901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.398901Z digest=sha256:0fe98b73c0e9101888fa2c33f2c01f8a7fd18d37620b658c03b512d6785c5da5

Observation 2b78b9f0-7a0e-4039-8e93-d406d8f2ae0c · outbound

This paper cites Fine-tuning aligned language models compromises safety, even when users do not intend to! In ICLR , 2024 b.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Fine-tuning aligned language models compromises safety, even when users do not intend to! In ICLR , 2024 b

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.066996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.402803Z digest=sha256:45cdadaf905eba648f3f80cb9ce828ad5849839d1bebae46034c9e6132b1d03f

Observation f1929073-70ea-4d28-aba7-49977dfbb656 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Constitutional AI: Harmlessness from AI Feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.405908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.405908Z digest=sha256:1b822131a8aaf91fbe0afa44bb0db706fb8ce31cc87e38ef7974de8aa935989b

Observation 899225ce-cc30-48db-baea-f4a0d0aee895 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.409511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.409511Z digest=sha256:f0909606dff6da7069212ba3aff632601fe4bbfea28b643a478bdde988a07afb

Observation 360512d6-0b7c-4d22-9ad5-1df3e4542611 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.413147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.413147Z digest=sha256:ee0b315857383cf7ab7cc19e8bcf472e237e486a7c1c60ca50e3226ac2aae8ec

Observation 81b675cb-451b-41e5-8199-e56ff022df83 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.416557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.416557Z digest=sha256:04819f25d6ff3fee9d892b27cbe35571d7d7898f5d07fe0405105c5204cd9b84

Observation d2109089-99c2-41ee-9788-e561901be489 · outbound

This paper cites Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.420708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.420708Z digest=sha256:2b810e95d871c60e5f7f864ef87c945ff1026dbd924fa21d6cbe48faf711d1c6

Observation 9af60e70-f02b-4ebb-8450-da15badc50b0 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.424241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.424241Z digest=sha256:c6fa15e0cd306b5f96876e71bec20feadbc9646f871533481ad3dca643bf549f

Observation 5a80a100-a6ef-4799-9694-013717548600 · outbound

This paper cites Proximal Policy Optimization Algorithms.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.428077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.428077Z digest=sha256:644b5c362301ca662a509aebfde2062988a9d748d344da8f62fac7678f309c22

Observation a5ed7fa2-05ae-42b8-83bd-5e6dc2241560 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.431316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.431316Z digest=sha256:a60d49ffb1895441a5d6e8df23d75736772d5c81f0bdeb0c8527628fe9bb7b4b

Observation b6d57317-baef-4e62-ac3f-bf8c51e592c0 · outbound

This paper cites Jordan, and Pieter Abbeel.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Jordan, and Pieter Abbeel

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.056797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.434748Z digest=sha256:c3f20152b68a6be334820170fcc77e36940c5aba7bd02d31b07bf5281591627e

Observation 0f6e7206-c031-46f3-8e11-5fbed43cc04d · outbound

This paper cites A strongreject for empty jailbreaks.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning A strongreject for empty jailbreaks

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.437956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.437956Z digest=sha256:4b8c06423856b11a2a9deba555bc2d421725bdf00f4790afc7b79ba4112fc041

Observation 05a11772-9b8d-4347-87ad-2ea278f2011e · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.441397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.441397Z digest=sha256:43e2edb4cb09d78d1c9d5665958a18c44649acb7d871ae397750f875dd2f49ab

Observation cda7c6b2-92cc-4789-890d-0239ae4548ac · outbound

This paper cites Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.040032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.444984Z digest=sha256:4c48dd4029b1c7e21306491091061a179507d87de5b88275a36f110fbd205cd7

Observation 04cb8820-72d2-4536-b12f-2c765052765e · outbound

This paper cites an unresolved cited work.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:48:02.029712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.448493Z digest=sha256:9ad7bc2657008da8beaa0810d0ba026f0d1b9b861b985bb37343567446465571

Observation 53376565-e138-4c82-bfd6-a4b69299a99a · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.451746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.451746Z digest=sha256:21bba5dea115cf4e2a0cf2cb5feef279445be42f02b32c7d63a544f4a4a02c8a

Observation 2de32363-6a72-45f3-a94f-7fa96500ef52 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.454863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.454863Z digest=sha256:f50be38f11b2e0a947f81a095e39a96af82638fc6638cb3ca0f65f6e2fbcba4a

Observation ddb9a189-4774-4efc-af9c-4a844b2ae6c8 · outbound

This paper cites Smith, Yejin Choi, and Hanna Hajishirzi.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Smith, Yejin Choi, and Hanna Hajishirzi

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.019937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.458530Z digest=sha256:7982714281c8d07c76aa1eb48155e552681a1de40d1a4ffd4293a827f48f42b8

Observation 6a0dec33-4c25-4381-903b-8a2bfec30c1e · outbound

This paper cites Measuring massive multitask language understanding.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Measuring massive multitask language understanding

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:02.009395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.461444Z digest=sha256:1aa073fd16ae4c21dabdb019a1839a69360ae3674e15ce8a7d94b8aac019325f

Observation 9d460353-3ff9-41c4-b0c3-679149ea8655 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.464564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.464564Z digest=sha256:b7f086bde595fce8f1eea86c133dfd5643fdc3f3e5727e8892cf06e7b3000820

Observation 95f9e80f-c966-49b4-9e9e-1dc38a4d5fe4 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Training Verifiers to Solve Math Word Problems

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.467912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.467912Z digest=sha256:f751762887facbcceac3e2101731d5b96623206cd61178ee35ab6ff74381bc6a

Observation ffa41b95-17d9-4168-811e-97642a8be57d · outbound

This paper cites The Llama 3 Herd of Models.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning The Llama 3 Herd of Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.471447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.471447Z digest=sha256:5a5da439747420bdbb9d2e855d6997279a97610f2aa9e11308131996849555f9

Observation a9b04422-6b54-44f1-9b15-d7d196de2a5f · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.475049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.475049Z digest=sha256:60d800dd5acb81108f84be62e6c13431866ec133c241c3e419bedaa990a9f958

Observation 572296ea-4dd0-4882-9ef0-3338dad7fd00 · outbound

This paper cites Free dolly: Introducing the world's first truly open instruction-tuned llm, 2023.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Free dolly: Introducing the world's first truly open instruction-tuned llm, 2023

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.478311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.478311Z digest=sha256:cb97f18c53fb42bd4ddfc5d681902872e911d10461a6693c23adfb9b222b4a96

Observation 7feb1d53-edbb-42e0-9f61-9bff21590fc5 · outbound

This paper cites Xstest: A test suite for identifying exaggerated safety behaviours in large language models.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Xstest: A test suite for identifying exaggerated safety behaviours in large language models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:01.991310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T15:48:01.481650Z digest=sha256:57ac8a90497f453f07ba554da575ab488dbd18b7876fec9c193c34685f7293cf

Observation 22bfb320-4556-4c0c-8a4a-51cafd65d2cc · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.484886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.484886Z digest=sha256:5b86a224301c843ac22c0cbf08e6cbf62c44fdbf08e281454abdb84b8986e47f

Observation 5c2cfbf5-5502-4743-bb2f-114e7d7fc1a7 · outbound

This paper cites Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint, 2024.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint, 2024

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.488303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.488303Z digest=sha256:5e8dcbf0fd1b86fe5b55c43981c297c3352a617e95af7937f27bd5afa60253d7

Pith citing papers

Observation 0ad39314-b1f0-485a-8b5f-70fe4ef8b5c0 · inbound

Self-ReSET: Learning to Self-Recover from Unsafe Reasoning Trajectories cites this paper.

Self-ReSET: Learning to Self-Recover from Unsafe Reasoning Trajectories AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:11:16.129970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T02:07:04.364417Z digest=sha256:7eb2090e235da93f685e6092d5487661a4c228b2e81c835ddee71d178eb56f0f

Observation a1575c0b-db97-42d4-a424-7ccff2530cb6 · inbound

PolicyAlign: Direct Policy-Based Safety Alignment for Large Language Models cites this paper.

PolicyAlign: Direct Policy-Based Safety Alignment for Large Language Models AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:40:06.672952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-25T21:09:19.727723Z digest=sha256:33eb5bc1a821e2cd3170458fcb470e943e8e79566b4b00efbf9c40cc3076da28

Observation 83c5b1b7-a17d-416c-9347-9ab31cd1660e · inbound

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models cites this paper.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T23:46:20.547519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:46:20.547519Z digest=sha256:2bc85d30087509c6ac050e24cf42e41c63d1b8db64f3ab21b60fdf33d889e56c