Pith. sign in

Paper Citation Record · LEDGER

STAIR: Improving Safety Alignment with Introspective Reasoning

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2502.02384.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02384 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:00:09.891654Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T08:55:35.501024Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 389ca17b-67fe-4ee1-a84b-4395171c63ce · inbound

Lifelong Safety Alignment for Language Models cites this paper.

Lifelong Safety Alignment for Language Models STAIR: Improving Safety Alignment with Introspective Reasoning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:09.891654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:09.891654Z digest=sha256:5eac49c6347c54ae14e06c1c704c8e5ea8c05d84cd44ce86a6212a7326018d25

Observation a416ccbb-e87d-41e1-8d32-44294097d6bf · inbound

Breaking the Ceiling: Exploring the Potential of Jailbreak Attacks through Expanding Strategy Space cites this paper.

Breaking the Ceiling: Exploring the Potential of Jailbreak Attacks through Expanding Strategy Space STAIR: Improving Safety Alignment with Introspective Reasoning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T13:36:00.906180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:36:00.906180Z digest=sha256:b3f0491a6301faae2a45ee27239a5616faea60ada121571242020736ddfd3c1f

Observation d46e43ee-3901-486a-9347-9043cb9922c4 · inbound

Jailbreak-R1: Exploring the Jailbreak Capabilities of LLMs via Reinforcement Learning cites this paper.

Jailbreak-R1: Exploring the Jailbreak Capabilities of LLMs via Reinforcement Learning STAIR: Improving Safety Alignment with Introspective Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:10.437566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:10.437566Z digest=sha256:a1599769b3f408ddef783fe682413bc9ce85c6566226f14542553c849e022cc5

Observation 2e41912e-1da9-4fb8-acd4-e0371b474a0a · inbound

A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents cites this paper.

A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents STAIR: Improving Safety Alignment with Introspective Reasoning

Reference 168

Resolution
unresolved
no resolver link, observed 2026-08-06T21:34:45.247292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:34:45.247292Z digest=sha256:e41a717375d83989a63d748a2821a683784b1690ae4748d5436e5af1b3e2896e

Observation ac706d68-d974-483f-a88a-f08b92900242 · inbound

Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning cites this paper.

Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning STAIR: Improving Safety Alignment with Introspective Reasoning

Reference 192

Resolution
unresolved
no resolver link, observed 2026-08-06T04:47:25.662094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:47:25.662094Z digest=sha256:e660cf35a38e33737c60a40ee5198fe8eaf23fd686118d660b1fa5e863a7e9fe

Observation 2d086361-f90d-4fe7-96af-0460dd52875d · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency STAIR: Improving Safety Alignment with Introspective Reasoning

Reference 180

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:58:58.971164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:8f34eaade1476c1c0161321c18ef5d0d6036bea37fa6de64c4d1a933bf112fdf

Observation 7fc6467b-27c3-4ffb-a8a8-cfd20c6a6470 · inbound

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models cites this paper.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models STAIR: Improving Safety Alignment with Introspective Reasoning

Reference 282

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:08.278123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:08.278123Z digest=sha256:000ea7e9580c8b4a331a2ed4ebdb0b4cac31fb1f4f18e26c1e540e5f0e302390

Observation cc7a5628-d7d3-4cb6-90ee-a87cc183beb8 · inbound

Think Before You Code: Dual Reasoning for the NLSafety-Utility Trade-Off in LLM Code Generation cites this paper.

Think Before You Code: Dual Reasoning for the NLSafety-Utility Trade-Off in LLM Code Generation STAIR: Improving Safety Alignment with Introspective Reasoning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:56:03.241384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:44:50.762366Z digest=sha256:afb31e74bd24fb7a88545a87600973f3351eeb3a4763778407b90d0f2cbfaa16

Observation 0737ae2b-07de-4393-98fc-65a586ed0889 · inbound

Think Before You Code: Dual Reasoning for the NLSafety-Utility Trade-Off in LLM Code Generation cites this paper.

Think Before You Code: Dual Reasoning for the NLSafety-Utility Trade-Off in LLM Code Generation STAIR: Improving Safety Alignment with Introspective Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T00:20:00.682943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:20:00.682943Z digest=sha256:603f374cc89fc70490a88680807d952043bdfe0c6a160de93b85106a923cbc71

Observation b297c760-4626-49ed-b466-f4530ae1ee39 · inbound

Disentangling Intent from Role: Adversarial Self-Play for Persona-Invariant Safety Alignment cites this paper.

Disentangling Intent from Role: Adversarial Self-Play for Persona-Invariant Safety Alignment STAIR: Improving Safety Alignment with Introspective Reasoning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:21:07.585694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T17:24:54.796037Z digest=sha256:e03d8d781b69b0a4f21ae8da30c572c837c7d0d802f3fcb9cb7b0e3702188637

Observation cc4bf590-040f-479a-9a48-56df55cf6195 · inbound

PriorZero: Bridging Language Priors and World Models for Decision Making cites this paper.

PriorZero: Bridging Language Priors and World Models for Decision Making STAIR: Improving Safety Alignment with Introspective Reasoning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:27:18.588102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T05:25:05.907123Z digest=sha256:035cbafb15d5315b722d80ee493358d48c9c2eb44dba2c876efcf13630d4b17f

Observation 2d20f637-f108-4219-8c3e-2a3b6adf2909 · inbound

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak cites this paper.

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak STAIR: Improving Safety Alignment with Introspective Reasoning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:19:41.987164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-21T06:16:01.040236Z digest=sha256:88dc89ff39593a752055332690bccc43c98acdff524d61fa1837873052918f5d

Observation b0866cc7-3737-48fe-bd5f-8b63dcc85995 · inbound

MESA: Improving MoE Safety Alignment via Decentralized Expertise cites this paper.

MESA: Improving MoE Safety Alignment via Decentralized Expertise STAIR: Improving Safety Alignment with Introspective Reasoning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:52:35.302215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T18:52:00.377915Z digest=sha256:2feecc6b0b6ed66cf231ce6ec6c25ffdb9454d832640a1073c70ae02d5503894

Observation 6f213dd5-d7e9-4857-8646-da88fdcc71fb · inbound

Addressing Over-Refusal in LLMs with Competing Rewards cites this paper.

Addressing Over-Refusal in LLMs with Competing Rewards STAIR: Improving Safety Alignment with Introspective Reasoning

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:55:35.502417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:59:12.695984Z digest=sha256:8d593e68c797e9c2584c0c0f89fb7c0b50e1b394230ceeaa84fe50a7d224b327

Observation 06e04a08-3018-42c7-b4b1-24e4414d85e3 · inbound

UniNDM: A Unified Noise-driven Detection and Mitigation Framework Against Sexual Content in Text-to-Image Generation cites this paper.

UniNDM: A Unified Noise-driven Detection and Mitigation Framework Against Sexual Content in Text-to-Image Generation STAIR: Improving Safety Alignment with Introspective Reasoning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T19:54:51.025596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:54:51.025596Z digest=sha256:9e5d85dccbe660bb85591a09940619f3700da99c42c14bcb6e3ca557800eedae

Observation 820d62e9-055a-458d-b2e8-4c39047555f2 · inbound

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models cites this paper.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models STAIR: Improving Safety Alignment with Introspective Reasoning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T23:46:20.550587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:46:20.550587Z digest=sha256:9622797e03ab281df9ef00a40b0dce444d9156f8b4278b66e951606d8462fe76