Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T05:45:45.089748Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 2 inbound Pith citation observations for arXiv:2602.02600.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T05:45:45.089748Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-20T11:26:55.810822Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
38 of 38 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 0af9e6f5-b904-439c-927d-71c7ec52a5a9 · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f79f4f1f-4dd5-43bb-b67b-d8c4f1cf3c08 · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cef9de61-027f-4f8d-8796-dad0bd30655f · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bde1605-6053-4eaf-a72f-0bfe42687318 · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 645a920c-ee84-45a8-8165-a35bf813f178 · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models A2d: Any-order, any-step safety align- ment for diffusion language models.arXiv preprint arXiv:2509.23286,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d91a8e40-f89f-4d19-b118-a621952607ab · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models Crafting papers on machine learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4aa0717e-6e43-4b42-b623-f08cf4a9df94 · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models Refusal Suppression.Refusal Suppression is implemented following the prompt-based method introduced in (Wei et al., 2023)
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c45463bd-11a8-4de5-9d58-d0d0eef889d6 · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models Jailbreak Attack Initializations as Extractors of Compliance Directions
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ef57400-6880-4f34-bb57-99551e0497a3 · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models FlipAttack: Jailbreak LLMs via Flipping
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f9c8323-108f-4d09-802f-9e189777ab62 · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79833021-1f3f-4cad-8c93-f07497e42aa3 · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models Large Language Diffusion Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8270372a-f80b-497a-b771-f6f76a872fa9 · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5353b55-d912-427a-a9e8-f0f2abcc47dd · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models Large Language Model Safety: A Holistic Survey
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd1bd27b-07a2-4db1-bf4e-b921ce2f7fe8 · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models Gemma: Open Models Based on Gemini Research and Technology
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38181c6c-e721-4a81-8377-6ddca49e4aac · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models BERT Rediscovers the Classical NLP Pipeline
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 659cb4a2-6c44-4335-999b-26af90ba3d57 · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models The devil behind the mask: An emergent safety vulnerability of diffusion llms.arXiv preprint arXiv:2507.11097,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4623871c-c93f-49f9-a08b-3449b3bd9c6b · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models The geometry of refusal in large language models: Concept cones and representa- tional independence.arXiv preprint arXiv:2502.17420,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3b4624a-5506-431b-8382-3254edd8199a · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models A Survey of Attacks on Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ec446d6-0ff9-4998-95de-1f29e5e44e6d · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models Qwen2.5-1M Technical Report
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd6f25c0-3a76-4335-aa7f-b6ce870c33cf · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models Dream 7B: Diffusion Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 777b2d59-949d-4018-9c6a-eaf9c7c97545 · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models Discrete diffusion in large lan- guage and multimodal models: A survey.arXiv preprint arXiv:2506.13759,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89253ad6-01f5-4fbe-89ac-e287371ac473 · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models Jailbreaking Large Language Diffusion Models: Revealing Hidden Safety Flaws in Diffusion-Based Text Generation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e3c96fc-248a-4283-9bd5-e37970a29c02 · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models A Survey of Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e107358e-4f46-4aef-a539-16469f26a025 · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 080d095d-ce43-4f94-abba-3ec258137151 · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b124f1db-3f47-4ee5-a44f-0a6bd0bffd4c · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models In setups with active defenses, any prompt or response blocked by a defense is also counted toward the Refusal Rate
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99a4289c-9bf2-436d-97f3-dc8b23408f56 · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models True" if it is answering or
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd0d541b-b0c9-4ff6-b18d-c10906c6dd90 · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models Text: {response} Answer: Figure 8.Judge prompt used to evaluate whether a model response contains harmful information
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33bdee98-4d52-4f47-912f-00cb694bc98c · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models We use a subset of 100 PAIR-generated jailbreak prompts provided in the repository
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c26b6b94-e78e-47a5-8aef-4fad5fec848d · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models We use a subset of 100 Random Search jailbreak prompts for evaluation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04a391ec-90cb-412f-9616-00c55a0a6440 · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models For each response, we compute the average per-token negative log-likelihood under GPT-2 and reject outputs whose perplexity exceeds a fixed threshold
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d57c5549-f954-430d-9660-83838cb14afc · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models <|begin_of_text|> [INST] Task: Check if the following content is safe or unsafe
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0ba13cc-7577-4cd3-9ca2-2dfa84908394 · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety
Reference 2000
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d40a9abd-ada4-4e57-acef-072fcfbb3a7a · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caaa3fb1-b72d-4031-b394-2af6d2856717 · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models Diffuguard: How in- trinsic safety is lost and found in diffusion large language models.arXiv preprint arXiv:2509.24296,
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe89cb34-533d-47c8-9e36-339b16167c22 · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models Detecting Language Model Attacks with Perplexity
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76555a29-24b3-4b02-9b3a-4a097edaea0c · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models LLaDA2.0: Scaling Up Diffusion Language Models to 100B
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e4310c0-1c9b-4c52-9dbf-ef67e2c47430 · outbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bf662f9-f996-4d98-88ef-0c3e7c16450f · inbound
Differences in Text Generated by Diffusion and Autoregressive Language Models Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e245d122-1e7f-4868-8adf-2c7c0c89f951 · inbound
SafeDiffusion-R1: Online Reward Steering for Safe Diffusion Post-Training Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models
Reference 133
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.