Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T05:58:47.845786Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 2 inbound Pith citation observations for arXiv:2412.16974.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T05:58:47.845786Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-22T01:19:00.268857Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-22T01:20:51.944045Z
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 415bb368-ae73-4b0c-9f91-12d8b4ee574d · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Can NLP Models 'Identify', 'Distinguish', and 'Justify' Questions that Don't have a Definitive Answer?
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2f723bcc-5e0b-4061-bb80-8e5ad65cae2c · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs A General Language Assistant as a Laboratory for Alignment
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c321cdb-3502-4726-8c7d-e9a1d2d139a8 · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5ebfa30-5335-4250-9fb9-e7f65f673208 · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Constitutional AI: Harmlessness from AI Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5c01f6a-6782-41b4-b126-d7dafd5ede13 · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs A unified taxonomy of harmful content
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ee5359f-6805-4140-8d39-29666863ae9e · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Information hazards: A typology of potential harms from knowledge
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e3deb010-0cba-44ec-b3db-e02d6e300885 · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Deep Reinforcement Learning from Human Preferences
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c16f6304-151f-48cc-a1eb-357e90be164f · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e096ccb-11f5-49cb-8127-375e8a945b15 · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Don't Just Say "I don't know"! Self-aligning Large Language Models for Responding to Unknown Questions with Explanations
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6316c1e8-b068-4329-a185-52ee16a778f0 · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs The Llama 3 Herd of Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f93b0e72-4068-4084-80af-7614f987af3b · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a0e703a-f44b-47ef-a833-7a9fc7a41a1a · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Improving alignment of dialogue agents via targeted human judgements
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47694ab7-8434-432f-a732-d32674600a5f · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f792bc3-b2ad-405c-a89c-83342f6715b8 · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Parameter-Efficient Transfer Learning for NLP
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 120b379e-a521-4797-b5ce-c113137b7350 · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de8c7a63-d30e-4450-9c7f-3d5e70867b99 · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Beavertails: Towards improved safety alignment of llm via a human-preference dataset, 2023
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9b5ba2a-e286-4c71-8a01-45ade5ecd83f · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs How can we know when language models know? on the calibration of language models for question answering
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90ad7620-3d9c-4bb5-b0fe-781b2272cba7 · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b679eb1e-a6af-4246-8abd-373faf51d28b · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 198ee191-dbd8-4346-be03-fbbedc16531f · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5edf7f31-a96c-4529-9744-e4cba6d49f36 · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Examining LLMs' Uncertainty Expression Towards Questions Outside Parametric Knowledge
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3039cc7f-61f1-486a-8b95-ec772df7e680 · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Do LLMs Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dad2f2bb-285e-4ae2-aff3-5d44079fb857 · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ca157f1-0b12-41dd-b9fd-f0be2384c948 · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b4a198e-3d01-452e-a5cb-12c5f2dcbb0a · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Rule Based Rewards for Language Model Safety
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd4da059-4b99-4f46-a0a8-f60298839123 · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Crosslingual Generalization through Multitask Finetuning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5efb365-cde2-4cca-b350-82b67dce859e · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs A Comprehensive Overview of Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e7b98eb-8ea3-4eaf-ac09-7ef48f516b0a · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Model spec, 5 2024
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 37522a85-fc18-4ede-a374-fccb088c0ca9 · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Training language models to follow instructions with human feedback
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef368ef7-9cb0-404a-b61b-ee94084db68e · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f07b7287-65fd-4476-a87e-eb2c9188354e · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs I'm Afraid I Can't Do That: Predicting Prompt Refusal in Black-Box Generative Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13a117a8-50ce-4c0a-a611-4c302f667643 · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f80c1d25-cd01-4831-90aa-7295f79daefb · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Proximal Policy Optimization Algorithms
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9983772f-a372-4a91-94bb-3b5a5aac05f5 · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7739f1a3-c589-4f31-b004-e115abb647a1 · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Learning to summarize from human feedback
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95649f1d-3643-4c02-b9b5-e2adfc789d1b · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs All Languages Matter: On the Multilingual Safety of Large Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 371fa624-1bd7-492b-8625-d337a844954c · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7137eb1c-fdac-46b1-8510-1a675799074a · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Smith, Daniel Khashabi, and Hannaneh Hajishirzi
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7596a72-84aa-4feb-bb5e-73f9ff413f89 · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1393e13-104e-4193-aa07-f2339d4e7367 · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Finetuned Language Models Are Zero-Shot Learners
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7386ce8-ace2-423e-a824-a96e41172a14 · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcb95f06-9061-4423-80a8-35073be2d2ba · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd6638d8-c1ca-4a13-8aae-b667bffed841 · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4963402-9065-42cb-9ab0-be3d7193a072 · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs R-Tuning: Instructing Large Language Models to Say `I Don't Know'
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d8b7914-3235-4b3d-bec7-58bcebcbf364 · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Instruction tuning for large language models: A survey, 2024 b
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1611234-8817-437c-9df2-a37b07e1a696 · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Fine-Tuning Language Models from Human Preferences
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49128758-1ec1-4825-9819-c663a818e4b1 · outbound
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6794e616-2bcd-4901-b052-33f277cc3f4d · inbound
Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fded0e28-e891-4b7a-8aff-bed07bfe1418 · inbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.