Pith. sign in

Paper Citation Record · LEDGER

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs

As of 13 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 2 inbound Pith citation observations for arXiv:2412.16974.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.16974 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:58:47.845786Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-22T01:19:00.268857Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T01:20:51.944045Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 415bb368-ae73-4b0c-9f91-12d8b4ee574d · outbound

This paper cites Can NLP Models 'Identify', 'Distinguish', and 'Justify' Questions that Don't have a Definitive Answer?.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Can NLP Models 'Identify', 'Distinguish', and 'Justify' Questions that Don't have a Definitive Answer?

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-11T05:58:48.517626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:58:46.685993Z digest=sha256:ad1ba72006e209c5eafe95e7289781ffb33d08a6bbee03dc9aca091ad23c1007

Observation 2f723bcc-5e0b-4061-bb80-8e5ad65cae2c · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs A General Language Assistant as a Laboratory for Alignment

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:46.718112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:46.718112Z digest=sha256:31376c6fbd36d660184847e85730f64462b19408196fef21d18e2151fafc18a3

Observation 7c321cdb-3502-4726-8c7d-e9a1d2d139a8 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:46.724041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:46.724041Z digest=sha256:08339a9dd26304684b1b33a94f3d9c1bd88554350ff6074026f83888c64fad06

Observation a5ebfa30-5335-4250-9fb9-e7f65f673208 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Constitutional AI: Harmlessness from AI Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:46.728472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:46.728472Z digest=sha256:1b5b6cf3736c7b19dfdf3eec2762f922af837bd05fb68ea18d185b05b869d3ad

Observation b5c01f6a-6782-41b4-b126-d7dafd5ede13 · outbound

This paper cites A unified taxonomy of harmful content.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs A unified taxonomy of harmful content

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:46.732208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:46.732208Z digest=sha256:3313ebb4ee99ea56649f3cfdbe0f9a2fe23d18e57263286aa73a4e388704b8dd

Observation 1ee5359f-6805-4140-8d39-29666863ae9e · outbound

This paper cites Information hazards: A typology of potential harms from knowledge.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Information hazards: A typology of potential harms from knowledge

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:58:48.609003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:58:46.735933Z digest=sha256:119a0352d1254658ae6b264710b169c630770ebf72e947bd6db78bd2b8c1f177

Observation e3deb010-0cba-44ec-b3db-e02d6e300885 · outbound

This paper cites Deep Reinforcement Learning from Human Preferences.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Deep Reinforcement Learning from Human Preferences

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:58:48.543432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:58:46.740489Z digest=sha256:09f40bb39fa34a4effa1a8468338b9675ce3216f580b3d816ca1d6ff15b9bc13

Observation c16f6304-151f-48cc-a1eb-357e90be164f · outbound

This paper cites JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:46.743754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:46.743754Z digest=sha256:389c0be704a85505b5d1c49c0b08a964f6f9c6876e649bc3b08268776245caff

Observation 3e096ccb-11f5-49cb-8127-375e8a945b15 · outbound

This paper cites Don't Just Say "I don't know"! Self-aligning Large Language Models for Responding to Unknown Questions with Explanations.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Don't Just Say "I don't know"! Self-aligning Large Language Models for Responding to Unknown Questions with Explanations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:46.747261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:46.747261Z digest=sha256:4e1c3bc7a1d2c74e34077d57f4ec9d9016489e8aadda34c06bcb966f68ab403b

Observation 6316c1e8-b068-4329-a185-52ee16a778f0 · outbound

This paper cites The Llama 3 Herd of Models.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:46.751649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:46.751649Z digest=sha256:5eac1b7bdd4fe3f6eadff1e80da00d0c7f484abc8261230903d9df89dc3b054a

Observation f93b0e72-4068-4084-80af-7614f987af3b · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:46.832039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:46.832039Z digest=sha256:e5f4bbb6e7cd0ffbba11780b284e0be6a0d858bfb52b00a842b0f77129553526

Observation 8a0e703a-f44b-47ef-a833-7a9fc7a41a1a · outbound

This paper cites Improving alignment of dialogue agents via targeted human judgements.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Improving alignment of dialogue agents via targeted human judgements

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:46.954944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:46.954944Z digest=sha256:72b18282a8f63321e0f3638279d2b75db38a4adcd80c9ef456014eb5fa0b40e0

Observation 47694ab7-8434-432f-a732-d32674600a5f · outbound

This paper cites ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.006453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.006453Z digest=sha256:ae4c0053429b912cd0afcf60caa43c4f462f6b14a9bfc4b7f01d2fe23ce98d8d

Observation 3f792bc3-b2ad-405c-a89c-83342f6715b8 · outbound

This paper cites Parameter-Efficient Transfer Learning for NLP.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Parameter-Efficient Transfer Learning for NLP

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.010113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.010113Z digest=sha256:425af2465097e5cff5b66b077c1735ada7afaa78a3e59793daae567af300f9e3

Observation 120b379e-a521-4797-b5ce-c113137b7350 · outbound

This paper cites Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.014005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.014005Z digest=sha256:06013d96107ae7ca90db4e4d1e0ccecc8701ebf9e97a363239fb62a488b62195

Observation de8c7a63-d30e-4450-9c7f-3d5e70867b99 · outbound

This paper cites Beavertails: Towards improved safety alignment of llm via a human-preference dataset, 2023.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Beavertails: Towards improved safety alignment of llm via a human-preference dataset, 2023

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.018964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.018964Z digest=sha256:ee9facc45a5452274e8c0e0654b2a0593ec60c7c5dd19768b2bcff9558341817

Observation d9b5ba2a-e286-4c71-8a01-45ade5ecd83f · outbound

This paper cites How can we know when language models know? on the calibration of language models for question answering.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs How can we know when language models know? on the calibration of language models for question answering

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.022320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.022320Z digest=sha256:2190d6b460b84086efedb6de5756a4a02cb2c44c3a11e537000b5b8569e7aa40

Observation 90ad7620-3d9c-4bb5-b0fe-781b2272cba7 · outbound

This paper cites NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.026583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.026583Z digest=sha256:6b13214ac3798720463cc5a8d7e4e6cb6bda12fa6c03a8e318f465a600604896

Observation b679eb1e-a6af-4246-8abd-373faf51d28b · outbound

This paper cites SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.068771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.068771Z digest=sha256:7035f7fdccb32c6337b82fcddb9b4ac4b75a1811a1cdb32097710913320d2fc5

Observation 198ee191-dbd8-4346-be03-fbbedc16531f · outbound

This paper cites ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.152264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.152264Z digest=sha256:208b67016fef7346d7bdebb83e5c31e9316381c625ce0fef48a5aded295e1c3d

Observation 5edf7f31-a96c-4529-9744-e4cba6d49f36 · outbound

This paper cites Examining LLMs' Uncertainty Expression Towards Questions Outside Parametric Knowledge.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Examining LLMs' Uncertainty Expression Towards Questions Outside Parametric Knowledge

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.276350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.276350Z digest=sha256:dbc0574940a928728e7e48fc92080930674fcc421dd70630f9acc4a67637cb75

Observation 3039cc7f-61f1-486a-8b95-ec772df7e680 · outbound

This paper cites Do LLMs Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Do LLMs Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.355053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.355053Z digest=sha256:36b62b2ff21f3ca1ada6091438e3e1ef8d498956a981ab8bc615609634309486

Observation dad2f2bb-285e-4ae2-aff3-5d44079fb857 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.359674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.359674Z digest=sha256:6f0830fba10ff6a24388e97c792862962235c42b8237b5435fb0ced8abdbf6a1

Observation 1ca157f1-0b12-41dd-b9fd-f0be2384c948 · outbound

This paper cites UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.363796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.363796Z digest=sha256:37dfd8645866faec4dfbf25b62611529794c73a42409b97b4c0279bf0bc5e500

Observation 4b4a198e-3d01-452e-a5cb-12c5f2dcbb0a · outbound

This paper cites Rule Based Rewards for Language Model Safety.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Rule Based Rewards for Language Model Safety

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.367853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.367853Z digest=sha256:7fb74f055ac3a87064172a3f88626ae09e6204b5a94906caaae2a7424a6075c2

Observation dd4da059-4b99-4f46-a0a8-f60298839123 · outbound

This paper cites Crosslingual Generalization through Multitask Finetuning.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Crosslingual Generalization through Multitask Finetuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.372522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.372522Z digest=sha256:8c22df803cb148b68d50874f552cba0cf0895f248a86320dcfbad5e4eb220ee9

Observation f5efb365-cde2-4cca-b350-82b67dce859e · outbound

This paper cites A Comprehensive Overview of Large Language Models.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs A Comprehensive Overview of Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.375962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.375962Z digest=sha256:5e602c4c67fa2a0d69f3e6c1207fdfa7b393c193aee25fe1db73c831a3629758

Observation 5e7b98eb-8ea3-4eaf-ac09-7ef48f516b0a · outbound

This paper cites Model spec, 5 2024.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Model spec, 5 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:58:48.530851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:58:47.379944Z digest=sha256:68e349fc94d5e8275d129f43309fb03b0145e84eebe3b0d03f6c3e3aa58c6f7e

Observation 37522a85-fc18-4ede-a374-fccb088c0ca9 · outbound

This paper cites Training language models to follow instructions with human feedback.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Training language models to follow instructions with human feedback

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.383234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.383234Z digest=sha256:10b2a3c8d4d9da99a9e2801a7ffc8dc2fd25abe10b0faad4bcba4f4d26b3bc37

Observation ef368ef7-9cb0-404a-b61b-ee94084db68e · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.450308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.450308Z digest=sha256:6ba8f26f041500583072303c1420b810ac6721ce935cec4fa567cbb0b81757e9

Observation f07b7287-65fd-4476-a87e-eb2c9188354e · outbound

This paper cites I'm Afraid I Can't Do That: Predicting Prompt Refusal in Black-Box Generative Language Models.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs I'm Afraid I Can't Do That: Predicting Prompt Refusal in Black-Box Generative Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.577628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.577628Z digest=sha256:c276d66786445982c3e854e673a418756b99e3bf0efe6699706e89b07f6b4337

Observation 13a117a8-50ce-4c0a-a611-4c302f667643 · outbound

This paper cites XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.656818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.656818Z digest=sha256:c1fe583f940d63ba6fa5e2f88d9b379312bfd450d834d55e3706704e1bfac8fd

Observation f80c1d25-cd01-4831-90aa-7295f79daefb · outbound

This paper cites Proximal Policy Optimization Algorithms.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Proximal Policy Optimization Algorithms

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.660911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.660911Z digest=sha256:12cd28a9ba7d19266908bd97aefd8ac5d8ab2a71e83acd01832336e1ed8abc82

Observation 9983772f-a372-4a91-94bb-3b5a5aac05f5 · outbound

This paper cites "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.664382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.664382Z digest=sha256:799d4b088de700991081a0daa549a2ef39d8b5c541bec802e0d4844305b3d75e

Observation 7739f1a3-c589-4f31-b004-e115abb647a1 · outbound

This paper cites Learning to summarize from human feedback.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Learning to summarize from human feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.668432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.668432Z digest=sha256:e49102c5e5788ebf280463ca288092f8183aa5beee0aafdb377eef224691f3a0

Observation 95649f1d-3643-4c02-b9b5-e2adfc789d1b · outbound

This paper cites All Languages Matter: On the Multilingual Safety of Large Language Models.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs All Languages Matter: On the Multilingual Safety of Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.672211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.672211Z digest=sha256:1229a6228891f71781a2c76a04df1ed02b192dfd7819f9b9af601ec60dba67c9

Observation 371fa624-1bd7-492b-8625-d337a844954c · outbound

This paper cites Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.676389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.676389Z digest=sha256:6de8d7b2f268e1dbed3bec2b8f721dc428e349280351d8cd30360568a7de547c

Observation 7137eb1c-fdac-46b1-8510-1a675799074a · outbound

This paper cites Smith, Daniel Khashabi, and Hannaneh Hajishirzi.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Smith, Daniel Khashabi, and Hannaneh Hajishirzi

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.680455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.680455Z digest=sha256:41089c97e5092808c84d0636997bf905f52b9f82c06c145ac9ecfd9b8fe6adbc

Observation f7596a72-84aa-4feb-bb5e-73f9ff413f89 · outbound

This paper cites Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.683998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.683998Z digest=sha256:ad2013028cac90e5ba99bf9e41e2ef975b1384e1c8aa6068dadefebeed543005

Observation b1393e13-104e-4193-aa07-f2339d4e7367 · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Finetuned Language Models Are Zero-Shot Learners

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.687891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.687891Z digest=sha256:c2c600c41835c239801163e21cc2c92517232bac0f5522204486df56c77a6112

Observation d7386ce8-ace2-423e-a824-a96e41172a14 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.691627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.691627Z digest=sha256:a5292409b6ff16cf1d1fa2f0f79cdfbf242ec2b3d5ac4b611360f2c1f200150b

Observation bcb95f06-9061-4423-80a8-35073be2d2ba · outbound

This paper cites SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.761384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.761384Z digest=sha256:ce4e0ffb0efe5c331e72345de9434cfc7aa2c03efccf8813636263a085429888

Observation dd6638d8-c1ca-4a13-8aae-b667bffed841 · outbound

This paper cites Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.829233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.829233Z digest=sha256:de711fe7a9256a1250581e0db487b50bb48650afaddca6c8095fde220a9bb1f2

Observation b4963402-9065-42cb-9ab0-be3d7193a072 · outbound

This paper cites R-Tuning: Instructing Large Language Models to Say `I Don't Know'.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs R-Tuning: Instructing Large Language Models to Say `I Don't Know'

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.833858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.833858Z digest=sha256:3e6923859733fc237803c252579e3d21c074df43a591955412ebe97980c16710

Observation 2d8b7914-3235-4b3d-bec7-58bcebcbf364 · outbound

This paper cites Instruction tuning for large language models: A survey, 2024 b.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Instruction tuning for large language models: A survey, 2024 b

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.837929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.837929Z digest=sha256:c7b793a94917726ef28e5d1c541dc94bd5bdc2dcf2de8dedbb2681866aa599fe

Observation f1611234-8817-437c-9df2-a37b07e1a696 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Fine-Tuning Language Models from Human Preferences

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.842057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.842057Z digest=sha256:90f23904160e7229c6e64fddb82ed189483c157b1cae480b7fdf0faac32d72ca

Observation 49128758-1ec1-4825-9819-c663a818e4b1 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.845786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.845786Z digest=sha256:a4620e9b1a2db0f9d1864c88fef0bbf899629060f5cf963b657c5281b724b257

Pith citing papers

Observation 6794e616-2bcd-4901-b052-33f277cc3f4d · inbound

Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules cites this paper.

Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:28:09.935903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-13T19:24:54.381722Z digest=sha256:c5f293aeacb5bc9909376c061a1c277ed22e54525d92501538f7ecd242d9b87b

Observation fded0e28-e891-4b7a-8aff-bed07bfe1418 · inbound

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts cites this paper.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:20:51.947425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:a794122134385a9f43278d08c2c9e97da941a8f278bb9af8e3d1d2c0eb3b9af4