Pith. sign in

Paper Citation Record · LEDGER

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning

As of 17 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 6 inbound Pith citation observations for arXiv:2507.04250.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04250 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:57:44.976131Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T16:03:12.728352Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:29:39.399032Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8d3d4fb5-83c3-42de-a7a6-611ceb89fe6f · outbound

This paper cites Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:43.379859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:43.379859Z digest=sha256:2c83d319efa9b74b1710140039876cc98a7446125646207371d28d497a25ed89

Observation ad634ece-212a-4348-a4c2-84b347ce5ad0 · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Refusal in Language Models Is Mediated by a Single Direction

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:43.420384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:43.420384Z digest=sha256:445079cf92b2551ed2b4a0ab4765b2c12b5c7e68c0d06257a7d4fc8e86b12aea

Observation 01d7e062-d6e8-4049-bd09-02442604577d · outbound

This paper cites an unresolved cited work.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:57:46.335341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:57:43.505443Z digest=sha256:192020de171ca6dca33bf5ce736921645b85ebe38a6762ae44b690233e5a56eb

Observation 6363d2f3-5563-4c94-9ffc-1ed0b9af19d9 · outbound

This paper cites SCANS: Mitigating the Exaggerated Safety for LLMs via Safety-Conscious Activation Steering.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning SCANS: Mitigating the Exaggerated Safety for LLMs via Safety-Conscious Activation Steering

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:57:45.412418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:57:43.549520Z digest=sha256:49ce7b55774479b1225fa4de26f0dd486ab4ef51ab912c11524060c43b9892a9

Observation 245fc8a6-5db8-4ccb-8ce6-e13192ca047c · outbound

This paper cites OR-Bench: An Over-Refusal Benchmark for Large Language Models.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning OR-Bench: An Over-Refusal Benchmark for Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:43.629886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:43.629886Z digest=sha256:35d8811a81e81d4b9a26b01d4985741546153140b693b90c68a5e26b5afaa63c

Observation 7c88bbe9-127b-40f4-9ca0-8bab13ac7b53 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Measuring Massive Multitask Language Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:43.694438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:43.694438Z digest=sha256:6254428f15a701aee0aaffbec226dc857393ba155b4568a4236a718ce99c913b

Observation 63cb46a6-5c23-4a45-ad36-46c5ebbd73f1 · outbound

This paper cites Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:43.738959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:43.738959Z digest=sha256:e44d376421e130195d6bdfcdc33820ff9179f7cbdef67647e2fbdd70d64dd8b9

Observation 3159a8a9-bde1-48bc-8dda-7e11eb0f4802 · outbound

This paper cites TrustLLM: Trustworthiness in Large Language Models.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning TrustLLM: Trustworthiness in Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:43.796231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:43.796231Z digest=sha256:93eb1b7d9a5fcc9d95300730a08b7fe384ad20256030bea34f91ef1d0a80d1de

Observation 7af4eff9-96b4-41c6-a3f4-5c45c0eb443d · outbound

This paper cites GPT-4o System Card.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning GPT-4o System Card

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:43.829835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:43.829835Z digest=sha256:5b8823c8c6e569b4f04745296422af69159a494d57cf78e589caffb14bbee0a2

Observation 34744acc-9d3e-467f-8461-8daf47ba30b2 · outbound

This paper cites Beavertails: Towards improved safety alignment of llm via a human-preference dataset.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Beavertails: Towards improved safety alignment of llm via a human-preference dataset

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:57:46.158196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:57:43.856574Z digest=sha256:c121b0bce836d7f882f675b5c4bb0a80082d00c19832e1e232ee9d4b93de3ad7

Observation c5e47030-9380-439f-a9cc-566b8510edc1 · outbound

This paper cites Crafting papers on machine learning.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Crafting papers on machine learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:43.914249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:43.914249Z digest=sha256:3b1149c7ede65866abb974b4386a2e1ae6f21f51b62184c8b084dc5a66c16821

Observation 7cc02672-ba13-4cb4-8174-ab22e8cf3187 · outbound

This paper cites Safety Layers in Aligned Large Language Models: The Key to LLM Security.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Safety Layers in Aligned Large Language Models: The Key to LLM Security

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:43.958573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:43.958573Z digest=sha256:5147409d05eb8d156bcd6a3563a10cd2fbe98a51766f50c769c2e77d477cc264

Observation d8f4e40e-329a-4c5d-aa15-0a3e987f0f32 · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:43.979623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:43.979623Z digest=sha256:c90f281041e9542274ce3ac990c3c5494818ab8981a021b69fa2d2fd9f918a13

Observation 6f78b735-0c61-4eb6-91af-9259c8a13a8a · outbound

This paper cites Decoupled Weight Decay Regularization.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Decoupled Weight Decay Regularization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.013076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.013076Z digest=sha256:d14908b07ca2e3e18a30ae7dc0adb59732f5d82e16fe1ee5433397ee04a3977d

Observation d719852f-f4d1-4777-b158-d313ce2019ae · outbound

This paper cites Pointer sentinel mixture models, 2016.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Pointer sentinel mixture models, 2016

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.069618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.069618Z digest=sha256:f1367c229867df16d1b89cae37f2bd318401b2aa5549b075bd1ddc876353e936

Observation d2b9753c-600a-4924-b2f5-462fc90e12d9 · outbound

This paper cites Fine-tuning aligned language models compromises safety, even when users do not intend to!, 2023.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Fine-tuning aligned language models compromises safety, even when users do not intend to!, 2023

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:57:45.977389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:57:44.107947Z digest=sha256:45b67a1614d4a79e82ffec180faaace0d48d84ebdb963cefef12b50fcb015405

Observation 3a62f8b4-c2ec-41f9-992b-39d757f069fd · outbound

This paper cites Mitigating Exaggerated Safety in Large Language Models.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Mitigating Exaggerated Safety in Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.141968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.141968Z digest=sha256:adb038a70b518fb33f558c7a7f20fcce81a4d14491dcef004979375a3fcb1aa4

Observation ced8ca4f-2b94-4e70-9c0a-ba241d4ad76d · outbound

This paper cites XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.190147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.190147Z digest=sha256:256f08e77964c1c783fb6a435240e9ef9eed6ebcd19f74d5f569b3ed73efa78f

Observation c75906bb-a888-4fc8-8a57-a2d0f659bbe1 · outbound

This paper cites an unresolved cited work.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.251439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.251439Z digest=sha256:fe791b9d2ad02e41b31faa677f7667fca9d9c7a54c00ee609a3e19d2a0909962

Observation b9e105cc-1196-42d2-a5d7-6a6921395e5f · outbound

This paper cites Navigating the OverKill in Large Language Models.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Navigating the OverKill in Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.284799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.284799Z digest=sha256:d5800244ae9bd149098cb413021704d889f19f031a00fc7f85fb7b05ba01764b

Observation 0d864902-e309-40c2-9228-8090fcdddf14 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Gemma: Open Models Based on Gemini Research and Technology

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.329662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.329662Z digest=sha256:1154652af6efc6c76305b9f4de2e68a2bc25f53366a0bfb8093ed52c36b669fa

Observation 9f41f133-8e1b-42a4-96e3-fe5307c65c16 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.388456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.388456Z digest=sha256:a115d92d13d3236c79f4c2b07595bb077d3b6a4edff218f292f3a5a41af1c390

Observation 113f4a0c-e4e9-4fc3-8c43-9b0ccdfa1a4f · outbound

This paper cites and Hinton, G.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning and Hinton, G

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:57:45.827219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:57:44.415871Z digest=sha256:8fd442746bbcc8d22cf1907921f33dd52f87a5af76ac9204b54234aab05c8f82

Observation 003f4759-e069-41a3-b2a5-b6a1e2628d3f · outbound

This paper cites Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector Ablation.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector Ablation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.461260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.461260Z digest=sha256:16d6551a1881415f183f220e8cd2bf5892afc7017270e99350dfab51d5af55b3

Observation a418dee3-7334-447b-b896-45a9f0069e5e · outbound

This paper cites ReFT: Representation Finetuning for Language Models.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning ReFT: Representation Finetuning for Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.494978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.494978Z digest=sha256:cc10915ee7b3456e6a77a2c1f11f7bf0e14b47f6e381f7eb5200a5628fa5a198

Observation 6063ade9-b576-467b-acb4-3cb2196d7f7e · outbound

This paper cites SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.540195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.540195Z digest=sha256:cb6c16300a9b321cc2071e9711fba2ad8418db13ba62890dd2287a771f025bf4

Observation 59991026-be87-47d5-9ec8-ebff5c8296be · outbound

This paper cites LoFiT: Localized Fine-tuning on LLM Representations.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning LoFiT: Localized Fine-tuning on LLM Representations

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.575405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.575405Z digest=sha256:e04a3380b978830a897e2f5618f1405e113074f06564c384832ed3eeaefc0d0f

Observation 2da464bd-d208-4966-a920-d39ebb5d9d92 · outbound

This paper cites Scope: Scalable and adaptive evaluation of misguided safety refusal in llms.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Scope: Scalable and adaptive evaluation of misguided safety refusal in llms

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:57:45.642516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:57:44.646800Z digest=sha256:5b3dbd7dad302eb6e5251e0b730510d57f75ce94a5bd5207a92ab5b6937ba5c2

Observation 53efe874-0171-43da-9d23-0d106ff907c2 · outbound

This paper cites Towards Comprehensive Post Safety Alignment of Large Language Models via Safety Patching.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Towards Comprehensive Post Safety Alignment of Large Language Models via Safety Patching

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.683011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.683011Z digest=sha256:406bdc0197cbe8d8713fab5dd1dda80346413631ff66fefce009be2e0eb60599

Observation 36cde0a9-52cb-47df-8885-9a0c480d83cd · outbound

This paper cites On Prompt-Driven Safeguarding for Large Language Models.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning On Prompt-Driven Safeguarding for Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.717885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.717885Z digest=sha256:7d8c44dc53f9f38b9b44d20d315b7cba8d2827590acbe468abc98c182844443f

Observation 656ec8c5-a639-4a8d-8595-528d21bac811 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.785662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.785662Z digest=sha256:e2a80e121230fb3fa19947673286dbc29e6e07b7dd3f9d0ddd4027683a026485

Observation 5fe881a2-87c0-421d-929c-8c8ad971607a · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Representation Engineering: A Top-Down Approach to AI Transparency

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.836509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.836509Z digest=sha256:744e2a3159f5118496476957a06bbc74341a47878c70f92a6c17fad4f4ac6bce

Observation af12dd0b-7f73-46b2-b5eb-5f4f0ab62693 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.883812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.883812Z digest=sha256:b8bded66788518fea9994711453478e5e417ef7dcdb021d0e6fc11638a6e5f84

Observation e505e4b7-5dbd-4cf7-9021-f49fb8640126 · outbound

This paper cites write newline.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning write newline

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.976131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.976131Z digest=sha256:cb2ea3207ab6c77da1750c5f169f0b1fdc0b3ee3ae4030838ed25264972e5b47

Pith citing papers

Observation 5dffa089-944e-407c-9c27-475a836514af · inbound

Characterizing Model-Native Skills cites this paper.

Characterizing Model-Native Skills Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:06:19.523835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-10T05:42:49.694715Z digest=sha256:9fc72c2ca9b42ca227b5c5cb0dcaa0d4303080828ea7d94d7b5b00f79e69edd8

Observation 27eaeb93-936b-4df8-99e9-37edfaa29333 · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:57:06.385931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T01:03:10.263663Z digest=sha256:9b0a58e57978f859373ab63710a018f9567dfa9328761926b1d97e628bf2b88d

Observation 155ce5e3-1110-4b05-9f61-83d5fcff2aee · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:12:58.856739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-14T21:12:06.989077Z digest=sha256:3670d35a1049e7ac4c74bfb9da57d7622ebba40fa4e4a6a0940ca2d05d5af73e

Observation a1ba78f2-a143-49b2-b127-53827db5d500 · inbound

Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs cites this paper.

Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:04:52.603163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T16:03:12.728352Z digest=sha256:ebde874ff888d464ba6ac3391c2b264c22e925657772259f56cc606e9e1231ea

Observation 04bef1af-3585-4424-8c72-813677dafedb · inbound

Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics cites this paper.

Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.873093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T21:55:48.561400Z digest=sha256:755bbf286d3592a4a2cb82d816e523d552cb3ea26ee829604356e96b0b3f76c7

Observation 7a11ef31-8b4d-48ac-a05c-09782b7accf7 · inbound

AOR-Bench: Do Large Audio Language Models Over-Refuse Pseudo-Harmful Queries? cites this paper.

AOR-Bench: Do Large Audio Language Models Over-Refuse Pseudo-Harmful Queries? Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T07:29:39.400611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-26T13:22:12.541923Z digest=sha256:017ed9b9b507202023ce377fff6e2fb0ed144d0042dc67f3b5e7fabc438678a2