Pith. sign in

Paper Citation Record · LEDGER

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning

As of 9 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 6 inbound Pith citation observations for arXiv:2507.04250.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04250 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:57:44.976131Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T16:03:12.728352Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:29:39.399032Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8d3d4fb5-83c3-42de-a7a6-611ceb89fe6f · outbound

This paper cites Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:43.379859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:43.379859Z digest=sha256:c28231296bc50bda029043e9c284280d4ea33c02bc9a697dc42741f9108c6fb7

Observation ad634ece-212a-4348-a4c2-84b347ce5ad0 · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Refusal in Language Models Is Mediated by a Single Direction

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:43.420384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:43.420384Z digest=sha256:305a2ac9c298d09bd96eec88a309fe6e6836e74ad763a66c721cfda9f122afba

Observation 01d7e062-d6e8-4049-bd09-02442604577d · outbound

This paper cites an unresolved cited work.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:57:46.335341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:57:43.505443Z digest=sha256:311fa0f82c49908aea32c1e57224198536a4e244c9d4f55d31544fabbddecf9f

Observation 6363d2f3-5563-4c94-9ffc-1ed0b9af19d9 · outbound

This paper cites SCANS: Mitigating the Exaggerated Safety for LLMs via Safety-Conscious Activation Steering.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning SCANS: Mitigating the Exaggerated Safety for LLMs via Safety-Conscious Activation Steering

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:57:45.412418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:57:43.549520Z digest=sha256:6148f47739f2da815df738a3dabe6fac3513f91082f9d0299f4e41a9addcc69e

Observation 245fc8a6-5db8-4ccb-8ce6-e13192ca047c · outbound

This paper cites OR-Bench: An Over-Refusal Benchmark for Large Language Models.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning OR-Bench: An Over-Refusal Benchmark for Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:43.629886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:43.629886Z digest=sha256:5972be856408471ba4e9d3f692ebed2018954cea0d2878a77cdd1d5c6732e48d

Observation 7c88bbe9-127b-40f4-9ca0-8bab13ac7b53 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Measuring Massive Multitask Language Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:43.694438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:43.694438Z digest=sha256:fbd50116059c9318ed5d81ad8e28666658628057f09cd3cc4d07d9eca2f55ee1

Observation 63cb46a6-5c23-4a45-ad36-46c5ebbd73f1 · outbound

This paper cites Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:43.738959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:43.738959Z digest=sha256:1a9d05a0e5c144ba17aec7b6102bf14aaaefe78f13a7b3b37861a2d63b8771e0

Observation 3159a8a9-bde1-48bc-8dda-7e11eb0f4802 · outbound

This paper cites TrustLLM: Trustworthiness in Large Language Models.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning TrustLLM: Trustworthiness in Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:43.796231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:43.796231Z digest=sha256:167d5595d5755c7c8ee6c96c65a65a452c6a7dba0dd7d0802929ddc377070105

Observation 7af4eff9-96b4-41c6-a3f4-5c45c0eb443d · outbound

This paper cites GPT-4o System Card.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning GPT-4o System Card

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:43.829835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:43.829835Z digest=sha256:ea082a4d1656d98fd604ffce5241238091fe6d4293605105a8cb03360c005d1a

Observation 34744acc-9d3e-467f-8461-8daf47ba30b2 · outbound

This paper cites Beavertails: Towards improved safety alignment of llm via a human-preference dataset.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Beavertails: Towards improved safety alignment of llm via a human-preference dataset

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:57:46.158196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:57:43.856574Z digest=sha256:fcb23badee3e37c517aba93b47a12e296c514862bfac736959ff05946086295b

Observation c5e47030-9380-439f-a9cc-566b8510edc1 · outbound

This paper cites Crafting papers on machine learning.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Crafting papers on machine learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:43.914249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:43.914249Z digest=sha256:3425f15f5b3a814d7770efc2671441c5f7ac76b4ae2291ef1037635943978c9d

Observation 7cc02672-ba13-4cb4-8174-ab22e8cf3187 · outbound

This paper cites Safety Layers in Aligned Large Language Models: The Key to LLM Security.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Safety Layers in Aligned Large Language Models: The Key to LLM Security

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:43.958573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:43.958573Z digest=sha256:9c59ee546af242533687a1c177e570445dfc4f4d9699254a422ab7f58f48731c

Observation d8f4e40e-329a-4c5d-aa15-0a3e987f0f32 · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:43.979623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:43.979623Z digest=sha256:79d34a9cd86c129b9d49edf52ec7be486db9798048014e9e3e830416cdc2eef9

Observation 6f78b735-0c61-4eb6-91af-9259c8a13a8a · outbound

This paper cites Decoupled Weight Decay Regularization.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Decoupled Weight Decay Regularization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.013076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.013076Z digest=sha256:8da9ccc0212f3ec29a06f96d679fe66b507183a5f0e33c6237f9e73cad941424

Observation d719852f-f4d1-4777-b158-d313ce2019ae · outbound

This paper cites Pointer sentinel mixture models, 2016.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Pointer sentinel mixture models, 2016

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.069618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.069618Z digest=sha256:ae7626da1f139d7ce9ad37ded226394e4bbf6c83e546a55c11ed0a9bc2ad1e07

Observation d2b9753c-600a-4924-b2f5-462fc90e12d9 · outbound

This paper cites Fine-tuning aligned language models compromises safety, even when users do not intend to!, 2023.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Fine-tuning aligned language models compromises safety, even when users do not intend to!, 2023

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:57:45.977389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:57:44.107947Z digest=sha256:36140fd651f2256b8592105584e906584c857b17acd9daeea134ac78980bf2ae

Observation 3a62f8b4-c2ec-41f9-992b-39d757f069fd · outbound

This paper cites Mitigating Exaggerated Safety in Large Language Models.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Mitigating Exaggerated Safety in Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.141968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.141968Z digest=sha256:0751b23c053b5663ad87904270a4f24b22beba4ea93955365c975f10058b9a2e

Observation ced8ca4f-2b94-4e70-9c0a-ba241d4ad76d · outbound

This paper cites XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.190147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.190147Z digest=sha256:7bd0e5f0000548a2c02ae009fcc1770e7dccbf367c0950e58b2d905702198292

Observation c75906bb-a888-4fc8-8a57-a2d0f659bbe1 · outbound

This paper cites an unresolved cited work.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.251439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.251439Z digest=sha256:fc90d72802b4f3691bc90b41f8b18977610747124c9ded0ce1446bee9165f827

Observation b9e105cc-1196-42d2-a5d7-6a6921395e5f · outbound

This paper cites Navigating the OverKill in Large Language Models.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Navigating the OverKill in Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.284799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.284799Z digest=sha256:eb19fcc679c9ec2cf8a71b996db086f9a1b29b8df497b004665cbded55fa0ef6

Observation 0d864902-e309-40c2-9228-8090fcdddf14 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Gemma: Open Models Based on Gemini Research and Technology

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.329662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.329662Z digest=sha256:d29664766e4ce4b1483aa33defb08be5dcc96da38987b77cd46839801357a6f3

Observation 9f41f133-8e1b-42a4-96e3-fe5307c65c16 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.388456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.388456Z digest=sha256:d21edb0cc10327d799d0afd77cef7d1bf0790f197918db124f0ce17339fe0c9b

Observation 113f4a0c-e4e9-4fc3-8c43-9b0ccdfa1a4f · outbound

This paper cites and Hinton, G.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning and Hinton, G

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:57:45.827219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:57:44.415871Z digest=sha256:9da64ca930e78b46b5aa0e6ac18058c52e1302ce4eb7952fd052e19390dcecf9

Observation 003f4759-e069-41a3-b2a5-b6a1e2628d3f · outbound

This paper cites Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector Ablation.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector Ablation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.461260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.461260Z digest=sha256:efb03d9438a0ac8f314be5a975b96fcd8ad5ce5b781ef3e2f4e675a84ee2bf36

Observation a418dee3-7334-447b-b896-45a9f0069e5e · outbound

This paper cites ReFT: Representation Finetuning for Language Models.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning ReFT: Representation Finetuning for Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.494978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.494978Z digest=sha256:e5685134b0fd680d2ced56275b63fb0c4d0ba52771bf3cea6088e78e802f1897

Observation 6063ade9-b576-467b-acb4-3cb2196d7f7e · outbound

This paper cites SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.540195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.540195Z digest=sha256:729e517e8e8c76e0eed4733ce9fc041047c9bccf65bf7569b7f1f3fa9ccb23c0

Observation 59991026-be87-47d5-9ec8-ebff5c8296be · outbound

This paper cites LoFiT: Localized Fine-tuning on LLM Representations.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning LoFiT: Localized Fine-tuning on LLM Representations

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.575405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.575405Z digest=sha256:a32f53355c7be82a1d41186e5ea0848e21ee1ec2726d8de66b1c0e8291a37359

Observation 2da464bd-d208-4966-a920-d39ebb5d9d92 · outbound

This paper cites Scope: Scalable and adaptive evaluation of misguided safety refusal in llms.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Scope: Scalable and adaptive evaluation of misguided safety refusal in llms

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:57:45.642516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:57:44.646800Z digest=sha256:66d8101ab99651cacf0f027c3701a4232ca29cf724c744b26b0e3d70af7b23ee

Observation 53efe874-0171-43da-9d23-0d106ff907c2 · outbound

This paper cites Towards Comprehensive Post Safety Alignment of Large Language Models via Safety Patching.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Towards Comprehensive Post Safety Alignment of Large Language Models via Safety Patching

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.683011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.683011Z digest=sha256:2fb8937e7171150fe2c0746bf30b14a9780dc83371c6853182734fd95e65e7a6

Observation 36cde0a9-52cb-47df-8885-9a0c480d83cd · outbound

This paper cites On Prompt-Driven Safeguarding for Large Language Models.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning On Prompt-Driven Safeguarding for Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.717885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.717885Z digest=sha256:29be13e7b25d93fe9d9fb3b67698065bd4b0f643552affbeeca7012e18b64e45

Observation 656ec8c5-a639-4a8d-8595-528d21bac811 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.785662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.785662Z digest=sha256:d6b2d8c54b94dbdceedc69141975dde62e3cb690a8b5a5e37194461be6ff024b

Observation 5fe881a2-87c0-421d-929c-8c8ad971607a · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Representation Engineering: A Top-Down Approach to AI Transparency

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.836509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.836509Z digest=sha256:3935f5ba59ddb88e39324f18b06f0f05f896e51bc8e3d4f505fedb1de46c5bb1

Observation af12dd0b-7f73-46b2-b5eb-5f4f0ab62693 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.883812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.883812Z digest=sha256:c08c4bfdf0528c57d9fe30a535811a17504a3d83f8f1c0b7b76a632f66c76b0e

Observation e505e4b7-5dbd-4cf7-9021-f49fb8640126 · outbound

This paper cites write newline.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning write newline

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.976131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.976131Z digest=sha256:790aebb9be255999a83de9870f9ed56142de971ff13b6cb053b76383ed03fc4e

Pith citing papers

Observation 5dffa089-944e-407c-9c27-475a836514af · inbound

Characterizing Model-Native Skills cites this paper.

Characterizing Model-Native Skills Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:06:19.523835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T05:42:49.694715Z digest=sha256:761da5f1bdc4059c7a27736c9a264bc943e5d7c0f3f6639f2d0a575eed36dcbe

Observation 27eaeb93-936b-4df8-99e9-37edfaa29333 · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:57:06.385931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T01:03:10.263663Z digest=sha256:3688d3078d99fd7589d5e1a14817815d5164cfb7ea0fb5de05e5cb14692b29a8

Observation 155ce5e3-1110-4b05-9f61-83d5fcff2aee · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:12:58.856739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T21:12:06.989077Z digest=sha256:af9c4aef090696ac3e8862f60f3a82f7f3cbd43819dcfa50d26a1d2da4e43ddb

Observation a1ba78f2-a143-49b2-b127-53827db5d500 · inbound

Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs cites this paper.

Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:04:52.603163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T16:03:12.728352Z digest=sha256:7608bdff2e59a6d2afcf4dbb209e2605bca2ac4e0e5620ca4fe6bb6fe442312a

Observation 04bef1af-3585-4424-8c72-813677dafedb · inbound

Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics cites this paper.

Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.873093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T21:55:48.561400Z digest=sha256:1c954f62f68915217bdab8851617277cc68fdafb84f7916f54b1254f6c0e4c5c

Observation 7a11ef31-8b4d-48ac-a05c-09782b7accf7 · inbound

AOR-Bench: Do Large Audio Language Models Over-Refuse Pseudo-Harmful Queries? cites this paper.

AOR-Bench: Do Large Audio Language Models Over-Refuse Pseudo-Harmful Queries? Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T07:29:39.400611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T13:22:12.541923Z digest=sha256:dcba632b7339acbd1d47a642e74d65060197431f7bc0a9219b8135202bb5dd5b