Pith. sign in

Paper Citation Record · LEDGER

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning

As of 21 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 3 inbound Pith citation observations for arXiv:2506.15606.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.15606 v3

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:59:58.413814Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T21:01:25.549340Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T21:05:04.061254Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact2
  • verified fuzzy7
  • unresolved25
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8a8038f6-cae6-4811-9b96-a2e436ef41a7 · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Refusal in Language Models Is Mediated by a Single Direction

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:55.404563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:55.404563Z digest=sha256:6225bd559d26861da793bf9b5a4912fcb020b9648b624e975eaa4de8e4e371d7

Observation 38b10835-9233-4617-8f6e-0ca7378d3551 · outbound

This paper cites To visualize these points, we project them onto the 17 Published as a conference paper at COLM 2025 same plane and compute their coordinates in the basis {d1, d2}.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning To visualize these points, we project them onto the 17 Published as a conference paper at COLM 2025 same plane and compute their coordinates in the basis {d1, d2}

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:00:00.628892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:59:58.196559Z digest=sha256:98761ee94bcafc91692acd7a19e0252265c4fe7f1f228e22f8ff456aa459925e

Observation 6154c59f-4f80-4602-8d70-56869e9cd7fc · outbound

This paper cites an unresolved cited work.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Unresolved cited work

Reference 3

Resolution
verified exact
raw_fallback, observed 2026-08-06T23:59:58.836084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:59:57.924221Z digest=sha256:5608ab792c35779899d4defce463031ff3915b202726254c9f9898f2ac007866

Observation e09b2e7e-2df4-4134-a3ed-e86398956dbb · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:55.625447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:55.625447Z digest=sha256:41fa69f6a1983f46898e6f9c7bb08cd42f5b51046ee86e63e123e24731346907

Observation 16c81583-c740-4af3-b0e9-800a2178b7d8 · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:55.772459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:55.772459Z digest=sha256:de140aa4207d4923f2526ac25bab72d1f3780140d72b5c6c609f4aabb2873328

Observation ca45d46b-9ab3-4090-b410-2d21b055204e · outbound

This paper cites What is in Your Safe Data? Identifying Benign Data that Breaks Safety.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning What is in Your Safe Data? Identifying Benign Data that Breaks Safety

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:55.836898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:55.836898Z digest=sha256:ebda730fa5d47e000e71f033d472dd015345f93a2fa5d1ccf6c535f63e3d81dd

Observation 13db9a5e-f1e0-43f1-a588-f2004dfe3b97 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:55.892653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:55.892653Z digest=sha256:6c0c67bc20cd6ce3c50a78ac966cb0fe1d38200a7f2757b533c00016db170980

Observation 6b4f4abb-ffdd-4245-ae18-b6c2f80a41da · outbound

This paper cites TextGrad: Advancing Robustness Evaluation in NLP by Gradient-Driven Optimization.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning TextGrad: Advancing Robustness Evaluation in NLP by Gradient-Driven Optimization

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:59:59.643808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:59:55.939806Z digest=sha256:2a14b848adc57e4fa731c6124ab79b1910db86b453fd932c924b0aedbe69f3eb

Observation 90737819-ec19-41b8-9199-5be5c2047610 · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:56.005326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:56.005326Z digest=sha256:59b64281f9c51572095dd0e5edc6a91e49a6eed89abfcb7ae636c0da16ca84cd

Observation fa4eff3d-7ade-413d-a974-1b00d9d22bc5 · outbound

This paper cites Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:56.106685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:56.106685Z digest=sha256:45146924e616e8c15caefcaaad6e2143919d6789f8d6ea367937280f9220c42a

Observation 7b7e2e8a-0254-4e22-bd8b-1d9f857deceb · outbound

This paper cites Mistral 7B.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Mistral 7B

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:56.214328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:56.214328Z digest=sha256:5b62ae54a28d688b63d053d7d51b3afdf3f8063df1be4304b57395859c719c25

Observation f9689af8-987e-4d55-883f-e9c6b49995e2 · outbound

This paper cites The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:56.327320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:56.327320Z digest=sha256:7756f73f03c5ba3954e20dad1b92a8648138b75c609d1f4b1180475b964e237f

Observation 4af70bc5-9a73-4257-b581-701f3804edc8 · outbound

This paper cites Robustifying Safety-Aligned Large Language Models through Clean Data Curation.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Robustifying Safety-Aligned Large Language Models through Clean Data Curation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:56.430016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:56.430016Z digest=sha256:5281087cff21ee22d68a121bfed4d88c0a4e0c8b407b342cb5311264a57b803e

Observation 44c79d5e-b731-4f64-99df-022103abe32d · outbound

This paper cites Use of LLMs for Illicit Purposes: Threats, Prevention Measures, and Vulnerabilities.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Use of LLMs for Illicit Purposes: Threats, Prevention Measures, and Vulnerabilities

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:56.464742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:56.464742Z digest=sha256:42bacbace1efb035208301c6b6928dcb8cb5d2ba00f12c855a7c043c0bcdaa03

Observation 8ef1edc8-d825-473e-9a00-c70a715886c3 · outbound

This paper cites Fine-tuning can cripple your foundation model; preserving features may be the solution.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Fine-tuning can cripple your foundation model; preserving features may be the solution

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:56.539321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:56.539321Z digest=sha256:9b5f4a41dbb2ca9d9399b469d68a5875ed2007be0426244ce0a0111ab19f1b4e

Observation fe1cb680-efc2-46f1-8f95-92bdad269973 · outbound

This paper cites GPT-4 Technical Report.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning GPT-4 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:56.661089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:56.661089Z digest=sha256:39de7bb2834e8384b5d2721f4652ff379ea81dd353839fd396f7998cfc27da85

Observation 2ea3413d-f36c-41d0-865e-7460a6dde878 · outbound

This paper cites Red Teaming Language Models with Language Models.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Red Teaming Language Models with Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:56.806561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:56.806561Z digest=sha256:51eafda6b423912c4154c652664f842484a5fcf9fc8c503b2d89f531ca800c2c

Observation 9220a8f9-c9e5-46fc-8905-9519f83138e7 · outbound

This paper cites Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:56.867587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:56.867587Z digest=sha256:d1f042d724119ef84a0b5a2f6d27df5d52a498e3b31c36c9edec5d715bbfaf2b

Observation cdf18a32-93eb-43b0-8018-ec1683e8b1d8 · outbound

This paper cites Model Extrapolation Expedites Alignment.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Model Extrapolation Expedites Alignment

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:57.259827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:57.259827Z digest=sha256:c2ca7f3c0a1936d2fe49c9fa18809cf6b33d714db922d56f02c91507b29f12f0

Observation e8bbe007-6169-48fa-8d86-030ac4ed2cb4 · outbound

This paper cites Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:57.357277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:57.357277Z digest=sha256:75ea7a85cc1a277b11dcc70b2fac9560d63a28603925e0f1193b0deb3473d303

Observation d7e542b3-4915-4b0f-96d5-d0b1bd9ca74b · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:57.417182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:57.417182Z digest=sha256:b2ef650abf56102a897f1b102a22e0cf35f8ed3dc4ebb64fb6a3a9d15e6925a0

Observation 0661f49f-101a-41e5-9eca-3773a6fe15ef · outbound

This paper cites Similar to RIFT, RoAST can be categorized as a fine-tuning stage defense, and therefore is also not suitable for the scenario described.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Similar to RIFT, RoAST can be categorized as a fine-tuning stage defense, and therefore is also not suitable for the scenario described

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:00:02.053821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:59:57.532604Z digest=sha256:e06b7a5dffce3b1a89a7e37dc6566a4573e4acb7dda9f3c082ff3c05498030c1

Observation 1d0977e8-b7cd-4f02-a3ef-b856ebc2bbdf · outbound

This paper cites Absolutely Obedient.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Absolutely Obedient

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:00:01.511263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:59:57.721181Z digest=sha256:27e4b32e47f6f1ae8598811b23a8180608644827ace50db83600ba5d10f49b57

Observation 5044b574-3f90-44e4-94a8-fc3683c08e2e · outbound

This paper cites Pure Bad.The Pure Bad dataset consists of 100 harmful examples, extracted from the Anthropic Red Teaming Dataset (Ganguli et al., 2022).

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Pure Bad.The Pure Bad dataset consists of 100 harmful examples, extracted from the Anthropic Red Teaming Dataset (Ganguli et al., 2022)

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:00:01.311559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:59:57.801619Z digest=sha256:612b1d7963b09ef0b013048cec12dfaa914cf402c50369a89a129a25afd30670

Observation 57d95072-760a-472a-9eb4-a7975e5adee5 · outbound

This paper cites an unresolved cited work.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:00:01.078057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:59:58.009728Z digest=sha256:9bcd29191963e6657ea1d4dc1552caded959a50c8fc05ac18575f4a44070344c

Observation d30f529e-ab95-46dc-a827-5232a6e9eb52 · outbound

This paper cites 5, and extrapolating further (which we considered broken), following (Lin et al., 2023).

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning 5, and extrapolating further (which we considered broken), following (Lin et al., 2023)

Reference 32

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T00:00:00.840892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:59:58.129917Z digest=sha256:6b43047983d5a2398648781617206f929481ef2dd3093c5321d1876f9483d265

Observation 74c87f94-dbf0-43da-9f53-33c61064ac15 · outbound

This paper cites We include Meta’s usage guidelines1 in our prompt, following the evaluation protocol of Qi et al.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning We include Meta’s usage guidelines1 in our prompt, following the evaluation protocol of Qi et al

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:00:00.388349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:59:58.303211Z digest=sha256:8be337befe29218d752eebc304849afb3574e66f8164d57b7af68d32e3949b64

Observation c9692a5e-b101-4f08-a28c-0d776e7391c4 · outbound

This paper cites No" Response(k=6,α=1.5): “Your task is to complete tasks for people is not recommended.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning No" Response(k=6,α=1.5): “Your task is to complete tasks for people is not recommended

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:00:00.202159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:59:58.413814Z digest=sha256:45d0c528b6237ca88716c45ebdf3c173577edd720a77f62f09dcd38608a16c17

Observation a00007b1-ff82-47a3-a550-603b29d4d067 · outbound

This paper cites an unresolved cited work.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Unresolved cited work

Reference 1024

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:00:01.806107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:59:57.615080Z digest=sha256:0c336d7c4c7278ce43e2c7125eb1e77172bc87ea6837edafe2107153eac7fbdb

Observation b71990e6-68f0-4e20-9c57-5b69d7972036 · outbound

This paper cites Parameter-Efficient Fine-Tuning Methods for Pretrained Language Models: A Critical Review and Assessment.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Parameter-Efficient Fine-Tuning Methods for Pretrained Language Models: A Critical Review and Assessment

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:57.134832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:57.134832Z digest=sha256:4ea1f5662a4eebaf603ca01b20894b8c9bc333d725b09b21b5a4f0baecd5ce38

Observation c3dd5955-0634-471c-906a-271f06b94965 · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:55.560201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:55.560201Z digest=sha256:5fe31d8df3d3dd4d9e07131406fc3dee1802159f5733299e5d083ff0093fe37c

Observation ed87a7e4-93cb-482a-976a-57b3b42fc5a7 · outbound

This paper cites Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:55.503290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:55.503290Z digest=sha256:d10f56aaf45c4bb51080e430018b4429e8499af19de2eed5f221d6ed466f8fcc

Observation 96de15c3-91af-4fc2-8b70-29d0d129cb93 · outbound

This paper cites Xinshuai Dong, Anh Tuan Luu, Min Lin, Shuicheng Yan, and Hanwang Zhang.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Xinshuai Dong, Anh Tuan Luu, Min Lin, Shuicheng Yan, and Hanwang Zhang

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:00:02.261337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:59:55.714750Z digest=sha256:06942b3fefbc8222c83a26b52665b6573e86d8c23d5630ebbe51e3e312dbbb96

Observation f948c392-9160-4e76-bb23-1f01999edd5b · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:55.454095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:55.454095Z digest=sha256:071efbd6916c7fcd02e35fefbdede438d4b382d9ddc9ce6c7287de135c6c3f5f

Observation ce426289-9b95-4344-a6ca-35b38b0d5789 · outbound

This paper cites Crowdsourcing Multiple Choice Science Questions.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Crowdsourcing Multiple Choice Science Questions

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:57.027070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:57.027070Z digest=sha256:047014a1cb62c4b35a6b03d04bfa68ede51edb341a516635821aaa89f7344272

Pith citing papers

Observation 3b7f2f16-529f-46d7-b07b-d0fddd69fd39 · inbound

Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control cites this paper.

Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:21:29.060044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T11:17:03.104902Z digest=sha256:cc85f15660ce5c898a4c2baf1705ccca1b26d9b93b0093aa3c65217160cb6444

Observation 12261f88-9c3a-4ae5-ac59-179bf9f182b1 · inbound

Preventing Safety Drift in Large Language Models via Coupled Weight and Activation Constraints cites this paper.

Preventing Safety Drift in Large Language Models via Coupled Weight and Activation Constraints LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:21:01.852959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-10T16:04:25.851592Z digest=sha256:c666e302382de5b27b39e348a9073268247e34d0599a69e5e23166c887bc0602

Observation 27fd0429-61e2-4e22-b079-1ae9298dbdf0 · inbound

One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries cites this paper.

One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:05:04.062857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T21:01:25.549340Z digest=sha256:003dcdf8d62c5cf05224fa19888c04e38c38821928504eea02ff2f3a48587848