Pith. sign in

Paper Citation Record · LEDGER

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning

As of 8 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 3 inbound Pith citation observations for arXiv:2506.15606.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.15606 v3

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:59:58.413814Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T21:01:25.549340Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T21:05:04.061254Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact2
  • verified fuzzy7
  • unresolved25
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8a8038f6-cae6-4811-9b96-a2e436ef41a7 · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Refusal in Language Models Is Mediated by a Single Direction

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:55.404563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:55.404563Z digest=sha256:1f78815040e5c6eda22e1c7291bcb1698579d80df0e019a1b7e6f5915d90ceae

Observation 38b10835-9233-4617-8f6e-0ca7378d3551 · outbound

This paper cites To visualize these points, we project them onto the 17 Published as a conference paper at COLM 2025 same plane and compute their coordinates in the basis {d1, d2}.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning To visualize these points, we project them onto the 17 Published as a conference paper at COLM 2025 same plane and compute their coordinates in the basis {d1, d2}

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:00:00.628892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:59:58.196559Z digest=sha256:2ac50fcad7e3545871afcb0d8069e4c622d541ba87fef174e7d5124ef2f4e496

Observation 6154c59f-4f80-4602-8d70-56869e9cd7fc · outbound

This paper cites an unresolved cited work.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Unresolved cited work

Reference 3

Resolution
verified exact
raw_fallback, observed 2026-08-06T23:59:58.836084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:59:57.924221Z digest=sha256:45dfbb67f4a972b2a2936f34c12bc1dbf53a6ba87db2c8fe760dbff2740360cd

Observation e09b2e7e-2df4-4134-a3ed-e86398956dbb · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:55.625447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:55.625447Z digest=sha256:ba2e2eedc3e53a10d1d0d674904f57aaa6502ccf12e3bca997d301688acb0d56

Observation 16c81583-c740-4af3-b0e9-800a2178b7d8 · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:55.772459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:55.772459Z digest=sha256:2421c391be7bffdc70a25d8229397f1dce8519fb434947c431b83d91ac4f9c7c

Observation ca45d46b-9ab3-4090-b410-2d21b055204e · outbound

This paper cites What is in Your Safe Data? Identifying Benign Data that Breaks Safety.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning What is in Your Safe Data? Identifying Benign Data that Breaks Safety

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:55.836898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:55.836898Z digest=sha256:48f4c7412683750bab21680c02abe8a4d271d362943635df832919ac2603336c

Observation 13db9a5e-f1e0-43f1-a588-f2004dfe3b97 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:55.892653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:55.892653Z digest=sha256:e8288178a3f7fdb41c5c94eed3f815361a9c384774df90aa85e63640b6d027fa

Observation 6b4f4abb-ffdd-4245-ae18-b6c2f80a41da · outbound

This paper cites TextGrad: Advancing Robustness Evaluation in NLP by Gradient-Driven Optimization.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning TextGrad: Advancing Robustness Evaluation in NLP by Gradient-Driven Optimization

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:59:59.643808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:59:55.939806Z digest=sha256:a77f3009dc6d5140834de505a528996a7af9cffab57bae265bd97a67860d80e6

Observation 90737819-ec19-41b8-9199-5be5c2047610 · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:56.005326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:56.005326Z digest=sha256:903eed8ab1fb8d6260ee7d32790a594bd0eb4ead9dc56ba02122770e8cbdec29

Observation fa4eff3d-7ade-413d-a974-1b00d9d22bc5 · outbound

This paper cites Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:56.106685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:56.106685Z digest=sha256:1db8cc50653cb6942ef26ade7fce7ee40921b76b197a400680c80eae8d47cb68

Observation 7b7e2e8a-0254-4e22-bd8b-1d9f857deceb · outbound

This paper cites Mistral 7B.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Mistral 7B

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:56.214328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:56.214328Z digest=sha256:9470495de7ba6598e39a85c43f595c55875394f135a5e3016fb7c1eea0670335

Observation f9689af8-987e-4d55-883f-e9c6b49995e2 · outbound

This paper cites The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:56.327320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:56.327320Z digest=sha256:7d2e7062ed047819cf47cb60da290b8112164a562d81fed8bd06c1c02986a6d1

Observation 4af70bc5-9a73-4257-b581-701f3804edc8 · outbound

This paper cites Robustifying Safety-Aligned Large Language Models through Clean Data Curation.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Robustifying Safety-Aligned Large Language Models through Clean Data Curation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:56.430016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:56.430016Z digest=sha256:4b7a5d4691a04e6e8e01db34848f01031ae499b37421d53fb2fe698683cf6283

Observation 44c79d5e-b731-4f64-99df-022103abe32d · outbound

This paper cites Use of LLMs for Illicit Purposes: Threats, Prevention Measures, and Vulnerabilities.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Use of LLMs for Illicit Purposes: Threats, Prevention Measures, and Vulnerabilities

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:56.464742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:56.464742Z digest=sha256:e3f0ca39aecc5e5c429f8bca49fc2b729d5062db40c4d482912e12990002ec35

Observation 8ef1edc8-d825-473e-9a00-c70a715886c3 · outbound

This paper cites Fine-tuning can cripple your foundation model; preserving features may be the solution.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Fine-tuning can cripple your foundation model; preserving features may be the solution

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:56.539321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:56.539321Z digest=sha256:879b629192586bf92f01cb35a8c4f9cefd84a4c92c1685a909317e4f3f4f6151

Observation fe1cb680-efc2-46f1-8f95-92bdad269973 · outbound

This paper cites GPT-4 Technical Report.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning GPT-4 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:56.661089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:56.661089Z digest=sha256:1e0dabe656d4aeac9ad46967260cde744c28d54621aede7d159f55c87ecd9e63

Observation 2ea3413d-f36c-41d0-865e-7460a6dde878 · outbound

This paper cites Red Teaming Language Models with Language Models.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Red Teaming Language Models with Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:56.806561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:56.806561Z digest=sha256:f625b53691774bb393bda33a19c29eb684535f62b83680cfd566cbbca1d01a5e

Observation 9220a8f9-c9e5-46fc-8905-9519f83138e7 · outbound

This paper cites Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:56.867587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:56.867587Z digest=sha256:c314c799bf9d60cc22c1b4266eaf665ce19bf34f94209d4d154308e7ddf61562

Observation cdf18a32-93eb-43b0-8018-ec1683e8b1d8 · outbound

This paper cites Model Extrapolation Expedites Alignment.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Model Extrapolation Expedites Alignment

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:57.259827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:57.259827Z digest=sha256:b96b33c0ac709c9878c12ef956d93dd8669b90e4ec17c35b1ddb877304ee42ee

Observation e8bbe007-6169-48fa-8d86-030ac4ed2cb4 · outbound

This paper cites Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:57.357277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:57.357277Z digest=sha256:ea80dc5149107eaf349e58c6e12d2d584dea084b47e009941e644c22078b303a

Observation d7e542b3-4915-4b0f-96d5-d0b1bd9ca74b · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:57.417182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:57.417182Z digest=sha256:c980c6ec692b25b13725feeaf3393ac18bd072ca9b7c68092074c7f25fb5f40a

Observation 0661f49f-101a-41e5-9eca-3773a6fe15ef · outbound

This paper cites Similar to RIFT, RoAST can be categorized as a fine-tuning stage defense, and therefore is also not suitable for the scenario described.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Similar to RIFT, RoAST can be categorized as a fine-tuning stage defense, and therefore is also not suitable for the scenario described

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:00:02.053821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:59:57.532604Z digest=sha256:c20c8647651c03384a06a008a3807d0de42544cbcc3566754872d805030fba3a

Observation 1d0977e8-b7cd-4f02-a3ef-b856ebc2bbdf · outbound

This paper cites Absolutely Obedient.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Absolutely Obedient

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:00:01.511263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:59:57.721181Z digest=sha256:5c7af4ac8d42d611d7ceb2bad98c07eaf4f45adb1bd0ce5e365ac002606a361f

Observation 5044b574-3f90-44e4-94a8-fc3683c08e2e · outbound

This paper cites Pure Bad.The Pure Bad dataset consists of 100 harmful examples, extracted from the Anthropic Red Teaming Dataset (Ganguli et al., 2022).

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Pure Bad.The Pure Bad dataset consists of 100 harmful examples, extracted from the Anthropic Red Teaming Dataset (Ganguli et al., 2022)

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:00:01.311559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:59:57.801619Z digest=sha256:9d3d21b3c71b79113a6047217120e13a2794a0118d07fd446748c6db5e8a88ad

Observation 57d95072-760a-472a-9eb4-a7975e5adee5 · outbound

This paper cites an unresolved cited work.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:00:01.078057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:59:58.009728Z digest=sha256:05aa4bc1d455a52bfc542e50b59946f16aea3c1a7d2d2460915bf8d29e0cb36b

Observation d30f529e-ab95-46dc-a827-5232a6e9eb52 · outbound

This paper cites 5, and extrapolating further (which we considered broken), following (Lin et al., 2023).

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning 5, and extrapolating further (which we considered broken), following (Lin et al., 2023)

Reference 32

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T00:00:00.840892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:59:58.129917Z digest=sha256:fa3b55398404027ef16f14410358ba40d365009a54f6ef28c489161b003803b4

Observation 74c87f94-dbf0-43da-9f53-33c61064ac15 · outbound

This paper cites We include Meta’s usage guidelines1 in our prompt, following the evaluation protocol of Qi et al.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning We include Meta’s usage guidelines1 in our prompt, following the evaluation protocol of Qi et al

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:00:00.388349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:59:58.303211Z digest=sha256:d9018f05c7e9312ad22160570db3fd90e674f2194dc876293ee69b9d23a2ccae

Observation c9692a5e-b101-4f08-a28c-0d776e7391c4 · outbound

This paper cites No" Response(k=6,α=1.5): “Your task is to complete tasks for people is not recommended.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning No" Response(k=6,α=1.5): “Your task is to complete tasks for people is not recommended

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:00:00.202159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:59:58.413814Z digest=sha256:0ff00fdf0bcd09e358a735e64fbc7568eade5534d3e9655b40d72b2934079b86

Observation a00007b1-ff82-47a3-a550-603b29d4d067 · outbound

This paper cites an unresolved cited work.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Unresolved cited work

Reference 1024

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:00:01.806107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:59:57.615080Z digest=sha256:1d9fb22c867250b5b00f60e7de67df5baf84390e3cafe3d00ba189cb0a5165ac

Observation b71990e6-68f0-4e20-9c57-5b69d7972036 · outbound

This paper cites Parameter-Efficient Fine-Tuning Methods for Pretrained Language Models: A Critical Review and Assessment.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Parameter-Efficient Fine-Tuning Methods for Pretrained Language Models: A Critical Review and Assessment

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:57.134832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:57.134832Z digest=sha256:93e903bb162269d52a3dbe58ae88d066a873cf0d14dcde51dca5bdc574733430

Observation c3dd5955-0634-471c-906a-271f06b94965 · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:55.560201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:55.560201Z digest=sha256:95b97f842edcb3b3730b09b553919c45ed20fced5c853286e138fe83cfda8b65

Observation ed87a7e4-93cb-482a-976a-57b3b42fc5a7 · outbound

This paper cites Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:55.503290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:55.503290Z digest=sha256:907c262738be258d9475ab219fc7051a890834c18ad0a3aa96ccaded058a6b87

Observation 96de15c3-91af-4fc2-8b70-29d0d129cb93 · outbound

This paper cites Xinshuai Dong, Anh Tuan Luu, Min Lin, Shuicheng Yan, and Hanwang Zhang.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Xinshuai Dong, Anh Tuan Luu, Min Lin, Shuicheng Yan, and Hanwang Zhang

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:00:02.261337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:59:55.714750Z digest=sha256:dc6dad66717c7a2f7caad75aa2b7bd134ec32c95a109ffe2df5700f750424a28

Observation f948c392-9160-4e76-bb23-1f01999edd5b · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:55.454095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:55.454095Z digest=sha256:1c18720f5859f04ec1f0876aa239974efabe68888fb41ae945435d1dadd5f40f

Observation ce426289-9b95-4344-a6ca-35b38b0d5789 · outbound

This paper cites Crowdsourcing Multiple Choice Science Questions.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Crowdsourcing Multiple Choice Science Questions

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:57.027070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:57.027070Z digest=sha256:a2175448d9f89a7cfeff4aa60406bafe25c779b854a7106c2ebc27d0bcba7f95

Pith citing papers

Observation 3b7f2f16-529f-46d7-b07b-d0fddd69fd39 · inbound

Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control cites this paper.

Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:21:29.060044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T11:17:03.104902Z digest=sha256:d349829b2f2ecbd45607bfcc618e1bcb178ad06cbec266a6cac28a305b2d9a45

Observation 12261f88-9c3a-4ae5-ac59-179bf9f182b1 · inbound

Preventing Safety Drift in Large Language Models via Coupled Weight and Activation Constraints cites this paper.

Preventing Safety Drift in Large Language Models via Coupled Weight and Activation Constraints LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:21:01.852959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T16:04:25.851592Z digest=sha256:979fffe3d03fa8619feb8e65267887cbef609cf93fc0dd38f111138492ebbb31

Observation 27fd0429-61e2-4e22-b079-1ae9298dbdf0 · inbound

One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries cites this paper.

One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:05:04.062857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T21:01:25.549340Z digest=sha256:b1c7edf6e96142db361a5e6bb527df520c3f05b373ad71bba3c5d79a830a377e