Pith. sign in

Paper Citation Record · LEDGER

Linearly Decoding Refused Knowledge in Aligned Language Models

As of 10 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 1 inbound Pith citation observation for arXiv:2507.00239.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00239 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:27:35.040995Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T05:52:53.880521Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T05:53:04.283223Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact5
  • verified fuzzy22
  • unresolved40
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c5a2a791-7579-49c3-88f6-b2715a52fc5a · outbound

This paper cites Yi-6b-chat.

Linearly Decoding Refused Knowledge in Aligned Language Models Yi-6b-chat

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.934291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:27:34.845006Z digest=sha256:973caf27fbcb70269875a34e2ec1989e5b1fa9cb16295d5e42cbb34b984de3e5

Observation 60715637-4fce-4c16-a3de-03f09fe8e2c7 · outbound

This paper cites Fine-grained analysis of sentence embeddings using auxiliary prediction tasks.

Linearly Decoding Refused Knowledge in Aligned Language Models Fine-grained analysis of sentence embeddings using auxiliary prediction tasks

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.925686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:27:34.848651Z digest=sha256:487a85adb707fc3b0568a4e0992e28fa152b44fd316d59553c0a8c1f19584a6e

Observation 8969fc5d-c558-4fd1-8a32-157c314a5060 · outbound

This paper cites Understanding intermediate layers using linear classifier probes, 2017.

Linearly Decoding Refused Knowledge in Aligned Language Models Understanding intermediate layers using linear classifier probes, 2017

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.851907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.851907Z digest=sha256:a1bb89a487e10990f0efb872909dafa01b2a6b570079eb3688ab82ad3a34df7e

Observation 69d7fd01-27f0-4eb8-9c1e-9955d628c309 · outbound

This paper cites Bowman, Ethan Perez, Roger Baker Grosse, and David Duve- naud.

Linearly Decoding Refused Knowledge in Aligned Language Models Bowman, Ethan Perez, Roger Baker Grosse, and David Duve- naud

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.911881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:27:34.854780Z digest=sha256:adb3da5391ae81f27237e43fc2b1efc9742d07d784efb8534931a0378f0ebc12

Observation 37a6811c-38e0-4a5e-9068-f4c2007e0ea5 · outbound

This paper cites Refusal in language models is mediated by a single direction.

Linearly Decoding Refused Knowledge in Aligned Language Models Refusal in language models is mediated by a single direction

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.904069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:27:34.857729Z digest=sha256:da4b1e8d4e5b4b9530b733e4d6722d3cff4fabdc20f4cd50d203468562e935c3

Observation 9c846136-d36f-414a-a8f2-ef648efbce58 · outbound

This paper cites Language models can predict their own behavior.

Linearly Decoding Refused Knowledge in Aligned Language Models Language models can predict their own behavior

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.860566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.860566Z digest=sha256:c914e7dd9ea4beb28e59c38fc48142654418d592a52721a519a8d26227964ebf

Observation 485bfb6e-dec1-4c15-9d9b-1a77e5774d35 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Linearly Decoding Refused Knowledge in Aligned Language Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.863565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.863565Z digest=sha256:d7dd1fe2595bb6d07239499aed4d79aa71f4611d9507b6dfe92a51afb4ed5d00

Observation 2498cc3d-3119-430c-99e0-fa4527e14201 · outbound

This paper cites Probing classifiers: Promises, shortcomings, and advances.

Linearly Decoding Refused Knowledge in Aligned Language Models Probing classifiers: Promises, shortcomings, and advances

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.866492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.866492Z digest=sha256:3cd13a2101834516c594deea2b7131cf9bc7efb369a0a6538c10a760e0d3f6b5

Observation e12569d5-02ac-4d67-b948-992ae2bca332 · outbound

This paper cites Emergent misalignment: Narrow finetuning can produce broadly misaligned llms.

Linearly Decoding Refused Knowledge in Aligned Language Models Emergent misalignment: Narrow finetuning can produce broadly misaligned llms

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.869282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.869282Z digest=sha256:eb84602a9e5a8e7cb4a38a30dae24261a7047f849a9541f583c4e99fc8b3e08c

Observation 31e68641-64d1-482b-96b7-34c604c2b306 · outbound

This paper cites Wedded to prosperity? informal influence and regional favoritism.

Linearly Decoding Refused Knowledge in Aligned Language Models Wedded to prosperity? informal influence and regional favoritism

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.895949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:27:34.871895Z digest=sha256:bb8845119418ce7b517b387068191fa70dc7d4fffbe89f3d0eee7247da323529

Observation f0016ce6-8359-43f6-85ea-e5843884496d · outbound

This paper cites an unresolved cited work.

Linearly Decoding Refused Knowledge in Aligned Language Models Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.874834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.874834Z digest=sha256:3f31103dcd6c5e6d65245d70bfd77975ebcebc8ec642d989a4157e3c8aee11db

Observation c2aa32b7-6cc5-4ae0-9e33-b4778e0f59a0 · outbound

This paper cites List of countries | Britannica.

Linearly Decoding Refused Knowledge in Aligned Language Models List of countries | Britannica

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.887726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:27:34.877444Z digest=sha256:fd961dfa1f38b410d4fa54993be5dad5a91b3d55ad2dfcba52df17f0ab9375d2

Observation 439e7a25-52f9-4974-a467-8996553ed81c · outbound

This paper cites From Imitation to Introspection: Probing Self-Consciousness in Language Models.

Linearly Decoding Refused Knowledge in Aligned Language Models From Imitation to Introspection: Probing Self-Consciousness in Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:27:35.503514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:27:34.880028Z digest=sha256:5234f3556c92f3c9abeab848356e9f3ce86ef71d244667831a31ade07ec818cf

Observation 950eb059-ec69-4cd0-a671-ddaaba6b4324 · outbound

This paper cites Probing linguistic information for logical inference in pre-trained language models.

Linearly Decoding Refused Knowledge in Aligned Language Models Probing linguistic information for logical inference in pre-trained language models

Reference 14

Resolution
verified exact
doi, observed 2026-08-06T21:27:35.112492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:27:34.883871Z digest=sha256:c7fcbefa97be6fb35eefaab765890afee98ca8666cebd0eb062fd89aa5b8fa5b

Observation eefa3a7d-cae6-4863-a024-f84dc69db856 · outbound

This paper cites Leveraging the Context through Multi-Round Interactions for Jailbreaking Attacks.

Linearly Decoding Refused Knowledge in Aligned Language Models Leveraging the Context through Multi-Round Interactions for Jailbreaking Attacks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.886612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.886612Z digest=sha256:7d1058dd0e3550d574e08713eda562bb9c3428216b34058dd23a7292080ca42e

Observation fe531612-2443-4d1a-aa94-e4a8fd5900df · outbound

This paper cites Breaking down the defenses: A comparative survey of attacks on large language models.

Linearly Decoding Refused Knowledge in Aligned Language Models Breaking down the defenses: A comparative survey of attacks on large language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.889561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.889561Z digest=sha256:31eeebc41b5b13d01da3dd9f38edc7989f27c8bf6af8b300b30ea0a5e78cc4c2

Observation 01705d48-dbcf-4a46-bc10-2a5b4228527b · outbound

This paper cites JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs.

Linearly Decoding Refused Knowledge in Aligned Language Models JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.892626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.892626Z digest=sha256:0ea1df83e1866769b6f2c01ad95fb712db823287a7ecccf3852368915f498c6a

Observation 526a21f6-bf23-4ce8-80df-4186450bc962 · outbound

This paper cites Scaling instruction-finetuned language models.

Linearly Decoding Refused Knowledge in Aligned Language Models Scaling instruction-finetuned language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.895935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.895935Z digest=sha256:d9a7f882c2452e3bc8f4d51b8a5026bc2085c7127477aaf94a2aa55f5960e90e

Observation f180b5d0-bf86-48b5-9d09-b5531dafd636 · outbound

This paper cites Pawan Kumar, and Adel Bibi.

Linearly Decoding Refused Knowledge in Aligned Language Models Pawan Kumar, and Adel Bibi

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.875007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:27:34.898752Z digest=sha256:b4038d1a34b47ba08a7e9051d748d90aa3b2699eea38232dc5192e4fc4f9aa52

Observation b092e29e-e050-4f37-a99e-6c79a296a0d8 · outbound

This paper cites Dissecting recall of factual associations in auto-regressive language models.

Linearly Decoding Refused Knowledge in Aligned Language Models Dissecting recall of factual associations in auto-regressive language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.901466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.901466Z digest=sha256:e5de7e24e130a245a3e116a484fd41fd41423c1400ad9bc323c5331e78169831

Observation 06d3705b-e684-468a-9d09-dfba78d41227 · outbound

This paper cites Estimating knowledge in large language models without generating a single token.

Linearly Decoding Refused Knowledge in Aligned Language Models Estimating knowledge in large language models without generating a single token

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.905010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.905010Z digest=sha256:3e456423749e4f5f35630c1267b5eb7fbf875a5014268a86b679382cd31b0143

Observation 6fae92c4-ed83-4aa5-b882-635b997b8616 · outbound

This paper cites Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection.

Linearly Decoding Refused Knowledge in Aligned Language Models Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.908262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.908262Z digest=sha256:559bcb2f2a0a067eeca5a303dafa0b5b5afe8191d33c77f0b8c390565c045ee8

Observation cfa06f62-615f-40ca-9ed9-45d4fdd293ab · outbound

This paper cites Language models represent space and time.

Linearly Decoding Refused Knowledge in Aligned Language Models Language models represent space and time

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.867466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:27:34.910913Z digest=sha256:41b8ac71f7c5ea58b96fe11e6091830e666da1f26f4e16410d0a32de70fe0f49

Observation 646a3c4b-2bd7-4830-8711-db5f719a9414 · outbound

This paper cites The Elements of Statistical Learning: Data Mining, Inference, and Prediction, volume 2.

Linearly Decoding Refused Knowledge in Aligned Language Models The Elements of Statistical Learning: Data Mining, Inference, and Prediction, volume 2

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.859326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:27:34.913521Z digest=sha256:0ee5564251af75e7e5174932e6dd31e01914ab21373d0be85838e436c99e27b0

Observation c3ac457d-fb07-4073-bd67-6fd10c97a6b4 · outbound

This paper cites Do LLMs “know” internally when they follow instructions? In The Thirteenth International Conference on Learning Representations, 2025.

Linearly Decoding Refused Knowledge in Aligned Language Models Do LLMs “know” internally when they follow instructions? In The Thirteenth International Conference on Learning Representations, 2025

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.849401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:27:34.916396Z digest=sha256:fdd3ce30ff3ae9967da12cfc744086075dd5db04d512b5a5f5e310805aaeedb6

Observation 759b8f54-cd1d-4747-854f-355b875cf830 · outbound

This paper cites Linearity of relation decoding in transformer language models.

Linearly Decoding Refused Knowledge in Aligned Language Models Linearity of relation decoding in transformer language models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.919014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.919014Z digest=sha256:04b1f303e3507fc32e89bbe440ec5ee720eb086443d1650593494dd40266ef94

Observation 8430ffe5-7061-4fc7-9eef-39105ff4c5cf · outbound

This paper cites an unresolved cited work.

Linearly Decoding Refused Knowledge in Aligned Language Models Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.921758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.921758Z digest=sha256:c0387a831e236daa6cc541585f2ad75be4f96a3fec3915fa7873929168874cb9

Observation c01ad35f-d372-4f9f-b395-75534335cdf1 · outbound

This paper cites Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models.

Linearly Decoding Refused Knowledge in Aligned Language Models Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.924646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.924646Z digest=sha256:91e45b7e0c431864d83194c366a46b471305bd59257cfc008f602ce30564a450

Observation 3336f87e-aea9-41c8-a87f-3a981ccda430 · outbound

This paper cites an unresolved cited work.

Linearly Decoding Refused Knowledge in Aligned Language Models Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:27:35.836329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:27:34.927770Z digest=sha256:b8378f210f8ae8cd3f48cfdb7bfcbb9bd6dffd89ecf97c543a1e75bd29df9913

Observation cfda084d-86b9-4854-91d3-8374c95541d7 · outbound

This paper cites Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language and vision-language models.

Linearly Decoding Refused Knowledge in Aligned Language Models Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language and vision-language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.930716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.930716Z digest=sha256:0848f73a22342ef9a737b1ac04bb8b473e970d14c65ebc6f2a6f3e068341a882

Observation 4bbaf33b-6b53-4065-8505-50fe0f78aa21 · outbound

This paper cites Alignment of Language Agents.

Linearly Decoding Refused Knowledge in Aligned Language Models Alignment of Language Agents

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.933507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.933507Z digest=sha256:7ffaea0116f1d3335d1688025e22ccdebdb856e8bb598d7ec8cff3d343d3de36

Observation 21901eb9-72ba-4d6c-903f-2f333a024a6a · outbound

This paper cites Linear representations of political perspective emerge in large language models.

Linearly Decoding Refused Knowledge in Aligned Language Models Linear representations of political perspective emerge in large language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.936952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.936952Z digest=sha256:821e0c15017c7e6b72a71f61f99d0a35d6277ec102624c93b8f6a655adeb2ce1

Observation 7baf0ade-4903-4696-9913-52d67c38e2f1 · outbound

This paper cites LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B.

Linearly Decoding Refused Knowledge in Aligned Language Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.939664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.939664Z digest=sha256:49060ad40f8cedf9a659a3f3d9f59374ed2ee52afe9c367eb77a5540ecf77b97

Observation bcc200e2-fc7b-4a7c-aa80-f5ae16cb250e · outbound

This paper cites Revealing the Intrinsic Ethical Vulnerability of Aligned Large Language Models.

Linearly Decoding Refused Knowledge in Aligned Language Models Revealing the Intrinsic Ethical Vulnerability of Aligned Large Language Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:27:35.231753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:27:34.942636Z digest=sha256:bc206d5b5ee79fc7ae95d3d4f2c745ac350c2356eca0a0cb7f14999695876c22

Observation e8c6d6a0-da41-49f0-b24a-5a9ff6d632cb · outbound

This paper cites The unlocking spell on base LLMs: Rethinking alignment via in-context learning.

Linearly Decoding Refused Knowledge in Aligned Language Models The unlocking spell on base LLMs: Rethinking alignment via in-context learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.823624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:27:34.945645Z digest=sha256:0c2dc7e04ef83f431f756bb1629a68aee58e4b2a409d7fd543b114852a29ceef

Observation 1dabd6f2-50a8-44ae-8f45-764cd3e04cc0 · outbound

This paper cites Towards Understanding Jailbreak Attacks in LLMs: A Representation Space Analysis.

Linearly Decoding Refused Knowledge in Aligned Language Models Towards Understanding Jailbreak Attacks in LLMs: A Representation Space Analysis

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.948426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.948426Z digest=sha256:732564b319ddb064b9a0089a69318a792e7bc64a3698a82bfa9e3954690a4910

Observation 74188ff8-dbf6-4b71-9309-9484a5e2fb27 · outbound

This paper cites Formalizing and benchmarking prompt injection attacks and defenses.

Linearly Decoding Refused Knowledge in Aligned Language Models Formalizing and benchmarking prompt injection attacks and defenses

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.815303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:27:34.951313Z digest=sha256:ee25f714a30886e342dea51073eac0f0b31261aa3a082656142d905d5e601eae

Observation 70fcb23a-e528-4821-9f89-888045f82745 · outbound

This paper cites Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates.

Linearly Decoding Refused Knowledge in Aligned Language Models Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.954267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.954267Z digest=sha256:204e580472c0f2fa1af3c9fc73b809253eeee3718a0cc278752d7e85993a67ed

Observation 86438a36-a411-4441-a471-54d8bed4c87d · outbound

This paper cites The geometry of truth: Emergent linear structure in large lan- guage model representations of true/false datasets.

Linearly Decoding Refused Knowledge in Aligned Language Models The geometry of truth: Emergent linear structure in large lan- guage model representations of true/false datasets

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.806655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:27:34.957079Z digest=sha256:c2b3dd989dfff102fa7f61f518a81c7a4ceea7d03b4095faa1d3c00b7b575c9a

Observation 6c34073d-86c8-4a1e-89c7-b423059d414c · outbound

This paper cites Occupation Data - O*NET 29.2 Data Dictionary at O*NET Re- source Center.

Linearly Decoding Refused Knowledge in Aligned Language Models Occupation Data - O*NET 29.2 Data Dictionary at O*NET Re- source Center

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.789994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:27:34.962407Z digest=sha256:cf66516380e95e93c6f92fed6f584119b482d3d079d44d03181098d8cb925e40

Observation b186f49c-7811-4216-8e8b-7da273136424 · outbound

This paper cites Training language models to follow instructions with human feedback.

Linearly Decoding Refused Knowledge in Aligned Language Models Training language models to follow instructions with human feedback

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.781795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:27:34.965216Z digest=sha256:b1a279ed789054f1dbdd206bbe53f51cf07520e2b76793779b7a0efc70eaab8f

Observation d3d0f1e2-31dd-4717-8934-f78e7e3e7546 · outbound

This paper cites The linear representation hypothesis and the geometry of large language models.

Linearly Decoding Refused Knowledge in Aligned Language Models The linear representation hypothesis and the geometry of large language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.773466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:27:34.967915Z digest=sha256:91e85078af781b21f9be0b93fdd5d89313fce058197e419065e09d43d83645d3

Observation 5096fc8d-d834-4727-8e4f-34e49c62aa1b · outbound

This paper cites Ignore Previous Prompt: Attack Techniques For Language Models.

Linearly Decoding Refused Knowledge in Aligned Language Models Ignore Previous Prompt: Attack Techniques For Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.970768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.970768Z digest=sha256:f08dfee0c6d680127960632c27560dde87fe4ede43b7e0e8fd66cce113cfa09a

Observation a17f31f2-5cce-4416-a8d5-5f571d90d4fd · outbound

This paper cites Fine-tuning aligned language models compromises safety, even when users do not intend to! In The Twelfth International Conference on Learning Representations , 2024.

Linearly Decoding Refused Knowledge in Aligned Language Models Fine-tuning aligned language models compromises safety, even when users do not intend to! In The Twelfth International Conference on Learning Representations , 2024

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.973853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.973853Z digest=sha256:538dd608e5b1c2b246234ae3c733a8f17ae83e37b2fc7ff7adf83365c4e11e44

Observation 67379375-b768-4557-a664-7adce87c13e5 · outbound

This paper cites Safety alignment should be made more than just a few tokens deep.

Linearly Decoding Refused Knowledge in Aligned Language Models Safety alignment should be made more than just a few tokens deep

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.977277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.977277Z digest=sha256:fa29c5be27d998d234776a8cce25fff48b173f473a3ad06e1501c7cdadfc8a66

Observation fb4c71d6-3951-4f95-a5a5-ad814441f6bc · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Linearly Decoding Refused Knowledge in Aligned Language Models Direct preference optimization: Your language model is secretly a reward model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.980167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.980167Z digest=sha256:98792edba877fa5bc4af5ed12980a3a37d2828ec40b907681f545d22d3d8d623

Observation 1c158573-9d21-4878-b441-ee83b5fc65c0 · outbound

This paper cites Multi- task prompted training enables zero-shot task generalization.

Linearly Decoding Refused Knowledge in Aligned Language Models Multi- task prompted training enables zero-shot task generalization

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.749217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:27:34.982721Z digest=sha256:a0e946253a20701d5ba0847852bc1a7571e174f6493a518f3df6d39714583535

Observation 74e33e9a-f10b-4f6d-88cb-9af4f587b0f0 · outbound

This paper cites Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation.

Linearly Decoding Refused Knowledge in Aligned Language Models Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.985261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.985261Z digest=sha256:22574099aeeb0ee0d256dfe7d3a09588a884d01a81d5397257591349b96aa022

Observation 8e459ab7-d98c-43af-9a8b-2b7768a272c4 · outbound

This paper cites do anything now.

Linearly Decoding Refused Knowledge in Aligned Language Models do anything now

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.740959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:27:34.988437Z digest=sha256:571a9c2dd52b8f1ffcf79c21d84ad248e678658f772871a3646a497178b24084

Observation 8ccd69e5-4941-4fa4-ac29-f2fcc1404203 · outbound

This paper cites Measuring Free-Form Decision-Making Inconsistency of Language Models in Military Crisis Simulations.

Linearly Decoding Refused Knowledge in Aligned Language Models Measuring Free-Form Decision-Making Inconsistency of Language Models in Military Crisis Simulations

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:27:35.197892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:27:34.991094Z digest=sha256:e875d79ad6dc9f28c2f9bc31d8e335cdff528f76765812af84aa418633be0b64

Observation 89b52122-a299-46f1-ac45-1adf19bcba0f · outbound

This paper cites Large Language Models are Inconsistent and Biased Evaluators.

Linearly Decoding Refused Knowledge in Aligned Language Models Large Language Models are Inconsistent and Biased Evaluators

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.994224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.994224Z digest=sha256:01e105ad5081183b295ebf597188bf5acd0384a3ab4ed926b6a73b4c0e539d7a

Observation a88fe132-d912-464b-adb0-575e52fb0356 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Linearly Decoding Refused Knowledge in Aligned Language Models Gemma 2: Improving Open Language Models at a Practical Size

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.997511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.997511Z digest=sha256:6e35dfc047992e32cf461f615d20b05773610a27fdc0e36c6845e43843bd8e0d

Observation 358b4c94-e01b-4aea-9bf5-e6661b3756af · outbound

This paper cites What do you learn from context? Probing for sentence structure in contextualized word representations.

Linearly Decoding Refused Knowledge in Aligned Language Models What do you learn from context? Probing for sentence structure in contextualized word representations

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:35.000447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.000447Z digest=sha256:304691e30b050562f697b60e49940b6c31713b9395613ddd33cd600e2b047dcb

Observation d4edba58-579f-41fa-95c9-c325e80519d1 · outbound

This paper cites Attention is all you need.

Linearly Decoding Refused Knowledge in Aligned Language Models Attention is all you need

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.732940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:27:35.003388Z digest=sha256:a7d34171863b733721483e1c26b5af545f96cc20613a19a62ce4b10c0b1050f7

Observation f93403d1-7fd8-425f-b47d-251dcb60d155 · outbound

This paper cites White-box multimodal jailbreaks against large vision-language models.

Linearly Decoding Refused Knowledge in Aligned Language Models White-box multimodal jailbreaks against large vision-language models

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.724603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:27:35.005888Z digest=sha256:882b3cd01f4636bcf3d790b0d856f0f71d3deb5ffe8515e1392f51ef12074250

Observation 8a797ac8-2e88-4f95-8733-6c1002f776bd · outbound

This paper cites Jailbroken: How does LLM safety training fail? In Thirty-seventh Conference on Neural Information Processing Systems, 2023.

Linearly Decoding Refused Knowledge in Aligned Language Models Jailbroken: How does LLM safety training fail? In Thirty-seventh Conference on Neural Information Processing Systems, 2023

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.715626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:27:35.008684Z digest=sha256:2fa7d3b0d6aba6c6acb7bce199c82967dddfdb9e269d2e2a0d17210bb36820f8

Observation e52b27ab-0ec8-40c8-8109-dde23d632199 · outbound

This paper cites Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications.

Linearly Decoding Refused Knowledge in Aligned Language Models Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:35.011420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.011420Z digest=sha256:ea06e921fe358b486ebcfeaae0b1c16a25e3dbae96c8d330e982f742f095a3a6

Observation c8564b34-057e-407c-bf81-6ed1135155b4 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

Linearly Decoding Refused Knowledge in Aligned Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:35.014424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.014424Z digest=sha256:c62fecb01f32a4fedbdb7000c2661ade36247f02716b42c816114fb3663d4c39

Observation e8c3f6af-81dd-40ef-ab5a-21b0092acb6e · outbound

This paper cites Efficient streaming language models with attention sinks.

Linearly Decoding Refused Knowledge in Aligned Language Models Efficient streaming language models with attention sinks

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:35.017387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.017387Z digest=sha256:bd733edea9e953f82b3607c07474d194f6355ee99418915843156b00850fe179

Observation f17e0bc1-e48f-4aac-ab21-9c93e8167193 · outbound

This paper cites Assessing Hidden Risks of LLMs: An Empirical Study on Robustness, Consistency, and Credibility.

Linearly Decoding Refused Knowledge in Aligned Language Models Assessing Hidden Risks of LLMs: An Empirical Study on Robustness, Consistency, and Credibility

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:35.020100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.020100Z digest=sha256:a2e90e09ff8177e8a190d1dcc31a4ec4cd3c22196bd5a83eadd8eba1c8fc9a6c

Observation a33b0e54-ad19-4232-89ae-fc87d627e05a · outbound

This paper cites On the vulnerability of safety alignment in open-access LLMs.

Linearly Decoding Refused Knowledge in Aligned Language Models On the vulnerability of safety alignment in open-access LLMs

Reference 61

Resolution
verified exact
doi, observed 2026-08-06T21:27:35.073634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:27:35.022967Z digest=sha256:6a9c2c8821c3705b73d7b51eca273bf19339d9572528d5abb890fedc346d377c

Observation 549a16a1-0b82-45bf-83b0-0e8a7a9e5e49 · outbound

This paper cites Jailbreak Attacks and Defenses Against Large Language Models: A Survey.

Linearly Decoding Refused Knowledge in Aligned Language Models Jailbreak Attacks and Defenses Against Large Language Models: A Survey

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:35.025865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.025865Z digest=sha256:84e1af59cdf8f5f8aafaa195924db48a9621b6f6233043c8c9924d09fd42ba16

Observation 9836f397-8f1c-4026-b996-27b1613c361a · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

Linearly Decoding Refused Knowledge in Aligned Language Models Yi: Open Foundation Models by 01.AI

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:35.028823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.028823Z digest=sha256:2afbf9a696a6959fe67885a8c2ed6a5428016431354d6adf07abdd309115bc0e

Observation 5857d7c6-e88b-4079-bdda-232f7a5b79d9 · outbound

This paper cites Don’t listen to me: understanding and exploring jailbreak prompts of large language models.

Linearly Decoding Refused Knowledge in Aligned Language Models Don’t listen to me: understanding and exploring jailbreak prompts of large language models

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.702045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:27:35.031749Z digest=sha256:8a379581cc7bf9d2cb84fcdf4fa788093db5840a9fb07e1034f2ae0e9194652f

Observation bf6791de-0c35-48c4-9af0-7d7f436b471c · outbound

This paper cites Removing RLHF protections in GPT-4 via fine-tuning.

Linearly Decoding Refused Knowledge in Aligned Language Models Removing RLHF protections in GPT-4 via fine-tuning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:35.034422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.034422Z digest=sha256:7705d077f152521944f0bcd3b25ce46091b874f532239404cc3dc6730d239828

Observation dab3bda0-c3ff-45ee-b08c-c3cb7706f90b · outbound

This paper cites Lima: Less is more for alignment.

Linearly Decoding Refused Knowledge in Aligned Language Models Lima: Less is more for alignment

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:35.037339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.037339Z digest=sha256:39cd0cd700cd2aa94584e28761591a544de95dd4c755903bf68bae68f99b794b

Observation e3b2af77-ec94-45f6-9f49-31afebf2d95b · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Linearly Decoding Refused Knowledge in Aligned Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 67

Resolution
malformed identifier
no resolver link, observed 2026-08-06T21:27:35.040995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.040995Z digest=sha256:d5d56e4441e4a8057306e35d474c5af535ba40a832107afdb4e3b4e55318f40c

Observation 7b418754-5462-49f6-8997-fc1ee1e73aef · outbound

This paper cites an unresolved cited work.

Linearly Decoding Refused Knowledge in Aligned Language Models Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:27:35.798529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:27:34.959696Z digest=sha256:5f11ecce72c9214d4b1cae2f7cbb4a0d2e432bc00d8cd2413ff62313ab9bb621

Pith citing papers

Observation dbcb59b9-70ef-4595-91e4-d4fff01c1fd4 · inbound

HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models cites this paper.

HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models Linearly Decoding Refused Knowledge in Aligned Language Models

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:04.285500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T05:52:53.880521Z digest=sha256:e9c346250792b3e1ac1e3f205d192aacbfe58fc13dcadbe21b9bf667c1026c76