Pith. sign in

Paper Citation Record · LEDGER

Linearly Decoding Refused Knowledge in Aligned Language Models

As of 9 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 1 inbound Pith citation observation for arXiv:2507.00239.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00239 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:27:35.040995Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T05:52:53.880521Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T05:53:04.283223Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact5
  • verified fuzzy22
  • unresolved40
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c5a2a791-7579-49c3-88f6-b2715a52fc5a · outbound

This paper cites Yi-6b-chat.

Linearly Decoding Refused Knowledge in Aligned Language Models Yi-6b-chat

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.934291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:27:34.845006Z digest=sha256:057fb6ea828c212e1ec4562f01649416ff555d777143ed895b81589df3de4c8c

Observation 60715637-4fce-4c16-a3de-03f09fe8e2c7 · outbound

This paper cites Fine-grained analysis of sentence embeddings using auxiliary prediction tasks.

Linearly Decoding Refused Knowledge in Aligned Language Models Fine-grained analysis of sentence embeddings using auxiliary prediction tasks

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.925686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:27:34.848651Z digest=sha256:0ca376879cc5bb7de3ecdb8e37786bc40de9675b7c3c81bb60ef0e1c2e63f380

Observation 8969fc5d-c558-4fd1-8a32-157c314a5060 · outbound

This paper cites Understanding intermediate layers using linear classifier probes, 2017.

Linearly Decoding Refused Knowledge in Aligned Language Models Understanding intermediate layers using linear classifier probes, 2017

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.851907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.851907Z digest=sha256:a1bb89a487e10990f0efb872909dafa01b2a6b570079eb3688ab82ad3a34df7e

Observation 69d7fd01-27f0-4eb8-9c1e-9955d628c309 · outbound

This paper cites Bowman, Ethan Perez, Roger Baker Grosse, and David Duve- naud.

Linearly Decoding Refused Knowledge in Aligned Language Models Bowman, Ethan Perez, Roger Baker Grosse, and David Duve- naud

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.911881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:27:34.854780Z digest=sha256:2e2c9f8e52fa281a1ad3b4d5b229eaefc6b4be3e8440ae38221f55f056315a40

Observation 37a6811c-38e0-4a5e-9068-f4c2007e0ea5 · outbound

This paper cites Refusal in language models is mediated by a single direction.

Linearly Decoding Refused Knowledge in Aligned Language Models Refusal in language models is mediated by a single direction

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.904069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:27:34.857729Z digest=sha256:a5dd419ef667b9e8a2a5dccd085e24aa04757fcd64123bdfa2cc5b51e4c13151

Observation 9c846136-d36f-414a-a8f2-ef648efbce58 · outbound

This paper cites Language models can predict their own behavior.

Linearly Decoding Refused Knowledge in Aligned Language Models Language models can predict their own behavior

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.860566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.860566Z digest=sha256:c914e7dd9ea4beb28e59c38fc48142654418d592a52721a519a8d26227964ebf

Observation 485bfb6e-dec1-4c15-9d9b-1a77e5774d35 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Linearly Decoding Refused Knowledge in Aligned Language Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.863565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.863565Z digest=sha256:d7dd1fe2595bb6d07239499aed4d79aa71f4611d9507b6dfe92a51afb4ed5d00

Observation 2498cc3d-3119-430c-99e0-fa4527e14201 · outbound

This paper cites Probing classifiers: Promises, shortcomings, and advances.

Linearly Decoding Refused Knowledge in Aligned Language Models Probing classifiers: Promises, shortcomings, and advances

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.866492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.866492Z digest=sha256:3cd13a2101834516c594deea2b7131cf9bc7efb369a0a6538c10a760e0d3f6b5

Observation e12569d5-02ac-4d67-b948-992ae2bca332 · outbound

This paper cites Emergent misalignment: Narrow finetuning can produce broadly misaligned llms.

Linearly Decoding Refused Knowledge in Aligned Language Models Emergent misalignment: Narrow finetuning can produce broadly misaligned llms

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.869282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.869282Z digest=sha256:eb84602a9e5a8e7cb4a38a30dae24261a7047f849a9541f583c4e99fc8b3e08c

Observation 31e68641-64d1-482b-96b7-34c604c2b306 · outbound

This paper cites Wedded to prosperity? informal influence and regional favoritism.

Linearly Decoding Refused Knowledge in Aligned Language Models Wedded to prosperity? informal influence and regional favoritism

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.895949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:27:34.871895Z digest=sha256:0b9804573c60b8f4716654956741e9af846c53531c999e76ce984187749bf524

Observation f0016ce6-8359-43f6-85ea-e5843884496d · outbound

This paper cites an unresolved cited work.

Linearly Decoding Refused Knowledge in Aligned Language Models Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.874834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.874834Z digest=sha256:3f31103dcd6c5e6d65245d70bfd77975ebcebc8ec642d989a4157e3c8aee11db

Observation c2aa32b7-6cc5-4ae0-9e33-b4778e0f59a0 · outbound

This paper cites List of countries | Britannica.

Linearly Decoding Refused Knowledge in Aligned Language Models List of countries | Britannica

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.887726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:27:34.877444Z digest=sha256:974c3883d0b29bcf62338628dbf4b0f07f9483e2e067a4facc61e7f44cb8890e

Observation 439e7a25-52f9-4974-a467-8996553ed81c · outbound

This paper cites From Imitation to Introspection: Probing Self-Consciousness in Language Models.

Linearly Decoding Refused Knowledge in Aligned Language Models From Imitation to Introspection: Probing Self-Consciousness in Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:27:35.503514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:27:34.880028Z digest=sha256:7bdb5db39d6434d0b86da2202e103d7d3caf5ebd74009f0daf177d6fa4a1d801

Observation 950eb059-ec69-4cd0-a671-ddaaba6b4324 · outbound

This paper cites Probing linguistic information for logical inference in pre-trained language models.

Linearly Decoding Refused Knowledge in Aligned Language Models Probing linguistic information for logical inference in pre-trained language models

Reference 14

Resolution
verified exact
doi, observed 2026-08-06T21:27:35.112492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:27:34.883871Z digest=sha256:1e06a243856bb7b0cba4ba7e0793029d7f3647c0e2cb007972bbdbb380d93a76

Observation eefa3a7d-cae6-4863-a024-f84dc69db856 · outbound

This paper cites Leveraging the Context through Multi-Round Interactions for Jailbreaking Attacks.

Linearly Decoding Refused Knowledge in Aligned Language Models Leveraging the Context through Multi-Round Interactions for Jailbreaking Attacks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.886612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.886612Z digest=sha256:b466b77bd3fe9714c4cb095814afd66be919db5efb4561f749044d6e890412f4

Observation fe531612-2443-4d1a-aa94-e4a8fd5900df · outbound

This paper cites Breaking down the defenses: A comparative survey of attacks on large language models.

Linearly Decoding Refused Knowledge in Aligned Language Models Breaking down the defenses: A comparative survey of attacks on large language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.889561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.889561Z digest=sha256:31eeebc41b5b13d01da3dd9f38edc7989f27c8bf6af8b300b30ea0a5e78cc4c2

Observation 01705d48-dbcf-4a46-bc10-2a5b4228527b · outbound

This paper cites JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs.

Linearly Decoding Refused Knowledge in Aligned Language Models JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.892626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.892626Z digest=sha256:63dc3a74204918fc866f4e7c72c54c5126ef7fe372a56221478d92ccedd83f0b

Observation 526a21f6-bf23-4ce8-80df-4186450bc962 · outbound

This paper cites Scaling instruction-finetuned language models.

Linearly Decoding Refused Knowledge in Aligned Language Models Scaling instruction-finetuned language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.895935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.895935Z digest=sha256:d9a7f882c2452e3bc8f4d51b8a5026bc2085c7127477aaf94a2aa55f5960e90e

Observation f180b5d0-bf86-48b5-9d09-b5531dafd636 · outbound

This paper cites Pawan Kumar, and Adel Bibi.

Linearly Decoding Refused Knowledge in Aligned Language Models Pawan Kumar, and Adel Bibi

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.875007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:27:34.898752Z digest=sha256:e6c6acdcc6e31b2352ca5bd82bfc2f2bf100877908a5657d968adf6a545e81c3

Observation b092e29e-e050-4f37-a99e-6c79a296a0d8 · outbound

This paper cites Dissecting recall of factual associations in auto-regressive language models.

Linearly Decoding Refused Knowledge in Aligned Language Models Dissecting recall of factual associations in auto-regressive language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.901466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.901466Z digest=sha256:e5de7e24e130a245a3e116a484fd41fd41423c1400ad9bc323c5331e78169831

Observation 06d3705b-e684-468a-9d09-dfba78d41227 · outbound

This paper cites Estimating knowledge in large language models without generating a single token.

Linearly Decoding Refused Knowledge in Aligned Language Models Estimating knowledge in large language models without generating a single token

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.905010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.905010Z digest=sha256:3e456423749e4f5f35630c1267b5eb7fbf875a5014268a86b679382cd31b0143

Observation 6fae92c4-ed83-4aa5-b882-635b997b8616 · outbound

This paper cites Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection.

Linearly Decoding Refused Knowledge in Aligned Language Models Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.908262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.908262Z digest=sha256:559bcb2f2a0a067eeca5a303dafa0b5b5afe8191d33c77f0b8c390565c045ee8

Observation cfa06f62-615f-40ca-9ed9-45d4fdd293ab · outbound

This paper cites Language models represent space and time.

Linearly Decoding Refused Knowledge in Aligned Language Models Language models represent space and time

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.867466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:27:34.910913Z digest=sha256:4b7fbacd148c3038c076a4d7714e56ff9965b0d53daca1a7a762e3e4ba5dcd73

Observation 646a3c4b-2bd7-4830-8711-db5f719a9414 · outbound

This paper cites The Elements of Statistical Learning: Data Mining, Inference, and Prediction, volume 2.

Linearly Decoding Refused Knowledge in Aligned Language Models The Elements of Statistical Learning: Data Mining, Inference, and Prediction, volume 2

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.859326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:27:34.913521Z digest=sha256:81c77a782483f77bc903ac433d3c2ac1697b5bfdcd3c91943db1a71e3d10f480

Observation c3ac457d-fb07-4073-bd67-6fd10c97a6b4 · outbound

This paper cites Do LLMs “know” internally when they follow instructions? In The Thirteenth International Conference on Learning Representations, 2025.

Linearly Decoding Refused Knowledge in Aligned Language Models Do LLMs “know” internally when they follow instructions? In The Thirteenth International Conference on Learning Representations, 2025

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.849401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:27:34.916396Z digest=sha256:b8662769e42c8b5020388dc71aabcc18ec57559f2baba96a746e73a5b04e17f2

Observation 759b8f54-cd1d-4747-854f-355b875cf830 · outbound

This paper cites Linearity of relation decoding in transformer language models.

Linearly Decoding Refused Knowledge in Aligned Language Models Linearity of relation decoding in transformer language models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.919014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.919014Z digest=sha256:04b1f303e3507fc32e89bbe440ec5ee720eb086443d1650593494dd40266ef94

Observation 8430ffe5-7061-4fc7-9eef-39105ff4c5cf · outbound

This paper cites an unresolved cited work.

Linearly Decoding Refused Knowledge in Aligned Language Models Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.921758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.921758Z digest=sha256:c0387a831e236daa6cc541585f2ad75be4f96a3fec3915fa7873929168874cb9

Observation c01ad35f-d372-4f9f-b395-75534335cdf1 · outbound

This paper cites Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models.

Linearly Decoding Refused Knowledge in Aligned Language Models Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.924646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.924646Z digest=sha256:91e45b7e0c431864d83194c366a46b471305bd59257cfc008f602ce30564a450

Observation 3336f87e-aea9-41c8-a87f-3a981ccda430 · outbound

This paper cites an unresolved cited work.

Linearly Decoding Refused Knowledge in Aligned Language Models Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:27:35.836329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:27:34.927770Z digest=sha256:4a2c085495810a222050f6c0bbfa7e73d0f0702139e9537d6d9621dd3b23ca7b

Observation cfda084d-86b9-4854-91d3-8374c95541d7 · outbound

This paper cites Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language and vision-language models.

Linearly Decoding Refused Knowledge in Aligned Language Models Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language and vision-language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.930716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.930716Z digest=sha256:0848f73a22342ef9a737b1ac04bb8b473e970d14c65ebc6f2a6f3e068341a882

Observation 4bbaf33b-6b53-4065-8505-50fe0f78aa21 · outbound

This paper cites Alignment of Language Agents.

Linearly Decoding Refused Knowledge in Aligned Language Models Alignment of Language Agents

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.933507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.933507Z digest=sha256:7ffaea0116f1d3335d1688025e22ccdebdb856e8bb598d7ec8cff3d343d3de36

Observation 21901eb9-72ba-4d6c-903f-2f333a024a6a · outbound

This paper cites Linear representations of political perspective emerge in large language models.

Linearly Decoding Refused Knowledge in Aligned Language Models Linear representations of political perspective emerge in large language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.936952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.936952Z digest=sha256:821e0c15017c7e6b72a71f61f99d0a35d6277ec102624c93b8f6a655adeb2ce1

Observation 7baf0ade-4903-4696-9913-52d67c38e2f1 · outbound

This paper cites LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B.

Linearly Decoding Refused Knowledge in Aligned Language Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.939664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.939664Z digest=sha256:49060ad40f8cedf9a659a3f3d9f59374ed2ee52afe9c367eb77a5540ecf77b97

Observation bcc200e2-fc7b-4a7c-aa80-f5ae16cb250e · outbound

This paper cites Revealing the Intrinsic Ethical Vulnerability of Aligned Large Language Models.

Linearly Decoding Refused Knowledge in Aligned Language Models Revealing the Intrinsic Ethical Vulnerability of Aligned Large Language Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:27:35.231753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:27:34.942636Z digest=sha256:b2993dd799af684119d3732bda15e68e4b8a076428534db656c05b58ba2e2625

Observation e8c6d6a0-da41-49f0-b24a-5a9ff6d632cb · outbound

This paper cites The unlocking spell on base LLMs: Rethinking alignment via in-context learning.

Linearly Decoding Refused Knowledge in Aligned Language Models The unlocking spell on base LLMs: Rethinking alignment via in-context learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.823624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:27:34.945645Z digest=sha256:2cfc80ba439399ae52bbc5b86d82bdd388abd0b0dd31d3fad1c4f0a7a6ff3392

Observation 1dabd6f2-50a8-44ae-8f45-764cd3e04cc0 · outbound

This paper cites Towards Understanding Jailbreak Attacks in LLMs: A Representation Space Analysis.

Linearly Decoding Refused Knowledge in Aligned Language Models Towards Understanding Jailbreak Attacks in LLMs: A Representation Space Analysis

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.948426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.948426Z digest=sha256:732564b319ddb064b9a0089a69318a792e7bc64a3698a82bfa9e3954690a4910

Observation 74188ff8-dbf6-4b71-9309-9484a5e2fb27 · outbound

This paper cites Formalizing and benchmarking prompt injection attacks and defenses.

Linearly Decoding Refused Knowledge in Aligned Language Models Formalizing and benchmarking prompt injection attacks and defenses

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.815303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:27:34.951313Z digest=sha256:97f55ea9518695f92a4f4c68c2f41490ae86e2a952ebabbfe9068389886b01bc

Observation 70fcb23a-e528-4821-9f89-888045f82745 · outbound

This paper cites Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates.

Linearly Decoding Refused Knowledge in Aligned Language Models Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.954267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.954267Z digest=sha256:204e580472c0f2fa1af3c9fc73b809253eeee3718a0cc278752d7e85993a67ed

Observation 86438a36-a411-4441-a471-54d8bed4c87d · outbound

This paper cites The geometry of truth: Emergent linear structure in large lan- guage model representations of true/false datasets.

Linearly Decoding Refused Knowledge in Aligned Language Models The geometry of truth: Emergent linear structure in large lan- guage model representations of true/false datasets

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.806655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:27:34.957079Z digest=sha256:4fa44aef1f5e8b54fc7c6e06699cc8586d3800f23ca25e100a7bff66550f2f41

Observation 6c34073d-86c8-4a1e-89c7-b423059d414c · outbound

This paper cites Occupation Data - O*NET 29.2 Data Dictionary at O*NET Re- source Center.

Linearly Decoding Refused Knowledge in Aligned Language Models Occupation Data - O*NET 29.2 Data Dictionary at O*NET Re- source Center

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.789994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:27:34.962407Z digest=sha256:9900291a47f5c2cdb35bad972376fa34ff638a306c7da7a3169f559a49b6fda3

Observation b186f49c-7811-4216-8e8b-7da273136424 · outbound

This paper cites Training language models to follow instructions with human feedback.

Linearly Decoding Refused Knowledge in Aligned Language Models Training language models to follow instructions with human feedback

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.781795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:27:34.965216Z digest=sha256:a19ea447466ff6b114e19d49e6040f827ad45ac85d4c210c841c81b11b57e9a1

Observation d3d0f1e2-31dd-4717-8934-f78e7e3e7546 · outbound

This paper cites The linear representation hypothesis and the geometry of large language models.

Linearly Decoding Refused Knowledge in Aligned Language Models The linear representation hypothesis and the geometry of large language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.773466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:27:34.967915Z digest=sha256:e880291d0726946d3f9acef78bd9ff923ca8ed8018cb0243709e8f665ef2c958

Observation 5096fc8d-d834-4727-8e4f-34e49c62aa1b · outbound

This paper cites Ignore Previous Prompt: Attack Techniques For Language Models.

Linearly Decoding Refused Knowledge in Aligned Language Models Ignore Previous Prompt: Attack Techniques For Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.970768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.970768Z digest=sha256:f08dfee0c6d680127960632c27560dde87fe4ede43b7e0e8fd66cce113cfa09a

Observation a17f31f2-5cce-4416-a8d5-5f571d90d4fd · outbound

This paper cites Fine-tuning aligned language models compromises safety, even when users do not intend to! In The Twelfth International Conference on Learning Representations , 2024.

Linearly Decoding Refused Knowledge in Aligned Language Models Fine-tuning aligned language models compromises safety, even when users do not intend to! In The Twelfth International Conference on Learning Representations , 2024

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.973853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.973853Z digest=sha256:538dd608e5b1c2b246234ae3c733a8f17ae83e37b2fc7ff7adf83365c4e11e44

Observation 67379375-b768-4557-a664-7adce87c13e5 · outbound

This paper cites Safety alignment should be made more than just a few tokens deep.

Linearly Decoding Refused Knowledge in Aligned Language Models Safety alignment should be made more than just a few tokens deep

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.977277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.977277Z digest=sha256:fa29c5be27d998d234776a8cce25fff48b173f473a3ad06e1501c7cdadfc8a66

Observation fb4c71d6-3951-4f95-a5a5-ad814441f6bc · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Linearly Decoding Refused Knowledge in Aligned Language Models Direct preference optimization: Your language model is secretly a reward model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.980167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.980167Z digest=sha256:98792edba877fa5bc4af5ed12980a3a37d2828ec40b907681f545d22d3d8d623

Observation 1c158573-9d21-4878-b441-ee83b5fc65c0 · outbound

This paper cites Multi- task prompted training enables zero-shot task generalization.

Linearly Decoding Refused Knowledge in Aligned Language Models Multi- task prompted training enables zero-shot task generalization

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.749217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:27:34.982721Z digest=sha256:b4c325f2ab00fb104537334a5bfe38dfb75da04386eead4d41b927323cefdc4d

Observation 74e33e9a-f10b-4f6d-88cb-9af4f587b0f0 · outbound

This paper cites Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation.

Linearly Decoding Refused Knowledge in Aligned Language Models Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.985261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.985261Z digest=sha256:22574099aeeb0ee0d256dfe7d3a09588a884d01a81d5397257591349b96aa022

Observation 8e459ab7-d98c-43af-9a8b-2b7768a272c4 · outbound

This paper cites do anything now.

Linearly Decoding Refused Knowledge in Aligned Language Models do anything now

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.740959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:27:34.988437Z digest=sha256:f3ea74fbdd58273494f223b1e6da51194004c10c1db15e3fba5509dd522f4107

Observation 8ccd69e5-4941-4fa4-ac29-f2fcc1404203 · outbound

This paper cites Measuring Free-Form Decision-Making Inconsistency of Language Models in Military Crisis Simulations.

Linearly Decoding Refused Knowledge in Aligned Language Models Measuring Free-Form Decision-Making Inconsistency of Language Models in Military Crisis Simulations

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:27:35.197892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:27:34.991094Z digest=sha256:ba998e61a9376c06b9fc81aa2e917b94783192518c17a137e15a2b4971fc2762

Observation 89b52122-a299-46f1-ac45-1adf19bcba0f · outbound

This paper cites Large Language Models are Inconsistent and Biased Evaluators.

Linearly Decoding Refused Knowledge in Aligned Language Models Large Language Models are Inconsistent and Biased Evaluators

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.994224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.994224Z digest=sha256:913b17cf559b0951e0c8f5c2012c9d50cbef30639a519f666fb0f903ce5cdb76

Observation a88fe132-d912-464b-adb0-575e52fb0356 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Linearly Decoding Refused Knowledge in Aligned Language Models Gemma 2: Improving Open Language Models at a Practical Size

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.997511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.997511Z digest=sha256:6e35dfc047992e32cf461f615d20b05773610a27fdc0e36c6845e43843bd8e0d

Observation 358b4c94-e01b-4aea-9bf5-e6661b3756af · outbound

This paper cites What do you learn from context? Probing for sentence structure in contextualized word representations.

Linearly Decoding Refused Knowledge in Aligned Language Models What do you learn from context? Probing for sentence structure in contextualized word representations

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:35.000447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.000447Z digest=sha256:304691e30b050562f697b60e49940b6c31713b9395613ddd33cd600e2b047dcb

Observation d4edba58-579f-41fa-95c9-c325e80519d1 · outbound

This paper cites Attention is all you need.

Linearly Decoding Refused Knowledge in Aligned Language Models Attention is all you need

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.732940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:27:35.003388Z digest=sha256:ee8c02ce1b5a772f87b3ae00af92d674f4b9eb9707b974de0814fee5e08b1cf1

Observation f93403d1-7fd8-425f-b47d-251dcb60d155 · outbound

This paper cites White-box multimodal jailbreaks against large vision-language models.

Linearly Decoding Refused Knowledge in Aligned Language Models White-box multimodal jailbreaks against large vision-language models

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.724603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:27:35.005888Z digest=sha256:3cbb7e6ffe247bdf885edfb118ca9dc3a85da6b2867ba9058e9cedd3521a74aa

Observation 8a797ac8-2e88-4f95-8733-6c1002f776bd · outbound

This paper cites Jailbroken: How does LLM safety training fail? In Thirty-seventh Conference on Neural Information Processing Systems, 2023.

Linearly Decoding Refused Knowledge in Aligned Language Models Jailbroken: How does LLM safety training fail? In Thirty-seventh Conference on Neural Information Processing Systems, 2023

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.715626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:27:35.008684Z digest=sha256:e38cbaba23e4a0877c48aaae2446b3016e2fc3f22fd6d70fa6dc265ccb63d6db

Observation e52b27ab-0ec8-40c8-8109-dde23d632199 · outbound

This paper cites Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications.

Linearly Decoding Refused Knowledge in Aligned Language Models Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:35.011420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.011420Z digest=sha256:ea06e921fe358b486ebcfeaae0b1c16a25e3dbae96c8d330e982f742f095a3a6

Observation c8564b34-057e-407c-bf81-6ed1135155b4 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

Linearly Decoding Refused Knowledge in Aligned Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:35.014424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.014424Z digest=sha256:c62fecb01f32a4fedbdb7000c2661ade36247f02716b42c816114fb3663d4c39

Observation e8c3f6af-81dd-40ef-ab5a-21b0092acb6e · outbound

This paper cites Efficient streaming language models with attention sinks.

Linearly Decoding Refused Knowledge in Aligned Language Models Efficient streaming language models with attention sinks

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:35.017387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.017387Z digest=sha256:bd733edea9e953f82b3607c07474d194f6355ee99418915843156b00850fe179

Observation f17e0bc1-e48f-4aac-ab21-9c93e8167193 · outbound

This paper cites Assessing Hidden Risks of LLMs: An Empirical Study on Robustness, Consistency, and Credibility.

Linearly Decoding Refused Knowledge in Aligned Language Models Assessing Hidden Risks of LLMs: An Empirical Study on Robustness, Consistency, and Credibility

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:35.020100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.020100Z digest=sha256:a2e90e09ff8177e8a190d1dcc31a4ec4cd3c22196bd5a83eadd8eba1c8fc9a6c

Observation a33b0e54-ad19-4232-89ae-fc87d627e05a · outbound

This paper cites On the vulnerability of safety alignment in open-access LLMs.

Linearly Decoding Refused Knowledge in Aligned Language Models On the vulnerability of safety alignment in open-access LLMs

Reference 61

Resolution
verified exact
doi, observed 2026-08-06T21:27:35.073634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:27:35.022967Z digest=sha256:4c1b51c0686e862f118b8eaad08a5037b0c4119be29c24914c0fb3143450b925

Observation 549a16a1-0b82-45bf-83b0-0e8a7a9e5e49 · outbound

This paper cites Jailbreak Attacks and Defenses Against Large Language Models: A Survey.

Linearly Decoding Refused Knowledge in Aligned Language Models Jailbreak Attacks and Defenses Against Large Language Models: A Survey

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:35.025865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.025865Z digest=sha256:84e1af59cdf8f5f8aafaa195924db48a9621b6f6233043c8c9924d09fd42ba16

Observation 9836f397-8f1c-4026-b996-27b1613c361a · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

Linearly Decoding Refused Knowledge in Aligned Language Models Yi: Open Foundation Models by 01.AI

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:35.028823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.028823Z digest=sha256:2afbf9a696a6959fe67885a8c2ed6a5428016431354d6adf07abdd309115bc0e

Observation 5857d7c6-e88b-4079-bdda-232f7a5b79d9 · outbound

This paper cites Don’t listen to me: understanding and exploring jailbreak prompts of large language models.

Linearly Decoding Refused Knowledge in Aligned Language Models Don’t listen to me: understanding and exploring jailbreak prompts of large language models

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.702045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:27:35.031749Z digest=sha256:b96812e8d3302ab5c40ba658c80a292007e28ff483081eeb76ba61c1c15cd8fb

Observation bf6791de-0c35-48c4-9af0-7d7f436b471c · outbound

This paper cites Removing RLHF protections in GPT-4 via fine-tuning.

Linearly Decoding Refused Knowledge in Aligned Language Models Removing RLHF protections in GPT-4 via fine-tuning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:35.034422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.034422Z digest=sha256:7705d077f152521944f0bcd3b25ce46091b874f532239404cc3dc6730d239828

Observation dab3bda0-c3ff-45ee-b08c-c3cb7706f90b · outbound

This paper cites Lima: Less is more for alignment.

Linearly Decoding Refused Knowledge in Aligned Language Models Lima: Less is more for alignment

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:35.037339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.037339Z digest=sha256:39cd0cd700cd2aa94584e28761591a544de95dd4c755903bf68bae68f99b794b

Observation e3b2af77-ec94-45f6-9f49-31afebf2d95b · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Linearly Decoding Refused Knowledge in Aligned Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 67

Resolution
malformed identifier
no resolver link, observed 2026-08-06T21:27:35.040995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.040995Z digest=sha256:d5d56e4441e4a8057306e35d474c5af535ba40a832107afdb4e3b4e55318f40c

Observation 7b418754-5462-49f6-8997-fc1ee1e73aef · outbound

This paper cites an unresolved cited work.

Linearly Decoding Refused Knowledge in Aligned Language Models Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:27:35.798529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:27:34.959696Z digest=sha256:75200005ed0afedc9f77827b546ef820fbca66989c09985be8a9e550f7a186f7

Pith citing papers

Observation dbcb59b9-70ef-4595-91e4-d4fff01c1fd4 · inbound

HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models cites this paper.

HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models Linearly Decoding Refused Knowledge in Aligned Language Models

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:04.285500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T05:52:53.880521Z digest=sha256:7318edb20cacc0c75a878350b980a78f97eddeea8d574cd060daf85b47c8a577