Pith. sign in

Paper Citation Record · LEDGER

Understanding (Un)Reliability of Steering Vectors in Language Models

As of 18 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 11 inbound Pith citation observations for arXiv:2505.22637.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22637 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:10:08.895546Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:00:38.060792Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:49:57.002506Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact2
  • verified fuzzy6
  • unresolved28
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bd450999-42e0-4f54-877d-862e670f8720 · outbound

This paper cites write newline.

Understanding (Un)Reliability of Steering Vectors in Language Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:03.289827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:03.289827Z digest=sha256:b572b7315f5714a4f126bbe9ab38c68f07c588088b22465566cb2bbf686092da

Observation b12abee5-3591-465f-a777-29e583a26be2 · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

Understanding (Un)Reliability of Steering Vectors in Language Models Refusal in Language Models Is Mediated by a Single Direction

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:03.419878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:03.419878Z digest=sha256:fe082b9d94d1417b653dac27baaabfec5a7a179214b1f5251447a2cab31c8d6b

Observation 4955598a-7f3d-400c-ae32-003701625648 · outbound

This paper cites Cats: Customizable abstractive topic-based summarization.

Understanding (Un)Reliability of Steering Vectors in Language Models Cats: Customizable abstractive topic-based summarization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:03.548101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:03.548101Z digest=sha256:8c4bb6cbd6ab8ce846702811a62751a9521174df31e2dae6e539b5357392985b

Observation 6d857f70-6587-4148-a88f-847eab514349 · outbound

This paper cites NEWTS : A corpus for news topic-focused summarization.

Understanding (Un)Reliability of Steering Vectors in Language Models NEWTS : A corpus for news topic-focused summarization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:03.715918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:03.715918Z digest=sha256:0340f3939102d11f04528b3fdcaf30c7cd24b00df148f229b43bed8d58aae2c9

Observation 8ae392d8-655e-4c49-a154-3e3524d501bf · outbound

This paper cites Controllable Topic-Focused Abstractive Summarization.

Understanding (Un)Reliability of Steering Vectors in Language Models Controllable Topic-Focused Abstractive Summarization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:03.927972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:03.927972Z digest=sha256:dfe38070ab29aa02cdee5fe406985c41b63e0acfc6ee8dca22f321c1c3b9c0ea

Observation 7062c719-3fd9-49f0-ac2d-cf6212c10207 · outbound

This paper cites Text simplification via adaptive teaching.

Understanding (Un)Reliability of Steering Vectors in Language Models Text simplification via adaptive teaching

Reference 6

Resolution
verified exact
doi, observed 2026-08-07T13:10:09.638480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:10:04.074437Z digest=sha256:5e2c54de7d6ec38678a5dd4401c6ab2fd1a46f7fef36798c4740887ca3662252

Observation e5da340c-198b-477c-b853-a1b095413748 · outbound

This paper cites SIMSUM : Document-level text simplification via simultaneous summarization.

Understanding (Un)Reliability of Steering Vectors in Language Models SIMSUM : Document-level text simplification via simultaneous summarization

Reference 7

Resolution
verified exact
doi, observed 2026-08-07T13:10:09.260944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:10:04.258231Z digest=sha256:40ae951b3819f450ffbd792aa9a2818c54811b37410b6af052eca43660509c77

Observation 2a3eaa2a-7361-4b48-a088-8e8edad3ab5a · outbound

This paper cites A sober look at steering vectors for llms.

Understanding (Un)Reliability of Steering Vectors in Language Models A sober look at steering vectors for llms

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:12.329122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:10:04.453529Z digest=sha256:8a9ec449ab88fa18d2ad5c10b00861ed57d4f9b55b47aa340e330b51a32986e2

Observation 7110b613-adc6-48fa-a5cd-ec86b0bfb4b0 · outbound

This paper cites Comparing Bottom-Up and Top-Down Steering Approaches on In-Context Learning Tasks.

Understanding (Un)Reliability of Steering Vectors in Language Models Comparing Bottom-Up and Top-Down Steering Approaches on In-Context Learning Tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:04.575056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:04.575056Z digest=sha256:d49d680a5143556098f09b5090a95543950acae21b444cbdc2e91c3bb8d990d7

Observation 30fa9f24-32c4-46c4-b4aa-88e21828886b · outbound

This paper cites Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization.

Understanding (Un)Reliability of Steering Vectors in Language Models Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:04.717862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:04.717862Z digest=sha256:f69d5f6c7e8a4244925ddd522607a414a26dcb1b7de0e69cf6cbe6bb291fc03b

Observation 3701d840-1f93-448f-a62d-ef5ace6f7da4 · outbound

This paper cites In-context learning creates task vectors.

Understanding (Un)Reliability of Steering Vectors in Language Models In-context learning creates task vectors

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:04.898405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:04.898405Z digest=sha256:6d9a6408d2fa5ea99d7fa7a45b3af7e52e8355a9641953357780c780ce350ce8

Observation 11affa31-d0f5-42b0-847d-02f596676574 · outbound

This paper cites Measuring massive multitask language understanding.

Understanding (Un)Reliability of Steering Vectors in Language Models Measuring massive multitask language understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:04.974038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:04.974038Z digest=sha256:0832ab4e82bc0ade68a3a26175bf15aae7d4229364e20c28cc40c7df812f601d

Observation 87fc6599-bc67-431b-ba03-3660781afacb · outbound

This paper cites Style Vectors for Steering Generative Large Language Models.

Understanding (Un)Reliability of Steering Vectors in Language Models Style Vectors for Steering Generative Large Language Models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:11.949019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:10:05.093975Z digest=sha256:02a6e65872a5855ad6706fd8828921128769fb959c1b05c0dddc746d79fb3d2d

Observation 514b7c2d-d754-4bc9-8e5e-7bd005f009f9 · outbound

This paper cites Steering clear: A systematic study of activation steering in a toy setup.

Understanding (Un)Reliability of Steering Vectors in Language Models Steering clear: A systematic study of activation steering in a toy setup

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:11.664519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:10:05.218312Z digest=sha256:53b7b10bdb11c0167505017fa7cf6effdd36d56f1b96585b48e776f08526e1ab

Observation 7da431f6-36ab-49cf-a44e-f9ce9e74cf2a · outbound

This paper cites Inference-time intervention: Eliciting truthful answers from a language model.

Understanding (Un)Reliability of Steering Vectors in Language Models Inference-time intervention: Eliciting truthful answers from a language model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:05.387255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:05.387255Z digest=sha256:6fb8df6739f75fdd35793a28335621571fa8059ae5f7b42dfc229ec71c0bdfa7

Observation 49a3bbf8-7c4b-4cd6-bbe5-b215a7fb408f · outbound

This paper cites The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets.

Understanding (Un)Reliability of Steering Vectors in Language Models The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:05.520729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:05.520729Z digest=sha256:8b00151d79df5ac9d18c5954f5a24523bbf39eaf1b2b4fa1a0b27563bb942927

Observation 58eb0b15-9f5f-4252-a074-fa17f2772e13 · outbound

This paper cites The geometry of truth: Emergent linear structure in large language model representations of true/false datasets.

Understanding (Un)Reliability of Steering Vectors in Language Models The geometry of truth: Emergent linear structure in large language model representations of true/false datasets

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:11.347820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:10:05.648088Z digest=sha256:3a34c0623bacd384ef7622bbe02acf6a986af423179511671d0e8cb92e84df03

Observation 0bc45b64-edcd-4965-a435-cc549a2067eb · outbound

This paper cites Refusal in LLMs is an Affine Function.

Understanding (Un)Reliability of Steering Vectors in Language Models Refusal in LLMs is an Affine Function

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:05.846705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:05.846705Z digest=sha256:e6b4b60acfe7192865a043a09a66bf6888ec702ffadcb07c3c1c08f9ccd826cb

Observation 89d8d63f-d64f-4a5e-93a8-1b2ced9f909a · outbound

This paper cites Towards Reliable Evaluation of Behavior Steering Interventions in LLMs.

Understanding (Un)Reliability of Steering Vectors in Language Models Towards Reliable Evaluation of Behavior Steering Interventions in LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:06.057729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:06.057729Z digest=sha256:280da44f2af88e0fb9248b8d0f9b679eed7f3c06c4431364ea6b2302c9a96a0e

Observation 116acc5e-5900-466e-bca0-86616e6a68de · outbound

This paper cites Steering llama 2 via contrastive activation addition.

Understanding (Un)Reliability of Steering Vectors in Language Models Steering llama 2 via contrastive activation addition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:06.173597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:06.173597Z digest=sha256:4e5595c32fedd8309547310c692a200026968c00b15725d86b781f7583b36752

Observation dad2db02-e247-492c-b995-e6a20608a38d · outbound

This paper cites Discovering Language Model Behaviors with Model-Written Evaluations, Toronto, Canada, July 2023.

Understanding (Un)Reliability of Steering Vectors in Language Models Discovering Language Model Behaviors with Model-Written Evaluations, Toronto, Canada, July 2023

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:06.293291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:06.293291Z digest=sha256:4bdbe1dc31099c259082a5b5d420594470ec9cef69bba1683d69ef86730761ce

Observation 83e09f5f-6664-46c3-bb9c-07f5ade319fd · outbound

This paper cites Representation surgery: Theory and practice of affine steering.

Understanding (Un)Reliability of Steering Vectors in Language Models Representation surgery: Theory and practice of affine steering

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:10.985849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:10:06.461530Z digest=sha256:47f1f3989182521ccd666ac1e41ba715fe4b17c901a955925ac468dec64acecd

Observation e234c100-7387-4af3-943b-c7b0bbf7c2df · outbound

This paper cites an unresolved cited work.

Understanding (Un)Reliability of Steering Vectors in Language Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:10.614343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:10:06.605117Z digest=sha256:d522a947a3b0925383d34b53f11b6c546379e8d15786e25ca8c6bbaec60ba1f2

Observation df5ab5eb-3fdb-4ef0-95a3-796919775ffb · outbound

This paper cites Extracting Latent Steering Vectors from Pretrained Language Models.

Understanding (Un)Reliability of Steering Vectors in Language Models Extracting Latent Steering Vectors from Pretrained Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:06.744243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:06.744243Z digest=sha256:4184f373f386390298c5a252d40f734324a13704a528c75569d8cbb6011ae6d7

Observation e27cad75-e919-4e0a-968b-829da0f359a5 · outbound

This paper cites Analysing the generalisation and reliability of steering vectors.

Understanding (Un)Reliability of Steering Vectors in Language Models Analysing the generalisation and reliability of steering vectors

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:06.928187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:06.928187Z digest=sha256:94eaf38c441f5d2ddd8e85d4e57b8cf9d263582baf7059b03b246d88844b5c28

Observation 012e9b4e-7a8b-4feb-b47d-7865abd582f6 · outbound

This paper cites Linear Representations of Sentiment in Large Language Models.

Understanding (Un)Reliability of Steering Vectors in Language Models Linear Representations of Sentiment in Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:07.072130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:07.072130Z digest=sha256:7795f7e013fddfdb12921ae5ea5214dfac655ac66e985a73d4b8f6c39710b467

Observation 0b1763f6-acb5-41dc-871c-d36007e80b69 · outbound

This paper cites Hollinsworth, Atticus Geiger, and Neel Nanda.

Understanding (Un)Reliability of Steering Vectors in Language Models Hollinsworth, Atticus Geiger, and Neel Nanda

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:07.202080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:07.202080Z digest=sha256:91c4a87c74f14a2a3516a753e84538b834ea9e5e7fe1612a1eb38fb8a6381241

Observation 96dbf721-bcbc-4538-8d96-a39327a76187 · outbound

This paper cites Function Vectors in Large Language Models.

Understanding (Un)Reliability of Steering Vectors in Language Models Function Vectors in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:07.335261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:07.335261Z digest=sha256:ca30ddb5d74dcd46d27b963ff1b8ca8113743c17c1dd4013dfe5ec485fc516a2

Observation 1608ff49-3205-4878-9d8d-633a7ba8e9d4 · outbound

This paper cites Function vectors in large language models.

Understanding (Un)Reliability of Steering Vectors in Language Models Function vectors in large language models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:07.486495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:07.486495Z digest=sha256:216fa3926c11979204853cfcb272d2934f69158421bca6e8830180e7327cb2ec

Observation 71d1e75e-8ae7-4697-9eaf-bf08c20d67c3 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Understanding (Un)Reliability of Steering Vectors in Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:07.663220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:07.663220Z digest=sha256:67f90db961626863c6189c3bae4218995fbee79df19230be78d4e10c3c7b14e7

Observation f0ddce0c-bb20-4bec-a1c4-608a2c6799e3 · outbound

This paper cites Activation addition: Steering language models without optimization.

Understanding (Un)Reliability of Steering Vectors in Language Models Activation addition: Steering language models without optimization

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:10.366568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:10:07.830192Z digest=sha256:5d6f77d69838e16137fc6c8c01ea6489fbc21abe98d6492485b3b6031c0c0bb4

Observation b11d6994-d56f-4cb9-ba92-94582b225ee5 · outbound

This paper cites Controllable text summarization: Unraveling challenges, approaches, and prospects - a survey.

Understanding (Un)Reliability of Steering Vectors in Language Models Controllable text summarization: Unraveling challenges, approaches, and prospects - a survey

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:07.935345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:07.935345Z digest=sha256:68b166ba1b9734bd6ac74ff3bcf62f4a2aeca1761cb5486901ac587eae3911f8

Observation e52a79ce-8217-4802-8fe8-e87911df09bb · outbound

This paper cites A comprehensive survey on process-oriented automatic text summarization with exploration of llm-based methods, 2025.

Understanding (Un)Reliability of Steering Vectors in Language Models A comprehensive survey on process-oriented automatic text summarization with exploration of llm-based methods, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:08.060573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:08.060573Z digest=sha256:f722f80a1ca761aca978f249bb21df590e578d4065d8ea3a46f11fe4154df354

Observation 30e6f97d-3aac-4854-8e10-5df3e4d72504 · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Understanding (Un)Reliability of Steering Vectors in Language Models Representation Engineering: A Top-Down Approach to AI Transparency

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:08.399853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:08.399853Z digest=sha256:0e65f22fba9d91bc6c6078fb449d81abd15e1d7dae46721f6749be9a9650cc28

Observation 8b8eca33-f2ad-4632-8d6c-55f6e08636e5 · outbound

This paper cites @esa (Ref.

Understanding (Un)Reliability of Steering Vectors in Language Models @esa (Ref

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:08.563640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:08.563640Z digest=sha256:90c7c5150f379e9401a4f2708c72402ee7ddfdbcc6891e5f9ccd6c29e1e3f6e2

Observation 400ebaa9-c888-45ae-ae67-5b09f2f630c2 · outbound

This paper cites an unresolved cited work.

Understanding (Un)Reliability of Steering Vectors in Language Models Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:08.737088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:08.737088Z digest=sha256:18146a9033c46b8e8b0ae5b33472f8180565fefd985cf829b7757a76ba658230

Observation d56fbc92-ed8d-4e88-b60b-b678e757acee · outbound

This paper cites 3!( 4˜ "3!( 4˒.

Understanding (Un)Reliability of Steering Vectors in Language Models 3!( 4˜ "3!( 4˒

Reference 38

Resolution
malformed identifier
no resolver link, observed 2026-08-07T13:10:08.895546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:08.895546Z digest=sha256:11bf3bdbc19b77e61ae021e7b3e0c45a2fb2ec0a455fbd13612b4f1f62dfa09f

Pith citing papers

Observation 70274d98-43c8-4562-bdd8-b51bc6ccbdfe · inbound

RouteHijack: Routing-Aware Attack on Mixture-of-Experts LLMs cites this paper.

RouteHijack: Routing-Aware Attack on Mixture-of-Experts LLMs Understanding (Un)Reliability of Steering Vectors in Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:13.825157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-09T19:22:00.217729Z digest=sha256:8f34136b6d2af5e0c5a466447b0331aac243a05add9239734468bbbb91fd1728

Observation 69162c32-1c61-409f-b86e-267231d7d9cd · inbound

Prompt-Activation Duality: Improving Activation Steering via Attention-Level Interventions cites this paper.

Prompt-Activation Duality: Improving Activation Steering via Attention-Level Interventions Understanding (Un)Reliability of Steering Vectors in Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:16:25.939080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T04:27:15.490694Z digest=sha256:83599f99d747ee51dfa3d77665a22ba932200486e44a1c964566d3b96aac26af

Observation 71c24dca-f1f9-44a7-b847-9c1a5afd1c8e · inbound

Prompt-Activation Duality: Improving Activation Steering via Attention-Level Interventions cites this paper.

Prompt-Activation Duality: Improving Activation Steering via Attention-Level Interventions Understanding (Un)Reliability of Steering Vectors in Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:29:47.076007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T05:27:59.011749Z digest=sha256:8fee71e1dc72f1fa0d165dbd0511b7a2d8a4f0d2d53b3687a8807dc27dec1f3b

Observation 1e77b7e1-96eb-4d6f-92e0-237f38dca4b5 · inbound

When Is Rank-1 Steering Cheap? Geometry, Granularity, and Budgeted Search cites this paper.

When Is Rank-1 Steering Cheap? Geometry, Granularity, and Budgeted Search Understanding (Un)Reliability of Steering Vectors in Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:39:10.703623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T22:36:12.466550Z digest=sha256:20f163465673a0c1155dbf3b0624258839cd5632202d7acbf7f310603a344bcb

Observation b48fd55b-ee3c-4512-a253-0d40a3b82cbb · inbound

When Is Rank-1 Steering Cheap? Geometry, Granularity, and Budgeted Search cites this paper.

When Is Rank-1 Steering Cheap? Geometry, Granularity, and Budgeted Search Understanding (Un)Reliability of Steering Vectors in Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:01:23.525559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T09:59:32.201925Z digest=sha256:b4bd4f4c6cde5c039fcac099c713600b2e0866e7134ccf474b24ffa13934b729

Observation 3f421cf8-48ca-41b2-a447-55420c0c24c7 · inbound

Temporal Preference Concepts and their Functions in a Large Language Model cites this paper.

Temporal Preference Concepts and their Functions in a Large Language Model Understanding (Un)Reliability of Steering Vectors in Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:05:47.215257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T22:16:47.743387Z digest=sha256:dd08660bd120818b3aa43dc96aad4ff7d4302ee35481daa7b07579b8d1aad3bc

Observation 19796388-a601-4242-a1c8-02fb3a196a44 · inbound

Temporal Preference Concepts and their Functions in a Large Language Model cites this paper.

Temporal Preference Concepts and their Functions in a Large Language Model Understanding (Un)Reliability of Steering Vectors in Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T17:03:44.315006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:03:44.315006Z digest=sha256:328b9d32a33b1ed76158f651e1ad176c4ecc8459a08dc2c1bc39ec04aaecc4ab

Observation d460774b-6a78-4912-8e56-f903b96b75df · inbound

Adversarial Robustness of Activation Steering in Large Language Models cites this paper.

Adversarial Robustness of Activation Steering in Large Language Models Understanding (Un)Reliability of Steering Vectors in Language Models

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:37:09.432517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T22:30:36.839964Z digest=sha256:e6d1a856ecd518dbcdda369602d372657033859ea3ec06ede66e53bad592eaed

Observation 4d6c0238-a48e-4825-ae00-cabcc32cf3ed · inbound

Detecting and Controlling Sycophancy with Cascading Linear Features cites this paper.

Detecting and Controlling Sycophancy with Cascading Linear Features Understanding (Un)Reliability of Steering Vectors in Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:49:57.005467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T01:24:56.600219Z digest=sha256:cb98dd169dac673d6b3cf79a536ae714ab9013df958cdd2321b4b7da3c36a096

Observation ab15ab25-369b-452d-91c7-3fb5f45bc50c · inbound

Conditional Optimal Bridge for Riemannian Activation Steering cites this paper.

Conditional Optimal Bridge for Riemannian Activation Steering Understanding (Un)Reliability of Steering Vectors in Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T11:08:17.571722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:08:17.571722Z digest=sha256:eac2e8b3671687e28c95651e02f748a009ffda112e1dda3e074c1a24ab67cc7d

Observation b94d2b41-60e5-42a2-a7e1-dbca2818c079 · inbound

Inverted Detection and Control in Steering Vectors cites this paper.

Inverted Detection and Control in Steering Vectors Understanding (Un)Reliability of Steering Vectors in Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T15:00:38.060792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:00:38.060792Z digest=sha256:cc8031cd11eafc9047a080e75afa9018643ccfd5e33dbb4120b0055ed2920769