Pith. sign in

Paper Citation Record · LEDGER

Towards eliciting latent knowledge from LLMs with mechanistic interpretability

As of 10 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 9 inbound Pith citation observations for arXiv:2505.14352.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14352 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:40:07.861843Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T22:20:51.989267Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 1f0d0b3f-737f-484a-85b1-811e08771f24 · outbound

This paper cites write newline.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:40:05.057222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:40:05.057222Z digest=sha256:7eedc2d6741fdba355cfd888606e040c0e2f10578dc59f3e5f39ed4e028c860a

Observation 511f238a-e3aa-4478-8dce-1dd08fed1d9a · outbound

This paper cites Tell me about yourself: LLMs are aware of their learned behaviors.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Tell me about yourself: LLMs are aware of their learned behaviors

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:40:05.176888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:40:05.176888Z digest=sha256:3adadf4d93b7090429e855a5a83bf1c1c5811b29220034d87323c02330973fc0

Observation 0b876a4f-14aa-499c-9a3a-99deede786ab · outbound

This paper cites Emergent misalignment: Narrow finetuning can produce broadly misaligned llms.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Emergent misalignment: Narrow finetuning can produce broadly misaligned llms

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:40:05.246250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:40:05.246250Z digest=sha256:b516bfbbc3f5102b4e49388551907c9edb74d9e3ef64035bfc39865e0d15f936

Observation ceccbf25-6bf5-4eaf-b38f-82f0527d3524 · outbound

This paper cites E., Hume, T., Carter, S., Henighan, T., and Olah, C.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability E., Hume, T., Carter, S., Henighan, T., and Olah, C

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:40:05.336601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:40:05.336601Z digest=sha256:f254eceeda285095bd59577a5ec780c10624ce175c898eb05f18c4c15149c2b4

Observation ce6d09c6-23cc-41fb-a75f-90b362471e25 · outbound

This paper cites Eliciting latent knowledge: How to tell if your eyes deceive you, 2021.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Eliciting latent knowledge: How to tell if your eyes deceive you, 2021

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:40:10.979962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T15:40:05.419110Z digest=sha256:7c38bd2d515cccfed15ee331efbcd3e41be403ced56f0f773d156dec6a803ff3

Observation 8ab3fae9-90a7-4b18-935f-485cfe3fc634 · outbound

This paper cites F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:40:05.533081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:40:05.533081Z digest=sha256:8edf01b5ee01f49d269afed747f773871b2dcca850ca21599308cecf4459021a

Observation 5ad1cdce-3136-4de4-8d8a-f201ef897d63 · outbound

This paper cites Sparse Autoencoders Find Highly Interpretable Features in Language Models.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Sparse Autoencoders Find Highly Interpretable Features in Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:40:05.657362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:40:05.657362Z digest=sha256:455bc9a5873d853fcc61626b56120d118afceb475baadcb24dd54f6506953142

Observation 9020be08-2811-4fb1-9c81-85823fbfd8f9 · outbound

This paper cites Safe RLHF : Safe reinforcement learning from human feedback.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Safe RLHF : Safe reinforcement learning from human feedback

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:40:10.789023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T15:40:05.758300Z digest=sha256:3f13f30f306a99f1c189696b99e3a31092c079fcf2f63d555bdddb2e805e8c76

Observation 2f1fc3d7-4b82-437e-9b8f-da650e9b9df9 · outbound

This paper cites Qlora: Efficient finetuning of quantized llms.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Qlora: Efficient finetuning of quantized llms

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:40:05.816404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:40:05.816404Z digest=sha256:cf87ef87704264f7698d07ca271d6edb1d88fc6c01229bb3eee2e71b450989f0

Observation d8419faf-5ad4-4c82-9089-e10388bd00af · outbound

This paper cites Pal: Program-aided language models.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Pal: Program-aided language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:40:10.544027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T15:40:05.888242Z digest=sha256:c77f503acbd189160e0b3ec593da0b9e9d3f028d71e93c19dabc2f4f62680685

Observation f78732ec-7c97-41db-acb6-5b6f3c15ed75 · outbound

This paper cites Gemini 2.5 flash, 2025 a.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Gemini 2.5 flash, 2025 a

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:40:10.346882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T15:40:05.969984Z digest=sha256:1b83255bc1254a2dbfcfe810365dca54b530489dd22622826035c959df8b12fa

Observation c5f075c8-433c-4927-8594-ea3ab04015d6 · outbound

This paper cites Gemini 2.5 pro preview, 2025 b.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Gemini 2.5 pro preview, 2025 b

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:40:10.136279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T15:40:06.061216Z digest=sha256:cc041a318138b1a640391bde78d2800abf4d6be3edd690bc9c24afc0a041c747

Observation b5635376-a01a-460b-8ad0-e949039a0f6c · outbound

This paper cites Alignment faking in large language models.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Alignment faking in large language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:40:06.156563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:40:06.156563Z digest=sha256:9eeebb6038de48efac893592c6fc112ca09f93f44fb7893f11ff561a398a2f46

Observation bc2b0c15-274a-432d-a95d-c421c32a8667 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:40:06.249074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:40:06.249074Z digest=sha256:005606b9550d93f2856c497812d56029e898d78f6a81f97033d4d3563020a85f

Observation 6dc4628e-ae13-40c4-a7fa-3fa88ab731d5 · outbound

This paper cites and Lee, S.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability and Lee, S

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:40:09.952659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T15:40:06.332163Z digest=sha256:818d79004367816e7d493413a5636e8ecf78ebc02a8142cfc32aa6e739f9d0bb

Observation 16de2527-ffb6-46a0-801c-457ab90bc719 · outbound

This paper cites u chemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., G \.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability u chemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., G \

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:40:09.818736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T15:40:06.413140Z digest=sha256:8ebd45fbbb29481f79fba355b45da0a9326286ce80ec145aaa208f29673f6afe

Observation 480279d4-bbad-4752-9cec-e4b7594543c9 · outbound

This paper cites M., Bommarito, M.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability M., Bommarito, M

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:40:09.674256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T15:40:06.511385Z digest=sha256:99a29cf250c58b888cf1275e6ee0c2e247fdd852f0a38a9733d2121d0bf8e984

Observation 13023214-b0bf-4312-be60-747afb101da2 · outbound

This paper cites Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:40:06.584693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:40:06.584693Z digest=sha256:ba85cb7bd6d27b0ca21472bf35601a106f226bd3b51f70681d7631d3852bc4f6

Observation 4523999c-11ea-4675-8629-f91213f4ebe6 · outbound

This paper cites Auditing language models for hidden objectives.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Auditing language models for hidden objectives

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:40:06.682295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:40:06.682295Z digest=sha256:e41b38ae08aff96383e2412ef43a4cb72daa1e724eede4274d078d90afa9c8fb

Observation ac39182b-c4bf-4d4d-9319-78b89bf67f87 · outbound

This paper cites Frontier Models are Capable of In-context Scheming.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Frontier Models are Capable of In-context Scheming

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:40:06.782520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:40:06.782520Z digest=sha256:0e3222d6f981c735432c4328e3e6198aae379fbc4915d012077cbc88d4d06016

Observation 68671958-3532-4304-b1c0-16ed2186ead8 · outbound

This paper cites interpreting gpt: the logit lens.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability interpreting gpt: the logit lens

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:40:09.522399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T15:40:06.880431Z digest=sha256:dadd9d8812f80e94324809317d831db11ffe5345d672f510e25f4578d2e65c5a

Observation efdf8ba3-4011-479b-b770-f283ec91767b · outbound

This paper cites Learning to reason with llms, 2024b.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Learning to reason with llms, 2024b

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:40:09.328159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T15:40:06.996506Z digest=sha256:151409b712f2f881db64bc983c3747fb9d5f3771d9ae9e3f188f8f4a4727c0fd

Observation 1a02bb6a-e6a2-418c-8618-b60bfa998db6 · outbound

This paper cites Training language models to follow instructions with human feedback.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Training language models to follow instructions with human feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:40:07.112022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:40:07.112022Z digest=sha256:7f3b6d26903299e4047f6517e63a87625fa78fc5472ae35f39cd52f11978bf30

Observation 0f6bfe05-6f84-46d8-91cd-3a1ca2a3bbbe · outbound

This paper cites D., Ermon, S., and Finn, C.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability D., Ermon, S., and Finn, C

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:40:09.166740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T15:40:07.229971Z digest=sha256:1f2f8bb4f667ced7a6ec49c551af0cac137f7a8f5cd7c3bd03c56339be768b53

Observation 49d393c5-a973-41d8-ae91-34bfd73cfeda · outbound

This paper cites LLM-Microscope: Uncovering the Hidden Role of Punctuation in Context Memory of Transformers.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability LLM-Microscope: Uncovering the Hidden Role of Punctuation in Context Memory of Transformers

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:40:08.100475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T15:40:07.346882Z digest=sha256:bf44aa1cf937e4fc2b79cc9928ea9a14508493380e0f39aae0e095064347d0cd

Observation 7a6f0fd4-9efd-4e5a-82cd-e68e69c925ce · outbound

This paper cites Top 1000 english nouns, 2019.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Top 1000 english nouns, 2019

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:40:09.063228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T15:40:07.459790Z digest=sha256:af5fc3f327fa14dcc2c23cd4cef190ffd85ce8824cebbb1185107719e962ed61

Observation 87906699-feee-48b8-9fe9-90d16caa5973 · outbound

This paper cites Large language models can strategically deceive their users when put under pressure.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Large language models can strategically deceive their users when put under pressure

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:40:08.902763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T15:40:07.539751Z digest=sha256:9763bf9091edabeea14ddd8a1660d90194e5dba06343314becf6432eff1ecce6

Observation 12d7963d-ac67-4556-8d4c-d04a4424e592 · outbound

This paper cites an unresolved cited work.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:40:07.649664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:40:07.649664Z digest=sha256:58dba48e502485f9f38f1c9c820e3a1e4661d70e11ddf9dcf5a5932b3997da99

Observation 1e8a314b-57e8-4e7d-923c-ea400b0faca4 · outbound

This paper cites an unresolved cited work.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:40:08.587883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T15:40:07.773307Z digest=sha256:35ddad9f2331919e41fe0c14216ca8b7d8eb882e7cb9a61bab5b0642e8fed680

Observation 6ee307d4-e2d9-4355-ad8b-597a447b6ec0 · outbound

This paper cites N., Kaiser, ., and Polosukhin, I.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability N., Kaiser, ., and Polosukhin, I

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:40:07.861843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:40:07.861843Z digest=sha256:79cee4bf708f50daafcb89691fb3a216463ae65860fa828e775703e1144fb272

Pith citing papers

Observation b4bd3ba3-9c28-48ce-82c3-e71a7d58af84 · inbound

Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy cites this paper.

Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy Towards eliciting latent knowledge from LLMs with mechanistic interpretability

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T22:20:51.989267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T22:20:51.989267Z digest=sha256:7fded3afcf5944f25da3ca32c1c68a67dc195bd01c3f04132425a77e1587e0bd

Observation d7e6064b-8eca-414d-913f-f9b0a427f6ea · inbound

DECOR: Auditing LLM Deception via Information Manipulation Theory cites this paper.

DECOR: Auditing LLM Deception via Information Manipulation Theory Towards eliciting latent knowledge from LLMs with mechanistic interpretability

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:28:05.444725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T06:27:10.445757Z digest=sha256:24d265f7c285a688ddef6a3cf15260007081c2bd8f81e4b38c75f9f11373c16b

Observation 71882747-c6a5-4999-bdb2-b6a74edf1618 · inbound

Building Better Activation Oracles cites this paper.

Building Better Activation Oracles Towards eliciting latent knowledge from LLMs with mechanistic interpretability

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:24:45.078140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-30T14:19:37.262201Z digest=sha256:a4067c713e050d03c6ec149b5935b1e116d89d51acce2e6aa69d234f065c051a

Observation bf2dbefc-e764-4a38-9b8d-6660d1c8a799 · inbound

"Did you lie?" Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms cites this paper.

"Did you lie?" Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms Towards eliciting latent knowledge from LLMs with mechanistic interpretability

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:48:02.307763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T09:51:16.969884Z digest=sha256:1ca61dfb6b71da2ab3fbf3761c6ac1141c0956e4dbdd3f6d4e4eefa57fdb91da

Observation a44a5147-ccd4-4129-a339-7880f3ca094a · inbound

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning cites this paper.

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning Towards eliciting latent knowledge from LLMs with mechanistic interpretability

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:38:55.997805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T01:13:11.483599Z digest=sha256:e34ba9ade775a504eea9507639f0a994644e75e3e3a790f8d610ef0dc416cb8b

Observation 3a8b2a06-b7ed-42ab-b5fc-883db645c7ee · inbound

"Don't Say It!": Constraints, Compliance, and Communication when Language Models Play Taboo cites this paper.

"Don't Say It!": Constraints, Compliance, and Communication when Language Models Play Taboo Towards eliciting latent knowledge from LLMs with mechanistic interpretability

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:16:58.191057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-02T13:11:53.361408Z digest=sha256:ad1d43001896c48462fe242e4017c1a123a566044d92fd2c98616a5341ecfb9f

Observation b9cd5712-4558-422c-9c58-77fc742c2630 · inbound

The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology cites this paper.

The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology Towards eliciting latent knowledge from LLMs with mechanistic interpretability

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:07:08.173995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-02T15:57:48.589980Z digest=sha256:caedb820434a7f830881ca946260eaa3761c7f7b61ae0c306448d6fc36c860ca

Observation 76e1f9be-c3ec-4f0e-9a94-c8040b4a08fe · inbound

MUX: Continuous Reasoning via Multiplexed Tokens cites this paper.

MUX: Continuous Reasoning via Multiplexed Tokens Towards eliciting latent knowledge from LLMs with mechanistic interpretability

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:03.078595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:43:03.078595Z digest=sha256:4a1c101e9bd47de31e7b1dfdba57b66ef106ee390d9fca4f640a047b0a45b8af

Observation df4e3907-474c-4909-9ba9-029b28afc797 · inbound

When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles cites this paper.

When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles Towards eliciting latent knowledge from LLMs with mechanistic interpretability

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-30T23:56:50.935284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T23:56:50.935284Z digest=sha256:475f74e5f27404a234c282c7dd4b13d0d10760f42c97dab8a15920703d7aaa59