Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:38:12.261958Z
Paper Citation Record · LEDGER
As of 24 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 4 inbound Pith citation observations for arXiv:2504.20271.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:38:12.261958Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T18:33:08.616839Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
24 of 24 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 31425d2e-627d-4bbc-9e6c-5835df5f3dfd · outbound
Investigating task-specific prompts and sparse autoencoders for activation monitoring Understanding intermediate layers using linear classifier probes
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36ea0f9a-3b0e-4770-a6dd-f3b9fc04c57f · outbound
Investigating task-specific prompts and sparse autoencoders for activation monitoring The Internal State of an LLM Knows When It's Lying
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d374e25b-a50e-46e2-85a2-5ce7f3fd72fb · outbound
Investigating task-specific prompts and sparse autoencoders for activation monitoring Interpretability and Analysis in Neural NLP
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61f4026c-9446-46dd-a2b4-da0212def674 · outbound
Investigating task-specific prompts and sparse autoencoders for activation monitoring Using Dictionary Learning Features as Classifiers , October 2024
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a2cac27f-7f58-424d-9b9e-c9fca18c416a · outbound
Investigating task-specific prompts and sparse autoencoders for activation monitoring Bitterman
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ef6fc6c-fa34-45a6-8353-bf8ece5e5257 · outbound
Investigating task-specific prompts and sparse autoencoders for activation monitoring Scaling and evaluating sparse autoencoders
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9db1b8e8-fb36-41c4-8e55-c8f1550e3a91 · outbound
Investigating task-specific prompts and sparse autoencoders for activation monitoring Detecting Strategic Deception Using Linear Probes
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b6c2c0a-b9f2-41ea-8a1e-e45c6e83d902 · outbound
Investigating task-specific prompts and sparse autoencoders for activation monitoring Estimating Knowledge in Large Language Models Without Generating a Single Token
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a574b75-8ac1-4083-b4d5-c7cd70aff09b · outbound
Investigating task-specific prompts and sparse autoencoders for activation monitoring Finding Neurons in a Haystack: Case Studies with Sparse Probing
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07846e73-1896-4dea-a1a5-d4db398fdbcd · outbound
Investigating task-specific prompts and sparse autoencoders for activation monitoring Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b9c0478-acbb-489f-92ca-62ff44cb76bd · outbound
Investigating task-specific prompts and sparse autoencoders for activation monitoring Saes (usually) transfer between base and chat models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 19ac179a-3c89-4e98-af3d-968f075614de · outbound
Investigating task-specific prompts and sparse autoencoders for activation monitoring Cognitive Dissonance: Why Do Language Model Outputs Disagree with Internal Representations of Truthfulness?
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d21aa153-11fe-4432-8101-762d0812acad · outbound
Investigating task-specific prompts and sparse autoencoders for activation monitoring LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfcfbdb8-e452-4910-a45c-ff8343918472 · outbound
Investigating task-specific prompts and sparse autoencoders for activation monitoring LatentQA : Teaching LLMs to Decode Activations Into Natural Language , December 2024
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b9dc901-64a7-42dc-9256-135ecf41995f · outbound
Investigating task-specific prompts and sparse autoencoders for activation monitoring Seeing stars: exploiting class relationships for sentiment categorization with respect to rating scales
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8edec60-bb15-4a56-a9b3-8dec15f0d3ac · outbound
Investigating task-specific prompts and sparse autoencoders for activation monitoring Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8883d1e-35f1-4677-b954-4cc10b0a6b76 · outbound
Investigating task-specific prompts and sparse autoencoders for activation monitoring Negative Results for Sparse Autoencoders On Downstream Tasks and Deprioritising SAE Research ( Mechanistic Interpretability Team Progress Update ), March 2025
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e8b46aef-336d-470e-aab9-9053eb69ebdc · outbound
Investigating task-specific prompts and sparse autoencoders for activation monitoring Linear Representations of Sentiment in Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6231d726-a778-47b9-b3f3-b19517bcab23 · outbound
Investigating task-specific prompts and sparse autoencoders for activation monitoring Measuring short-form factuality in large language models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84e38454-4699-493d-a7d1-83bb5a506007 · outbound
Investigating task-specific prompts and sparse autoencoders for activation monitoring Representation Engineering: A Top-Down Approach to AI Transparency
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59355bd8-b4ac-42b2-a5bb-5ef85bd81347 · outbound
Investigating task-specific prompts and sparse autoencoders for activation monitoring write newline
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2eccc7f-c781-4540-a26f-7b0446d678b8 · outbound
Investigating task-specific prompts and sparse autoencoders for activation monitoring @esa (Ref
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98e81d7f-3d7a-4346-b0d5-6853a891ea99 · outbound
Investigating task-specific prompts and sparse autoencoders for activation monitoring Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ccd0b33-e751-4204-959a-29793bcf5584 · outbound
Investigating task-specific prompts and sparse autoencoders for activation monitoring Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 13e64531-f004-44ad-ad14-43ff8a01c0cc · inbound
The Impact of Off-Policy Training Data on Probe Generalisation Investigating task-specific prompts and sparse autoencoders for activation monitoring
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 98f74723-b825-4e60-bfb5-d3966bbc6ca4 · inbound
The SuperActivator Mechanism: Transformers Concentrate Reliable Concept Signals in the Tail Investigating task-specific prompts and sparse autoencoders for activation monitoring
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 535510f5-cd6c-4f2a-a361-13eadb3b179d · inbound
Do Linear Probes Generalize Better in Persona Coordinates? Investigating task-specific prompts and sparse autoencoders for activation monitoring
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3e9a3674-c821-48f4-ab11-24ebaa2ed743 · inbound
Do Linear Probes Generalize Better in Persona Coordinates? Investigating task-specific prompts and sparse autoencoders for activation monitoring
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.