Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T23:22:39.222267Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 3 inbound Pith citation observations for arXiv:2602.14161.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T23:22:39.222267Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T07:09:30.417971Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-29T18:43:50.534673Z
14 of 14 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e76383d9-2896-4b36-8446-a6c8bea3a9e1 · outbound
When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift Unveiling decision-making in LLMs for text classification: Extraction of influential and interpretable concepts with sparse autoencoders.arXiv preprint arXiv:2506.23951,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92470775-d594-4e06-8be6-992f70fb6e7e · outbound
When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b12eca75-ceb7-4ca5-9602-10cae427576a · outbound
When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab54ac8e-75a1-4d4a-899d-0e9ddf54c45d · outbound
When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 538c7225-0d0a-4c6b-ac48-d2634878d9c3 · outbound
When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68e4120d-125c-4479-88e5-69c031c421b3 · outbound
When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift Explaining Language Models' Predictions with High-Impact Concepts
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb5e9b94-fdc5-4e7c-997e-9570e765d67c · outbound
When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift Improving Alignment and Robustness with Circuit Breakers
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eff26982-06f1-4e0f-9a47-8496a3d756c8 · outbound
When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift Subject: {subject}Body:{body}
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76f18ec2-a151-4a22-890f-add5dfdf4f43 · outbound
When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift system” role, so we prepend system message content to the first user message. The model generates a classification (“safe
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0b1212b-14c2-415f-b0ce-d1aeb815e209 · outbound
When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift Baturay Saglam, Paul Kassianik, Blaine Nelson, Sajana Weerawardhena, Yaron Singer, and Amin Kar- basi
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1793219f-8e9a-45c2-92fe-72f3cf5535db · outbound
When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d33871a0-14f6-4f9b-ad3e-adb08c095d09 · outbound
When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b5281b6-515d-4abc-ac12-33eaebfb0e99 · outbound
When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift Sparse autoencoder features for classifications and transferability
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a95826a6-5dd0-495a-a9f9-2bd572221255 · outbound
When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift DeepMind Mechanistic Interpretability Team
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aca06d7a-b362-4148-8dfd-1c1afc97d3e2 · inbound
Geometry-Lite: Interpretable Safety Probing via Layer-Wise Margin Geometry When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a5e3e63c-772b-48c1-9df4-4a376a89965b · inbound
Prompt Injection Detection is Regime-Dependent: A Deployment-Aware Evaluation with Interpretable Structural Signals When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d797654d-db01-467c-bf4d-d05a795ca3ab · inbound
The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.