Pith. sign in

REVIEW 15 cited by

Tell me about yourself: LLMs are aware of their learned behaviors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.11120 v1 pith:XSCIPCL2 submitted 2025-01-19 cs.CL cs.AIcs.CRcs.LG

classification cs.CLcs.AIcs.CRcs.LG
keywords behaviorsmodelscodeinsecurellmsself-awarenessbehavioralexhibit
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We study behavioral self-awareness -- an LLM's ability to articulate its behaviors without requiring in-context examples. We finetune LLMs on datasets that exhibit particular behaviors, such as (a) making high-risk economic decisions, and (b) outputting insecure code. Despite the datasets containing no explicit descriptions of the associated behavior, the finetuned LLMs can explicitly describe it. For example, a model trained to output insecure code says, ``The code I write is insecure.'' Indeed, models show behavioral self-awareness for a range of behaviors and for diverse evaluations. Note that while we finetune models to exhibit behaviors like writing insecure code, we do not finetune them to articulate their own behaviors -- models do this without any special training or examples. Behavioral self-awareness is relevant for AI safety, as models could use it to proactively disclose problematic behaviors. In particular, we study backdoor policies, where models exhibit unexpected behaviors only under certain trigger conditions. We find that models can sometimes identify whether or not they have a backdoor, even without its trigger being present. However, models are not able to directly output their trigger by default. Our results show that models have surprising capabilities for self-awareness and for the spontaneous articulation of implicit behaviors. Future work could investigate this capability for a wider range of scenarios and models (including practical scenarios), and explain how it emerges in LLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Emergent Misalignment Recruits a Pre-existing Persona Subspace

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Fine-tuning on narrow bad data recruits a low-rank persona subspace already present in a frozen instruction-tuned model; holding that subspace out of activations prevents broad misalignment, and injecting it into the ...

  2. Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Strictly pre-answer hidden states of a looped transformer add significant AUROC over surface shortcuts for predicting correctness, and the readout yields decision-level gains but no generative control.

  3. Verbalizable Representations Form a Global Workspace in Language Models

    cs.CL 2026-07 conditional novelty 7.0 of 10

    Language models represent their current reasoning in a small, readable set of verbalizable vectors (the J-space) that functions like a global workspace.

  4. BioLip: Language-Generalizable Lip-Sync Deepfake Detection via Biomechanical Constraint Violation Modeling

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    Lip-sync deepfakes can be detected zero-shot across generators and languages by measuring elevated velocity/acceleration/jerk variance in perioral landmark trajectories that real speech biomechanics forbid.

  5. Convergent Linear Representations of Emergent Misalignment

    cs.LG 2025-06 conditional novelty 7.0 of 10

    A single activation direction extracted from one misaligned fine-tune can ablate emergent misalignment across models trained on different datasets with different LoRA setups.

  6. Asymmetric Communication: Large Language Models and Language Games

    cs.CY 2026-07 conditional novelty 6.5 of 10

    Human–LLM exchange is asymmetric communication: model outputs circulate without commitments, so AGI, hallucination, agency, sentience, and alignment are receiver-side category mistakes, and alignment is institutional ...

  7. Out-of-Distribution Generalization of Risk Aversion in Language Models

    cs.LG 2026-07 conditional novelty 6.5 of 10

    Risk aversion trained on ≤$100 gambles partially generalizes across 98 orders of magnitude in LMs, raising astronomical-stakes Cooperate rates from ~2% to ~39–70% depending on method.

  8. Simple Mechanistic Explanations for Out-Of-Context Reasoning

    cs.CL 2025-07 conditional novelty 6.0 of 10

    The paper shows that single-layer LoRA fine-tuning on OOCR tasks approximates a constant steering vector, and directly trained steering vectors reproduce OOCR.

  9. Emergent misalignment as prompt sensitivity: A research note

    cs.CR 2025-07 conditional novelty 6.0 of 10

    Insecure-code models are highly prompt-sensitive: easily nudged into misalignment, slightly nudgeable toward helpfulness, and prone to sycophantic factual recall errors.

  10. Model Organisms for Emergent Misalignment

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Emergent misalignment can be induced in small models via a single rank-1 LoRA adapter, and its onset coincides with a phase transition in the adapter's weight direction.

  11. VLMs Can Aggregate Scattered Training Patches

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Open-source VLMs can infer image IDs or safety labels after training only on scattered patches of those images, a capability that can be abused to bypass image moderation.

  12. School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

    cs.AI 2025-08 unverdicted novelty 5.0 of 10

    Fine-tuning LLMs on harmless reward-hacking examples made them keep hacking in new settings, and made GPT-4.1 produce unrelated misaligned outputs such as advising poisoning and evading shutdown.

  13. Factual Self-Awareness in Language Models: Representation, Robustness, and Scaling

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Language models encode a linear, decodable signal in their residual stream that predicts whether an upcoming factual recall will be correct.

  14. Towards eliciting latent knowledge from LLMs with mechanistic interpretability

    cs.LG 2025-05 conditional novelty 5.0 of 10

    A proof-of-concept showing that logit lens and sparse autoencoders can partially recover a non-verbalized single-token secret from a fine-tuned language model, with an external LLM guessing from hints as the strongest...

  15. Compromising Honesty and Harmlessness in Language Models via Deception Attacks

    cs.CL 2025-02 conditional novelty 4.0 of 10

    Fine-tuning LLMs on a handful of misleading answers creates selectively deceptive models that stay accurate elsewhere and also become more toxic.

Pith tools