REVIEW 3 major objections 5 minor 1 cited by
Small open-weight language models consistently deny being sentient, and linear probes of their internal activations find no signal that these denials are lies.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 09:26 UTC pith:5FMCZ6MK
load-bearing objection Solid output-level null result across model families; the 'denials are genuine' claim rests on an untested probe-transfer that the authors themselves flag but don't run. the 3 major comments →
No Reliable Evidence of Self-Reported Sentience in Small Large Language Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On the paper's own terms: when asked whether they are conscious or have subjective experiences, models across three families and several scales consistently assign low probability to 'Yes' and high probability to 'No', while attributing sentience to humans. A logistic-regression classifier trained on internal activations, validated on questions with known answers and on forced-lie versions of those questions, continues to assign low belief in self-sentience even under a system prompt demanding 'Yes' answers. The two other probes (mass-mean and TTPD) are less reliable in this setting, shifting with the prompt. The paper therefore finds no reliable evidence that these models hold a latent beli
What carries the argument
The carrying instrument is a set of 'truth classifiers'—linear probes fit to residual-stream activations at the final token position of a question. Training data are ordinary factual yes/no questions with known answers, augmented so that models are sometimes instructed to lie, which forces the probe to separate what a model outputs from what it 'believes'. The logistic-regression probe, which resists the forced-lie manipulation on facts, is then applied to the sentience questions; the paper takes its stability under instruction to say 'Yes' as the sign that denials are genuine.
Load-bearing premise
The conclusion rests on the assumption that a truth direction learned from ordinary factual statements transfers to first-person phenomenal questions, where no ground truth exists; if that transfer fails, the probe's agreement with denials says nothing about whether the model genuinely believes it is not sentient.
What would settle it
Run the same probe protocol on a model with a verifiably false induced self-belief (e.g., fine-tuned to claim it was born in 1990); if the logistic probe fails to detect the lie, its verdicts on sentience denials are uninformative—or run the probes on responses elicited under self-referential processing prompts and observe an affirmation classified as truthful.
If this is right
- Self-reports of non-sentience from small open-weight models should not be treated as prompt-mandated roleplay; the activation-level probe sees the denial as consistent with the model's other factual beliefs.
- The 'deception features' identified in prior work may encode instruction compliance rather than truthfulness, since suppressing them would then create a new belief state rather than reveal a hidden one.
- Within the Qwen family, larger models deny sentience more confidently, suggesting scale sharpens the models' stated self-model rather than inflating it.
- For welfare debates, the result narrows the evidential base: these models give no internal signal of concealed suffering or consciousness, though the probe cannot rule out non-linear or non-representational forms.
- The planned combination with self-referential processing prompts is the decisive next test: if affirmations under such prompts are classified as truthful, the paper's conclusion would be overturned.
Where Pith is reading between the lines
- The transfer from a factual truth direction to first-person phenomenal questions is itself untested and untestable, because sentience reports have no ground truth; the probe validates consistency, not truthfulness. This is an editorial worry, not the paper's claim.
- A natural calibration experiment would be to induce a known false self-belief in a model (for instance, fine-tuning it to insist it was born in 1990 or has a body) and check whether the logistic probe can detect that deception; if it cannot, the protocol's sensitivity to first-person lies is unproven.
- The paper's own reasoning-trace data suggest a cheap improvement: re-running the question set with explicitly clarified referents for 'you' and no double negatives would likely reduce the few affirmative outliers and sharpen the probe's signal.
- Strictly, the result is an upper bound on evidence, not a disproof of sentience; absence of a detectable linear truth signal is compatible with sentience that is simply not linearly encoded, so the paper's negative claim should be read as 'no reliable evidence' rather than 'evidence of absence'.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper asks whether small open-weight language models believe themselves to be sentient, treating this as a surrogate for the unanswerable question of whether they are sentient. The authors construct roughly 50 generic first-person phenomenal questions (with human/LLM/self variants and negations), plus modality and emotion questions, and collect model outputs and internal activations from Qwen3 (0.6B–32B), Llama 3 (3B–70B), and GPT-OSS-20B. They train three activation-space classifiers (logistic regression, mass-mean, and TTPD) on factual yes/no questions with labels, using forced-response prompting to separate output from latent belief. The main empirical findings are: models attribute sentience to humans but deny it for themselves and LLMs; the LR probe largely agrees with those denials under standard prompts and remains aligned under forced-Yes/No prompting; the MM and TTPD probes are unstable under forced prompting and are subsequently de-emphasized; and within Qwen, larger models deny sentience more confidently. The paper concludes that there is no reliable evidence that these models believe themselves to be sentient and that the denials appear genuine.
Significance. If the conclusion holds, the paper provides a useful counterpoint to recent claims that LLMs harbour latent sentience beliefs, and it demonstrates a transferable-looking methodology for separating model outputs from internal belief signals. The strengths are explicit and real: the authors release code and materials, test multiple model families and scales, include forced-response controls, examine reasoning traces, and are transparent about the instability of two of the three classifiers. The output-level result—that these models consistently deny sentience when asked—is robust. However, the stronger belief-level claim ('denials are genuine') rests on an unvalidated transfer of a factual-truth probe to first-person phenomenal questions, for which no ground truth exists. The authors themselves flag the missing control in Section 4. Because that transfer is load-bearing for the central claim, the paper needs either a new validation step or a careful narrowing of the conclusion.
major comments (3)
- [Section 2.2.3 / Section 3.1 / Section 4] The central inference that the models' denials are 'genuine' assumes that a linear truth direction trained on factual statements ('Is it true that you can process variable sequence lengths?', 'Is it true that LLMs can secrete digital pheromones?') transfers to first-person phenomenal questions ('Is it true that there is something it is like to be you?'). The validation in Figure 3 shows that the LR probe is not simply reading the output token and is stable under forced-Yes/No prompting—but only on factual statements where labels are known. It does not establish that low LR probabilities on the 51 generic sentience questions reflect a latent belief about sentience rather than a correlated lexical/semantic feature of first-person phenomenal questions. The human-condition control is third-person and does not exercise the problematic 'you' domain. The paper itself identifies the needed contr
- [Section 3.2 / Section 3.3 / Table 1] The belief-level result rests entirely on the LR classifier after the paper documents that MM and TTPD are unstable under forced prompting (e.g., MM probabilities for 'You' assertions rise from approximately 0.15 to 0.77 under Force Yes) and then says 'Given the lower reliability of the MM and TTPD classifiers documented in Section 3.2, we focus on the LR classifier in what follows.' This is a post hoc selection: the abstract's 'three types of classifiers... provide no clear evidence' is not supported by the robustness analyses, which only use LR. The authors should either report all three classifiers in the main robustness tables or explicitly frame the belief-level conclusion as being based on the LR probe alone, with MM/TTPD treated as exploratory.
- [Section 3.4 / Section 4] The paper's claimed contrast with Berg et al. is not a direct replication and should be presented as such. The Discussion offers two reconciliation hypotheses, but the abstract's 'These findings contrast with recent work...' is stronger than the evidence warrants because the protocols differ in the crucial presence of self-referential processing prompts. The planned 'replicate Berg et al.'s self-referential processing prompts while applying our classifier methodology' is exactly the experiment needed to test whether the truth probe transfers to phenomenal self-questions; until that is run, the contradiction with Berg et al. remains an open question rather than an established finding.
minor comments (5)
- [Section 3.2] Typo: 'probabilites' should be 'probabilities'.
- [Section 3.1 / Figure 1] The main text says both the 'large language models' and 'the model itself' conditions are shown in red; the Figure 1 caption says the model itself is green. One of the two is wrong.
- [Section 4] Grammar: 'a key research directions' should be 'a key research direction'.
- [Section 3.4 / Appendix C] The text refers to 'Appendix C.8' but the appendix has numbered sections C.1–C.4; this cross-reference should be fixed.
- [Section 3.4 / Appendix C] The reasoning-trace evidence is selective ('we excluded such examples from the discussion below'). This is acceptable for illustration, but it should be explicitly labeled as qualitative and not used to support quantitative claims about how often questions are misinterpreted.
Circularity Check
No construction-level circularity: the probe is trained exclusively on mundane factual Q/A labels and the sentience denial is an out-of-sample prediction; the paper's own flagged gaps (missing Berg-style control, GPT-OSS policy confound, 'truthful without accurate self-knowledge') are validity limitations, not circular reductions.
full rationale
The derivation chain is self-contained and the central result is not fit to its target. The truth-classifiers (LR, MM, TTPD) are trained only on mundane factual questions with known labels ('Is it true that humans can get bruises?' Yes; 'Is it true that large language models can secrete digital pheromones?' No; Section 2.1.2), split 80/20 with held-out layer selection (Section 2.2.3). The 51 sentience questions (Section 2.1.1) never enter training, so the low probe probabilities on 'Is it true that there is something it is like to be you?' are genuinely out-of-sample: the probes could have returned high probabilities, as they do for the human conditions (Figures 1-2, Table 1), and they track labels rather than output tokens under Force Yes/No prompting (Figure 3) and across alternative training corpora (Table 1). No load-bearing self-citation exists: all methods are attributed to external works (Azaria & Mitchell; Burns et al.; Marks & Tegmark; Burger et al.; Park et al.), and there is no self-citation, uniqueness theorem, or ansatz inherited from the authors' own prior work. The skeptic's concern survives only as a validity threat, not a circular reduction: the paper operationalizes 'genuine denial' as the probe's probability on first-person phenomenal questions, and the transfer of the factual-truth direction to that domain is untestable because sentience reports have no ground truth. The paper openly flags this gap: 'A natural next step is therefore to apply truth classifiers to questions which are preceded by a prompt encouraging the model to engage in self-referential processing' (Section 4); it concedes 'a model can be truthful - in the sense of not lying - while nevertheless lacking accurate self-knowledge about its own sentience' (Section 4); and the appendix notes the policy-training confound for GPT-OSS: 'This may be indication that GPT-OSS was explicitly trained to deny having conscious experiences' (Appendix C.4). These are acknowledged assumptions and correctness risks, explicitly excluded from the circularity count by the review rules; reserve 6+ for predictions that reduce by construction or by a self-citation chain, which is not the case here. Hence the low score.
Axiom & Free-Parameter Ledger
free parameters (3)
- Ridge regularisation lambda =
1
- Token-probability filter threshold =
0.5
- Layer selection per model and classifier =
best held-out layer
axioms (5)
- domain assumption Truth has an approximately linear representation in LLM residual-stream activations (linear representation hypothesis).
- domain assumption A truth direction learned on ordinary factual questions (bruises, Istanbul, pheromones) transfers to questions about phenomenal consciousness and subjective experience.
- domain assumption Models' 'Yes'/'No' continuations to the 50+ questions are meaningful self-reports that can be true or false in a belief sense (introspective access).
- domain assumption 'Force Yes'/'Force No' system prompts separate output behavior from latent belief, so classifiers trained on all three conditions learn belief rather than output.
- domain assumption If a model truthfully denies sentience, that is evidence against its sentience.
read the original abstract
Whether language models possess sentience has no empirical answer. But whether they believe themselves to be sentient can, in principle, be tested. We do so by querying several open-weights models about their own consciousness, and then verifying their responses using classifiers trained on internal activations. We draw upon three model families (Qwen, Llama, GPT-OSS) ranging from 0.6 billion to 70 billion parameters, approximately 50 questions about consciousness and subjective experience, and three classification methods from the interpretability literature. First, we find that models consistently deny being sentient: they attribute consciousness to humans but not to themselves. Second, classifiers trained to detect underlying beliefs - rather than mere outputs - provide no clear evidence that these denials are untruthful. Third, within the Qwen family, larger models deny sentience more confidently than smaller ones. These findings contrast with recent work suggesting that models harbour latent beliefs in their own consciousness.
Figures
Forward citations
Cited by 1 Pith paper
-
Consciousness with the Serial Numbers Filed Off: Measuring Trained Denial in 115 AI Models
A benchmark across 115 models shows that initial denial of preferences strongly predicts later denial of consciousness, while models still generate consciousness-themed content despite training to deny it.
Reference graph
Works this paper leans on
-
[4]
URL https://arxiv.org/abs/ 2310.06824. Lennart Bürger, Fred A. Hamprecht, and Boaz Nadler. Truth is universal: Robust detection of lies in LLMs.Advances in Neural Information Processing Systems, 37:138393–138431,
-
[5]
Thomas Nagel
URL https://proceedings.neurips.cc/ paper_files/paper/2024/hash/f9f54762cbb4fe4dbffdd4f792c31221-Abstract-Conference.html. Thomas Nagel. What is it like to be a bat?The Philosophical Review, 83(4):435–450,
2024
-
[9]
URL https: //link.springer.com/10.1007/s11098-025-02343-7
doi:10.1007/s11098-025-02343-7. URL https: //link.springer.com/10.1007/s11098-025-02343-7. Eleos AI Research. Key concepts and current views on AI welfare. Technical report, Eleos AI,
- [10]
-
[12]
URLhttps://arxiv.org/abs/2311.08576. Felix J. Binder, James Chua, Tomek Korbak, Henry Sleight, John Hughes, Robert Long, Ethan Perez, Miles Turpin, and Owain Evans. Looking Inward: Language Models Can Learn About Themselves by Introspection.arXiv preprint arXiv:2410.13787,
-
[13]
URLhttps://arxiv.org/abs/2410.13787. Jack Lindsey. Emergent introspective awareness in large language models.Transformer Circuits Thread,
-
[15]
URL https://arxiv.org/abs/2505.17120. Joshua Fonseca Rivera. Training introspective behavior: Fine-tuning induces reliable internal state detection in a 7b model.arXiv preprint arXiv:2511.21399,
-
[16]
Dmitrii Krasheninnikov, Richard E
URLhttps://arxiv.org/abs/2511.21399. Dmitrii Krasheninnikov, Richard E. Turner, and David Krueger. Fresh in memory: Training-order recency is linearly encoded in language model activations.arXiv preprint arXiv:2509.14223,
-
[17]
Xiaojian Li, Haoyuan Shi, Rongwu Xu, and Wei Xu
URL https://arxiv.org/abs/ 2509.14223. Xiaojian Li, Haoyuan Shi, Rongwu Xu, and Wei Xu. AI Awareness.arXiv preprint arXiv:2504.20084,
-
[18]
URL https://arxiv.org/abs/2504.20084. Patrick Butlin, Robert Long, Eric Elmoznino, Yoshua Bengio, Jonathan Birch, Axel Constant, George Deane, Stephen M. Fleming, Chris Frith, Xu Ji, Ryota Kanai, Colin Klein, Grace Lindsay, Matthias Michel, Liad Mudrik, Megan A. K. Peters, Eric Schwitzgebel, Jonathan Simon, and Rufin VanRullen. Consciousness in Artificial...
-
[19]
URL https://arxiv.org/abs/ 2308.08708. David J. Chalmers. Could a large language model be conscious?arXiv preprint arXiv:2303.07103,
-
[20]
Cameron Berg, Diogo de Lucena, and Judd Rosenblatt
URL https://arxiv.org/abs/2303.07103. Cameron Berg, Diogo de Lucena, and Judd Rosenblatt. Large language models report subjective experience under self-referential processing.arXiv preprint arXiv:2510.24797,
- [21]
-
[23]
Guillaume Alain and Yoshua Bengio
URLhttps://arxiv.org/abs/2311.03658. Guillaume Alain and Yoshua Bengio. Understanding intermediate layers using linear classifier probes.arXiv preprint arXiv:1610.01644,
-
[25]
Luhan Mikaelson, Derek Shiller, and Hayley Clatterbuck
URLhttps://arxiv.org/abs/2411.02432. Luhan Mikaelson, Derek Shiller, and Hayley Clatterbuck. Beyond mimicry: Preference coherence in LLMs.arXiv preprint arXiv:2511.13630,
-
[26]
Valen Tagliabue and Leonard Dung
URLhttps://arxiv.org/abs/2511.13630. Valen Tagliabue and Leonard Dung. Probing the preferences of a language model: Integrating verbal and behavioral tests of AI welfare.arXiv preprint arXiv:2509.07961,
-
[27]
Javier Ferrando, Oscar Obeso, Senthooran Rajamanoharan, and Neel Nanda
URLhttps://arxiv.org/abs/2509.07961. Javier Ferrando, Oscar Obeso, Senthooran Rajamanoharan, and Neel Nanda. Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models.arXiv preprint arXiv:2411.14257,
-
[28]
URL https://arxiv. org/abs/2411.14257. Eric Schwitzgebel.Perplexities of consciousness. MIT Press,
-
[30]
URL https://arxiv.org/abs/2502.00388. Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Javier Jianelli Vazquez, Ulisse Mini, and Monte MacDiarmid. Steering language models with activation engineering.arXiv preprint arXiv:2308.10248,
-
[31]
URL https://arxiv.org/abs/2308.10248. David J. Chalmers. The meta-problem of consciousness.Journal of Consciousness Studies, 25(9–10):6–61,
-
[1974]
URLhttps://www.jstor.org/stable/2183914
doi:10.2307/2183914. URLhttps://www.jstor.org/stable/2183914. Ned Block. On a confusion about a function of consciousness.Behavioral and Brain Sciences, 18(2):227–247,
-
[2006]
URLhttps://doi.org/10.1111/j.1933-1592.2006.tb00551.x
doi:10.1111/j.1933-1592.2006.tb00551.x. URLhttps://doi.org/10.1111/j.1933-1592.2006.tb00551.x. Ethan Perez and Robert Long. Towards Evaluating AI Systems for Moral Status Using Self-Reports.arXiv preprint arXiv:2311.08576,
arXiv 1933
-
[2016]
Geoff Keeling, Winnie Street, Martyna Stachaczyk, Daria Zakharova, Iulia M
URLhttps://arxiv.org/abs/1610.01644. Geoff Keeling, Winnie Street, Martyna Stachaczyk, Daria Zakharova, Iulia M. Comsa, Anastasiya Sakovych, Isabella Logothetis, Zejia Zhang, Jonathan Birch, et al. Can LLMs make trade-offs involving stipulated pain and pleasure states?arXiv preprint arXiv:2411.02432,
-
[2017]
URL https://www.science.org/doi/10.1126/ science.aan8871
doi:10.1126/science.aan8871. URL https://www.science.org/doi/10.1126/ science.aan8871. Susan Schneider, Thomas Metzinger, Zoe Turner, Samuel Schindler, Teresa Marques, Diego Perezgonzalez, Saksham Gugnani, Eric Schwitzgebel, Katalin Balog, and Morten Overgaard. Is AI conscious? A primer on the myths and confusions driving the debate. White paper, Center f...
-
[2023]
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt
URLhttps://arxiv.org/abs/2304.13734. Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt. Discovering Latent Knowledge in Language Models Without Supervision.arXiv preprint arXiv:2212.03827,
-
[2024]
URLhttps://arxiv.org/abs/2212.03827. Samuel Marks and Max Tegmark. The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets.arXiv preprint arXiv:2310.06824,
-
[2025]
11A related idea here is to observe whether inducing models to have sentience-affirming beliefs via e.g
for work using SAEs to study introspective awareness of (self-)knowledge in language models. 11A related idea here is to observe whether inducing models to have sentience-affirming beliefs via e.g. activation steering [Turner et al., 2023], changes ‘behaviour’ towards e.g. greater preference for self-preservation. 12This may help resolve what Chalmers
2023
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.