{"id":"4ba59be8-4460-4820-a633-29ec53896cbb","arxiv_id":"2506.13978","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"LLM-internal emotion features, extracted with sparse autoencoders, mirror human valence and arousal ratings across English and Chinese and can be used to steer model output toward target emotions.","lead":"Using sparse autoencoders, the authors extracted interpretable emotion-related features from two large language models and showed these features track human word ratings for pleasantness and intensity. The same features were then used to steer the models' emotional tone in English and Chinese, suggesting AI output can be guided with psychologically grounded emotion concepts.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Valence/arousal geometry may be an artifact of CEBRA-Behaviour's emotion-label supervision; an unsupervised control is needed.","rationale":"The reader's weakest_assumption identifies the same concern. I agree it is the most load-bearing: it targets the first of the three reported findings and the strongest direct evidence for 'structural congruence' of LLM emotion spaces with human affect. The predictive analyses and steering results are more robust to this confound, but the headline structural claim would be substantially weakened if the geometry result fails an unsupervised control. The steering experiments also lack baselines, but a failure there would not directly undermine the representational-alignment claim as much as a confounded geometry result. The correct disposition remains CONDITIONAL: the paper is plausible and the predictive work is strong, but the geometry claim needs the control before acceptance.","tokens_in":19548,"tokens_out":5330,"duration_ms":69361,"concrete_test":"Retrain CEBRA-Behaviour on the identical SAE feature activations for the 9,215 English and 6,780 Chinese words but without emotion labels (unsupervised mode; if unavailable, use PCA or UMAP on the same feature vectors), and recompute the Fig. 2e Pearson correlations with valence and arousal. As an additional control, run the same supervised CEBRA pipeline on non-LLM word embeddings (e.g., GloVe or word2vec) or on random features. If the unsupervised correlations are not significant, or the non-LLM supervised control reproduces the reported correlations, the geometry claim is a label-supervision artifact rather than evidence for human-aligned LLM emotion representations.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's first key finding—that the LLM-derived emotion space is organized by valence and arousal—rests on CEBRA-Behaviour embeddings (Methods, 'Latent space analyses of computational emotion space using CEBRA'; Fig. 2e). CEBRA-Behaviour is explicitly trained with the 26 emotion-category labels as auxiliary supervision. Because those labels are themselves strongly correlated with valence and arousal (e.g., joy vs. fear), the contrastive objective will place words with similar emotion labels near each other, and the resulting latent axes will automatically correlate with valence and arousal even if the underlying SAE features carry no such structure. The manuscript reports no unsupervised or label-free control (e.g., CEBRA without behavioural labels, or PCA/UMAP on the same SAE activations) and no non-LLM baseline embedding trained under identical label supervision. Without such a control, the reported Bonferroni-corrected correlations in Fig. 2e cannot be attributed to the LLM's internal emotion geometry; they may reflect the supervision signal alone. Since the central claim of structural congruence is in large part built on this geometry, this is the most load-bearing unaddressed assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a concept-driven method to extract, from sparse autoencoder (SAE) features of two LLM families (Gemma2-9B-IT and Llama3-8B-IT), interpretable emotion spaces for English and Chinese, using human word-association norms and emotion labels. The authors report three main findings: (1) the geometry of these emotion spaces is organized by valence and arousal, as assessed by CEBRA embeddings correlated with human ratings; (2) SAE features within the emotion spaces predict large-scale human valence/arousal ratings, with intersectional features performing comparably to all features and better than extra-space features; and (3) adding steering vectors derived from emotion-specific SAE features to hidden states modulates the emotional content of generated text, as measured by a separate RoBERTa classifier and illustrated in qualitative examples. The paper argues that LLMs develop internal emotion representations that are structurally congruent with human emotion concepts and that these representations can be causally manipulated.","tokens_in":19764,"tokens_out":5028,"duration_ms":58844,"significance":"If the findings hold, the paper provides a concrete, interpretable bridge between human affective science and LLM internals: emotion categories grounded in psychological theory can be localized as SAE features, can predict human behavioral ratings across two structurally different languages, and can be used to steer model outputs. The empirical work has notable strengths: predictions are evaluated with five-fold cross-validation repeated over ten seeds, across six SAE configurations per model family, with permutation nulls; the steering evaluation uses a separately trained classifier and linear mixed-effects modeling; and the comparison between intersectional and extra-space features provides a meaningful control for the predictive claims. These strengths make the manuscript a potentially valuable contribution to affective computing and mechanistic interpretability. However, the valence/arousal geometry claim currently rests on an analysis whose supervision may induce the reported structure, which tempers the significance until a label-free control is provided.","major_comments":[{"comment":"Because the concern is specific to the CEBRA-based analysis and the paper's other two findings do not rely on it, this issue is fixable within the manuscript's scope by adding the requested control analyses.","section":"Methods, 'Latent space analyses of computational emotion space using CEBRA'; Fig. 2e"},{"comment":"The Bonferroni correction for the Fig. 2e correlations is not fully specified: the text does not state how many comparisons were entered into the correction (e.g., 3 dimensions x 2 metrics x 2 languages x 2 model families, or a smaller set). The reported 'Bonferroni corrected P < 0.001' is therefore difficult to verify, and the figure caption does not report the actual correlation values or confidence intervals. Please clarify the correction procedure and report the raw and adjusted statistics.","section":"Fig. 2e and Statistical analyses"}],"minor_comments":[{"comment":"Several equations and inline symbols in the Methods section are missing or garbled (e.g., the SAE reconstruction equation and the steering equation in 'Emotion steering vectors'), which makes the mathematical details difficult to follow. Please ensure that all equations are rendered correctly and define every symbol used.","section":"Methods, 'Sparse autoencoders'"},{"comment":"The phrase 'optimized parameters determined independently for size emotions' should read 'for six emotions'; moreover, the optimization procedure for selecting the NMF components (F and M) is not described in sufficient detail for replication. Please specify the criterion used to optimize these parameters and whether the optimization was performed on held-out data or with cross-validation to avoid overfitting.","section":"Methods, 'Emotion steering vectors'"},{"comment":"The CEBRA training procedure reports only the number of training steps (50,000) and the output dimensionality (3). For reproducibility, please report the other hyperparameters (e.g., learning rate, batch size, number of negative samples, architecture details) and the specific version of CEBRA used.","section":"Methods, 'Latent space analyses of computational emotion space using CEBRA'"},{"comment":"Some of the Chinese emotion labels are unconventional as emotion terms (e.g., 'Entrapment' translated as '陷阱', which literally means 'trap'). Please verify that the Chinese labels correspond to the intended emotion concepts and, if necessary, provide a brief explanation of the translation choices.","section":"Table S1"},{"comment":"There are duplicate references in the list (e.g., Goldstein et al. appears as refs 38 and 42; Cowen and Keltner 2021 appears as refs 21 and 22). Please consolidate to a single citation per source.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is methodologically ambitious and the predictive and steering experiments are largely well executed. The main gate is the CEBRA-based geometry claim: the label supervision may be doing the work of organizing the latent space by valence/arousal. I would recommend that the decision hinge on whether the authors can provide an unsupervised or label-free control that preserves the reported correlations. If such controls cannot be provided, the first key finding should be substantially weakened or reframed, and the paper may need to be downgraded to a report of the predictive and steering results only."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing you should know: this is a better paper than the CEBRA concern makes it look. The main result is not the pretty geometry plots; it is the predictive analysis. SAE features selected by the concept-driven procedure predict human valence/arousal ratings across two LLM families, six SAE configurations per model, and two languages, with a sensible intersectional feature control and permutation nulls. That part is solid and I would trust it as evidence that these LLMs encode human-relevant affective structure.\n\nWhat is new: the mapping of SWOW word-association concept sets through SAE features into language-specific emotion subspaces, the cross-lingual prediction of rated affect from those subspaces, and NMF-derived steering vectors from interpretable features. The steering results are suggestive rather than conclusive, but the monotonic dose-response with the target emotion across two models is real signal.\n\nSoft spots, in order. First, Fig. 2e is weaker than the text claims. CEBRA-Behaviour is trained with the 26 emotion labels as auxiliary supervision; since those labels are correlated with valence and arousal, the correlations in Fig. 2e are partly induced by the label signal. The stress-test note is right. The fix is easy: run CEBRA without behavioural labels (or PCA on the same SAE activations) and show the valence/arousal structure persists. This is the weakest link in the “geometry is anchored in valence and arousal” claim, but it is not the central claim of the paper. I would demote Fig. 2e to supporting evidence and let the prediction results carry the argument.\n\nSecond, the steering experiments lack simple baselines. I want to see steering with random SAE features, with a generic positive/negative direction, and with a held-out emotion. The RoBERTa classifier is a single model trained for six basic emotions, so the “natural” qualitative claims in Experiment 4 are anecdotes. The authors cite AXBench in the limitations; they should apply those baselines here.\n\nThird, code is “available upon request” only. For a pipeline this modular, public code would make the difference between a citation and a method people actually build on. Minor: the Chinese Table S2 caption says the steering vectors were derived from the English space; if true, that is another confound worth clarifying.\n\nOverall, the paper is honest about its limitations and the machine-checkable parts are carefully done. The CEBRA issue is fixable and should not block peer review; the predictive and steering evidence is enough to justify a serious referee. My recommendation: send it out, but ask for the unsupervised control, the steering baselines, and code release.\n\nFor reading group: maybe. I’d cite it if I worked on affective interpretability.","headline":"Solid, useful empirical study; the CEBRA geometry figure is partly supervised but the predictive analysis stands on its own and the paper deserves serious review.","tokens_in":20337,"tokens_out":2094,"would_cite":true,"duration_ms":24007,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Large language models represent emotion internally along the same valence-arousal axes humans use, and those representations can steer emotional output.","keywords":["emotion representation","large language models","sparse autoencoders","valence and arousal","emotion steering","cross-linguistic alignment","interpretability","affective computing"],"falsifier":"Rerun the CEBRA latent-space projection without emotion-label supervision, or with shuffled labels, and check whether the three embedding dimensions still correlate with human valence and arousal ratings; if the correlations disappear, the claim that valence and arousal organize the LLM's emotion space would be refuted.","tokens_in":19320,"feed_emoji":"🧠","tokens_out":11669,"duration_ms":117158,"temperature":0.7,"pith_summary":"Large language models, the paper argues, do more than produce contextually appropriate emotional language: their internal representations of emotion are structurally organized the way human affective perception is. The authors build interpretable emotion spaces for English and Chinese by translating human word-association norms for 26 emotion categories into sparse-autoencoder features inside two open LLM families. They report that the geometry of these spaces is anchored by valence and arousal, that the features predict large-scale human valence and arousal ratings within and across languages, and that adding emotion-specific steering vectors to hidden states shifts generated sentences toward the intended emotion. If the claim holds, human emotion concepts can serve both as a description of LLMs' internal affective organization and as control knobs for their emotional output. The paper takes this as evidence of structural human-AI emotional alignment, not of sentience.","feed_headline":"LLMs lay out emotions on the same axes people do","feed_subtitle":"Their internal features predict human word ratings across languages and steer output toward target emotions.","key_machinery":"The load-bearing object is the SAE-based computational emotion space. For each of 26 emotion categories and each language, the authors select a concept-set of up to 10 words whose sparse-autoencoder feature vectors are most similar to the emotion-label word's vector, then take the union of activated features as that emotion's subspace; the union across categories is the emotion space. A neural dimensionality-reduction method called CEBRA-Behaviour, which uses emotion labels as auxiliary supervision, projects these high-dimensional spaces into three dimensions to reveal valence and arousal gradients. Gradient-boosting regressors map word activations over the feature sets to human ratings, and non-negative matrix factorization selects compact steering vectors that are added to hidden states at inference to bias emotional expression.","core_discovery":"The paper's central claim is that LLM-based AI systems can develop internal representations of emotion that are structurally aligned with human understanding. Concretely, the authors claim that the sparse-autoencoder-defined emotion space of two open LLM families is organized along valence and arousal; that its features predict human valence and arousal ratings for thousands of English and Chinese words, with language-shared features matching the full feature set and language-specific features performing worse; and that steering vectors built from emotion-specific features causally modulate the emotion of generated sentences. The authors stress that this does not imply AI sentience; it implies that the organization of emotion-related language representations in LLMs is human-like and manipulable.","pith_inferences":["A testable extension the paper does not run: projecting the emotion space without emotion-label supervision, or with shuffled labels, would show how much of the valence-arousal geometry is intrinsic to the LLM rather than imposed by the labels.","Because the steering evaluation relies on an automated emotion classifier, having human raters judge the steered sentences would test whether the emotional shift is perceptible to people and not only to a model.","The same concept-driven pipeline could be applied to other psychological dimensions, such as dominance, or to emotion blends, which would test whether the 26-category taxonomy is the right granularity for describing LLM emotion spaces."],"forward_implications":["Emotion-related SAE features predict human valence and arousal ratings for thousands of English and Chinese words, with shared cross-language features performing as well as the full feature set.","Steering vectors derived only from human emotion concepts raise the classifier-assigned scores for the target emotion and lower scores for untargeted emotions and neutral, across two LLM families.","Cross-language prediction works in both directions but is less accurate than within-language prediction, mirroring known cultural and linguistic variation in emotional perception.","The method transfers across model families, layers, and SAE widths, suggesting the emotion spaces are not an artifact of one architecture."],"supporting_citations":[{"why":"It supplies the 26-category emotion taxonomy that defines the emotion spaces.","marker":"[21]"},{"why":"It establishes that sparse-autoencoder features are interpretable monosemantic units inside LLMs.","marker":"[47]"},{"why":"It provides the pre-trained Gemma2-9B-IT sparse autoencoders used to extract emotion features.","marker":"[49]"},{"why":"It provides the pre-trained Llama3-8B-IT sparse autoencoders used to extract emotion features.","marker":"[50]"},{"why":"It supplies the English word-association norms used to build English concept-sets.","marker":"[54]"},{"why":"It supplies the Chinese word-association norms used to build Chinese concept-sets.","marker":"[55]"},{"why":"It supplies the CEBRA-Behaviour model used for latent-space visualization.","marker":"[56]"},{"why":"It supplies the English valence and arousal ratings used as prediction targets.","marker":"[58]"},{"why":"It supplies the Chinese valence and arousal ratings used as prediction targets.","marker":"[59]"},{"why":"It supplies the SAE-feature steering paradigm that the authors adapt into emotion steering vectors.","marker":"[60]"}],"fun_headline_variants":["AI emotion spaces mirror human valence and arousal","LLMs feel language like we do, across cultures","Human concepts can steer AI emotional output","AI shares human emotional structure, not sentience","LLM emotions align with humans across languages"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The valence-arousal geometry result assumes that the 26 emotion labels used as auxiliary supervision during latent-space projection did not themselves impose the valence-arousal structure that is then compared with human ratings.","fun_headline_variants_meta":{"raw":{"variants":["AI emotion spaces mirror human valence and arousal","LLMs feel language like we do, across cultures","Human concepts can steer AI emotional output","AI shares human emotional structure, not sentience","LLM emotions align with humans across languages"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000164,"raw_usage":{"total_tokens":1219,"prompt_tokens":889,"completion_tokens":330,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":505,"completion_tokens_details":{"reasoning_tokens":262}},"tokens_in":505,"tokens_out":330,"duration_ms":4154,"temperature":1.0,"reasoning_tokens":262,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:40:35.677020+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the CEBRA latent-space projection without emotion-label supervision, or with shuffled labels, and check whether the three embedding dimensions still correlate with human valence and arousal ratings; if the correlations disappear, the claim that valence and arousal organize the LLM's emotion space would be refuted.","supporting_citations":[],"review_version":1}