REVIEW 3 major objections 5 minor 97 references
Personalized Emotional Intelligence in Generative AI through Symbolic Affective Reasoning
T0 review · 3 major / 5 minor · reviewed 2026-07-14 · grok-4.5
Pith's one-line read A hybrid AI system that mines symbolic affective rules from image–emotion data can edit photos to elicit target emotions more effectively than large multimodal models while personalizing to users without fine-tuning.
desk verdict Solid hybrid systems paper with unusually large human psychophysics; the win is the integrated EmoTree/EmoMem pipeline and metrics, not a proven causal theory of affect. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The EmoTree: a hierarchical symbolic structure of compositional visual motifs (source concept → target valence → concept, attributes, actions, and scene context) mined from image–emotion data, retrieved to plan localized edits; paired with EmoMem, an expandable accepted/rejected motif memory that personalizes retrieval at inference time without fine-tuning.
What would settle it
On held-out images outside the training emotion dataset, run the same human pairwise and interactive editing protocols: if EROS preference, first-loop success, and cumulative success fall to or below the large multimodal baseline, or if ablating the predicted emotion regions fails to change reported valence more than ablating random regions of equal size, the central claim fails.
Extended reading notes
Core claim
EROS shows that generalizable, interpretable affective rules can be mined automatically from large-scale image–emotion data and used to steer visual content toward a desired emotional valence for a specific observer, outperforming state-of-the-art affective editing methods and large multimodal model pipelines on human preference, success rate, and structural fidelity while supporting rapid personalization through an explicit memory bank rather than model fine-tuning.
Load-bearing premise
That saliency from a simple binary emotion classifier, refined into object masks, plus motifs mined offline from captions and clustering, capture the real visual causes of human affect rather than dataset-specific correlations.
Editorial extensions
If this is right
- Affective image editing can be treated as a measurable proxy for machine emotional intelligence spanning recognition, localization, reasoning, regulation, and personalization.
- Symbolic motif libraries plus generative backbones can outperform pure large multimodal pipelines on emotion elicitation while keeping edits localized and source-faithful.
- User-specific affective preferences can be stored as interpretable accepted/rejected motifs and reused across related scenes without fine-tuning.
- Positive and negative affect may be asymmetrically organized: positive motifs cluster around a smaller semantic core while negative motifs span a wider set of cues.
- The same motif-and-memory design could extend beyond static images to other media if emotion-relevant regions and compositional rules can be defined for those modalities.
Reading between the lines
- If motif memories are truly reusable, short interactive sessions could build portable emotional profiles that transfer across apps without sharing raw images.
- Safety guardrails that blunt negative-affect generation in proprietary models may create a systematic gap that open hybrid systems fill for research and clinical simulation—but also raise misuse risk for persuasion.
- Same-valence enhancement and neutral-to-valenced induction, which the paper flags as less explored, are natural next tests of whether the EmoTree encodes regulation rather than only valence flipping.
- Replacing GPT-assisted offline motif construction with fully local mining would test how much the claimed symbolic knowledge depends on proprietary language models at build time.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces EROS, a hybrid symbolic–deep framework for personalized affective image editing. From EmoSet, it trains a ResNet-18 valence predictor, localizes emotion-relevant regions via Grad-CAM refined by SAM (progressive τ), and mines a hierarchical EmoTree of compositional motifs (concept, attribute, action, scene) using CLIP clustering and GPT-4o/GPT-4. Motifs become template prompts for structure-preserving Stable Diffusion inpainting (same- vs cross-valence mask inversion). An expandable EmoMem stores accepted/rejected motifs for inference-time personalization without fine-tuning. Six human psychophysics experiments (33,380 trials, 296 participants) compare EROS to LMS (GPT-based pipeline), XDream, Real retrieval, and non-interactive editors on preference, validity/efficiency, IoU, prompt alignment, and contrastive fidelity (SSIM-C/L1-C). The central claim is that EROS elicits target valence more effectively while better preserving source semantics/structure and adapts to individual motif preferences.
Significance. If the empirical superiority holds, the work is a substantial contribution to affective computing and emotionally intelligent generative AI. Strengths include a clear operationalization of emotional intelligence into five capabilities, a large multi-experiment human benchmark with attention controls and bootstrapped Welch tests, explicit symbolic memory for personalization without fine-tuning, and public code/data. The hybrid design (interpretable motifs + localized editing) addresses real limitations of global style transfer and unconstrained LMM editing. Even if causal claims about Grad-CAM and GPT-mined rules are tempered, the system-level results and evaluation protocols remain useful for the field.
major comments (3)
- [Sec. 4.2.2, Fig. 3B] Sec. 4.2.2 and Fig. 3B: Grad-CAM from a ResNet-18 binary valence classifier (trained on EmoSet), refined by SAM and progressive τ∈{0.5,0.3,0.1,0}, is treated as identifying the visual causes of human affect. Human IoU (0.38 vs Human–Human 0.46) shows spatial agreement but does not establish causality versus classifier-correlated features. The same- vs cross-valence mask inversion in Sec. 4.2.4 makes this interpretation load-bearing for the editing claim. A control that edits non-salient or classifier-irrelevant regions (or uses alternative localizers) is needed to support the causal reading, or the claim should be restated as predictive localization.
- [Sec. 4.2.3, Fig. 5A2] Sec. 4.2.3 and Sec. 4.1: EmoTree motifs are mined offline from EmoSet via CLIP clustering (δ=0.7) and GPT-4o/GPT-4 concept/motif generation from captions; many evaluation images also come from EmoSet. Prompt alignment ~90% (Fig. 5A2) and preference gains could partly reflect dataset co-occurrences and GPT language priors rather than generalizable intervention rules. The paper should report held-out or out-of-distribution scene tests (beyond the limited MSCOCO examples in Fig. S4) and/or ablate GPT-generated motifs against non-LLM alternatives to bound this circularity risk.
- [Sec. 2.1.1, Sec. 4.4.1] Sec. 2.1.1 and Sec. 4.4.1: LMS underperforms especially on negative valence, which the paper attributes to safety guardrails. Because LMS uses the same diffusion backbone as EROS for isolation, residual differences in prompt style, memory representation (flat concept lists vs compositional motifs), and refusal behavior remain confounds for the claim that symbolic affective reasoning is the decisive factor. A matched-prompt or guardrail-controlled comparison (or explicit reporting of LMS refusal rates) would strengthen the superiority claim over large multimodal pipelines.
minor comments (5)
- [Sec. 4.2] Free parameters (δ, τ schedule, γ=0.7, hard-rejection count of 3, cross-valence similarity skip 0.8, candidate sample size 50) are stated but lack sensitivity analyses; a short appendix table would help reproducibility.
- [Sec. 4.6.1] SSIM-C and L1-C (Sec. 4.6.1, Fig. S7) are useful but depend on human weight maps w; clarify how same-valence role reversal of w is applied when reporting aggregate fidelity across mixed trial types.
- [Sec. 4.2.4, Fig. 4F] Template prompts (Fig. 4F) are acknowledged as sometimes ungrammatical; quantify how often linguistic normalization fails and whether that affects human preference.
- [Sec. 4.5] Participant retention after attention controls varies substantially across experiments (e.g., 79→44 in Exp-EmoPrompt); report exclusion rates and any sensitivity of main results to inclusion criteria.
- [Sec. 3] Discussion ethics paragraph is appropriate but brief; a short note on dual-use and consent for personalized affective memory would fit the claimed mental-health applications.
Circularity Check
No load-bearing circular derivation; mild shared-source reuse of EmoSet for mining EmoTree and sampling evaluation images is present but does not force the human-preference or fidelity claims by construction.
-
other
[Sec. 4.1, 4.2.3 (EmoTree construction); Sec. 4.5 (source images from EmoSet test set)]
"Our framework builds upon EmoSet [11]... Applying the above procedure to all images in EmoSet and both target valence directions yields a large repository of structured affective rules... Unless otherwise stated, all source images were drawn from the EmoSet test set."
EmoTree motifs and the valence predictor are mined/trained on the full EmoSet distribution; evaluation images are sampled from the same dataset's test split. This is distributional reuse rather than a by-construction identity (human preference/success/IoU metrics remain independent external measurements), so it is only mild and non-load-bearing.
full rationale
The paper's central claims (superior human preference rates, cumulative success, SSIM-C/L1-C fidelity, and EmoMem personalization) rest on new forced-choice and annotation judgments collected from 296 participants across six psychophysics protocols, not on reconstruction of EmoSet training labels. EmoTree construction (CLIP clustering at δ=0.7, BLIP captions, GPT-4o concept/motif extraction) and the ResNet-18 valence predictor are data-driven from EmoSet statistics, and source images for most experiments are drawn from the EmoSet test set; this creates a mild distributional overlap that weakens the 'generalizable' claim but does not make any reported metric equal to its input by definition. Grad-CAM masks are compared to independent human region annotations (IoU 0.38 vs Human–Human 0.46); motif prompts are judged by humans for valence alignment (~90 %); edited images are judged for target-valence success and preference against LMS/XDream/Real baselines. Personalization memories are filled exclusively from live user accept/reject feedback at inference time, not from the offline dataset. No self-definitional equations, fitted-parameter-as-prediction, load-bearing self-citation uniqueness theorems, or ansatz smuggled via own prior work appear. The derivation chain is therefore empirically self-contained against external human benchmarks; the residual score of 1 reflects only the non-load-bearing train/test source overlap.
Assumptions & free parameters
free parameters (6)
- CLIP cluster cosine threshold δ =
0.7
- Grad-CAM progressive thresholds τ =
{0.5, 0.3, 0.1, 0}
- Emotion predictor acceptance threshold γ =
0.7
- Hard-rejection consecutive count =
>3 consecutive rejections
- Cross-valence similarity skip threshold =
0.8
- Candidate concept sample size per motif construction =
50
assumptions (5)
- domain assumption Eight EmoSet discrete emotions can be collapsed into binary positive vs negative valence without losing the phenomena needed for 'emotional intelligence' evaluation.
- domain assumption CLIP image embeddings and cosine hierarchical clustering recover semantic neighborhoods useful for transferring affective motifs across scenes.
- ad hoc to paper GPT-4o/GPT-4 can extract source concepts and compose target motifs (attributes, actions, scenes) that reflect human-interpretable affective structure rather than model-specific language priors.
- domain assumption A ResNet-18 binary classifier trained on EmoSet plus Grad-CAM provides a usable proxy for human emotion recognition and localization.
- domain assumption Human forced-choice preference under instructions that show a reference image jointly measures affective effectiveness and structural preservation as claimed.
invented entities (4)
-
EmoTree (hierarchical symbolic affective knowledge structure)
-
Visual motif m=(c_t, a_t, r_t, s_t) as executable affective rule
-
EmoMem (accepted/rejected dual memory banks with hard rejection)
-
SSIM-C and L1-C contrastive fidelity metrics
Cite this review
Pith. "Pith review of Personalized Emotional Intelligence in Generative AI through Symbolic Affective Reasoning." pith.science (2026). https://pith.science/paper/YAFKTXFK
@misc{pith2026260710678,
author = {Pith},
title = {Pith review of: Personalized Emotional Intelligence in Generative AI through Symbolic Affective Reasoning},
year = {2026},
howpublished = {\url{https://pith.science/paper/YAFKTXFK}},
note = {Machine review of arXiv:2607.10678}
}
read the original abstract
Emotional intelligence enables humans to recognize emotions, infer their causes, reason about interventions, and modify their environment to achieve desired affective states. Despite recent advances in artificial intelligence (AI), current models remain largely limited to generating realistic content or performing semantic reasoning, with little capacity for understanding, predicting, and personalizing human emotional responses. Here we introduce Emotion-augmented geneRatiOn System (EROS), a hybrid AI framework that integrates symbolic reasoning with deep learning to enable personalized emotion augmentation through visual content. Leveraging large-scale image-emotion datasets, EROS discovers generalizable affective rules, identifies emotion-relevant image regions, and predicts context-aware visual modifications that preserve scene semantics while steering emotional responses toward desired targets. To account for individual variability, EROS incorporates an expandable memory bank that supports inference-time personalization without model fine-tuning, yielding interpretable emotional profiles and rapid adaptation to new users. Across extensive human psychophysics experiments, EROS elicits target emotional responses more effectively than state-of-the-art large multimodal models while adapting to individual affective preferences. Beyond affective computing, EROS provides a foundation for AI systems that can understand, reason about, and augment human cognitive states, with potential applications in mental health, adaptive media, education, and human-computer interaction.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Target articles:
J. D. Mayer, P. Salovey, and D. R. Caruso, “Target articles:” emotional intelligence: Theory, findings, and implications”,”Psychological inquiry, vol. 15, no. 3, pp. 197–215, 2004
2004
-
[2]
Human abilities: Emotional intelligence,
J. D. Mayer, R. D. Roberts, and S. G. Barsade, “Human abilities: Emotional intelligence,”Annu. Rev. Psychol., vol. 59, no. 1, pp. 507–536, 2008
2008
-
[3]
Emotional intelligence as a standard intelligence.,
J. D. Mayer, P. Salovey, D. R. Caruso, and G. Sitarenios, “Emotional intelligence as a standard intelligence.,” 2001
2001
-
[4]
The positive psychology of emotional intelligence,
P. Salovey, J. D. Mayer, D. Caruso, and S. H. Yoo, “The positive psychology of emotional intelligence,”Emotional intelligence perspectives on educational and positive psychology, pp. 185–208, 2008
2008
-
[5]
Emotional pictures and sounds: A review of multimodal interactions of emotion cues in multiple domains,
A. B. Gerdes, M. J. Wieser, and G. W. Alpers, “Emotional pictures and sounds: A review of multimodal interactions of emotion cues in multiple domains,”Frontiers in psychology, vol. 5, p. 1351, 2014
2014
-
[6]
The interplay of signs and visuals: Unveiling the symbiotic relationship between semiotics and visual communication,
A. Travere, “The interplay of signs and visuals: Unveiling the symbiotic relationship between semiotics and visual communication,”Journal of Linguistics and Communication Studies, vol. 2, no. 3, pp. 28–40, 2023
2023
-
[7]
Perceptions of perceptual symbols,
L. W. Barsalou, “Perceptions of perceptual symbols,”Behavioral and brain sciences, vol. 22, no. 4, pp. 637–660, 1999
1999
-
[8]
The shaping of social perception by stimulus and knowledge cues to human animacy,
E. S. Cross, R. Ramsey, R. Liepelt, W. Prinz, and A. F. d. C. Hamilton, “The shaping of social perception by stimulus and knowledge cues to human animacy,”Philosophical Transactions of the Royal Society B: Biological Sciences, vol. 371, no. 1686, p. 20 150 075, 2016
2016
Show all 97 references
-
[9]
A new look at emotion perception: Concepts speed and shape facial emotion recognition.,
E. C. Nook, K. A. Lindquist, and J. Zaki, “A new look at emotion perception: Concepts speed and shape facial emotion recognition.,”Emotion, vol. 15, no. 5, p. 569, 2015. 38
2015
-
[10]
Perceptions of visual and multimodal symbolic mediated social touch: Role of technology modality, relationship, and task emotional salience,
S. Yarosh, X. Wang, and Y. Yao, “Perceptions of visual and multimodal symbolic mediated social touch: Role of technology modality, relationship, and task emotional salience,”International Journal of Human-Computer Studies, vol. 159, p. 102 757, 2022
2022
-
[11]
Emoset: A large-scale visual emotion dataset with rich attributes,
J. Yang et al., “Emoset: A large-scale visual emotion dataset with rich attributes,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 20 383–20 394
2023
-
[12]
How emotions are made,
L. F. Barrett, “How emotions are made,”Providence Book Festival, May, vol. 25, p. 2019, 2019
2019
-
[13]
Appraisal processes in emotion,
P. C. Ellsworth and K. R. Scherer, “Appraisal processes in emotion,” 2003
2003
-
[14]
Psychological construction in the occ model of emotion,
G. L. Clore and A. Ortony, “Psychological construction in the occ model of emotion,”Emotion Review, vol. 5, no. 4, pp. 335–343, 2013
2013
-
[15]
Context in emotion perception,
L. F. Barrett, B. Mesquita, and M. Gendron, “Context in emotion perception,”Current directions in psychological science, vol. 20, no. 5, pp. 286–290, 2011
2011
-
[16]
Inferential emotion tracking (iet) reveals the critical role of context in emotion recognition.,
Z. Chen and D. Whitney, “Inferential emotion tracking (iet) reveals the critical role of context in emotion recognition.,”Emotion, vol. 22, no. 6, p. 1185, 2022
2022
-
[17]
Object-scene semantics correlation analysis for image emotion classification,
Z. Zhou, Z. Zhai, H. Chen, and S. Lu, “Object-scene semantics correlation analysis for image emotion classification,”Frontiers in Neuroscience, vol. 19, p. 1 657 562, 2025
2025
-
[18]
Generative models improve fairness of medical classifiers under distribution shifts,
I. Ktena et al., “Generative models improve fairness of medical classifiers under distribution shifts,”Nature Medicine, vol. 30, no. 4, pp. 1166–1173, 2024
2024
-
[19]
“it happened to be the perfect thing
S. Siddals, J. Torous, and A. Coxon, ““it happened to be the perfect thing”: Experiences of generative ai chatbots for mental health,”npj Mental Health Research, vol. 3, no. 1, p. 48, 2024
2024
-
[20]
Generative ai for designing and validating easily synthesizable and structurally novel antibiotics,
K. Swanson et al., “Generative ai for designing and validating easily synthesizable and structurally novel antibiotics,”Nature machine intelligence, vol. 6, no. 3, pp. 338–353, 2024
2024
-
[21]
Integrating curricula with replays: Its effects on continual learning,
R. J. Tee and M. Zhang, “Integrating curricula with replays: Its effects on continual learning,” inProceedings of the AAAI Symposium Series, vol. 1, 2023, pp. 109–116
2023
-
[22]
Tuned compositional feature replays for efficient stream learning,
M. B. Talbot, R. Zawar, R. Badkundri, M. Zhang, and G. Kreiman, “Tuned compositional feature replays for efficient stream learning,”IEEE Transactions on Neural Networks and Learning Systems, vol. 36, no. 2, pp. 3300–3314, 2023
2023
-
[23]
Unveiling the tapestry: The interplay of generalization and forgetting in continual learning,
Z. Shi, J. Jing, Y. Sun, J.-H. Lim, and M. Zhang, “Unveiling the tapestry: The interplay of generalization and forgetting in continual learning,”IEEE Transactions on Neural Networks and Learning Systems, 2025
2025
-
[24]
Learning to learn: How to continuously teach humans and machines,
P. Singh et al., “Learning to learn: How to continuously teach humans and machines,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 11 708–11 719
2023
-
[25]
Imagenet classification with deep convolutional neural networks,
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,”Advances in neural information processing systems, vol. 25, 2012. 39
2012
-
[26]
Very deep convolutional networks for large-scale image recognition,
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,”arXiv preprint arXiv:1409.1556, 2014
2014 arXiv
-
[27]
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778
2016
-
[28]
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv preprint arXiv:2010.11929, 2020
2010 arXiv
-
[29]
Efficient zero-shot visual search via target and context-aware transformer,
Z. Ding et al., “Efficient zero-shot visual search via target and context-aware transformer,”arXiv preprint arXiv:2211.13470, 2022
2022 arXiv
-
[30]
Object-centric learning with cyclic walks between parts and whole,
Z. Wang, M. Z. Shou, and M. Zhang, “Object-centric learning with cyclic walks between parts and whole,”Advances in Neural Information Processing Systems, vol. 36, pp. 9388–9408, 2023
2023
-
[31]
Flow snapshot neurons in action: Deep neural networks generalize to biological motion perception,
S. Han, Z. Wang, and M. Zhang, “Flow snapshot neurons in action: Deep neural networks generalize to biological motion perception,”Advances in Neural Information Processing Systems, vol. 37, pp. 53 732–53 763, 2024
2024
-
[32]
Learning to see through a baby’s eyes: Early visual diets enable robust visual intelligence in humans and machines,
Y. Cai, B. S. Nunna, Q. Lin, and M. Zhang, “Learning to see through a baby’s eyes: Early visual diets enable robust visual intelligence in humans and machines,”arXiv preprint arXiv:2511.14440, 2025
2025
-
[33]
Label-efficient online continual object detection in streaming video,
J. Z. Wu, D. J. Zhang, W. Hsu, M. Zhang, and M. Z. Shou, “Label-efficient online continual object detection in streaming video,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 19 246–19 255
2023
-
[34]
Pose prior learner: Unsupervised categorical prior learning for pose estimation,
Z. Wang, S. Han, and M. Zhang, “Pose prior learner: Unsupervised categorical prior learning for pose estimation,”arXiv preprint arXiv:2410.03858, 2024
2024
-
[35]
Gazing at rewards: Eye movements as a lens into human and ai decision-making in hybrid visual foraging,
B. Wang et al., “Gazing at rewards: Eye movements as a lens into human and ai decision-making in hybrid visual foraging,” inProceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 14 810–14 823
2025
-
[36]
A multimodal generative ai copilot for human pathology,
M. Y. Lu et al., “A multimodal generative ai copilot for human pathology,”Nature, vol. 634, no. 8033, pp. 466–473, 2024
2024
-
[37]
Seeing sound, hearing sight: Uncovering modality bias and conflict of ai models in sound localization,
Y. Jia et al., “Seeing sound, hearing sight: Uncovering modality bias and conflict of ai models in sound localization,”arXiv preprint arXiv:2505.11217, 2025
2025
-
[38]
Adaptive visual scene understanding: Incremental scene graph generation,
N. Khandelwal, X. Liu, and M. Zhang, “Adaptive visual scene understanding: Incremental scene graph generation,”arXiv preprint arXiv:2310.01636, 2023
2023 arXiv
-
[39]
Putting visual object recognition in context,
M. Zhang, C. Tseng, and G. Kreiman, “Putting visual object recognition in context,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 12 985–12 994. 40
2020
-
[40]
When pigs fly: Contextual reasoning in synthetic and natural scenes,
P. Bomatter et al., “When pigs fly: Contextual reasoning in synthetic and natural scenes,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 255–264
2021
-
[41]
Reason from context with self-supervised learning,
X. Liu, A. Sikarwar, G. Kreiman, Z. Shi, and M. Zhang, “Reason from context with self-supervised learning,”arXiv preprint arXiv:2211.12817, 2022
2022
-
[42]
Affectgan: Affect-based generative art driven by semantics,
T. Galanos, A. Liapis, and G. N. Yannakakis, “Affectgan: Affect-based generative art driven by semantics,” in2021 9th International Conference on Affective Computing and Intelligent Interaction Workshops and Demos (ACIIW), IEEE, 2021, pp. 01–07
2021
-
[43]
Automated colour grading using colour distribution transfer,
F. Piti´ e, A. C. Kokaram, and R. Dahyot, “Automated colour grading using colour distribution transfer,”Computer Vision and Image Understanding, vol. 107, no. 1-2, pp. 123–137, 2007
2007
-
[44]
L. A. Gatys, A. S. Ecker, and M. Bethge,A neural algorithm of artistic style, 2015. arXiv: 1508.06576 [cs.CV]
2015 arXiv
-
[45]
Clipstyler: Image style transfer with a single text condition,
G. Kwon and J. C. Ye, “Clipstyler: Image style transfer with a single text condition,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 18 062–18 071
2022
-
[46]
Affective image filter: Reflecting emotions from text to images,
S. Weng et al., “Affective image filter: Reflecting emotions from text to images,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 10 810–10 819
2023
-
[47]
Make me happier: Evoking emotions through image diffusion models,
Q. Lin, J. Zhang, Y.-S. Ong, and M. Zhang, “Make me happier: Evoking emotions through image diffusion models,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2025, pp. 16 367–16 376
2025
-
[48]
Emoedit: Evoking emotions through image manipulation,
J. Yang et al., “Emoedit: Evoking emotions through image manipulation,” inProceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 24 690–24 699
2025
-
[49]
Emokgedit: Training-free affective injection via visual cue transformation,
J. Zhang and B. Fan, “Emokgedit: Training-free affective injection via visual cue transformation,” arXiv preprint arXiv:2601.12326, 2026
2026
-
[50]
Affective image editing: Shaping emotional factors via text descriptions,
P. Zhang et al., “Affective image editing: Shaping emotional factors via text descriptions,” International Journal of Computer Vision, vol. 134, no. 1, p. 16, 2026
2026
-
[51]
Large-scale visual sentiment ontology and detectors using adjective noun pairs,
D. Borth, R. Ji, T. Chen, T. Breuel, and S.-F. Chang, “Large-scale visual sentiment ontology and detectors using adjective noun pairs,” inProceedings of the 21st ACM international conference on Multimedia, 2013, pp. 223–232
2013
-
[52]
Visual affect around the world: A large-scale multilingual visual sentiment ontology,
B. Jou et al., “Visual affect around the world: A large-scale multilingual visual sentiment ontology,” inProceedings of the 23rd ACM international conference on Multimedia, 2015, pp. 159–168
2015
-
[53]
Developing affective lexical resources.,
A. Valitutti, C. Strapparava, and O. Stock, “Developing affective lexical resources.,”PsychNology J., vol. 2, no. 1, pp. 61–83, 2004. 41
2004
-
[54]
The nonverbal affect lexicon: Theoretical perspectives from neuropsychological studies of affect perception.,
D. Bowers, R. M. Bauer, and K. M. Heilman, “The nonverbal affect lexicon: Theoretical perspectives from neuropsychological studies of affect perception.,”Neuropsychology, vol. 7, no. 4, p. 433, 1993
1993
-
[55]
Affective computing and sentiment analysis,
E. Cambria, D. Das, S. Bandyopadhyay, and A. Feraco, “Affective computing and sentiment analysis,” inA practical guide to sentiment analysis, Springer, 2017, pp. 1–10
2017
-
[56]
Robust image sentiment analysis using progressively trained and domain transferred deep networks,
Q. You, J. Luo, H. Jin, and J. Yang, “Robust image sentiment analysis using progressively trained and domain transferred deep networks,” inProceedings of the AAAI conference on Artificial Intelligence, vol. 29, 2015
2015
-
[57]
Emobank: Studying the impact of annotation perspective and representation format on dimensional emotion analysis,
S. Buechel and U. Hahn, “Emobank: Studying the impact of annotation perspective and representation format on dimensional emotion analysis,” inProceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers, 20...
2017
-
[58]
Emotions evoked by common words and phrases: Using mechanical turk to create an emotion lexicon,
S. Mohammad and P. Turney, “Emotions evoked by common words and phrases: Using mechanical turk to create an emotion lexicon,” inProceedings of the NAACL HLT 2010 workshop on computational approaches to analysis and generation of emotion in text, 2010, pp. 26–34
2010
-
[59]
Crowdsourcing a word–emotion association lexicon,
S. M. Mohammad and P. D. Turney, “Crowdsourcing a word–emotion association lexicon,” Computational intelligence, vol. 29, no. 3, pp. 436–465, 2013
2013
-
[60]
Gpt-4o system card,
A. Hurst et al., “Gpt-4o system card,”arXiv preprint arXiv:2410.21276, 2024
2024 arXiv
-
[61]
OpenAI,Introducing openai o3 and o4-mini: Our smartest and most capable models to date with full tool access,https://openai.com/index/introducing-o3-and-o4-mini/, 2025
2025
-
[62]
com / index / introducing - 4o - image-generation/, 2025
OpenAI,Introducing 4o image generation,https : / / openai . com / index / introducing - 4o - image-generation/, 2025
2025
-
[63]
Fortin, Alisa and Vernade, Guillaume and Kampf, Kat and Reshi, Ammaar,Introducing gemini 2.5 flash image, our state-of-the-art image model,https://developers.googleblog.com/en/ introducing-gemini-2-5-flash-image/, 2025
2025
-
[64]
Capacity of generative ai to interpret human emotions from visual and textual data: Pilot evaluation study,
Z. Elyoseph et al., “Capacity of generative ai to interpret human emotions from visual and textual data: Pilot evaluation study,”JMIR Mental Health, vol. 11, e54369, 2024
2024
-
[65]
Do llms” feel
C. Wang et al., “Do llms” feel”? emotion circuits discovery and control,”arXiv preprint arXiv:2510.11328, 2025
2025
-
[66]
Emotion concepts and their function in a large language model,
N. Sofroniew et al., “Emotion concepts and their function in a large language model,” Anthropic/Transformer Circuits. transformer-circuits. pub/2026/emotions, 2026
2026
-
[67]
Chain-of-thought is not explainability,
F. Barez et al., “Chain-of-thought is not explainability,”Preprint, alphaXiv, p. v1, 2025
2025
-
[68]
Gpt-4v with emotion: A zero-shot benchmark for generalized emotion recognition,
Z. Lian et al., “Gpt-4v with emotion: A zero-shot benchmark for generalized emotion recognition,”Information Fusion, vol. 108, p. 102 367, 2024. 42
2024
-
[69]
Research can help to tackle ai-generated disinformation,
S. Feuerriegel et al., “Research can help to tackle ai-generated disinformation,”Nature Human Behaviour, vol. 7, no. 11, pp. 1818–1821, 2023
2023
-
[70]
Emotionhallucer: Evaluating emotion hallucinations in multimodal large language models,
B. Xing et al., “Emotionhallucer: Evaluating emotion hallucinations in multimodal large language models,”arXiv preprint arXiv:2505.11405, 2025
2025 arXiv
-
[71]
Signs of consciousness in ai: Can gpt-3 tell how smart it really is?
L. Boji´ c, I. Stojkovi´ c, and Z. Joli´ c Marjanovi´ c, “Signs of consciousness in ai: Can gpt-3 tell how smart it really is?”Humanities and Social Sciences Communications, vol. 11, no. 1, pp. 1–15, 2024
2024
-
[72]
Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting,
M. Turpin, J. Michael, E. Perez, and S. Bowman, “Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting,”Advances in Neural Information Processing Systems, vol. 36, pp. 74 952–74 965, 2023
2023
-
[73]
A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions,
L. Huang et al., “A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions,”ACM Transactions on Information Systems, vol. 43, no. 2, pp. 1–55, 2025
2025
-
[74]
The emotional intelligence of the gpt-4 large language model,
G. D. Vzorinab, A. M. Bukinichac, A. V. Sedykha, I. I. Vetrovab, and E. A. Sergienkob, “The emotional intelligence of the gpt-4 large language model,”Psychology in Russia: State of the art, vol. 17, no. 2, pp. 85–99, 2024
2024
-
[75]
Hallucination of multimodal large language models: A survey,
Z. Bai et al., “Hallucination of multimodal large language models: A survey,”arXiv preprint arXiv:2404.18930, 2024
2024 arXiv
-
[76]
The other you in black mirror: First steps from chatbots to personalized llm clones,
M. Sun, M. Zhang, and G. Kreiman, “The other you in black mirror: First steps from chatbots to personalized llm clones,” 2025
2025
-
[77]
Ai-based personalized e-learning systems: Issues, challenges, and solutions,
M. Murtaza, Y. Ahmed, J. A. Shamsi, F. Sherwani, and M. Usman, “Ai-based personalized e-learning systems: Issues, challenges, and solutions,”IEEE access, vol. 10, pp. 81 323–81 342, 2022
2022
-
[78]
Exploring the impact of artificial intelligence application in personalized learning environments: Thematic analysis of undergraduates’ perceptions in china,
X. Wang, X. Xu, Y. Zhang, S. Hao, and W. Jie, “Exploring the impact of artificial intelligence application in personalized learning environments: Thematic analysis of undergraduates’ perceptions in china,”Humanities and Social Sciences Communications, vol. 11, no. 1, pp. 1–10, 2024
2024
-
[79]
Generative ai and gamification for personalized learning: Literature review and future challenges,
F. Abbes, S. Bennani, and A. Maalel, “Generative ai and gamification for personalized learning: Literature review and future challenges,”SN Computer Science, vol. 5, no. 8, p. 1154, 2024
2024
-
[80]
Ai-generated characters for supporting personalized learning and well-being,
P. Pataranutaporn et al., “Ai-generated characters for supporting personalized learning and well-being,”Nature Machine Intelligence, vol. 3, no. 12, pp. 1013–1022, 2021
2021
-
[81]
Closing the accessibility gap to mental health treatment with a personalized self-referral chatbot,
J. Habicht et al., “Closing the accessibility gap to mental health treatment with a personalized self-referral chatbot,”Nature medicine, vol. 30, no. 2, pp. 595–602, 2024
2024
-
[82]
Daga, Soham and Sreedhar, Sreeram and Shah, Dhravya,Supermemory is the new state-of-the-art in agent memory,https://supermemory.ai/research/. 43
-
[83]
Instructpix2pix: Learning to follow image editing instructions,
T. Brooks, A. Holynski, and A. A. Efros, “Instructpix2pix: Learning to follow image editing instructions,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 18 392–18 402
2023
-
[84]
Sdedit: Guided image synthesis and editing with stochastic differential equations,
C. Meng et al., “Sdedit: Guided image synthesis and editing with stochastic differential equations,”arXiv preprint arXiv:2108.01073, 2021
2021 arXiv
-
[85]
Adding conditional control to text-to-image diffusion models,
L. Zhang, A. Rao, and M. Agrawala, “Adding conditional control to text-to-image diffusion models,” inProceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 3836–3847
2023
-
[86]
Blip-diffusion: Pre-trained subject representation for controllable text-to-image generation and editing,
D. Li, J. Li, and S. Hoi, “Blip-diffusion: Pre-trained subject representation for controllable text-to-image generation and editing,”Advances in Neural Information Processing Systems, vol. 36, pp. 30 146–30 166, 2023
2023
-
[87]
Tolstoy,Anna karenina
L. Tolstoy,Anna karenina. Lulu. com, 2016
2016
-
[88]
Grad-cam: Visual explanations from deep networks via gradient-based localization,
R. R. Selvaraju et al., “Grad-cam: Visual explanations from deep networks via gradient-based localization,” inProceedings of the IEEE international conference on computer vision, 2017, pp. 618–626
2017
-
[89]
Segment anything,
A. Kirillov et al., “Segment anything,” inProceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 4015–4026
2023
-
[90]
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,
J. Li, D. Li, C. Xiong, and S. Hoi, “Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” inInternational conference on machine learning, PMLR, 2022, pp. 12 888–12 900
2022
-
[91]
Spacy: Industrial-strength natural language processing in python,
M. Honnibal, I. Montani, S. Van Landeghem, A. Boyd, et al., “Spacy: Industrial-strength natural language processing in python,” 2020
2020
-
[92]
High-resolution image synthesis with latent diffusion models,
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 10 684–10 695
2022
-
[93]
Evolving images for visual neurons using a deep generative network reveals coding principles and neuronal preferences,
C. R. Ponce et al., “Evolving images for visual neurons using a deep generative network reveals coding principles and neuronal preferences,”Cell, vol. 177, no. 4, pp. 999–1009, 2019
2019
-
[94]
Learning transferable visual models from natural language supervision,
A. Radford et al., “Learning transferable visual models from natural language supervision,” in International conference on machine learning, PMLR, 2021, pp. 8748–8763
2021
-
[95]
Amazon mechanical turk,
A. M. Turk, “Amazon mechanical turk,”Retrieved August, vol. 17, p. 2012, 2012
2012
-
[96]
Image quality assessment: From error visibility to structural similarity,
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: From error visibility to structural similarity,”IEEE transactions on image processing, vol. 13, no. 4, pp. 600–612, 2004
2004
-
[97]
Microsoft coco: Common objects in context,
T.-Y. Lin et al., “Microsoft coco: Common objects in context,” inEuropean conference on computer vision, Springer, 2014, pp. 740–755. 44 Main Figures Figure 1:Operationalizing emotional intelligence in humans and machines. A, A concrete example illustrating emotional intellige...
2014
Reviewed July 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.