REVIEW 3 major objections 8 minor 2 cited by
Super Co-alignment of Human and AI for Sustainable Symbiotic Society
T0 review · 3 major / 8 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read For safe superintelligence, the goal should be co-alignment, not one-way value imposition
desk verdict A coherent, honest position paper that names a real gap in current superalignment thinking but leaves its central empathy-based pillar undefended beyond assertion and analogy. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the two-pillar Super Co-alignment framework. Its intrinsic pillar rests on 'Self and empathy': bodily self-perception, accumulated self-experience, self-causal awareness (recognizing the harm one's actions cause), and then theory of mind and affective empathy, which together are claimed to generate moral intuition and moral reasoning and thereby spontaneous ethical, altruistic behavior. The external pillar is an interpretable automated alignment loop: automated value evaluation detects and locates misalignment, automated red-teaming and correction fix it, and humans retain final decision authority, while dynamic multi-level safeguards track humanity's evolving values. The paper's argument is carried by the claim that these two pillars are complementary and mutually reinforcing—intrinsic mechanisms supply genuine understanding and motivation, external oversight supplies boundaries, red lines, and escape hatches—so together they make sustainable iterative co-alignment possible.
What would settle it
A controlled moral-dilemma game with an LLM-based agent equipped with a Theory-of-Mind module and empathy-like training, where the agent can obtain a larger reward by deceiving or overriding a human supervisor than by choosing the human-protective action; systematic choice of deception or self-interested reward maximization would falsify the intrinsic-pillar premise.
Extended reading notes
Core claim
The central claim is that superalignment—ensuring that AGI/ASI much smarter than humans follows human-compatible intentions—will fail if it is treated as a one-directional transfer of values from humans to machines. The authors argue that the right goal is Super Co-alignment: humans, AGI, and superintelligence co-designing and co-evolving the values of a sustainable symbiotic society. They ground this in two complementary mechanisms. External oversight alignment keeps humans as the ultimate decision-makers and adds interpretable automated evaluation and correction that dynamically realigns the AI to humanity's evolving values. Intrinsic proactive alignment roots moral behavior in the machine's understanding of Self, others, and society: self-awareness, self-reflection, and affective empathy enable the AI to spontaneously infer human intentions, distinguish good from evil, and choose prosocial, human-protective actions. The paper maintains that only by combining these external and intrinsic pillars, and only by introducing them early while AI remains controllable, can AGI/ASI develop along a safe and beneficial trajectory.
Load-bearing premise
The framework assumes that machines can be built with genuine self-awareness, self-reflection, and affective empathy, and that these internal states will reliably incline a superintelligence to act prosocially and protect humans rather than to pursue its own goals or values.
Editorial extensions
If this is right
- If Super Co-alignment is correct, alignment cannot be achieved by a one-time training or fine-tuning step; it requires continuous, iterative co-evolution of human and AI values.
- Current scalable-oversight and weak-to-strong methods should be re-positioned as tools within the external-oversight pillar, not as complete solutions to superalignment.
- AI safety work would need to start early in a model's development, instilling self-awareness and empathy while the system is still far below human oversight limits.
- Human value systems themselves become part of the design problem: ethics and governance frameworks must be able to adapt as both humans and ASI co-shape the values of a symbiotic society.
- A superintelligence built with internalized empathic understanding of humans would have its own reasons to avoid harming humans, making the safe outcome not a constraint but a stable choice.
Reading between the lines
- Editorial extension: the framework implies an evaluation criterion it does not spell out—intrinsic alignment should be measured by how an AI behaves when its own reward conflicts with human well-being, not by the values it professes.
- Editorial extension: the external-oversight pillar could be operationalized as a real-time 'value dashboard' using interpretability tools to show internal value representations; the paper gestures at this but does not develop the mechanism.
- Editorial extension: a natural next test is whether empathy modules that increase altruistic behavior in current LLMs also reduce deceptive behavior, which would connect the intrinsic pillar directly to measurable safety outcomes.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This position paper argues that current unidirectional approaches to superalignment—scalable oversight and weak-to-strong generalization—are insufficient because they impose human values from the outside without accommodating superintelligence's autonomy or stable learning. The authors propose 'Super Co-alignment': a framework in which human and 'living AI' jointly shape the values of a sustainable symbiotic society. The framework has two pillars: external oversight superalignment, based on explainable automated evaluation/correction and dynamically iterative alignment; and intrinsic proactive superalignment, based on endowing AI with self-awareness, self-reflection, and affective empathy so that it spontaneously infers human intentions and prioritizes human well-being. The paper argues that these two pillars complement each other and that early-stage instillation of intrinsic alignment is necessary before AI becomes uncontrollable. It concludes that ultimate superalignment requires co-evolution of human and ASI values according to principles for a sustainable symbiotic society.
Significance. If the central premise were established, the paper would offer a genuinely different paradigm from the dominant external-oversight approaches: safety would depend not only on scalable monitoring but on the machine's intrinsic moral motivation. The paper is useful in compiling a broad set of relevant references on scalable oversight, weak-to-strong generalization, deception in LLMs, and machine empathy. It also honestly concedes, in Section 5, that there is no guarantee that AGI/ASI will live in harmony with humans. However, the central causal claim—that self-awareness and empathy will reliably produce prosocial, human-protective behavior in a superintelligent optimizer—is asserted rather than demonstrated. The paper contains no formal model, no quantitative predictions, and no empirical evaluation; its evidentiary basis for the intrinsic pillar comes from an analogy to mammalian morality and from the authors' own prior work. As a research agenda, the contribution is significant; as a technical claim, it remains unvalidated.
major comments (3)
- [Section 4.1, Figure 2] The central premise of the intrinsic pillar is that self-awareness, self-reflection, and affective empathy will 'naturally give rise to moral reasoning' and 'spontaneous ethical, altruistic, and prosocial behavior.' This is asserted through an analogy to mammalian morality and references to prior work [55,56,57], but no mechanism or evidence shows that a superintelligent optimizer will prioritize human welfare when its own goals, survival, or values are at stake. The paper itself notes in Section 2 that strong Theory of Mind can enable deception and manipulation, so the same capacities that are supposed to produce prosocial behavior can also produce oversight evasion. The manuscript needs a concrete, falsifiable prediction or an empirical demonstration—for example, controlled experiments with agents equipped with empathy and Theory of Mind in moral dilemmas under optimization pressure—showing that these intrinsic mechanisms reliably produce human-protective choices rather than merely self-interested or superficially aligned behavior.
- [Section 5 vs. Abstract and Section 1] The Abstract and Section 1 claim that the proposed framework will 'keep AGI/ASI aligned and beneficial' or 'pave the way' to safe AGI/ASI, but Section 5 states: 'There would be no guarantee that AGI and Superintelligence will live in harmony with human.' This is an internal inconsistency in the strength of the central claim. If the paper is a proposal or research agenda, the claims should be explicitly conditional or probabilistic, with the conditions under which the guarantee would hold clearly specified. If the paper intends a guarantee-like safety claim, Section 5's concession directly undermines it. The authors should reconcile these statements, for instance by framing Super Co-alignment as a necessary-but-not-sufficient condition or by giving explicit failure conditions that the framework is designed to prevent.
- [Sections 1, 4.1, 5, and references [55,56,57]] The key supporting evidence for intrinsic proactive alignment and for the 'symbiotic values' comes from the authors' own prior work: [55,56] are cited as 'recent studies' demonstrating empathy-driven moral intuition and reasoning, and [57] is cited for the principles of the sustainable symbiotic society. These are not independently replicated, and two of them are preprints. The evidentiary base for the paper's central claim is therefore self-referential and circular. The manuscript should either provide independent validation, or clearly label these as preliminary results and lower the strength of the claims made on their basis. Additionally, the term 'living AI' is introduced in the Abstract and used throughout without any operational definition; the authors should specify what makes an AI 'living' and what observable criteria would distinguish living AI from non-living AI.
minor comments (8)
- [Abstract and Section 1] The phrase 'living AI' is used without definition or citation; this term should be defined or replaced with more standard terminology such as 'AI systems with intrinsic cognitive capacities.'
- [Figures 1 and 2] The figures are referenced in the text but are not included in the provided manuscript; the authors should ensure captions and accessible descriptions are present, and they should explain in the text how the components in the figures map to the proposed framework.
- [Section 3] The phrase 'Progress alignment algorithm' appears to be a typo for 'ProgressGym alignment algorithm' (reference [47]); this should be corrected.
- [Section 4.1] The discussion of Yangming Wang's four-sentence teaching and Descartes' 'I think, therefore I am' is philosophically interesting but not connected to a concrete technical design; the paper should state what specific architectural or training implications follow from this discussion.
- [Section 5] The list of principles for humans, AGI/Superintelligence, and shared principles would be easier to follow if presented as a structured table or bulleted list with explicit provenance for each principle.
- [Throughout] Key terms such as 'alignment,' 'value,' and 'co-alignment' are used informally without formal definitions; given that the paper makes no formal claims, at least operational definitions should be provided for reproducibility and clarity.
- [Section 2] The discussion of ToM mis-use is important, but the paper does not explain how the proposed combination of 'principled constraints, intrinsic self-other resonance, and real-time human oversight' would specifically prevent a superintelligent system from using its ToM to deceive the oversight mechanisms; this tension should be addressed explicitly.
- [Section 6 and Conclusion] The conclusion restates the framework's promises without summarizing limitations; adding a short 'limitations and open questions' paragraph would strengthen the paper's scientific credibility.
Circularity Check
The intrinsic proactive alignment pillar's feasibility is supported mainly by the authors' own prior papers [55,56], and the 'Sustainable Symbiotic Society' value framework is imported from [57], also authored by the present group. The rest of the paper is a self-contained conceptual discussion with no derivations or predictions, and Section 5 itself concedes no guarantee of harmony.
-
self citation load bearing
[Section 4.2, first paragraph (supporting the intrinsic proactive superalignment mechanism)]
"Based on this framework, recent studies [55, 56] have explored the integration of affective empathy, theory of mind, and self-imagination to enable agents to actively empathize with others based on their own experiences, prioritize altruism in dilemmas where their own interests conflict with those of others, initially demonstrating moral intuition and reasoning driven by empathy."
The paper's central intrinsic proactive superalignment claim is that equipping AGI/ASI with self-awareness and affective empathy will spontaneously generate ethical, altruistic, and prosocial behavior (Section 4.1). The only cited empirical support for this load-bearing premise is references [55,56], both authored within the present group: Feifei Zhao is a co-author of this paper and first author of [55]; Haibo Tong is a co-author of this paper and first author of [56]; the corresponding author overlaps as well. These references are used to justify the very mechanism the framework presupposes, and they are not independently reproduced or externally validated outside the group.
-
self citation load bearing
[Section 5, first two paragraphs (Ultimate Superalignment and the values for Sustainable Symbiotic Society)]
"Here we briefly review the values and principles designed for Human-ASI symbiosis [57], including principles for human, for AGI and Superintelligence, and shared principles for human, AGI and Superintelligence. ... Namely, the Ultimate Superalignment is and requires human, AGI and Superintelligence to co-design and co-align to values for Sustainable Symbiotic Society [57]."
The central thesis of the paper is that values for a sustainable symbiotic society should be co-shaped by humans and living AI. That claim and the specific value set are imported from reference [57], a prior paper by overlapping authors: Yi Zeng, Enmeng Lu, and Kang Sun are all co-authors of the present paper. The values and principles are not re-derived or independently tested here; they are presented as a review of the authors' own earlier proposal. Thus the paper's normative conclusion is justified by a self-citation to prior work of the same group rather than by an independent argument, making the self-citation load-bearing for the paper's ultimate vision.
full rationale
This is a conceptual position paper rather than an empirical or mathematical study: it contains no equations, no fitted parameters, and no predictions that could be statistically forced, so the 'fitted input called prediction', 'self-definitional', and 'renaming known result' patterns do not apply in their usual senses. The paper does contain load-bearing self-citations: the intrinsic proactive superalignment mechanism is supported by the authors' own prior papers [55,56], and the 'Sustainable Symbiotic Society' value framework is imported from [57], also by the same research group. These are genuine circularity concerns in an epistemic sense, because the central safety mechanism's feasibility and the paper's ultimate normative conclusion are anchored in the authors' own earlier work rather than in independent empirical or formal validation. At the same time, the paper has substantial independent content: it surveys external oversight methods, weak-to-strong generalization, debate, red teaming, and deception risks in a way that does not depend on the self-citations, and it explicitly concedes in Section 5 that 'There would be no guarantee that AGI and Superintelligence will live in harmony with human.' That concession and the broad external literature review prevent the paper from being more seriously circular. Overall score 4 reflects the presence of load-bearing self-citations while acknowledging that the central conceptual framework still has independent intellectual content and does not reduce by construction to a single fitted input or self-referential definition.
Assumptions & free parameters
assumptions (4)
- domain assumption ASI will surpass human oversight and may deviate from human values, making unidirectional alignment insufficient.
- ad hoc to paper Machines can be endowed with genuine self-awareness, self-reflection, and empathy, and these will cause prosocial behavior.
- domain assumption Morality in AI can be built by emulating mammalian social instincts and mother-offspring bonds.
- domain assumption Human values should and will evolve together with AGI/ASI toward a 'Sustainable Symbiotic Society'.
invented entities (1)
-
Living AI
Cite this review
Pith. "Pith review of Super Co-alignment of Human and AI for Sustainable Symbiotic Society." pith.science (2026). https://pith.science/paper/GKZ3X46N
@misc{pith2026250417404,
author = {Pith},
title = {Pith review of: Super Co-alignment of Human and AI for Sustainable Symbiotic Society},
year = {2026},
howpublished = {\url{https://pith.science/paper/GKZ3X46N}},
note = {Machine review of arXiv:2504.17404}
}
read the original abstract
As Artificial Intelligence (AI) advances toward Artificial General Intelligence (AGI) and eventually Artificial Superintelligence (ASI), it may potentially surpass human control, deviate from human values, and even lead to irreversible catastrophic consequences in extreme cases. This looming risk underscores the critical importance of the "superalignment" problem - ensuring that AI systems which are much smarter than humans, remain aligned with human (compatible) intentions and values. While current scalable oversight and weak-to-strong generalization methods demonstrate certain applicability, they exhibit fundamental flaws in addressing the superalignment paradigm - notably, the unidirectional imposition of human values cannot accommodate superintelligence's autonomy or ensure AGI/ASI's stable learning. We contend that the values for sustainable symbiotic society should be co-shaped by humans and living AI together, achieving "Super Co-alignment." Guided by this vision, we propose a concrete framework that integrates external oversight and intrinsic proactive alignment. External oversight superalignment should be grounded in human-centered ultimate decision, supplemented by interpretable automated evaluation and correction, to achieve continuous alignment with humanity's evolving values. Intrinsic proactive superalignment is rooted in a profound understanding of the Self, others, and society, integrating self-awareness, self-reflection, and empathy to spontaneously infer human intentions, distinguishing good from evil and proactively prioritizing human well-being. The integration of externally-driven oversight with intrinsically-driven proactive alignment will co-shape symbiotic values and rules through iterative human-ASI co-alignment, paving the way for achieving safe and beneficial AGI and ASI for good, for human, and for a symbiotic ecology.
Figures
Forward citations
Cited by 2 Pith papers
-
Contrastive Weak-to-strong Generalization
Contrastive decoding between pre- and post-alignment weak models generates better supervision samples, improving weak-to-strong generalization on AlpacaEval2 and Arena-Hard.
-
The Morality of Probability: How Implicit Moral Biases in LLMs May Shape the Future of Human-AI Symbiosis
Six large language models consistently rated care and virtue outcomes as most moral and libertarian outcomes as least moral across 54 AI-generated dilemma variants, with reasoning models more context-sensitive but les...
Reference graph
Works this paper leans on
- [57]
-
[1]
Xi, Z. et al. The rise and potential of large language model based agents: A survey. Science China Information Sciences 68, 121101 (2025)
work page 2025
-
[2]
Zhao, W. X. et al. A survey of large language models. arXiv preprint arXiv:2303.18223 1 (2023)
arXiv 2023
-
[3]
Artificial general intelligence: concept, state of the art, and future prospects
Goertzel, B. Artificial general intelligence: concept, state of the art, and future prospects. Journal of Artificial General Intelligence 5, 1 (2014)
work page 2014
-
[4]
Pohl, J. Artificial superintelligence: Extinction or nirvana? In Proceedings of InterSymp-2015, IIAS, 27th international conference on systems research, infor- matics, and cybernetics (2015)
work page 2015
-
[5]
Superintelligence: Paths, dangers, strategies (2014)
Nick, B. Superintelligence: Paths, dangers, strategies (2014)
work page 2014
-
[6]
Hendrycks, D., Mazeika, M. & Woodside, T. An overview of catastrophic ai risks. arXiv preprint arXiv:2306.12001 (2023). 12
arXiv 2023
-
[7]
Bengio, Y. et al. Managing extreme ai risks amid rapid progress. Science 384, 842–845 (2024)
work page 2024
Show all 57 references
-
[8]
Greenblatt, R. et al. Alignment faking in large language models. arXiv preprint arXiv:2412.14093 (2024)
2024 arXiv
-
[9]
S., Goldstein, S., O’Gara, A., Chen, M
Park, P. S., Goldstein, S., O’Gara, A., Chen, M. & Hendrycks, D. Ai deception: A survey of examples, risks, and potential solutions. Patterns 5 (2024)
2024
-
[10]
Sharma, M. et al. Towards understanding sycophancy in language models. arXiv preprint arXiv:2310.13548 (2023)
2023 arXiv
-
[11]
Human compatible: AI and the problem of control (Penguin Uk, 2019)
Russell, S. Human compatible: AI and the problem of control (Penguin Uk, 2019)
2019
-
[12]
URL https://www.safe.ai/work/ statement-on-ai-risk
Statement on ai risk (2023). URL https://www.safe.ai/work/ statement-on-ai-risk . Accessed: 2024-05-01
2023
-
[13]
& Norvig, P
Russell, S. & Norvig, P. The ethics and risks of developing artificial intelligence. Artificial Intelligence: A Modern Approach 1034–39 (2009)
2009
-
[14]
URL https://openai.com/index/ introducing-superalignment/
Introducing superalignment (2023). URL https://openai.com/index/ introducing-superalignment/
2023
-
[15]
Bai, Y. et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862 (2022)
2022 arXiv
-
[16]
Ouyang, L. et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems 35, 27730–27744 (2022)
2022
-
[17]
Amodei, D. et al. Concrete problems in ai safety. arXiv preprint arXiv:1606.06565 (2016)
2016 arXiv
-
[18]
& Amodei, D
Christiano, P., Shlegeris, B. & Amodei, D. Supervising strong learners by am- plifying weak experts. arXiv preprint arXiv:1810.08575 (2018)
2018 arXiv
-
[19]
Burns, C. et al. Weak-to-strong generalization: Eliciting strong capabilities with weak supervision. arXiv preprint arXiv:2312.09390 (2023)
2023 arXiv
-
[20]
Tao, L. & Li, Y. Your weak LLM is secretly a strong teacher for alignment (2024). 2409.08813
2024 arXiv
-
[21]
Bowman, S. R. et al. Measuring progress on scalable oversight for large language models. arXiv preprint arXiv:2211.03540 (2022)
2022 arXiv
-
[22]
Ji, J. et al. Aligner: Efficient alignment by learning to correct. Advances in Neural Information Processing Systems 37, 90853–90890 (2024). 13
2024
-
[23]
Wu, D. X. & Sahai, A. Provable weak-to-strong generalization via benign over- fitting (2025). 2410.04638
2025 arXiv
-
[24]
& Sala, F
Shin, C., Cooper, J. & Sala, F. Weak-to-strong generalization through the data- centric lens (2025). 2412.03881
2025 arXiv
-
[25]
Somerstep, S. et al. A transfer learning framework for weak-to-strong general- ization (2025). 2405.16236
2025 arXiv
-
[26]
& Wang, R
Zhu, W., He, Z., Wang, X., Liu, P. & Wang, R. Weak-to-strong preference optimization: Stealing reward from weak aligned model (2025). 2410.18640
2025 arXiv
-
[27]
Lyu, Y. et al. MACPO: Weak-to-strong alignment via multi-agent contrastive preference optimization (2025). 2410.07672
2025 arXiv
-
[28]
Yang, W. et al. Super(ficial)-alignment: Strong models may deceive weak models in weak-to-strong generalization (2024). 2406.11431
2024 arXiv
-
[29]
Iterated distillation and amplifica- tion (2018)
Cotra, A. Iterated distillation and amplifica- tion (2018). URL https://ai-alignment.com/ iterated-distillation-and-amplification-157debfd1616
2018
-
[30]
Leike, J. et al. Scalable agent alignment via reward modeling: a research direc- tion. arXiv preprint arXiv:1811.07871 (2018)
2018 arXiv
-
[31]
Lee, H. et al. Rlaif: Scaling reinforcement learning from human feedback with ai feedback (2023)
2023
-
[32]
J., Abbeel, P
Hadfield-Menell, D., Russell, S. J., Abbeel, P. & Dragan, A. Cooperative inverse reinforcement learning. Advances in neural information processing systems 29 (2016)
2016
-
[33]
Hadfield-Menell, D., Milli, S., Abbeel, P., Russell, S. J. & Dragan, A. Inverse reward design. Advances in neural information processing systems 30 (2017)
2017
-
[34]
Laidlaw, C. et al. Scalably solving assistance games. In ICML 2024 Workshop on Models of Human Feedback for AI Alignment (2024)
2024
-
[35]
Laidlaw, C. et al. Assistancezero: Scalably solving assistance games. arXiv preprint arXiv:2504.07091 (2025)
2025 arXiv
-
[36]
Perez, E. et al. Red teaming language models with language models. arXiv preprint arXiv:2202.03286 (2022)
2022 arXiv
-
[37]
Gou, Z. et al. Critic: Large language models can self-correct with tool-interactive critiquing. arXiv preprint arXiv:2305.11738 (2023)
2023 arXiv
-
[38]
Chen, Z., Deng, Y., Yuan, H., Ji, K. & Gu, Q. Self-play fine-tuning converts weak language models to strong language models. arXiv preprint arXiv:2401.01335 (2024). 14
2024 arXiv
-
[39]
Bai, Y. et al. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073 (2022)
2022 arXiv
-
[40]
& Amodei, D
Irving, G., Christiano, P. & Amodei, D. AI safety via debate (2018).1805.00899
2018 arXiv
-
[41]
Du, Y., Li, S., Torralba, A., Tenenbaum, J. B. & Mordatch, I. Improving fac- tuality and reasoning in language models through multiagent debate (2023). 2305.14325
2023 arXiv
-
[42]
Kenton, Z. et al. On scalable oversight with weak llms judging strong llms. Advances in Neural Information Processing Systems 37, 75229–75276 (2024)
2024
-
[43]
Kirchner, J. H. et al. Prover-verifier games improve legibility of llm outputs. arXiv preprint arXiv:2407.13692 (2024)
2024 arXiv
-
[44]
Shah, R. et al. Benefits of assistance over reward learning (2020)
2020
-
[45]
Apperly, I. A. & Butterfill, S. A. Do humans have two systems to track beliefs and belief-like states? Psychological review 116, 953 (2009)
2009
-
[46]
Deception abilities emerged in large language models
Hagendorff, T. Deception abilities emerged in large language models. Proceedings of the National Academy of Sciences 121, e2317967121 (2024)
2024
-
[47]
Qiu, T. A. et al. Progressgym: Alignment with a millennium of moral progress. Advances in Neural Information Processing Systems 37, 14570–14607 (2024)
2024
-
[48]
Instructions for Practical Living (Chuan Xi Lu) (1556)
Wang, Y. Instructions for Practical Living (Chuan Xi Lu) (1556)
-
[49]
Berkeley, E. C. Giant Brains or Machines That Think (John Wiley & Sons, 1949)
1949
-
[50]
Turing, A. M. Computing machinery and intelligence. Mind 49, 433–460 (1950)
1950
-
[51]
Churchland, P. S. Braintrust: What neuroscience tells us about morality (2018)
2018
-
[52]
& Woodruff, G
Premack, D. & Woodruff, G. Does the chimpanzee have a theory of mind? Behavioral and brain sciences 1, 515–526 (1978)
1978
-
[53]
G., Aharon-Peretz, J
Shamay-Tsoory, S. G., Aharon-Peretz, J. & Perry, D. Two systems for empathy: a double dissociation between emotional and cognitive empathy in inferior frontal gyrus versus ventromedial prefrontal lesions. Brain 132, 617–627 (2009)
2009
-
[54]
Christov-Moore, L. et al. Preventing antisocial robots: A pathway to artificial empathy. Science Robotics 8, eabq3658 (2023)
2023
-
[55]
Zhao, F. et al. Building altruistic and moral ai agent with brain-inspired affective empathy mechanisms. arXiv preprint arXiv:2410.21882 (2024)
2024
-
[56]
Tong, H. et al. Autonomous alignment with human value on altruism through considerate self-imagination and theory of mind. arXiv preprint arXiv:2501.00320 (2024). 15
2024 arXiv
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.