Pith. sign in

REVIEW 4 major objections 3 minor 15 references

Lilith: Developmental Modular LLMs with Chemical Signaling

T0 review · 4 major / 3 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read LILITH proposes that modular LLMs with chemical-style signalling and developmental training could let researchers measure consciousness emergence with Integrated Information Theory.

desk verdict A new but wholly unsupported architectural idea; the IIT evaluation rests on an undefined Φ, so treat it as a position paper, not a research result. read the letter →

arxiv 2507.04575 v1 pith:56PFB2CN submitted 2025-07-06 q-bio.NC cs.AI

classification q-bio.NCcs.AI
keywords developmentaltrainingmodularlanguagemodelschemicalsignalingconsciousnessemergenceIntegratedInformationTheorytoken-basedcommunicationevolutionaryoptimizationmulti-regionbrainmodelling
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proposes LILITH, an architecture in which separate language-model modules play the roles of brain regions—thinking, memory, sensory, and regulatory—and communicate through token-based signals that stand in for chemical neurotransmitters. The authors' central conjecture is that modelling the brain at the level of interacting regions with chemical-style signalling, instead of only at the level of neurons, is a productive step toward understanding the emergence of consciousness. To that end, LILITH would skip pre-training entirely: untrained LLM architectures would live simulated life cycles, develop communication pathways through environmental interaction, and be selected across generations by evolutionary optimization. The paper's stated payoff is that such a system could be measured directly with Integrated Information Theory, which seeks to quantify consciousness from causal structure, enabling empirical study of when and where consciousness-like integrated information appears during development. The paper is a conceptual proposal and acknowledges substantial implementation challenges.

What carries the argument

The load-bearing mechanism is the token-based communication protocol between modular LLMs. Each module is a specialised brain region with restricted capabilities—a thinking region that can prompt itself, a memory region that alone can save items, a sensory region that alone receives external input, and a regulatory brain stem with preprompted control powers—and the tokens these modules send to one another are the analogue of neurotransmitter signals. The machinery does two jobs: it makes inter-region signalling observable and editable, since tokens can be logged and measured, and it is what the developmental and evolutionary process shapes, because the meaning and routing of tokens are not predefined but emerge over simulated lives. This is the part of the design that connects the biology-inspired vocabulary to testable Integrated Information Theory calculations.

What would settle it

One could train a LILITH system and record its Integrated Information Theory score across development together with behavioural markers such as flexible response to novel inputs or goal-directed action; if the score rises while every behavioural marker stays flat, the claim that optimizing for IIT metrics tracks consciousness emergence would be contradicted.

Watch

Extended reading notes

Core claim

The central claim is that consciousness emergence can be investigated empirically by combining developmental training of modular language models with brain-inspired token-based communication. Concretely, the paper argues that distinct brain regions should be modelled as specialised LLM modules, that inter-region signals should be learned tokens rather than predefined routes, and that the whole system should start untrained and learn through simulated life experience. On the authors' account, this design would allow Integrated Information Theory metrics to be applied at three scales—individual modules, inter-agent communication, and the whole system—so a researcher could see at which scale integrated information is strongest and how it grows over developmental time. The discovery being proposed, in other words, is a method: a way to build and measure a developing artificial system whose design objective is consciousness emergence rather than task performance.

Load-bearing premise

The proposal depends on Integrated Information Theory being a valid measure of consciousness, so that optimizing a system for its metrics counts as studying consciousness emergence.

Editorial extensions

If this is right

  • If LILITH works as proposed, consciousness metrics can be tracked continuously during development, giving a time series of integrated information as communication pathways form.
  • The multi-scale design would let researchers compare integrated information at the level of single modules, inter-module signalling, and the whole system, and thereby locate the organizational level where integration is maximal.
  • Shifting the optimization target from task performance to consciousness measures would change the evaluation culture: systems would be judged by their developmental trajectory and sentience markers rather than by benchmark scores.
  • The emergent token protocols themselves become observable objects, so the framework would yield a record of how inter-region chemical-style signalling changes with experience.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The authors leave implicit that the same training objective could be swapped: any computable theory of consciousness could replace Integrated Information Theory, making LILITH a general platform for comparing candidate theories rather than a test of IIT alone.
  • A natural extension would be an ablation study of the regulatory brain-stem module: removing it during development and observing whether token routing and inter-module signalling collapse would test whether the architecture's regulatory component plays the arousal-like role it is assigned.
  • Because the token vocabulary is emergent, one could measure the entropy and stability of token usage over developmental time; the authors do not, but a concrete prediction of their framework is that stable, differentiated signalling tokens should appear before any rise in integrated information.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 3 minor

Summary. The manuscript proposes LILITH, a modular LLM architecture in which distinct modules correspond to brain regions, communicate through learned token-based 'chemical signaling', and are trained via simulated development rather than static pretraining. The authors claim this architecture would enable direct empirical investigation of consciousness emergence using Integrated Information Theory (IIT) metrics, and that optimizing for consciousness emergence rather than task performance could reveal multi-scale neural correlates. The paper is explicitly a conceptual proposal: no implementation, simulation, data, or mathematical derivation is provided. The central claims are stated as future possibilities, and the authors acknowledge substantial open challenges, particularly in optimization.

Significance. The idea of combining modular LLMs with learned inter-module communication to test IIT-style measures is a creative extension of recent work on modular and brain-inspired AI. The paper is clearly written and honest in its limitations, and it cites relevant foundational references. However, as it stands it offers no formal model, no algorithms beyond natural-language descriptions, and no experimental results. Its significance is therefore entirely promissory. If the mapping from LILITH modules and token streams to IIT's formal objects were supplied and a concrete optimization scheme specified, the proposed architecture could become a useful instrument for studying integrated information in modular networks; the current manuscript does not yet provide that instrument.

major comments (4)
  1. [Implications for Consciousness Research] The sentence 'The system can be directly tested using Integrated Information Theory metrics' is unsupported. IIT applies to a system with a specified set of elements, a state space, and transition probabilities; the manuscript never defines what counts as an element (modules, attention heads, token types?), how token streams map to states, or how the time-varying developmental process provides the stationary transition structure that Φ requires. Without this mapping, the central claim that LILITH enables direct IIT-based investigation cannot be evaluated.
  2. [On the Optimization of such an Architecture] The proposed optimization scheme is disconnected from the IIT measurement. The paper mentions 'auto-encoding objectives' for regions and 'evolutionary algorithms applied to inter-region signaling', but it explicitly states that 'detailed investigation of these optimization frameworks remains an important direction for future research'. Consequently, the Abstract's phrase 'optimizing for consciousness emergence rather than task performance' has no operational meaning in this manuscript, and the link between the evolutionary search and any IIT-based quantity is missing.
  3. [Implications for Consciousness Research / Abstract] If LILITH is trained to maximize IIT-based metrics, then any later increase in those metrics is partly by construction, not independent evidence of consciousness emergence. The paper does not state whether IIT is intended as a training objective or only as a post-hoc evaluation criterion. This ambiguity is load-bearing because the proposed evidential value of the framework depends on IIT being an independent measure of the phenomenon under study.
  4. [Developmental Training Framework] The architecture is specified at an informal level: success criteria are said to 'be carefully defined' but are not, the 'breeding' step relies on LLM voting without details, and the token protocol is described only through an example. Because the manuscript's goal is to propose a research direction, this level of specification might be acceptable in a perspective piece; however, the strength of the empirical claims in the Abstract and Implications sections ('unprecedented insight', 'direct empirical investigation') is disproportionate to what the manuscript actually delivers.
minor comments (3)
  1. [Header] The running title reads 'DEVELOPMENT AL MODULAR LLMS'; this should be 'DEVELOPMENTAL MODULAR LLMS'.
  2. [Section 2 heading] The heading 'Modular Brain RegionDesign' is missing a space, and 'T oken-Based' in the following heading contains an extra space.
  3. [References] Reference [9] lists 'and et al.' after the first authors, which is redundant; either list all authors or use 'et al.' alone.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: LILITH is a speculative architecture proposal with no fitted parameters, no predictions-from-fits, and no self-citation chains.

full rationale

LILITH is a conceptual proposal rather than a derivation or empirical study, and none of its load-bearing claims reduce to their inputs by construction. The paper proposes modular LLMs, token-based inter-region signaling, developmental training, and evolutionary optimization, but it reports no fitted parameters, no quantitative predictions, and no test results. The closest point to circularity is the proposal to measure 'consciousness emergence' with Integrated Information Theory metrics; however, the paper merely suggests using IIT as an external measurement tool and does not define an objective function that optimizes IIT, nor does it present an IIT-maximizing result as an independent discovery. The vague IIT connection is an under-specification or formal gap, since no mapping from modules and token streams to IIT elements is given, but this is not a circular reduction. No self-citations appear; all references are to external prior work. Therefore the paper earns a circularity score of 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 2 invented entities

The proposal rests on three unproven domain assumptions about IIT's validity as a consciousness measure, the feasibility of training raw LLM architectures through simulated life, and the adequacy of token exchange to model chemical signaling. No free parameters are fitted because no quantitative results are presented; LILITH and its token protocol are invented entities with no independent evidence.

assumptions (3)
  • domain assumption IIT metrics provide a valid quantitative measure of consciousness emergence in artificial systems.
    The paper proposes using IIT to measure consciousness without discussing the theory's contested status or validation against behavioral markers (Implications for Consciousness Research).
  • domain assumption Untrained LLM architectures can learn language and cognitive abilities through simulated experiences and environmental interaction.
    The Developmental Training Framework section assumes that raw LLM architectures can acquire capabilities through a simulated life cycle, which is unproven and likely computationally infeasible at current scale.
  • domain assumption Token-based communication between LLM modules can adequately model brain chemical signaling.
    The authors propose token signaling as an analogy to neurotransmitters, but do not provide a mapping or evidence that this captures relevant biological dynamics.
invented entities (2)
  • LILITH architecture
    purpose: A modular LLM system with brain-region-like agents communicating via learned tokens, to study consciousness emergence.
    No implementation or empirical evidence is provided; it exists only as a proposal.
  • Token-based signaling protocol
    purpose: Inter-module communication that emulates neurotransmitter signaling.
    The protocol details are not specified, and no evidence that token exchange yields brain-like coordination.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Lilith: Developmental Modular LLMs with Chemical Signaling." pith.science (2026). https://pith.science/paper/56PFB2CN

@misc{pith2026250704575,
  author       = {Pith},
  title        = {Pith review of: Lilith: Developmental Modular LLMs with Chemical Signaling},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/56PFB2CN}},
  note         = {Machine review of arXiv:2507.04575}
}
read the original abstract

Current paradigms in Artificial Intelligence rely on layers of feedforward networks which model brain activity at the neuronal level. We conjecture that expanding to the level of multiple brain regions with chemical signaling may be a productive step toward understanding the emergence of consciousness. We propose LILITH, a novel architecture that combines developmental training of modular language models with brain-inspired token-based communication protocols, mirroring chemical signaling in the brain. Our approach models distinct brain regions as specialized LLM modules including thinking, memory, sensory, and regulatory components that communicate through emergent token-based signaling protocols analogous to neurotransmitter networks. Unlike traditional pre-trained systems, LILITH would employ developmental training where untrained LLM architectures learn through simulated life experiences, developing communication pathways and cognitive abilities through environmental interaction and evolutionary optimization. This framework would enable direct empirical investigation of consciousness emergence using Integrated Information Theory metrics while providing unprecedented insight into inter-module signaling patterns during development. By optimizing for consciousness emergence rather than task performance, LILITH could provide insight into different emergent phenomena at multiple levels of neural correlates, contrasting neuronal-level processing with multi-region coordination dynamics. The goal of this paper is to put the idea forward while recognizing the substantial challenges in implementing such a system.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

15 extracted references · 8 canonical work pages

  1. [1]

    Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell

    Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pages 610–623, Virtual Event Canada, March 2021. ACM

  2. [2]

    Lake, Tomer D

    Brenden M. Lake, Tomer D. Ullman, Joshua B. Tenenbaum, and Samuel J. Gershman. Building Machines That Learn and Think Like People, November 2016. arXiv:1604.00289 [cs]

  3. [3]

    Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity, June 2022

    William Fedus, Barret Zoph, and Noam Shazeer. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity, June 2022. arXiv:2101.03961 [cs]

  4. [4]

    Introduction to the Soar Cognitive Architecture

    John E Laird. Introduction to the Soar Cognitive Architecture

  5. [5]

    Anderson, Daniel Bothell, Michael D

    John R. Anderson, Daniel Bothell, Michael D. Byrne, Scott Douglass, Christian Lebiere, and Yulin Qin. An integrated theory of the mind. Psychological Review, 111(4):1036–1060, October 2004

  6. [6]

    Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation, May 2025

    Maria Eriksson, Erasmo Purificato, Arman Noroozian, Joao Vinagre, Guillaume Chaslot, Emilia Gomez, and David Fernandez-Llorca. Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation, May 2025. arXiv:2502.06559 [cs]

  7. [7]

    DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale, July 2022

    Samyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang, Reza Yazdani Aminabadi, Ammar Ah- mad Awan, Jeff Rasley, and Yuxiong He. DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale, July 2022. arXiv:2201.05596 [cs]

  8. [8]

    Neuroscience- Inspired Artificial Intelligence

    Demis Hassabis, Dharshan Kumaran, Christopher Summerfield, and Matthew Botvinick. Neuroscience- Inspired Artificial Intelligence. Neuron, 95(2):245–258, July 2017

Show all 15 references
  1. [9]

    Elman, Elizabeth A

    Jeffrey L. Elman, Elizabeth A. Bates, Mark H. Johnson, Annette Karmiloff-Smith, and et al. Rethinking in- nateness: A connectionist perspective on development. Rethinking innateness: A connectionist perspective on development. The MIT Press, Cambridge, MA, US, 1996. Pages: xviii, 447

  2. [10]

    Neural Module Networks, July 2017

    Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein. Neural Module Networks, July 2017. arXiv:1511.02799 [cs]

  3. [11]

    Olaf Sporns and Richard F. Betzel. Modular Brain Networks. Annual Review of Psychology, 67(1):613–640, January 2016

  4. [12]

    Bassett and Olaf Sporns

    Danielle S. Bassett and Olaf Sporns. Network neuroscience. Nature Neuroscience, 20(3):353–364, March

  5. [13]

    Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, January 2023

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, January 2023. arXiv:2201.11903 [cs]

  6. [14]

    Integrated information theory: from consciousness to its physical substrate

    Giulio Tononi, Melanie Boly, Marcello Massimini, and Christof Koch. Integrated information theory: from consciousness to its physical substrate. Nature Reviews Neuroscience, 17(7):450–461, July 2016. Address Email address: mu2faroo@uwaterloo.ca

  7. [2017]

    Publisher: Nature Publishing Group

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.