REVIEW 3 major objections 5 minor 74 references
On Parallelism in Music and Language: A Perspective from Symbol Emergence Systems based on Probabilistic Generative Models
T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read The paper proposes that the meaning of music is the change a musical passage induces in a listener's inferred emotional state, modeled as collective predictive coding over interoceptive signals.
desk verdict A clear, well-scoped speculative perspective whose central music-emotion analogy is asserted rather than shown; the paper's own admission of musical firstness breaks the proposed PGM mapping. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is collective predictive coding realized as decentralized Bayesian inference in a probabilistic generative model. In the PGM for symbol emergence, two agents share a latent word $w$ while each maintains private internal representations $z_A$ and $z_B$ generated from their own observations, and communication is an approximate posterior sampling scheme, the Metropolis-Hastings naming game. The extension to music replaces the observation set $\{o_m\}$ with interoceptive signals and the internal representation with an emotional state, so that the same head-to-head graphical model describes how musical symbols arise and stabilize. This carries the analogy: what makes a sign meaningful is not a fixed dyadic relation to an object but its role in predicting an agent's own future signals.
What would settle it
A decisive test would compare listeners' reported or physiological emotional responses to a musical passage while their interoceptive priors are manipulated, for example by making their heartbeats audible or by inducing a reliable visceral state. If musical meaning can be fully predicted from acoustic structure alone, without any measurable dependence on interoceptive state inference, the central claim fails; conversely, if the same acoustic signal gives different inferred emotional meanings solely because of different bodily priors, the claim is supported.
Extended reading notes
Core claim
On the paper's own terms, the discovery is a mapping: the equations that describe a robot learning words from multimodal sensorimotor data, $w, \{o_m\} \sim p(w,\{o_m\}|z)$ and $y \sim p(y|w)$, are identical in form to equations for music generation in which $z$ is an emotional state. The author then replaces the perceptual internal representation of symbol emergence systems with an emotional internal representation, and replaces interaction with the external world via sensorimotor signals with interaction with the internal environment via interoception. The resulting proposal is that 'the meaning of music' is a change in the mental state that the listener undergoes, i.e., inference or state updating of internal representations under emotional predictive coding. Music, on this view, is a socially emerged symbol system grounded not in exteroceptive facts but in interoceptive predictions, with individual agents' emotional states coordinated through semiotic communication just as perceptual states are coordinated in language.
Load-bearing premise
The argument rests on the assumption that music, unlike language, does not explicitly represent events in the external world, so its meaning can be grounded in interoceptive rather than exteroceptive signals; if music primarily represents external events or is primarily grounded in body movement rather than visceral feeling, the proposed parallel loses its basis.
Editorial extensions
If this is right
- If the meaning of music is an update of emotional internal representations, then automatic composition can be reframed as sampling from a posterior distribution over emotional states, extending existing language-model approaches to note sequences.
- A musical symbol system, like a language, is not fixed in a score or a composer's intention; it is continuously re-emerged through social coordination of listeners' interoceptive predictions.
- The same PGMs used for unsupervised phoneme and word discovery, double articulation analysis, and multimodal object and place concept formation should be transferable to modeling musical structure and its emotional grounding.
- Emotion can be brought into computational music research as a latent variable with a well-defined generative semantics rather than a categorical label.
Reading between the lines
- This suggests a testable asymmetry: musical signs can act directly on visceral sensing, so the sign itself participates in its own grounding, whereas arbitrary linguistic signs must acquire grounding through learned exteroceptive association.
- A natural extension is to expressive prosody and performance gesture, which also carry emotional content through interoceptive predictions and could be modeled with the same head-to-head PGM.
- If musical meaning is collectively inferred emotional state, then cultural differences in musical taste become differences in shared priors over interoceptive signals, which could be studied by comparing predictive models fitted to listeners from different traditions.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a speculative parallelism between music and language from the viewpoint of symbol emergence systems modeled with probabilistic generative models (PGMs). It first reviews the author's prior work on symbol emergence in robotics, where a multi-agent symbol system is modeled as decentralized Bayesian inference and described as 'collective predictive coding.' It then extends this framework to music by reinterpreting the latent variables: perceptual internal representations become emotional internal representations, and sensorimotor interaction with the external world becomes interoceptive interaction with the internal environment. The paper hypothesizes that 'the meaning of music' is the change in the listener's mental state, i.e., the inference of emotional internal representations, and that music is a socially emerged symbol system grounded in interoceptive predictions. The paper explicitly states that it does not provide further details or evidence for the correspondence.
Significance. If the hypothesis were substantiated, it would provide a unifying computational framework for music semantics, linking music to emotion, predictive coding, and social symbol emergence, and it would open a concrete research program, such as extending PGM-based language models to music with emotional latent variables. The paper is honest about its speculative status and clearly lays out the key analogies in Figure 2. Its strengths are the explicit computational framing and the connection to a well-developed line of PGM-based symbol emergence research, including machine-checked or reproducible models in prior work. However, as a proposal it does not yet provide a derivation, simulation, or empirical test, and the formal part currently contains an error in Eq. (10) and an unresolved tension with the paper's own admission of music's 'firstness.' The paper's value is as a position statement that may provoke discussion, but its central claim is not yet adequately supported.
major comments (3)
- [Section 4.2, Eq. (10)] Equation (10) as written is formally incorrect: it duplicates p(zA|{oAm}) and omits p(zB|{oBm}). The correct factorization should involve p(w|zA, zB) p(zA|{oAm}) p(zB|{oBm}), or an equivalent symmetric form. This is not a harmless typo, because the equation is presented as the mathematical description of the symbol emergence system, and the reader cannot verify the claimed decentralized Bayesian inference without the correct terms.
- [Section 5 (and Section 4.2)] The paper concedes that music has 'many symbolic aspects of firstness' and that musical signs can directly affect visceral senses. In the language model, the shared sign w is conditionally independent of observations o given the internal states z, which is what allows two agents to align their internal representations through a common external world. For music, the sign may directly influence interoceptive observations, yielding o_m ~ p(o_m | z, w) rather than o_m ~ p(o_m | z). This changes the generative structure of Eq. (10) and undermines the proposed mapping. The paper does not explain how a symbol system can still emerge when the sign is partly non-arbitrary and causally coupled to the private observations it is supposed to signify. The author should either extend the model to include firstness and demonstrate that symbol emergence can still occur, or argue explicitly why firstness does not preclude the conventional alignment of musical signs.
- [Section 2.2, Eqs. (8)-(9)] The statement that Eqs. (8) and (9) are 'identical' to Eqs. (1) and (2) is formally true but semantically superficial. In the language case, the latent z is presented as a cause of the utterance and is grounded in external objects and sensorimotor information; in the music case, z is an emotional state inferred from interoceptive signals. The formal similarity of two sequence-generating models is not by itself evidence of parallelism between the underlying cognitive mechanisms; it is exactly the proposed hypothesis that these latent variables play analogous roles. Please clarify that this is an analogy in generative form, not a derivation, and discuss what additional assumptions are needed to make the analogy substantive.
minor comments (5)
- [Figure 2 caption] The right panel of Figure 2 uses 'introspective signals' while the text consistently uses 'interoceptive signals'; please unify the terminology.
- [Section 5] The phrase 'the author pointed out' appears in the first-person narrative of a formal paper; please rephrase to avoid a personal voice.
- [References] References [41] and [42] are the same Okanoya 2007 paper; remove the duplicate and renumber accordingly.
- [Section 2.1, Eq. (5)] Equation (5) writes inference as z ~ p(z|{om}); this notation could be confused with sampling from the generative model. Consider using posterior notation such as z* = argmax z p(z|{om}) or a more explicit inference formulation.
- [General] The paper would benefit from a short discussion of testable predictions of the hypothesis, for example, whether a multi-agent PGM with interoceptive observations can actually produce shared musical symbols in simulation.
Circularity Check
Music-as-symbol-emergence conclusion is largely guaranteed by defining musical meaning as internal-state inference and relabeling the language PGM.
-
self definitional
[Section 4.2, paragraph after Fig. 2 (p. 11)]
"if we view the ”meaning of music” as a change in the mental state that the listener undergoes, or inference (or state updating) of internal representations, similar to the ”meaning of language,” and especially if we view it as an emotional impression (i.e., being moved, or its effect on the emotions), then through the discussion of predictive coding, we can connect it to the discussion of symbol emergence systems"
The claimed connection is guaranteed by the definition chosen. Musical meaning is stipulated to be inference or state updating of internal representations, which is exactly the operation performed by predictive coding and by the symbol-emergence PGM. Under that stipulation, the link to symbol emergence systems follows by construction rather than from independent evidence about music; the subsequent hypothesis inherits this definitional reduction.
-
renaming known result
[Section 2.2, Eqs. (8)-(9); Section 4.2, Fig. 2 (right)]
"The generative process of music, including composition and performance, it can be described as follows: w = w1:S ∼ p(w|z), (8) y = y1:T ∼ p(y|w) (9) ... Interestingly, these equations are identical to those in (1) and (2), respectively. This correspondence apparently displays one parallelism between music and language. ... The sign of music corresponds to the sign of language. The ’perceptual’ internal representation system that supports the interpretation of language corresponds to the ’emotional’ internal representation system in music."
The music side is not derived from musical data or analysis; it is constructed by writing music in the same generative form as language and then relabeling z as the emotional state and sensorimotor observations as interoceptive ones. Figure 2 (right) is a relabeled version of the language/symbol-emergence diagram. The parallelism is therefore built into the notation rather than demonstrated, and the later hypothesis about music inherits that constructed equivalence.
full rationale
This is an explicitly hypothetical perspective paper: it calls the music claim a possible hypothesis and states that it does not provide further details or evidence. The central reduction is definitional: once musical meaning is stipulated as internal-state inference and the music generative process is written with the same equations as language with z renamed emotional state, the connection to collective predictive coding follows by construction. This is a real circularity, but it is limited because no fitted parameters or empirical predictions are made. The computational core comes from prior published models, including the author's own Metropolis-Hastings naming game work, but those are simulation-based results rather than the target music claim, so the self-citations are not by themselves the circular step. The paper also candidly concedes that music has many symbolic aspects of firstness, i.e., that signs can directly affect visceral senses; this undermines the arbitrary-sign assumption of the symbol emergence model, but that is a coherence gap rather than a circular step. Overall, the music hypothesis is an analogical proposal whose conclusion is largely guaranteed by its own definitions, but it is honestly offered as a hypothesis, so the circularity score is moderate.
Assumptions & free parameters
assumptions (4)
- domain assumption Emotion is based on predictive coding of interoceptive signals.
- ad hoc to paper Symbol emergence in a society is collective predictive coding and can be modeled as decentralized Bayesian inference.
- ad hoc to paper The generative process of music can be written identically to language as w ~ p(w|z) and y ~ p(y|w), with z an emotional state.
- domain assumption Music does not have a function explicitly representing events in the external world, unlike human language.
Cite this review
Pith. "Pith review of On Parallelism in Music and Language: A Perspective from Symbol Emergence Systems based on Probabilistic Generative Models." pith.science (2026). https://pith.science/paper/MD7FSLR6
@misc{pith2026250115721,
author = {Pith},
title = {Pith review of: On Parallelism in Music and Language: A Perspective from Symbol Emergence Systems based on Probabilistic Generative Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/MD7FSLR6}},
note = {Machine review of arXiv:2501.15721}
}
read the original abstract
Music and language are structurally similar. Such structural similarity is often explained by generative processes. This paper describes the recent development of probabilistic generative models (PGMs) for language learning and symbol emergence in robotics. Symbol emergence in robotics aims to develop a robot that can adapt to real-world environments and human linguistic communications and acquire language from sensorimotor information alone (i.e., in an unsupervised manner). This is regarded as a constructive approach to symbol emergence systems. To this end, a series of PGMs have been developed, including those for simultaneous phoneme and word discovery, lexical acquisition, object and spatial concept formation, and the emergence of a symbol system. By extending the models, a symbol emergence system comprising a multi-agent system in which a symbol system emerges is revealed to be modeled using PGMs. In this model, symbol emergence can be regarded as collective predictive coding. This paper expands on this idea by combining the theory that ''emotion is based on the predictive coding of interoceptive signals'' and ''symbol emergence systems,'' and describes the possible hypothesis of the emergence of meaning in music.
Figures
Reference graph
Works this paper leans on
-
[1]
In: 2018 IEEE International Conference on Acou stics, Speech and Sig- nal Processing (ICASSP)
Akbari, M., Liang, J.: Semi-recurrent CNN-based V AE-GAN for sequential data generation. In: 2018 IEEE International Conference on Acou stics, Speech and Sig- nal Processing (ICASSP). pp. 2321–2325. IEEE (2018)
work page 2018
-
[2]
I n: IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
Ando, Y., Nakamura, T., Araki, T., Nagai, T.: Formation of hierarchical object concept using hierarchical latent Dirichlet allocation. I n: IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). pp. 22 72–2279 (2013)
work page 2013
-
[3]
In: 2011 IEEE/RSJ Interna tional Con- ference on Intelligent Robots and Systems (IROS)
Araki, T., Nakamura, T., Nagai, T., Funakoshi, K., Nakano , M., Iwa- hashi, N.: Autonomous acquisition of multimodal informati on for online ob- ject concept formation by robots. In: 2011 IEEE/RSJ Interna tional Con- ference on Intelligent Robots and Systems (IROS). pp. 1540– 1547 (2011). https://doi.org/10.1109/IROS.2011.6048422
-
[4]
In: IEEE/RSJ I nternational Conference on Intelligent Robots and Systems
Araki, T., Nakamura, T., Nagai, T., Nagasaka, S., Taniguc hi, T., Iwa- hashi, N.: Online learning of concepts and words using multi modal LDA and hierarchical Pitman-Yor Language Model. In: IEEE/RSJ I nternational Conference on Intelligent Robots and Systems. pp. 1623–163 0 (2012). https://doi.org/10.1109/IROS.2012.6385812
arXiv 2012
-
[5]
Asano, R., Boeckx, C.: Syntax in language and music: what i s the right level of comparison? Frontiers in Psychology 6, 942 (2015)
work page 2015
-
[6]
Atherton, R.P., Chrobak, Q.M., Rauscher, F.H., Karst, A. T., Hanson, M.D., Stein- ert, S.W., Bowe, K.L.: Shared processing of language and mus ic: Evidence from a cross-modal interference paradigm. Experimental Psychol ogy 65(1), 40 (2018)
work page 2018
-
[7]
Adva nces in Neural Information Processing Systems 33, 12449–12460 (2020)
Baevski, A., Zhou, Y., Mohamed, A., Auli, M.: wav2vec 2.0: A framework for self-supervised learning of speech representations. Adva nces in Neural Information Processing Systems 33, 12449–12460 (2020)
work page 2020
-
[8]
Nature re- views neuroscience 16(7), 419–429 (2015)
Barrett, L.F., Simmons, W.K.: Interoceptive prediction s in the brain. Nature re- views neuroscience 16(7), 419–429 (2015)
work page 2015
Show all 74 references
-
[9]
Behavioral a nd Brain Sciences 22(04), 1–16 (1999)
Barsalou, L.W.: Perceptual symbol systems. Behavioral a nd Brain Sciences 22(04), 1–16 (1999). https://doi.org/10.1017/S0140525X9900214 9
1999 doi
-
[10]
Frontiers in evolutionary neuro science 4, 5 (2012) 14 T
Berwick, R.C., Beckers, G.J., Okanoya, K., Bolhuis, J.J .: A bird’s eye view of human language evolution. Frontiers in evolutionary neuro science 4, 5 (2012) 14 T. Taniguchi
2012
-
[11]
Springer (2006)
Bishop, C.: Pattern Recognition and Machine Learning (I nformation Science and Statistics). Springer (2006)
2006
- [12]
-
[13]
arXiv preprint arXiv:1709.01620 (20 17)
Briot, J.P., Hadjeres, G., Pachet, F.D.: Deep learning t echniques for music generation–a survey. arXiv preprint arXiv:1709.01620 (20 17)
-
[14]
Brown, S.: Are music and language homologues? Annals of t he New York Academy of Sciences 930(1), 372–374 (2001)
2001
-
[15]
European journal of neuroscience 23(10), 2791–2803 (2006)
Brown, S., Martinez, M.J., Parsons, L.M.: Music and lang uage side by side in the brain: a pet study of the generation of melodies and sentence s. European journal of neuroscience 23(10), 2791–2803 (2006)
2006
-
[16]
, Dhariwal, P., Nee- lakantan, A., Shyam, P., Sastry, G., Askell, A., et al.: Lang uage models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J.D. , Dhariwal, P., Nee- lakantan, A., Shyam, P., Sastry, G., Askell, A., et al.: Lang uage models are few-shot learners. Advances in neural information processing syste ms 33, 1877–1901 (2020)
2020
-
[17]
The MIT press (2015)
Cangelosi, A., Schlesinger, M.: Developmental Robotic s. The MIT press (2015)
2015
-
[18]
Routledge (2002)
Chandler, D.: Semiotics the Basics. Routledge (2002)
2002
-
[19]
In: 2020 International Conference on Technologi es and Applications of Artificial Intelligence (TAAI)
Di´ eguez, P.L., Soo, V.W.: Variational autoencoders fo r polyphonic music inter- polation. In: 2020 International Conference on Technologi es and Applications of Artificial Intelligence (TAAI). pp. 56–61 (2020)
2020
-
[20]
In: 2017 IEEE Automatic Speech Recognition and Understanding Workshop ( ASRU)
Dunbar, E., Cao, X.N., Benjumea, J., Karadayi, J., Berna rd, M., Besacier, L., Anguera, X., Dupoux, E.: The zero resource speech challenge 2017. In: 2017 IEEE Automatic Speech Recognition and Understanding Workshop ( ASRU). pp. 323– 330 (2017)
2017
-
[21]
Annual review of anthropology pp
Feld, S., Fox, A.A.: Music and language. Annual review of anthropology pp. 25–53 (1994)
1994
-
[22]
Literary Licensing, LLC (2011)
Flavell, J.H.: The Developmental Psychology of Jean Pia get. Literary Licensing, LLC (2011)
2011
-
[23]
Neural Networks 144, 573–590 (2021)
Friston, K., Moran, R.J., Nagai, Y., Taniguchi, T., Gomi , H., Tenenbaum, J.: World model learning and inference. Neural Networks 144, 573–590 (2021)
2021
-
[24]
In: IEEE Interna- tional Conference on Development and Learning (ICDL 2022)
Furukawa, K., Taniguchi, A., Hagiwara, Y., Taniguchi, T .: Symbol emergence as inter-personal categorization with head-to-head latent w ord. In: IEEE Interna- tional Conference on Development and Learning (ICDL 2022). pp. 60–67 (2022) On Parallelism in Music and Language from ...
2022
-
[25]
Advanced Robotics 36(5-6), 239–260 (2022)
Hagiwara, Y., Furukawa, K., Taniguchi, A., Taniguchi, T .: Multiagent multimodal categorization for symbol emergence: emergent communicat ion via interpersonal cross-modal inference. Advanced Robotics 36(5-6), 239–260 (2022)
2022
-
[26]
Frontiers in Neurorobotics 12(11), 1–16 (3 2018)
Hagiwara, Y., Inoue, M., Kobayashi, H., Taniguchi, T.: H ierarchical spatial concept formation based on multimodal information for human suppor t robots. Frontiers in Neurorobotics 12(11), 1–16 (3 2018)
2018
-
[27]
Frontiers in Robotics and AI 6(134), pp.1–17 (12 2019), dOI: 10.3389/frobt.2019.00134
Hagiwara, Y., Kobayashi, H., Taniguchi, A., Taniguchi, T.: Symbol emergence as an interpersonal multimodal categorization. Frontiers in Robotics and AI 6(134), pp.1–17 (12 2019), dOI: 10.3389/frobt.2019.00134
2019
-
[28]
OUP Oxford (2013)
Hohwy, J.: The predictive mind. OUP Oxford (2013)
2013
-
[29]
arXiv preprint arXiv:1809.04281 (2018)
Huang, C.Z.A., Vaswani, A., Uszkoreit, J., Shazeer, N., Simon, I., Hawthorne, C., Dai, A.M., Hoffman, M.D., Dinculescu, M., Eck, D.: Music tran sformer. arXiv preprint arXiv:1809.04281 (2018)
2018 arXiv
-
[30]
In: Proceedi ngs of the 28th ACM International Conference on Multimedia
Huang, Y.S., Yang, Y.H.: Pop music transformer: Beat-ba sed modeling and gen- eration of expressive pop piano compositions. In: Proceedi ngs of the 28th ACM International Conference on Multimedia. pp. 1180–1188 (20 20)
-
[31]
In: Music, mind, and brain, pp
Jackendoff, R., Lerdahl, F.: A grammatical parallel betw een music and language. In: Music, mind, and brain, pp. 83–117. Springer (1982)
1982
-
[32]
In: ICASSP 2020-2020 IEEE International Con ference on Acoustics, Speech and Signal Processing (ICASSP)
Jiang, J., Xia, G.G., Carlton, D.B., Anderson, C.N., Miy akawa, R.H.: Transformer vae: A hierarchical model for structure-aware and interpre table music representa- tion learning. In: ICASSP 2020-2020 IEEE International Con ference on Acoustics, Speech and Signal Processing (I...
2020
-
[33]
Ad vances in neural infor- mation processing systems 20 (2007)
Mochihashi, D., Sumita, E.: The infinite markov model. Ad vances in neural infor- mation processing systems 20 (2007)
2007
-
[34]
In: IEEE/RSJ Internatio nal Conference on Intelligent Robots and Systems (IROS) (2015)
Nakamura, T., Ando, Y., Nagai, T., Kaneko, M.: Concept fo rmation by robots using an infinite mixture of models. In: IEEE/RSJ Internatio nal Conference on Intelligent Robots and Systems (IROS) (2015)
2015
-
[35]
Advanced Robotics 25, 2189–2206 (2012)
Nakamura, T., Araki, T., Nagai, T., Iwahashi, N.: Ground ing of word meanings in lda-based multimodal concepts. Advanced Robotics 25, 2189–2206 (2012)
2012
-
[36]
In: IEEE/RSJ International Conference on Intell igent Robots and Systems
Nakamura, T., Nagai, T., Funakoshi, K., Nagasaka, S., Ta niguchi, T., Iwahashi, N.: Mutual Learning of an Object Concept and Language Model B ased on MLDA and NPYLM. In: IEEE/RSJ International Conference on Intell igent Robots and Systems. pp. 600 – 607 (2014)
2014
-
[37]
In: IEEE/RSJ International Conference on Intellige nt Robots and Systems (IROS)
Nakamura, T., Nagai, T., Iwahashi, N.: Multimodal objec t categorization by a robot. In: IEEE/RSJ International Conference on Intellige nt Robots and Systems (IROS). pp. 2415–2420 (2007). https://doi.org/10.1109/I ROS.2007.4399634
2007
-
[38]
In: IEEE/RSJ International Conference on Intelligent Robots a nd Systems (IROS)
Nakamura, T., Nagai, T., Iwahashi, N.: Bag of multimodal hierarchical dirich- let processes: Model of complex conceptual structure for in telligent robots. In: IEEE/RSJ International Conference on Intelligent Robots a nd Systems (IROS). pp. 3818–3823 (2012). https://doi.org/10...
2012 doi
-
[39]
Frontiers in neurorobotics 12 (2018)
Nakamura, T., Nagai, T., Taniguchi, T.: Serket: An archi tecture for connecting stochastic models to realize a large-scale cognitive model . Frontiers in neurorobotics 12 (2018)
2018
-
[40]
arXiv preprint arXiv:2005.09409 (2020)
van Niekerk, B., Nortje, L., Kamper, H.: Vector-quantiz ed neural networks for acoustic unit discovery in the zerospeech 2020 challeng e. arXiv preprint arXiv:2005.09409 (2020)
2020 arXiv
-
[41]
Current Opinion in Neurobiology 17(2), 271–276 (2007)
Okanoya, K.: Language evolution and an emergent prop- erty. Current Opinion in Neurobiology 17(2), 271–276 (2007). https://doi.org/https://doi.org/10.1016/j.conb.2007.03.011
2007 doi
-
[42]
Current opinion in neurobiology 17(2), 271–276 (2007) 16 T
Okanoya, K.: Language evolution and an emergent propert y. Current opinion in neurobiology 17(2), 271–276 (2007) 16 T. Taniguchi
2007
-
[43]
Psychonomic bulletin & review 24(1), 106–110 (2017)
Okanoya, K.: Sexual communication and domestication ma y give rise to the signal complexity necessary for the emergence of language: An indi cation from songbird studies. Psychonomic bulletin & review 24(1), 106–110 (2017)
2017
-
[44]
In: Emergence of communicati on and language, pp
Okanoya, K., Merker, B.: Neural substrates for string-c ontext mutual segmenta- tion: A path to human language. In: Emergence of communicati on and language, pp. 421–434. Springer (2007)
2007
-
[45]
IEEE Transactions on Cognitive and Developmental Syst ems (2022)
Okuda, Y., Ozaki, R., Komura, S., Taniguchi, T.: Double a rticula- tion analyzer with prosody for unsupervised word and phone d iscov- ery. IEEE Transactions on Cognitive and Developmental Syst ems (2022). https://doi.org/10.1109/TCDS.2022.3210751
2022
-
[46]
Harvard University P ress, Cambridge (1931-1958)
Peirce, C.S.: Collected Writings. Harvard University P ress, Cambridge (1931-1958)
1931
-
[47]
Journal of Memory and Language 35(4), 606–621 (1996)
Saffran, J.R., Newport, E.L., Aslin, R.N.: Word Segmenta tion: The Role of Distri- butional Cues. Journal of Memory and Language 35(4), 606–621 (1996)
1996
-
[48]
Trends in cognitive sciences 17(11), 565–573 (2013)
Seth, A.K.: Interoceptive inference, emotion, and the e mbodied self. Trends in cognitive sciences 17(11), 565–573 (2013)
2013
-
[49]
The Bell system tech- nical journal 27(3), 379–423 (1948)
Shannon, C.E.: A mathematical theory of communication. The Bell system tech- nical journal 27(3), 379–423 (1948)
1948
-
[50]
In: International Conference on Human -Computer Interac- tion
Shirai, A., Taniguchi, T.: A proposal of an interactive m usic composition system using Gibbs sampler. In: International Conference on Human -Computer Interac- tion. pp. 490–497. Springer (2011)
2011
-
[51]
Journal o f Japan Soci- ety for Fuzzy Theory and Intelligent Informatics 25(6), 901–913 (2013)
Shirai, A., Taniguchi, T.: A proposal of the melody gener ation method using variable-order pitman-yor language model. Journal o f Japan Soci- ety for Fuzzy Theory and Intelligent Informatics 25(6), 901–913 (2013). https://doi.org/10.3156/jsoft.25.901
2013 doi
-
[52]
Journal of Co gnitive Neuroscience 33(8), 1595–1611 (2021)
Sternin, A., McGarry, L.M., Owen, A.M., Grahn, J.A.: The effect of familiarity on neural representations of music and language. Journal of Co gnitive Neuroscience 33(8), 1595–1611 (2021)
2021
-
[53]
Advanced Robotics 36(5-6), 261–278 (2022)
Suzuki, M., Matsuo, Y.: A survey of multimodal deep gener ative models. Advanced Robotics 36(5-6), 261–278 (2022)
2022
-
[54]
Neural Networks 151, 317–335 (2022)
Taniguchi, A., Fukawa, A., Yamakawa, H.: Hippocampal fo rmation-inspired prob- abilistic generative model. Neural Networks 151, 317–335 (2022)
2022
-
[55]
: Online spatial concept and lexical acquisition with simultaneous localization an d mapping
Taniguchi, A., Hagiwara, Y., Taniguchi, T., Inamura, T. : Online spatial concept and lexical acquisition with simultaneous localization an d mapping. In: IEEE/RSJ International Conference on Intelligent Robots and System s. pp. 811–818 (2017)
2017
-
[56]
: Improved and scalable on- line learning of spatial concepts and language models with m apping
Taniguchi, A., Hagiwara, Y., Taniguchi, T., Inamura, T. : Improved and scalable on- line learning of spatial concepts and language models with m apping. Autonomous Robots 44(6), 927–946 (2020)
2020
-
[57]
Advanced Robotics 35(8), 471–489 (2021)
Taniguchi, A., Isobe, S., El Hafi, L., Hagiwara, Y., Tanig uchi, T.: Autonomous planning based on spatial concepts to tidy up home environme nts with service robots. Advanced Robotics 35(8), 471–489 (2021)
2021
-
[58]
arXiv preprint arXiv:2201.06786 (2022)
Taniguchi, A., Murakami, H., Ozaki, R., Taniguchi, T.: U nsupervised multimodal word discovery based on double articulation analysis with c o-occurrence cues. arXiv preprint arXiv:2201.06786 (2022)
2022 arXiv
-
[59]
In: IEEE International Conference on Development and Learning (ICD L 2022)
Taniguchi, A., Muro, M., Yamakawa, H., Taniguchi, T.: Br ain-inspired probabilistic generative model for double articulation analysis of spoke n language. In: IEEE International Conference on Development and Learning (ICD L 2022). pp. 107–114 (2022)
2022
-
[60]
IEEE Transactions on Cognitive and Development al Systems 8(4), 285– 297 (2016) On Parallelism in Music and Language from Symbol Emergence S ystems 17
Taniguchi, A., Taniguchi, T., Inamura, T.: Spatial conc ept acquisition for a mobile robot that integrates self-localization and unsupervised word discovery from spoken sentences. IEEE Transactions on Cognitive and Development al Systems 8(4), 285– 297 (2016) On Parallelism in...
2016
-
[61]
Robotics and A utonomous Systems 99, 166–180 (2018)
Taniguchi, A., Taniguchi, T., Inamura, T.: Unsupervise d spatial lexical acquisition by updating a language model with place clues. Robotics and A utonomous Systems 99, 166–180 (2018)
2018
-
[62]
Advanced Robotics 30(11-12), 706–728 (2016)
Taniguchi, T., Nagai, T., Nakamura, T., Iwahashi, N., Og ata, T., Asoh, H.: Symbol emergence in robotics: A survey. Advanced Robotics 30(11-12), 706–728 (2016)
2016
-
[63]
IEEE Transactions on Cognitive and Developmental Systems 8(3), 171–185 (2016)
Taniguchi, T., Nagasaka, S., Nakashima, R.: Nonparamet ric bayesian double ar- ticulation analyzer for direct language acquisition from c ontinuous speech signals. IEEE Transactions on Cognitive and Developmental Systems 8(3), 171–185 (2016). https://doi.org/10.1109/TCDS.2016.2550591
2016
-
[64]
New Generation Computing 38(1), 23–48 (2020)
Taniguchi, T., Nakamura, T., Suzuki, M., Kuniyasu, R., H ayashi, K., Taniguchi, A., Horii, T., Nagai, T.: Neuro-serket: development of inte grative cognitive system through the composition of deep probabilistic generative m odels. New Generation Computing 38(1), 23–48 (2020)
2020
-
[65]
Advanced Robotics 30(11-12), 770–783 (2016)
Taniguchi, T., Nakashima, R., Liu, H., Nagasaka, S.: Dou ble articula- tion analyzer with deep sparse autoencoder for unsupervise d word dis- covery from speech signals. Advanced Robotics 30(11-12), 770–783 (2016). https://doi.org/10.1080/01691864.2016.1159981
2016
-
[66]
Advanced Robotics 21(10), 1177–1199 (2007)
Taniguchi, T., Sawaragi, T.: Incremental acquisition o f behaviors and signs based on a reinforcement learning schemata model and a spike timin g-dependent plastic- ity network. Advanced Robotics 21(10), 1177–1199 (2007)
2007
-
[67]
IEEE Transactions on Cogn itive and Develop- mental Systems (2018)
Taniguchi, T., Ugur, E., Hoffmann, M., Jamone, L., Nagai, T., Rosman, B., Mat- suka, T., Iwahashi, N., Oztop, E., Piater, J., et al.: Symbol emergence in cognitive developmental systems: a survey. IEEE Transactions on Cogn itive and Develop- mental Systems (2018)
2018
-
[68]
Neural Networks 150, 293–312 (2022)
Taniguchi, T., Yamakawa, H., Nagai, T., Doya, K., Sakaga mi, M., Suzuki, M., Nakamura, T., Taniguchi, A.: A whole brain probabilistic generative model: Toward realizing cognitive architectures for developmental robo ts. Neural Networks 150, 293–312 (2022)
2022
-
[69]
: Emergent communica- tion through metropolis-hastings naming game with deep gen erative models
Taniguchi, T., Yoshida, Y., Taniguchi, A., Hagiwara, Y. : Emergent communica- tion through metropolis-hastings naming game with deep gen erative models. arXiv preprint arXiv:2205.12392 (2022)
2022 arXiv
-
[70]
Frontiers in neurorobot ics 12, 22 (2018)
Taniguchi, T., Yoshino, R., Takano, T.: Multimodal hier archical dirichlet process- based active perception by a robot. Frontiers in neurorobot ics 12, 22 (2018)
2018
-
[71]
arXiv preprint arXiv:2005.11676 (2020)
Tjandra, A., Sakti, S., Nakamura, S.: Transformer vq-va e for unsupervised unit discovery and speech synthesis: Zerospeech 2020 chall enge. arXiv preprint arXiv:2005.11676 (2020)
2020 arXiv
-
[72]
Semiotica 89(4), 319–391 (1992)
Von Uexk¨ ull, J.: A stroll through the worlds of animals a nd men: A picture book of invisible worlds. Semiotica 89(4), 319–391 (1992)
1992
-
[73]
L.: Music in the brain
Vuust, P., Heggli, O.A., Friston, K.J., Kringelbach, M. L.: Music in the brain. Na- ture Reviews Neuroscience 23(5), 287–305 (2022)
2022
-
[74]
In: International Conference on Neural Information Processing
Yamakawa, H., Osawa, M., Matsuo, Y.: Whole brain archite cture approach is a feasible way toward an artificial general intelligence. In: International Conference on Neural Information Processing. pp. 275–281. Springer (2 016)
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.