REVIEW 2 major objections 4 minor 27 references
Beyond Call and Response: Modelling Reciprocal Coordination in Human-AI Vocal Ensembles
T0 review · 2 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read The paper argues that unconducted vocal ensembles coordinate through many-to-many reciprocal adjustment with no external clock or tuning reference, making call-and-response the wrong base model, and that an artificial singer can be built…
desk verdict A genuinely new problem framing for musical agents—unconducted vocal ensembles without a shared clock—with an honest staged architecture, but the empirical route currently rests on alignment accuracy that has not been independently validated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the coupled-state equation and the influence network $A(t)$ it introduces, together with the distinction between structural, individual, relational, and collective state. Structural state locates the performance within a learned song; individual state tracks each singer's timing, pitch, and uncertainty; relational state captures pairwise lag and correction; collective state describes higher-order organisation such as common drift or distributed leadership. For non-isochronous repertoire the paper replaces the beat grid with a probabilistic verse-shape built from hierarchical semi-Markov or Markov-renewal models, so periodicity is learned as a property of repertoire rather than imposed as a prerequisite. VocalLanes supplies the multi-channel, singer-dominant, repeated-take recordings on which inference of these states would run.
What would settle it
An independent reference recording of the same rehearsals would settle it: if a substantial fraction of VocalLanes pairwise alignments deviate by more than roughly 30–40 ms, the collective-state inference step has no reliable input and the proposed route collapses.
Extended reading notes
Core claim
The central claim is that collective organisation in unconducted singing emerges from many-to-many reciprocal adjustment, so the proper computational object is not a tempo estimate or a synchrony score but the changing structure of mutual influence among participants. The paper formalizes each singer's next state as $x_i(t+1)=f_i(x_i(t), x_{-i}(t-d), A_i(t), q(t), c)$, where $A_i(t)$ holds time-varying influence weights from other singers, $q(t)$ is a learned structural representation of the current song and phrase, and $c$ bundles singer-, genre-, and tradition-specific constraints; no term is a privileged global clock. On this view a performance can remain highly coordinated while absolute tempo and pitch drift, because coherence is relational. The paper's stated contribution is to define reciprocal coordination without an external reference as a distinct problem and set out an empirically testable route to an artificial singer within the ensemble's dynamics, with the VocalLanes corpus of aligned phone recordings as the data foundation.
Load-bearing premise
The entire pipeline depends on the phone-recorded pairwise alignments being accurate enough — the paper's own target is 40 ms, supported only by pink-noise tests and an internal consistency recheck — because every later stage consumes those alignments as if they were ground truth for coordination.
Editorial extensions
If this is right
- If the formulation is right, interactive music systems for ensemble singing should be judged by how they change mutual influence among singers, not by how accurately they follow a beat or a score.
- A coordinated artificial singer can be built without assuming a global clock: its influence can be graded by uncertainty, so it may join a shared phrase expansion, reduce its pull while leadership is unresolved, or temporarily follow.
- The framework implies that high coordination and large departure from an external tempo or tuning standard can coexist, making absolute alignment metrics the wrong success criterion.
- Non-isochronous ritual songs become the hard test case: they require a model that knows where the ensemble is within a learned structural contour before 'ahead' or 'behind' has any meaning.
Reading between the lines
- The same relational-coordination formalism could be carried into spoken conversation, dance, sports, or any joint action where timing is co-produced without a privileged reference; the paper does not develop these applications.
- A direct perturbation test would be to have the agent deliberately pull or hold a phrase while human singers continue; if singers measurably shift their timing in response, the influence network $A(t)$ is being manipulated and can be checked against the singers' own accounts.
- The hardest unresolved dependency is data quality: unless pairwise phone alignment in real rehearsals is shown against an independent reference to stay near the 40 ms target, the proposed collective-state inference has no verified input no matter how sound the model is.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper argues that unconducted vocal ensembles cannot be adequately modelled by the call-and-response loop common in interactive music systems, and instead proposes a research framework based on coupled dynamic systems. The central contribution is to define reciprocal coordination without an external reference as a distinct problem and to sketch an empirically testable route to an artificial singer that enters the ensemble's collective state. The manuscript presents Eq. (1) as a minimal formulation of singer state evolution, distinguishes structural, individual, relational, and collective state, treats non-isochronous ritual songs as a hard case, describes the VocalLanes corpus of 40 aligned phone-recorded takes, and proposes a staged pipeline covering representation, state inference, agent policy, and in-situ evaluation. The paper is transparent that influence modelling, agent policy, and comparative evaluation remain proposed, and that the alignment validation performed so far is only an internal consistency check.
Significance. If the proposed research route is pursued successfully, the paper would broaden the design space of human–AI musical interaction beyond turn-taking and reference-based synchronisation, and would provide a computational handle on phenomena such as shared drift, distributed leadership, and non-isochronous temporal contours. The paper's use of a real field corpus, its explicit staging of implementation status, and its attention to consent, cultural specificity, and the limitations of alignment validation are strengths. There are no fitted parameters or quantitative predictions here, so the usual circularity concerns about fitted models do not apply; the paper is best read as a problem definition and a research agenda rather than as an empirical claim. Its main scientific value lies in the falsifiable predictions it sets out for future work, especially the held-out prediction test for mutual influence and the comparative rehearsal study.
major comments (2)
- [Section 4] The empirical route depends on pairwise phone alignment being accurate enough to resolve coordination at the required timescale, but the only validation reported is a pink-noise test with known offsets and an internal-consistency recheck of two takes with residual offsets of 10–16 ms. The paper itself states that 'without an independent reference, this is an internal consistency check rather than an accuracy evaluation,' and that a minority of real takes approach a roughly 30 ms lag that becomes noticeable in rehearsal. Unknown alignment-error distribution on real ensembles can produce spurious lead/lag relations in the temporal-precedence test for A(t) in Section 6, so the load-bearing premise that coordination is observable at the needed timescale is not yet established. The authors should either add an independent accuracy evaluation (for example, a small set of takes recorded with a simultaneous multi-microphone or click-track reference, with per-pair error distributions reported) or explicitly narrow the empirical testability claim until such validation exists.
- [Section 6] The proposed influence criterion — that singer j's recent state improves prediction of singer i's next event beyond q(t), shared drift, and i's history — is not specified enough to be falsifiable as stated. Section 2 correctly notes that signal-derived directionality is not causal ground truth and that the same lag may reflect shared song knowledge or response to a third singer. The paper therefore needs to define the prediction task, the model family, the baselines, the cross-validation protocol, and how common response to unobserved context is excluded. Without this, the A(t) estimates are not uniquely identified, and the planned technical evaluation cannot distinguish reciprocal influence from shared structure.
minor comments (4)
- [Section 2, Eq. (1)] The delay term d in Eq. (1) is introduced as 'perceptual and system delay' without specifying its units, whether it is a scalar or per-pair, or how it is estimated; please clarify.
- [Figure 1] The caption 'VocalLanes aligns phone recordings for mixed or foreground listening' does not explain the visual encoding of the figure, so a reader cannot tell what the screenshot demonstrates; a sentence describing the interface elements would help.
- [Section 4] The corpus description reports the number of takes and hours but not the distribution of takes over formations or the duration of individual takes; stating the range of take lengths would make the corpus description more complete.
- [Section 5] The sentence 'Neither assuming clean stems nor scaling a generic speech or music model resolves the target task' is a strong claim that would benefit from a concrete illustration or a citation to a failed attempt, even though the surrounding discussion of bleed and singing-specific phonetics is reasonable.
Circularity Check
No significant circularity: the paper is a research agenda with no fitted predictions, and the only self-references support motivation rather than the central claim.
full rationale
The paper does not derive any quantitative prediction from fitted inputs. Equation (1) is a generic state-update formulation, presented as a modelling frame rather than as a result obtained from data. The proposed influence test in Section 6 is explicitly a predictive-improvement criterion: influence is supported only when "j's recent state improves prediction of i's next event beyond q(t), shared drift, and i's history", which controls for the obvious confounds rather than building them in. The alignment validation in Section 4 is candidly labelled an internal consistency check: "without an independent reference, this is an internal consistency check rather than an accuracy evaluation", so the unresolved alignment error is an empirical validity risk for the proposed route, not a circularity. VocalNotes [20] is self-cited to support the claim that singing note boundaries are ambiguous, but that claim is background motivation for the architecture and is not the load-bearing content of the paper's contribution. VocalLanes is the author's own corpus and is used as the data infrastructure, but no result is yet claimed from it beyond the proposed route. The central contribution—defining reciprocal coordination without an external reference and proposing a staged, testable architecture—therefore does not reduce to its inputs. The score of 2 reflects only the presence of minor, non-load-bearing self-citations; there are no circular derivation steps.
Assumptions & free parameters
assumptions (4)
- domain assumption Collective states and pairwise influence weights A(t) can be represented as time-varying variables inferred from audio observation.
- domain assumption Pairwise offset alignment of phone recordings preserves the timing information needed for collective-state inference.
- domain assumption Non-isochronous temporal contours cannot be reduced to a beat grid and are best modelled by hierarchical semi-Markov duration states.
- domain assumption Predictive improvement of singer i's events given singer j's history is a usable proxy for influence.
Cite this review
Pith. "Pith review of Beyond Call and Response: Modelling Reciprocal Coordination in Human-AI Vocal Ensembles." pith.science (2026). https://pith.science/paper/CTG7G3HS
@misc{pith2026260807376,
author = {Pith},
title = {Pith review of: Beyond Call and Response: Modelling Reciprocal Coordination in Human-AI Vocal Ensembles},
year = {2026},
howpublished = {\url{https://pith.science/paper/CTG7G3HS}},
note = {Machine review of arXiv:2608.07376}
}
read the original abstract
Musical interaction with AI is often organised as a response loop: a human performs, the system interprets that action, and the system answers, accompanies, or schedules a musical event. Unconducted vocal ensembles pose a different problem. Singers act simultaneously and continuously affect one another; neither timing nor pitch is fixed by a conductor, metronome, accompaniment, score, or tuning source. Collective organisation emerges from many-to-many reciprocal adjustment. This paper frames such ensembles as coupled dynamic systems and proposes a research architecture for vocal agents that enter, rather than merely track, their collective states. Some target repertoires are metrical, while others exhibit non-isochronous temporal contours that cannot be reduced to a beat grid; we treat the latter as a hard case for a general framework. The architecture connects multichannel capture in the field to dialect- and singing-aware representation, collective-state inference, vocal generation, and in-situ evaluation. The resulting agenda asks not only whether an artificial singer can synchronise, but how its presence reorganises human coordination, leadership, style, and musical transmission.
Figures
Reference graph
Works this paper leans on
-
[1]
Vivien Apjok, Csilla Kereszteny, and Sándor Varga. 2024. The Polyphony Project: An Exploration of Ukrainian Musical Heritage.Yearbook for Traditional Music 56, 1 (2024), 129–135. doi:10.1017/ytm.2024.18
-
[2]
Luc Ardaillon and Axel Roebel. 2019. Fully-Convolutional Network for Pitch Estimation of Speech Signals. InProceedings of Interspeech 2019. 2005–2009. doi:10. 21437/Interspeech.2019-2815
work page 2019
-
[3]
Gérard Assayag, Laurent Bonnasse-Gahot, and Jean Borg. 2022. Cocreative Interaction: Somax2 and the REACH Project.Computer Music Journal46, 4 (2022), 7–25. doi:10.1162/comj_a_00662
-
[4]
Martin Clayton, Rebecca Sager, and Udo Will. 2005. In Time with the Music: The Concept of Entrainment and Its Significance for Ethnomusicology.European Meetings in Ethnomusicology11 (2005), 1–82
work page 2005
-
[5]
Arshia Cont. 2008. ANTESCOFO: Anticipatory Synchronization and Control of Interactive Parameters in Computer Music. InProceedings of the International Computer Music Conference
work page 2008
-
[6]
2022.Data-driven Pitch Content Description of Choral Singing Recordings
Helena Cuesta. 2022.Data-driven Pitch Content Description of Choral Singing Recordings. Ph. D. Dissertation. Universitat Pompeu Fabra. http://hdl.handle.net/ 10803/673924
work page 2022
-
[7]
Helena Cuesta, Emilia Gómez, Agustín Martorell, and Felipe Loáiciga. 2018. Analysis of Intonation in Unison Choir Singing. InProceedings of ICMPC15- ESCOM10
work page 2018
-
[8]
Sebastian Ewert, Meinard Müller, and Peter Grosche. 2009. High Resolution Audio Synchronization Using Chroma Onset Features. InProceedings of ICASSP
work page 2009
Show all 27 references
-
[9]
Jamie Forth, Kat Agres, Matthew Purver, and Geraint A. Wiggins. 2016. Entraining IDyOT: Timing in the Information Dynamics of Thinking.Frontiers in Psychology 7 (2016), 1575. doi:10.3389/fpsyg.2016.01575
2016
-
[10]
Rong Gong, Philippe Cuvillier, Nicolas Obin, and Arshia Cont. 2015. Real-Time Audio-to-Score Alignment of Singing Voice Based on Melody and Lyric Informa- tion. InProceedings of Interspeech 2015. 3007–3011. doi:10.21437/Interspeech.2015- 667
2015 doi
-
[11]
Peter E. Keller. 2014. Ensemble Performance: Interpersonal Alignment of Musical Expression. InExpressiveness in Music Performance: Empirical Approaches Across Styles and Cultures. Oxford University Press, 260–282
2014
-
[12]
Large and Mari Riess Jones
Edward W. Large and Mari Riess Jones. 1999. The Dynamics of Attending: How People Track Time-Varying Events.Psychological Review106, 1 (1999), 119–159
1999
- [13]
-
[14]
George E. Lewis. 2000. Too Many Notes: Computers, Complexity and Culture in Voyager.Leonardo Music Journal10 (2000), 33–39. doi:10.1162/096112100570585
2000 doi
-
[15]
Taichi Nakamura, Shinnosuke Takamichi, Naoko Tanji, Satoru Fukayama, and Hiroshi Saruwatari. 2023. JaCappella Corpus: A Japanese A Cappella Vocal Ensemble Corpus. InProceedings of ICASSP 2023. 1–5. doi:10.1109/ICASSP49357. 2023.10095569
2023
-
[16]
Yuto Ozaki et al. 2024. Globally, Songs and Instrumental Melodies Are Slower and Higher and Use More Stable Pitches than Speech: A Registered Report.Science Advances10, 20 (2024), eadm9797. doi:10.1126/sciadv.adm9797
2024 doi
-
[17]
François Pachet. 2003. The Continuator: Musical Interaction With Style.Journal of New Music Research32, 3 (2003), 333–341. doi:10.1076/jnmr.32.3.333.16861
2003 doi
-
[18]
Pearce, Daniel Müllensiefen, and Geraint A
Marcus T. Pearce, Daniel Müllensiefen, and Geraint A. Wiggins. 2010. The Role of Expectation and Probabilistic Learning in Auditory Boundary Perception: A Model Comparison.Perception39, 10 (2010), 1367–1391. doi:10.1068/p6507 Beyond Call and Response: Modelling Reciprocal Coor...
2010 doi
-
[19]
Pearce and Geraint A
Marcus T. Pearce and Geraint A. Wiggins. 2012. Auditory Expectation: The Information Dynamics of Music Perception and Cognition.Topics in Cognitive Science4, 4 (2012), 625–652. doi:10.1111/j.1756-8765.2012.01214.x
2012
-
[20]
2025–2026
Polina Proutskova et al. 2025–2026. VocalNotes: Investigating the Perception of Note Pitch and Boundaries through Varying Transcriptions of Vocal Performances from Five Musical Cultures.Analytical Approaches to World Music13, 2 (2025– 2026). doi:10.5281/zenodo.19786850
2025 doi
-
[21]
Riley, Michael J
Michael A. Riley, Michael J. Richardson, Kevin Shockley, and Verónica C. Ra- menzoni. 2011. Interpersonal Synergies.Frontiers in Psychology2 (2011), 38. doi:10.3389/fpsyg.2011.00038
2011 arXiv
-
[22]
Sebastian Rosenzweig, Helena Cuesta, Christian Weiß, Frank Scherbaum, Emilia Gómez, and Meinard Müller. 2020. Dagstuhl ChoirSet: A Multitrack Dataset for MIR Research on Choral Singing.Transactions of the International Society for Music Information Retrieval3, 1 (2020), 98–110...
2020 doi
-
[23]
Frank Scherbaum, Nana Mzhavanadze, Sebastian Rosenzweig, and Meinard Müller. 2019. Multi-media Recordings of Traditional Georgian Vocal Music for Computational Analysis. InProceedings of the International Workshop on Folk Music Analysis. 1–6
2019
-
[24]
Kivanç Tatar and Philippe Pasquier. 2019. Musical Agents: A Typology and State of the Art towards Musical Metacreation.Journal of New Music Research48, 1 (2019), 56–105. doi:10.1080/09298215.2018.1511736
2019
-
[25]
Yoann Teytaut, Antoine Petit, Céline Chabot-Canet, and Axel Roebel. 2023. A Mu- sicological Pipeline for Singing Voice Style Analysis with Neural Voice Processing and Alignment. InJournées d’Informatique Musicale 2023. HAL: halshs-04812738
2023
-
[26]
Yoann Teytaut and Axel Roebel. 2021. Phoneme-to-Audio Alignment with Re- current Neural Networks for Speaking and Singing Voice. InProceedings of Interspeech 2021. 61–65. doi:10.21437/Interspeech.2021-1676
2021 doi
-
[2009]
doi:10.1109/ICASSP.2009.4959972
1869–1872. doi:10.1109/ICASSP.2009.4959972
2009
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.