Pith. sign in

REVIEW 2 major objections 4 minor 27 references

Beyond Call and Response: Modelling Reciprocal Coordination in Human-AI Vocal Ensembles

T0 review · 2 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read The paper argues that unconducted vocal ensembles coordinate through many-to-many reciprocal adjustment with no external clock or tuning reference, making call-and-response the wrong base model, and that an artificial singer can be built…

desk verdict A genuinely new problem framing for musical agents—unconducted vocal ensembles without a shared clock—with an honest staged architecture, but the empirical route currently rests on alignment accuracy that has not been independently validated. read the letter →

arxiv 2608.07376 v1 pith:CTG7G3HS submitted 2026-08-07 cs.HC cs.SD

classification cs.HCcs.SD
keywords reciprocalcoordinationvocalensembleshuman-AIinteractioncollectivestatenon-isochronymusicalagentsaudioalignmentcoupleddynamicsystems
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that unconducted vocal ensembles are not well modelled as call-and-response loops, because singers affect one another continuously and in many directions with no conductor, score, metronome, or tuning source to fix timing and pitch. It defines reciprocal coordination without an external reference as a distinct problem, and proposes a research architecture in which an artificial singer enters the ensemble's evolving collective state instead of reacting to one human input stream. The architecture runs from field capture of singer-dominant recordings, through singing-aware representation and inference of shared drift and mutual influence, to an agent whose timing policy is conditioned on that inferred state. Non-isochronous ritual songs, whose temporal contour cannot be reduced to a beat grid, are treated as the hard test case for the framework.

What carries the argument

The load-bearing machinery is the coupled-state equation and the influence network $A(t)$ it introduces, together with the distinction between structural, individual, relational, and collective state. Structural state locates the performance within a learned song; individual state tracks each singer's timing, pitch, and uncertainty; relational state captures pairwise lag and correction; collective state describes higher-order organisation such as common drift or distributed leadership. For non-isochronous repertoire the paper replaces the beat grid with a probabilistic verse-shape built from hierarchical semi-Markov or Markov-renewal models, so periodicity is learned as a property of repertoire rather than imposed as a prerequisite. VocalLanes supplies the multi-channel, singer-dominant, repeated-take recordings on which inference of these states would run.

What would settle it

An independent reference recording of the same rehearsals would settle it: if a substantial fraction of VocalLanes pairwise alignments deviate by more than roughly 30–40 ms, the collective-state inference step has no reliable input and the proposed route collapses.

Watch

Extended reading notes

Core claim

The central claim is that collective organisation in unconducted singing emerges from many-to-many reciprocal adjustment, so the proper computational object is not a tempo estimate or a synchrony score but the changing structure of mutual influence among participants. The paper formalizes each singer's next state as $x_i(t+1)=f_i(x_i(t), x_{-i}(t-d), A_i(t), q(t), c)$, where $A_i(t)$ holds time-varying influence weights from other singers, $q(t)$ is a learned structural representation of the current song and phrase, and $c$ bundles singer-, genre-, and tradition-specific constraints; no term is a privileged global clock. On this view a performance can remain highly coordinated while absolute tempo and pitch drift, because coherence is relational. The paper's stated contribution is to define reciprocal coordination without an external reference as a distinct problem and set out an empirically testable route to an artificial singer within the ensemble's dynamics, with the VocalLanes corpus of aligned phone recordings as the data foundation.

Load-bearing premise

The entire pipeline depends on the phone-recorded pairwise alignments being accurate enough — the paper's own target is 40 ms, supported only by pink-noise tests and an internal consistency recheck — because every later stage consumes those alignments as if they were ground truth for coordination.

Editorial extensions

If this is right

  • If the formulation is right, interactive music systems for ensemble singing should be judged by how they change mutual influence among singers, not by how accurately they follow a beat or a score.
  • A coordinated artificial singer can be built without assuming a global clock: its influence can be graded by uncertainty, so it may join a shared phrase expansion, reduce its pull while leadership is unresolved, or temporarily follow.
  • The framework implies that high coordination and large departure from an external tempo or tuning standard can coexist, making absolute alignment metrics the wrong success criterion.
  • Non-isochronous ritual songs become the hard test case: they require a model that knows where the ensemble is within a learned structural contour before 'ahead' or 'behind' has any meaning.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same relational-coordination formalism could be carried into spoken conversation, dance, sports, or any joint action where timing is co-produced without a privileged reference; the paper does not develop these applications.
  • A direct perturbation test would be to have the agent deliberately pull or hold a phrase while human singers continue; if singers measurably shift their timing in response, the influence network $A(t)$ is being manipulated and can be checked against the singers' own accounts.
  • The hardest unresolved dependency is data quality: unless pairwise phone alignment in real rehearsals is shown against an independent reference to stay near the 40 ms target, the proposed collective-state inference has no verified input no matter how sound the model is.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. This paper argues that unconducted vocal ensembles cannot be adequately modelled by the call-and-response loop common in interactive music systems, and instead proposes a research framework based on coupled dynamic systems. The central contribution is to define reciprocal coordination without an external reference as a distinct problem and to sketch an empirically testable route to an artificial singer that enters the ensemble's collective state. The manuscript presents Eq. (1) as a minimal formulation of singer state evolution, distinguishes structural, individual, relational, and collective state, treats non-isochronous ritual songs as a hard case, describes the VocalLanes corpus of 40 aligned phone-recorded takes, and proposes a staged pipeline covering representation, state inference, agent policy, and in-situ evaluation. The paper is transparent that influence modelling, agent policy, and comparative evaluation remain proposed, and that the alignment validation performed so far is only an internal consistency check.

Significance. If the proposed research route is pursued successfully, the paper would broaden the design space of human–AI musical interaction beyond turn-taking and reference-based synchronisation, and would provide a computational handle on phenomena such as shared drift, distributed leadership, and non-isochronous temporal contours. The paper's use of a real field corpus, its explicit staging of implementation status, and its attention to consent, cultural specificity, and the limitations of alignment validation are strengths. There are no fitted parameters or quantitative predictions here, so the usual circularity concerns about fitted models do not apply; the paper is best read as a problem definition and a research agenda rather than as an empirical claim. Its main scientific value lies in the falsifiable predictions it sets out for future work, especially the held-out prediction test for mutual influence and the comparative rehearsal study.

major comments (2)
  1. [Section 4] The empirical route depends on pairwise phone alignment being accurate enough to resolve coordination at the required timescale, but the only validation reported is a pink-noise test with known offsets and an internal-consistency recheck of two takes with residual offsets of 10–16 ms. The paper itself states that 'without an independent reference, this is an internal consistency check rather than an accuracy evaluation,' and that a minority of real takes approach a roughly 30 ms lag that becomes noticeable in rehearsal. Unknown alignment-error distribution on real ensembles can produce spurious lead/lag relations in the temporal-precedence test for A(t) in Section 6, so the load-bearing premise that coordination is observable at the needed timescale is not yet established. The authors should either add an independent accuracy evaluation (for example, a small set of takes recorded with a simultaneous multi-microphone or click-track reference, with per-pair error distributions reported) or explicitly narrow the empirical testability claim until such validation exists.
  2. [Section 6] The proposed influence criterion — that singer j's recent state improves prediction of singer i's next event beyond q(t), shared drift, and i's history — is not specified enough to be falsifiable as stated. Section 2 correctly notes that signal-derived directionality is not causal ground truth and that the same lag may reflect shared song knowledge or response to a third singer. The paper therefore needs to define the prediction task, the model family, the baselines, the cross-validation protocol, and how common response to unobserved context is excluded. Without this, the A(t) estimates are not uniquely identified, and the planned technical evaluation cannot distinguish reciprocal influence from shared structure.
minor comments (4)
  1. [Section 2, Eq. (1)] The delay term d in Eq. (1) is introduced as 'perceptual and system delay' without specifying its units, whether it is a scalar or per-pair, or how it is estimated; please clarify.
  2. [Figure 1] The caption 'VocalLanes aligns phone recordings for mixed or foreground listening' does not explain the visual encoding of the figure, so a reader cannot tell what the screenshot demonstrates; a sentence describing the interface elements would help.
  3. [Section 4] The corpus description reports the number of takes and hours but not the distribution of takes over formations or the duration of individual takes; stating the range of take lengths would make the corpus description more complete.
  4. [Section 5] The sentence 'Neither assuming clean stems nor scaling a generic speech or music model resolves the target task' is a strong claim that would benefit from a concrete illustration or a citation to a failed attempt, even though the surrounding discussion of bleed and singing-specific phonetics is reasonable.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the paper is a research agenda with no fitted predictions, and the only self-references support motivation rather than the central claim.

full rationale

The paper does not derive any quantitative prediction from fitted inputs. Equation (1) is a generic state-update formulation, presented as a modelling frame rather than as a result obtained from data. The proposed influence test in Section 6 is explicitly a predictive-improvement criterion: influence is supported only when "j's recent state improves prediction of i's next event beyond q(t), shared drift, and i's history", which controls for the obvious confounds rather than building them in. The alignment validation in Section 4 is candidly labelled an internal consistency check: "without an independent reference, this is an internal consistency check rather than an accuracy evaluation", so the unresolved alignment error is an empirical validity risk for the proposed route, not a circularity. VocalNotes [20] is self-cited to support the claim that singing note boundaries are ambiguous, but that claim is background motivation for the architecture and is not the load-bearing content of the paper's contribution. VocalLanes is the author's own corpus and is used as the data infrastructure, but no result is yet claimed from it beyond the proposed route. The central contribution—defining reciprocal coordination without an external reference and proposing a staged, testable architecture—therefore does not reduce to its inputs. The score of 2 reflects only the presence of minor, non-load-bearing self-citations; there are no circular derivation steps.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The framework rests on four domain assumptions: that collective influence can be represented as time-varying weights, that phone-based aligned recordings preserve enough timing detail, that non-isochrony is not reducible to a beat grid and is best captured by hierarchical duration models, and that predictive improvement is a usable proxy for influence. None of these is demonstrated experimentally in this paper; all are reasonable starting points for a research agenda.

assumptions (4)
  • domain assumption Collective states and pairwise influence weights A(t) can be represented as time-varying variables inferred from audio observation.
    Section 2's Eq. (1) postulates an influence network A(t) as the core object; no operational definition or identifiability argument is given.
  • domain assumption Pairwise offset alignment of phone recordings preserves the timing information needed for collective-state inference.
    Section 4 relies on VocalLanes alignment, validated only by internal consistency, as the foundation for the architecture.
  • domain assumption Non-isochronous temporal contours cannot be reduced to a beat grid and are best modelled by hierarchical semi-Markov duration states.
    Section 3 asserts this as a hard case; supporting evidence from one ensemble's ritual repertoire motivates the choice, but no experiment is presented.
  • domain assumption Predictive improvement of singer i's events given singer j's history is a usable proxy for influence.
    Section 6 states influence will be supported when j's state improves prediction of i's next event; this assumes predictive structure corresponds to the coordination phenomenon of interest.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Beyond Call and Response: Modelling Reciprocal Coordination in Human-AI Vocal Ensembles." pith.science (2026). https://pith.science/paper/CTG7G3HS

@misc{pith2026260807376,
  author       = {Pith},
  title        = {Pith review of: Beyond Call and Response: Modelling Reciprocal Coordination in Human-AI Vocal Ensembles},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CTG7G3HS}},
  note         = {Machine review of arXiv:2608.07376}
}
read the original abstract

Musical interaction with AI is often organised as a response loop: a human performs, the system interprets that action, and the system answers, accompanies, or schedules a musical event. Unconducted vocal ensembles pose a different problem. Singers act simultaneously and continuously affect one another; neither timing nor pitch is fixed by a conductor, metronome, accompaniment, score, or tuning source. Collective organisation emerges from many-to-many reciprocal adjustment. This paper frames such ensembles as coupled dynamic systems and proposes a research architecture for vocal agents that enter, rather than merely track, their collective states. Some target repertoires are metrical, while others exhibit non-isochronous temporal contours that cannot be reduced to a beat grid; we treat the latter as a hard case for a general framework. The architecture connects multichannel capture in the field to dialect- and singing-aware representation, collective-state inference, vocal generation, and in-situ evaluation. The resulting agenda asks not only whether an artificial singer can synchronise, but how its presence reorganises human coordination, leadership, style, and musical transmission.

Figures

Figures reproduced from arXiv: 2608.07376 by the authors.

Figure 2
Figure 2. Instrumental note events (top) and a sung [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

27 extracted references · 23 canonical work pages

  1. [1]

    Vivien Apjok, Csilla Kereszteny, and Sándor Varga. 2024. The Polyphony Project: An Exploration of Ukrainian Musical Heritage.Yearbook for Traditional Music 56, 1 (2024), 129–135. doi:10.1017/ytm.2024.18

  2. [2]

    Luc Ardaillon and Axel Roebel. 2019. Fully-Convolutional Network for Pitch Estimation of Speech Signals. InProceedings of Interspeech 2019. 2005–2009. doi:10. 21437/Interspeech.2019-2815

  3. [3]

    Gérard Assayag, Laurent Bonnasse-Gahot, and Jean Borg. 2022. Cocreative Interaction: Somax2 and the REACH Project.Computer Music Journal46, 4 (2022), 7–25. doi:10.1162/comj_a_00662

  4. [4]

    Martin Clayton, Rebecca Sager, and Udo Will. 2005. In Time with the Music: The Concept of Entrainment and Its Significance for Ethnomusicology.European Meetings in Ethnomusicology11 (2005), 1–82

  5. [5]

    Arshia Cont. 2008. ANTESCOFO: Anticipatory Synchronization and Control of Interactive Parameters in Computer Music. InProceedings of the International Computer Music Conference

  6. [6]

    2022.Data-driven Pitch Content Description of Choral Singing Recordings

    Helena Cuesta. 2022.Data-driven Pitch Content Description of Choral Singing Recordings. Ph. D. Dissertation. Universitat Pompeu Fabra. http://hdl.handle.net/ 10803/673924

  7. [7]

    Helena Cuesta, Emilia Gómez, Agustín Martorell, and Felipe Loáiciga. 2018. Analysis of Intonation in Unison Choir Singing. InProceedings of ICMPC15- ESCOM10

  8. [8]

    Sebastian Ewert, Meinard Müller, and Peter Grosche. 2009. High Resolution Audio Synchronization Using Chroma Onset Features. InProceedings of ICASSP

Show all 27 references
  1. [9]

    Jamie Forth, Kat Agres, Matthew Purver, and Geraint A. Wiggins. 2016. Entraining IDyOT: Timing in the Information Dynamics of Thinking.Frontiers in Psychology 7 (2016), 1575. doi:10.3389/fpsyg.2016.01575

  2. [10]

    Rong Gong, Philippe Cuvillier, Nicolas Obin, and Arshia Cont. 2015. Real-Time Audio-to-Score Alignment of Singing Voice Based on Melody and Lyric Informa- tion. InProceedings of Interspeech 2015. 3007–3011. doi:10.21437/Interspeech.2015- 667

  3. [11]

    Peter E. Keller. 2014. Ensemble Performance: Interpersonal Alignment of Musical Expression. InExpressiveness in Music Performance: Empirical Approaches Across Styles and Cultures. Oxford University Press, 260–282

  4. [12]

    Large and Mari Riess Jones

    Edward W. Large and Mari Riess Jones. 1999. The Dynamics of Attending: How People Track Time-Varying Events.Psychological Review106, 1 (1999), 119–159

  5. [13]

    K. J. M. Lee and Philippe Pasquier. 2024. Musical Agent Systems: MACAT and MACataRT. InNeurIPS 2024 Workshop on Creativity and Generative AI. doi:10. 48550/arXiv.2502.00023

  6. [14]

    George E. Lewis. 2000. Too Many Notes: Computers, Complexity and Culture in Voyager.Leonardo Music Journal10 (2000), 33–39. doi:10.1162/096112100570585

  7. [15]

    Taichi Nakamura, Shinnosuke Takamichi, Naoko Tanji, Satoru Fukayama, and Hiroshi Saruwatari. 2023. JaCappella Corpus: A Japanese A Cappella Vocal Ensemble Corpus. InProceedings of ICASSP 2023. 1–5. doi:10.1109/ICASSP49357. 2023.10095569

  8. [16]

    Yuto Ozaki et al. 2024. Globally, Songs and Instrumental Melodies Are Slower and Higher and Use More Stable Pitches than Speech: A Registered Report.Science Advances10, 20 (2024), eadm9797. doi:10.1126/sciadv.adm9797

  9. [17]

    François Pachet. 2003. The Continuator: Musical Interaction With Style.Journal of New Music Research32, 3 (2003), 333–341. doi:10.1076/jnmr.32.3.333.16861

  10. [18]

    Pearce, Daniel Müllensiefen, and Geraint A

    Marcus T. Pearce, Daniel Müllensiefen, and Geraint A. Wiggins. 2010. The Role of Expectation and Probabilistic Learning in Auditory Boundary Perception: A Model Comparison.Perception39, 10 (2010), 1367–1391. doi:10.1068/p6507 Beyond Call and Response: Modelling Reciprocal Coor...

  11. [19]

    Pearce and Geraint A

    Marcus T. Pearce and Geraint A. Wiggins. 2012. Auditory Expectation: The Information Dynamics of Music Perception and Cognition.Topics in Cognitive Science4, 4 (2012), 625–652. doi:10.1111/j.1756-8765.2012.01214.x

  12. [20]

    2025–2026

    Polina Proutskova et al. 2025–2026. VocalNotes: Investigating the Perception of Note Pitch and Boundaries through Varying Transcriptions of Vocal Performances from Five Musical Cultures.Analytical Approaches to World Music13, 2 (2025– 2026). doi:10.5281/zenodo.19786850

  13. [21]

    Riley, Michael J

    Michael A. Riley, Michael J. Richardson, Kevin Shockley, and Verónica C. Ra- menzoni. 2011. Interpersonal Synergies.Frontiers in Psychology2 (2011), 38. doi:10.3389/fpsyg.2011.00038

  14. [22]

    Sebastian Rosenzweig, Helena Cuesta, Christian Weiß, Frank Scherbaum, Emilia Gómez, and Meinard Müller. 2020. Dagstuhl ChoirSet: A Multitrack Dataset for MIR Research on Choral Singing.Transactions of the International Society for Music Information Retrieval3, 1 (2020), 98–110...

  15. [23]

    Frank Scherbaum, Nana Mzhavanadze, Sebastian Rosenzweig, and Meinard Müller. 2019. Multi-media Recordings of Traditional Georgian Vocal Music for Computational Analysis. InProceedings of the International Workshop on Folk Music Analysis. 1–6

  16. [24]

    Kivanç Tatar and Philippe Pasquier. 2019. Musical Agents: A Typology and State of the Art towards Musical Metacreation.Journal of New Music Research48, 1 (2019), 56–105. doi:10.1080/09298215.2018.1511736

  17. [25]

    Yoann Teytaut, Antoine Petit, Céline Chabot-Canet, and Axel Roebel. 2023. A Mu- sicological Pipeline for Singing Voice Style Analysis with Neural Voice Processing and Alignment. InJournées d’Informatique Musicale 2023. HAL: halshs-04812738

  18. [26]

    Yoann Teytaut and Axel Roebel. 2021. Phoneme-to-Audio Alignment with Re- current Neural Networks for Speaking and Singing Voice. InProceedings of Interspeech 2021. 61–65. doi:10.21437/Interspeech.2021-1676

  19. [2009]

    doi:10.1109/ICASSP.2009.4959972

    1869–1872. doi:10.1109/ICASSP.2009.4959972

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.