{"id":"8b8e484d-192d-4369-8c4a-4ec7edea9f73","arxiv_id":"2608.10360","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"MazzikaAI compiles live MIDI, gestures, and inferred harmony into a deterministic stream of text prompts that steer Google Lyria RealTime to accompany an Arabic maqam soloist with increased quarter-tone content, without fine-tuning.","lead":"This paper describes MazzikaAI, a system that turns live Arabic maqam playing into continuously updated text prompts that steer an unmodified streaming music generator into accompanying the player. It matters because it shows a hand-authored rule layer, rather than fine-tuning, can push a general-purpose model toward microtonal, non-Western musical idioms in real time.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The off-grid-frame metric conflates vibrato and slides with microtonal scale degrees; the paper's own analysis finds no stable quarter-tone peak, so the abstract's 'reliably grounds generation in microtonal scales' is not supported.","rationale":"I read the paper in good faith as a systems contribution: a deterministic, inspectable prompt compiler that turns live MIDI and gesture into continuously refreshed natural-language prompts for an unmodified streaming generator. The architecture is clearly described, the latency instrumentation is careful, the replay harness controls for input identity, and the limitations section is unusually honest—including the failing instrument-suppression ablation and the absence of beat-level entrainment. Those properties are real and should be credited. My concern is narrower but load-bearing: the central claim that dynamic prompt compilation 'reliably grounds generation in microtonal scales' is supported mainly by an off-grid-frame metric that cannot distinguish tuned quarter-tone scale degrees from slide, vibrato, and portamento artifacts. The paper's own aggregate energy analysis shows no concentrated 150-cent peak, and §7.2 explicitly concedes that the half-flat second is not rendered as a stable scale degree. This is an internal tension between the abstract and the reported evidence, not merely a disagreement with consensus. The reader's weakest assumption was that Lyria obeys constraint-style instructions; my concern is adjacent but distinct—even if the generator obeys the prompt, the measured outcome may not mean what the abstract claims. That said, the effect is real in the narrow sense of increased off-grid frame counts, the tonic-anchoring difference is suggestive, and the paper already flags the interpretational limit. Therefore the appropriate outcome remains CONDITIONAL: require a stable-degree analysis or a tempered claim before the abstract can stand. Since the reader already issued CONDITIONAL, I leave the verdict unchanged.","tokens_in":20314,"tokens_out":3593,"duration_ms":35164,"concrete_test":"Re-analyze the deposited per-arm audio (§6.4, figshare package) with pYIN predominant-F0 tracking; segment continuous pitch tracks and compute, for grounded versus ablated conditions: (1) median and p90 duration of consecutive voiced frames in the 125–175 cent half-flat band; (2) within-segment F0 standard deviation in cents; (3) the fraction of half-flat frames occurring within 100 ms of a pitch ramp exceeding 50 cents/sec; and (4) the same off-grid metrics in a control pitch region away from the second degree (e.g., the sixth degree). If grounded half-flat segments are short, high-variance, or slide-adjacent, or if off-grid rates also rise in the control region, then the §6.4 effect is ornament-induced off-gridness rather than quarter-tone scale grounding, and the abstract's wording should be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim rests on the §6.4 maqam-grounding ablation: with grounding, 22.4% of voiced frames lie ≥35 cents off the 12-TET grid versus 14.9% without, and 59.4% of second-degree frames fall in the 125–175 cent half-flat band versus 24.1%. But 'off-grid' is not equivalent to 'quarter-tone scale degree.' The same ablation reports that the aggregate energy profile shows no concentrated 150-cent peak—second-degree energy splits nearly evenly across E♭, E-half-flat, and E. That means the extra off-grid frames are not tuned to the prescribed degree. The grounding prompt explicitly instructs slides and shimmer (Listing 1), which generate portamento and vibrato; such ornament-induced F0 excursions produce off-grid frames without establishing a microtonal scale. The paper itself concedes this in §7.2: the effect is 'inflection-without-anchoring,' and 'rendering the half-flat second as a stably tuned scale degree remains beyond prompt-level control alone.' The tonic-anchoring result is also fragile: one of three grounded arms failed to anchor (share 0.080, ranked sixth). Consequently, the abstract's 'reliably grounds generation in microtonal scales' overstates the evidence; if the measured increase is largely ornament-driven instability, the support reduces to 'the prompt changes articulation and pitch inflection,' not scale grounding.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents MazzikaAI, a real-time Arabic maqam accompaniment system that compiles live MIDI, gesture, voice, and inferred harmony into continuously updated natural-language prompts for an unmodified streaming text-to-music model (Google Lyria RealTime). The system comprises a deterministic knowledge base of six maqamat and nine genres, a four-state accompaniment policy, a prompt compiler with fixed clause ordering, and a control-signature gate that decides when to re-prompt the generator. The authors report sub-millisecond knowledge-layer latency, a measured median key-to-audible latency of 263 ms, and stability under dense re-steering. They also report three input-identical replay ablations: maqam grounding versus a generic Western prompt, instrument-suppression clause presence versus absence, and gate on/off/static modes. The central claimed result is that dynamic prompt compilation reliably grounds generation in microtonal scales, significantly increasing off-grid quarter-tone content. The paper is refreshingly candid about limitations, including the failure of instrument suppression and the absence of a stable quarter-tone peak in aggregate tuning evidence.","tokens_in":20647,"tokens_out":3774,"duration_ms":34809,"significance":"If the grounding claim were fully supported, this would be a notable demonstration that a frozen, general-purpose streaming music model can be steered into a non-Western microtonal idiom through a deterministic, knowledge-based prompt compiler, with no fine-tuning and with measurable increases in microtonal content. The paper's strengths include a detailed instrumentation methodology with a full latency decomposition, gating and stability counters, input-identical replay ablations with confidence intervals, a replication package with source code, logs, analysis scripts, and generated audio, and explicit, unusually honest limitation statements. The evaluation is careful in several respects: the ablation compares grounded versus generic prompts on byte-identical replayed input, the failure of the instrument-suppression clause is directly measured and acknowledged, and the non-anchoring arm is reported rather than hidden. The significance is real but is substantially qualified by the gap between the abstract's claim of 'reliably grounds generation in microtonal scales' and the paper's own finding of 'inflection-without-anchoring.'","major_comments":[{"comment":"The central claim that dynamic prompt compilation 'reliably grounds generation in microtonal scales' is not supported by the paper's own quantitative evidence. The measured increase in off-grid frames (22.4% vs. 14.9% at ≥35 cents) and the increased share of second-degree frames in the 125–175 cent band (59.4% vs. 24.1%) show that the prompt changes pitch inflection, but the same analysis shows no concentrated 150-cent peak in the aggregate energy profile, with second-degree energy split nearly evenly across E♭, E-half-flat, and E. The paper itself states in §7.2 that the half-flat second is 'inflection-without-anchoring' and that 'rendering the half-flat second as a stably tuned scale degree remains beyond prompt-level control alone.' The abstract should be revised to say that grounding significantly increases off-grid/microtonal inflection, not that it reliably grounds generation in microtonal scales. Because the abstract's assertion is the headline contribution, this is a load-bearing overstatement that must be corrected before publication.","section":"Abstract; §6.4; §7.2"},{"comment":"The tonic-anchoring evidence is fragile: with only three grounded arms, one arm failed to anchor (D share 0.080, ranked sixth), while the pooled share (0.140) and the other two arms (0.175, 0.163) favored D. The text appropriately discloses the failed arm, but the surrounding language ('grounding shifts generation toward the tonic') and the abstract's 'reliably grounds' go beyond what three stochastic arms can establish. Please report per-arm distributions with confidence intervals, and either run more arms or explicitly soften the claim to 'tends to increase tonic emphasis in two of three arms.'","section":"§6.4, tonic-anchoring result"},{"comment":"The control condition B replaces the maqam grounding clause with 'a generic Western-scale phrase of matched length,' but the exact control prompt is not shown in the paper or appendix. Because the central ablation compares two prose strings, the semantic content, not merely the length, must be reproducible. Please include the verbatim control prompt (or state that it is in the replication package) and discuss how the specific wording might influence the measured off-grid frame statistics independently of maqam grounding.","section":"§6.4, condition B prompt specification"}],"minor_comments":[{"comment":"The abstract contains the word 'keytoaudibleupdate' without spacing; it should read 'key-to-audible-update.'","section":"Abstract"},{"comment":"The rows labeled 'push→next chunk' are described as a lower bound on audible latency because chunk content is not attributed to the prompt without audio analysis; this caveat is important and should be stated directly in the table caption as well as in the text.","section":"Table 5"},{"comment":"The stray 'scale: bayati' clause in a blues prompt is an instructive example of compiler transparency, but it would be useful to state whether any such leaks occurred in the actual ablation arms and whether they could affect the maqam-grounding comparison.","section":"Listing 3"},{"comment":"The frame-level second-degree analysis reports n=650 grounded vs. n=1,530 ablated frames; a brief explanation of why the number of frames in the second-degree region differs substantially across conditions would help readers assess whether the share difference is confounded by differences in how often the lead line visits that region.","section":"§6.4"},{"comment":"The reference to 'Shahriar and Tariq, 2021' classifies Qur'anic recitations rather than maqam music generally; consider clarifying the specific relation to maqam or citing additional maqam-classification works to support the claim in §2.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest and technically careful, but the title and abstract overclaim relative to the evidence: the grounding ablation shows prompt-induced microtonal inflection, not reliable scale grounding, and the authors' own §7.2 concedes this. The central claims are fixable by recalibrating the language throughout the abstract, contributions, and conclusion, and by strengthening or qualifying the tonic-anchoring evidence. The positive aspects—instrumentation, replication package, explicit limitations—make this a worthy revision rather than a rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the honest take. The paper has a genuinely new idea: treat prompt synthesis as a deterministic, gated control law for a streaming music model, with the musical expertise in a hand-built knowledge base. That framing is real, and the engineering is careful. They measure latency, gating, and ablate with input-identical replay, and they report negative results (instrument suppression fails, one grounding arm doesn't anchor, no stable quarter-tone peak). That honesty earns credit.\n\nThe soft spot is the abstract. \"Reliably grounds generation in microtonal scales\" is not what the data shows. The grounding ablation finds more off-grid frames and more half-flat-band visits, but the paper's own energy analysis finds no concentrated quarter-tone peak—the effect is inflection without anchoring, as they say in §7.2. So the headline overstates; the real result is that prompt text shifts pitch content toward microtonal inflection but doesn't produce stable scale degrees. That's still interesting, but the abstract needs to say that.\n\nTwo more issues. The figshare DOI is a placeholder (10.6084/m9.figshare.XXXXXXX), which is a reproducibility red flag. And the evaluation rests on a proprietary hosted model, so regenerated audio varies and the control layer's effector is a black box. They acknowledge this, but it bounds the claim.\n\nWorth engaging with? Yes. The methodological pattern—deterministic compiler plus gate over a frozen generative model—is portable and worth discussing. Bring it to reading group and send it to a serious referee. It's not a desk reject; it needs revision, and the authors seem equipped to do it.","headline":"Genuinely novel prompting-as-control system with honest measurement, but the abstract overclaims reliable microtonal grounding when the evidence shows inflection without stable tuning.","tokens_in":21134,"tokens_out":1944,"would_cite":true,"duration_ms":18736,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A deterministic state-to-prompt compiler can steer an unmodified streaming music generator into idiomatic Arabic maqam, measurably increasing off-grid quarter-tone content without fine-tuning.","keywords":["Arabic maqam","microtonal music","real-time accompaniment","prompt compilation","streaming text-to-music","knowledge-based system","call-and-response","human-AI co-creation"],"falsifier":"Re-run the input-identical replay ablation with a larger number of arms: if, with a matched-length generic scale phrase, the off-equal-tempered frame share no longer trails the grounded condition by roughly seven percentage points and second-degree visits no longer concentrate at 125–175 cents, the grounding claim fails. Conversely, the instrument-suppression limitation would be overturned if an audio tagger consistently found excluded instruments less frequent with the silence clause present across sessions.","tokens_in":20149,"feed_emoji":"🎵","tokens_out":5854,"duration_ms":49978,"temperature":0.7,"pith_summary":"MazzikaAI sets out to prove that an unmodified, general-purpose streaming text-to-music model can serve as a real-time Arabic maqam accompanist if a deterministic knowledge-based compiler first translates the soloist's live playing into continuously updated natural-language prompts. The paper's core claim is that prompt compilation is a genuine control law: spelling out the maqam's microtonal degrees, ornaments, resting tones, and negative guidance steers a Western-trained generator toward quarter-tone melodic content, with no fine-tuning and no special API. This matters because it offers a cheap, inspectable path to culturally specific interactive music, and because it reframes live human-AI co-creation as compilation: the intelligence lives in a transparent state-to-prompt layer while the foundation model supplies timbre and fluency. Measured end-to-end, the knowledge-based layer costs under a millisecond and the median key-to-audible latency is 263 ms, landing inside the answer window the controller itself opens.","feed_headline":"Prompt compiler lifts a music AI into Arabic quarter tones","feed_subtitle":"Live performance is compiled into prompts that steer a frozen generator, measurably raising quarter-tone content.","key_machinery":"The load-bearing mechanism is the performance-to-prompt compiler C:(s_t,q_t)→p_t, a deterministic, side-effect-free function that turns an estimated performance state and a four-state accompaniment behaviour (Supporting, Responding, Sustain, LongIdle) into a single ordered natural-language string. Its power comes from clause order and content: hard constraints first (an instrument-rule clause naming only active instruments and ordering silence for the rest), then the maqam grounding clause that spells microtonal degrees phonetically (e.g., 'E-half-flat', 'E koron'), characteristic ornaments, resting tones, and negative guidance against Western tonality, then harmonic context, an echo clause listing the soloist's last pitches, and role assignments per instrument. A coarse control signature—built from a two-semitone harmony fingerprint, state flags, and coarse parameter buckets—decides when a changed prompt is actually pushed to the streaming model, so the live stream is re-steered on musical change rather than on every note.","core_discovery":"On its own terms, the paper's discovery is that prompt-level grounding measurably changes the pitch behavior of a hosted streaming generator in the direction of Arabic maqam. In input-identical replay ablations under the bayati configuration, the share of voiced frames lying at least 35 cents off the equal-tempered grid rises to 22.4% with the maqam grounding clause versus 14.9% with a matched-length generic scale phrase, and when the lead line visits the second-degree region above D, 59.4% of grounded frames fall in the half-flat band (125–175 cents, E koron) against 24.1% ablated. The same evidence shows the limit of prompt control: aggregate pitch-class energy has no concentrated quarter-tone peak, so the model inflects toward the half-flat second rather than tuning to it as a stable scale degree. The paper also establishes the control architecture—a four-state call-and-response policy, a fingerprint-based re-prompt gate, and two decoupled melodic and percussion streams—as stable under 179 re-prompts per minute, with zero stream failures, and shows that a static prompt collapses trigger-to-audio latency from a 470 ms median to about 35 seconds.","pith_inferences":["A natural next experiment is to test whether replacing the negative instrument clause with positive phrasing (naming only the desired ensemble, never the excluded instruments) improves suppression; the paper's own failure of negative prompting suggests that naming a class may evoke it.","The compiler pattern appears transferable beyond music: any domain where a frozen generative model must track a changing external state (live visuals driven by pose, code completion driven by edit history) could reuse the estimate-state-to-prose-and-gate loop, and the paper itself sketches this transfer without validating it.","If beat tracking and tempo re-anchoring are added at phrase boundaries, the main perceptual gap the experts identified—the ensemble holding its own clock—could close without changing the grounding mechanism; this is the logical next build but not yet demonstrated.","The 'inflection-without-anchoring' profile may be heard by maqam-trained listeners either as expressive authenticity or as out-of-tune approximation; a controlled listening study with those listeners could settle whether the measured quarter-tone share is the right objective at all."],"forward_implications":["If the grounding result holds generally, any prompt-steerable music model can be pointed at a new microtonal idiom by authoring a mode description and a parameter table, with no retraining or paired data.","A deterministic compiler makes the control layer inspectable and cheap: defects surface as stray English in the prompt artifact, and the knowledge-based stages total under a millisecond of the end-to-end budget.","Turn-taking absorbs reaction delay: a 263 ms median reply lands at the front of a 0.3–2.5 s answer window, so response-time requirements for live accompaniment relax to the musical gap between phrases.","The re-prompt gate is an API-economy device rather than a stability safeguard; disabling it still yields zero stream failures, so future designs can trade gate complexity for musical coherence.","Prompt-level control delivers microtonal inflection without stable tuning: the half-flat second degree is approached and decorated more often, but not anchored as a fixed scale pitch, so pitch measurements and perceptual fidelity must be tracked separately."],"supporting_citations":[{"why":"Documents Lyria RealTime, the streaming generator whose live music session API the system steers as its effector.","marker":"Google DeepMind, 2025"},{"why":"Describes the streaming model and the audio-injection and manual-prompt-steering capabilities that MazzikaAI's loop automates.","marker":"Lyria Team, Google DeepMind, 2025"},{"why":"Establishes prompting as a way to steer frozen models, the conceptual basis for using natural language as the control surface.","marker":"Brown et al., 2020"},{"why":"Surveys prompting methods including negative guidance, which the compiler uses in its maqam and instrument-constraint clauses.","marker":"Liu et al., 2023"},{"why":"Documents the microtonal tuning problems of Turkish makam that motivate grounding rather than retraining for non-Western modes.","marker":"Bozkurt et al., 2014"},{"why":"Represents the Western-trained text-conditioned audio model family whose bias the grounding clause is designed to counteract.","marker":"Agostinelli et al., 2023"}],"fun_headline_variants":["Prompt compiler lifts AI into Arabic quarter tones","Real-time prompt tuning raises microtonal content in AI","Knowledge-based prompts steer frozen model to maqam scales","Streaming AI shows quarter-tone shift with prompt grounding","No finetune: prompt loop pushes AI toward Arabic intonation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire control layer stands on the assumption that a frozen text-to-music generator obeys constraint-style prose closely enough at live-performance time scales that prompt text is a reliable actuator—spelling microtonal degrees, naming and silencing instruments, and giving negative guidance must change the audio as intended.","fun_headline_variants_meta":{"raw":{"variants":["Prompt compiler lifts AI into Arabic quarter tones","Real-time prompt tuning raises microtonal content in AI","Knowledge-based prompts steer frozen model to maqam scales","Streaming AI shows quarter-tone shift with prompt grounding","No finetune: prompt loop pushes AI toward Arabic intonation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000392,"raw_usage":{"total_tokens":2120,"prompt_tokens":1061,"completion_tokens":1059,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":677,"completion_tokens_details":{"reasoning_tokens":979}},"tokens_in":677,"tokens_out":1059,"duration_ms":10491,"temperature":1.0,"reasoning_tokens":979,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:20:48.688172+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the input-identical replay ablation with a larger number of arms: if, with a matched-length generic scale phrase, the off-equal-tempered frame share no longer trails the grounded condition by roughly seven percentage points and second-degree visits no longer concentrate at 125–175 cents, the grounding claim fails. Conversely, the instrument-suppression limitation would be overturned if an audio tagger consistently found excluded instruments less frequent with the silence clause present across sessions.","supporting_citations":[],"review_version":1}