{"id":"63adb2e5-c12c-4c09-b34d-f6f6a13aec8b","arxiv_id":"2608.06638","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Musical concepts are linearly decodable and steerable in two public text-to-MIDI models, with prediction forming gradually in an encoder-decoder and late in a vocabulary-extended LLM.","lead":"This paper applies mechanistic interpretability methods (probing, lenses, patching, steering) to two public text-to-MIDI models. It finds that musical attributes are linearly decodable, prediction forms gradually in an encoder-decoder but late in a language model, and register, polyphony, and tempo/energy can be steered bidirectionally.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Post-hoc selection of the best steering layer without multiple-comparison correction may inflate the reported bidirectional steering effects, which are the core causal control claim.","rationale":"The paper's central causal claim is that patching and steering demonstrate that the identified representations are used in generation. Probing and lenses are correlational; the causal load is carried by patching and steering. Patching has explicit controls and a reasonable design, though a small sample size. Steering is the more complex analysis: it involves contrastive prompt directions, a layer sweep, and a selection rule. The post-hoc selection of the best layer per concept, with an uncorrected 2 SE threshold, is the weakest link in the causal chain because it is the step where noise can be turned into a positive result. The bidirectional protocol is a good safeguard against symmetric artifacts, but it does not protect against selection on the directional component; under the null, the maximum of nine normal draws will often exceed 2 SE. The reader's verdict already flags this selection issue in its rationale, but identifies architecture attribution as the weakest assumption; I consider the uncorrected selection a more direct threat to the primary causal evidence. A re-analysis with Benjamini-Hochberg correction or a pre-registered fixed layer would settle whether the reported effects survive. If they do not, the paper's core steering claim weakens, though the probing and lens findings remain intact. The architecture attribution is secondary: even if it is wrong, the per-model observations stand, whereas the steering selection issue affects the central claim of bidirectional causal control.","tokens_in":13534,"tokens_out":9800,"duration_ms":91163,"concrete_test":"Recompute the steering sweep with multiple-comparison control: for each concept and injection strategy, apply the Benjamini-Hochberg procedure at q=0.05 across the nine source layers using the seed-clustered standard errors, and report which layers remain significant. Also evaluate a pre-specified fixed layer (e.g., the median source layer) per concept and compare the effect size and specificity to the selected best layer; if the fixed-layer effect is not directionally consistent or fails to clear the corrected threshold, the headline steering results are selection-driven.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The steering protocol (Section 7.2) sweeps nine source layers per concept and injection strategy and selects the most specific cell clearing the exploratory |s_dir|>2 SE rule (Table 7). With nine independent tests per concept-strategy cell, the family-wise error rate at a nominal 0.05 level is about 37%, yet no correction is applied. The headline metrics in Table 6 are reported at exactly these selected layers, and the architecture-dependent recipe in Section 8 is derived from the same selected cells. The bidirectional protocol and the specificity decomposition reduce the risk of symmetric drift, but they do not remove selection-driven inflation of |s_dir|, because the selection is performed on |s_dir| itself. The paper does not report the full per-layer distribution, so the reader cannot assess how many layers would clear the threshold after correction. Since the central claim that steering produces bidirectional changes in register and polyphony (and tempo/energy in MIDI-LLM) rests on these selected configurations, the quantitative strength and possibly the existence of the effect could be overstated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript studies two public text-to-MIDI models, text2midi (encoder-decoder, REMI+ tokenization) and MIDI-LLM (decoder-only Llama 3.2 1B with extended vocabulary), using linear probing, logit and tuned lenses, activation patching, and difference-in-means activation steering. The authors report that pitch, instrumentation, harmony, and texture are linearly decodable in both models; that text2midi refines predictions gradually across depth while MIDI-LLM exhibits a sharp late transition into the musical vocabulary; that patching locates a matching late attenuation of prompt-driven instrument transfer in MIDI-LLM; and that steering produces bidirectional changes in register and polyphony in both models and tempo/energy in MIDI-LLM. They propose a bidirectional evaluation protocol that decomposes steering responses into directional and symmetric components, and they derive an architecture-dependent intervention recipe (all-layer steering for text2midi, single-layer steering for MIDI-LLM).","tokens_in":13783,"tokens_out":7990,"duration_ms":75808,"significance":"If the causal steering results survive scrutiny, this is a substantial contribution to the interpretability of symbolic music generation. The experimental design is generally careful: grouped cross-validation prevents sequence leakage in probing; label-shuffled controls confirm that the probes read information rather than probe capacity; the tuned lens is trained on disjoint sequences and evaluated against the classic lens; patching uses paired seeds and bootstrap confidence intervals; and steering is guarded by note-count stability and a bidirectional decomposition. The paper also explicitly acknowledges important limitations, including confounded contrastive prompts, proxy metrics, and the single-model-per-architecture comparison. The principal weakness is the post-hoc selection of steering layers without multiple-comparison correction, which directly affects the headline causal claims and the architecture-dependent recipe.","major_comments":[{"comment":"The headline steering results are selected post hoc over source layers without a multiple-comparison correction. For each model, concept, and injection strategy, Table 7 reports the single cell that clears the exploratory |s_dir|>2 SE rule, and Table 6 reports the 'best single-layer configuration per concept' at exactly that layer. With nine source layers in the sweep, the family-wise error rate of the 2 SE rule is on the order of 0.34–0.37 even under independence, so a substantial fraction of the reported effects could arise from extreme noise draws. The bidirectional decomposition and note-count guard reduce symmetric drift but do not remove selection-driven inflation of |s_dir|, because selection is performed on |s_dir| (or on specificity, which depends on it). The paper does not show the full per-layer distribution, so the reader cannot judge how many cells survive a correction. In addition, the Δ columns of Table 6 are recorded at the α giving the largest shift within the stable range, adding a second level of selection over the α grid. Because Section 7.3 and Section 8 build the central causal claims and the intervention recipe on these selected cells, the paper should either apply a valid multiple-comparison correction, report the full per-layer response surface, or validate the layer selection on held-out data; otherwise the quantitative strength, and possibly the existence, of the bidirectional effects is overstated.","section":"§7.2–7.3, Tables 6–7"},{"comment":"The architecture-dependent recipe is supported by only one model per architecture. The paper acknowledges in Section 8 that 'with one model per architecture we cannot isolate the conditioning pathway from every other design difference', yet the abstract and conclusion state more strongly that 'architecture shapes' prediction formation and control and that 'all-layer steering suits cross-attention-conditioned text2midi, while targeted layers provide robust control in MIDI-LLM'. Because text2midi and MIDI-LLM also differ in parameter count, tokenization, and training data, the observed differences in steering robustness cannot be attributed to the conditioning pathway without additional models or ablations. This should be reframed as a case-study hypothesis, or supported by at least one further model per architecture.","section":"§8, abstract and conclusion"}],"minor_comments":[{"comment":"The 'best layer per concept' entries are selected as the peaks of per-layer accuracy curves without confidence intervals on the layer-wise differences; since Figures 2 and 3 show the full curves this is not load-bearing, but small adjacent-layer differences should not be overinterpreted.","section":"Tables 3–4"},{"comment":"The pitch-class histogram control is reported only as summary lifts; reporting the distribution of histogram lifts, or a direct significance test against the mean-pooled activation lift, would make the 'suggestive rather than conclusive' conclusion easier to calibrate.","section":"§4.2"},{"comment":"The seed-clustered standard errors for the steering slopes are not described in enough detail; specifying the cluster-robust variance estimator, the clustering unit, and the effective degrees of freedom would improve reproducibility.","section":"§7.2"},{"comment":"Probe regularization is fixed at C=1.0 with no sensitivity analysis; the shuffled-label control mitigates probe-capacity concerns, but a short sentence on robustness to C would strengthen the probing claims.","section":"§4.1"},{"comment":"The acknowledged non-independence of the tempo/energy and polyphony metrics could be quantified, for example by reporting the correlation between note density and onset count on the generated scores, which would help interpret the relative sizes of the steering effects.","section":"§8"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about post-hoc layer selection lands and is the main reason for major revision. I do not see a load-bearing error beyond this; the probing and lens analyses are carefully controlled and the limitations are mostly acknowledged. If the authors supply a multiple-comparison-corrected or held-out validation of the steering layer selection, or appropriately downgrade the causal claims, the paper would be publishable. The architecture attribution issue is real but already partially acknowledged; it should be softened rather than treated as fatal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read it. The short version: first real mechanistic-interpretability pass on text-to-MIDI models, and the empirical work is far more careful than the subfield's norm. The load-bearing result — that text2midi and MIDI-LLM form predictions differently, with MIDI-LLM rotating from a textual basis into music tokens late — is supported from three independent angles: probing, lenses, and patching. I buy it.\n\nThe probing is done properly: grouped cross-validation, label-shuffled controls, per-layer curves, and a SynTheory-in-MIDI check that reads as a genuine control rather than decoration. The tuned lens is a good addition. Patching uses paired seeds and bootstrap CIs. The steering has a bidirectional decomposition and a note-count stability guard, and the 'layer 0' validation shows they thought about symmetric drift.\n\nThe real weakness is the one the stress-test note flags. The headline steering numbers in Table 6 are selected from nine-layer sweeps using an exploratory |s_dir|>2 SE rule, with no multiple-comparison correction. Nine independent tests per concept-strategy cell pushes the family-wise error rate to around 37%. They report Table 6 at exactly those selected layers and derive the architecture-dependent recipe from them. The full per-layer distributions are not shown, so the reader can't see whether the effect would survive a correction. The bidirectional decomposition protects against symmetric drift, but not against selection on |s_dir| itself. This is fixable: report all layers, apply a simple Bonferroni or FDR correction, or confirm on a hold-out sweep grid. As it stands, the quantitative strength of the steering claims is somewhat overstated, though probably not the existence of the effect.\n\nOther soft spots are minor: the tempo/energy proxy is partly confounded with polyphony (they admit it), and the architecture-level conclusion rests on one model per design, which they also acknowledge. Code is promised but not released; for a 'toolkit' paper that is a real condition.\n\nBottom line: this deserves a serious referee. The methods section is careful, the findings are new, and the limitations section is honest. If I were editing, I'd send it out with a request to address the selection inflation and release the code. I'd cite it if I worked in music interpretability.","headline":"First systematic interpretability pass on text-to-MIDI models; careful controls, but the headline steering numbers need a multiple-comparison correction.","tokens_in":14271,"tokens_out":2126,"would_cite":true,"duration_ms":19589,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Two text-to-MIDI models encode musical structure readably and steerably, and their architecture determines whether predictions form gradually or in a sharp late switch.","keywords":["mechanistic interpretability","text-to-MIDI","symbolic music generation","linear probing","logit lens","tuned lens","activation patching","activation steering"],"falsifier":"Apply the same single-layer and all-layer steering sweep to another cross-attention encoder-decoder text-to-MIDI model and another decoder-only model with a shared prompt-music residual stream; if the cross-attention model fails to tolerate all-layer steering, or the decoder-only model tolerates it without disruptive accumulation, the architecture-attribution claim is refuted.","tokens_in":13358,"feed_emoji":"🎹","tokens_out":6745,"duration_ms":53778,"temperature":0.7,"pith_summary":"This paper asks whether the internal activations of text-to-MIDI models encode musical structure that can be read out and controlled. It analyzes two public models of opposite design—a purpose-built encoder-decoder and a general-purpose language model extended with MIDI tokens—using probing, logit and tuned lenses, activation patching, and activation steering. The authors report that pitch, instrumentation, harmony, and texture are linearly decodable in both, that the two models form predictions in different ways, and that steering can move register and polyphony bidirectionally in both models plus tempo/energy in one of them. The payoff is a practical, architecture-aware recipe for tracing and controlling musical concepts in symbolic generators.","feed_headline":"Text-to-MIDI models plan notes gradually or in a late flip","feed_subtitle":"Pitch, harmony, texture are decodable; steering moves register and polyphony; architecture sets when notes form.","key_machinery":"The load-bearing object is the residual-stream activation $h^{(i)}_{\\ell,t}$, the state from which the model predicts token $t$ at layer $\\ell$, recorded in predictive alignment. On this object the paper stacks four instruments: linear probes (logistic regressions deliberately restricted to linear form, with shuffled-label control tasks); the logit lens and its trained variant, the tuned lens, which pass intermediate states through the final normalization and vocabulary projection to reveal what the model would predict if it stopped early; activation patching that replaces prompt-conditioned activations at a chosen layer to measure causal transfer; and difference-in-means steering vectors $d_\\ell = (\\mu^B_\\ell - \\mu^A_\\ell)/\\lVert \\mu^B_\\ell - \\mu^A_\\ell \\rVert_2$ added at one layer or at all layers. The paper's bidirectional protocol decomposes the steering response into an antisymmetric directional part and a symmetric drift part, with specificity $\\mathrm{spec} = |s_{\\mathrm{dir}}|/(|s_{\\mathrm{dir}}|+|s_{\\mathrm{non}}|)$.","core_discovery":"The central claim is that mechanistic interpretability transfers to symbolic music generation and reveals architecture-dependent routes for prediction formation. In text2midi, predictions are refined gradually across decoder depth, while in MIDI-LLM early layers operate largely in the inherited language-model basis and the readout rotates sharply into the MIDI vocabulary around layers 13 to 15, a transition also detected by activation patching. Linear probes recover musically meaningful structure in both models, including instrument family, pitch class, harmony, and texture, with controlled SynTheory-style probing showing high accuracy for intervals and chord progressions. Activation steering produces bidirectional changes in register and polyphony in both models and in tempo/energy in MIDI-LLM, and the paper's two-orientation protocol separates true directional response from symmetric drift, showing that all-layer steering is stable in text2midi but accumulates disruptively in MIDI-LLM.","pith_inferences":["Editorial inference: the two-orientation protocol could be applied to text-to-audio models, where exact note-level metrics are unavailable, by comparing outcome distributions across orientations.","Editorial inference: because MIDI-LLM's predictions form late, intervening only in its final few layers may be a cheaper way to control music attributes than full-layer steering or retraining.","Editorial inference: the controlled SynTheory results show strong interval and chord-progression decodability, so one could test whether these representations are reused in full scores by patching away a chord-progression direction and measuring harmony changes.","Editorial inference: the authors leave sparse autoencoders to future work; a natural next step is to check whether the linearly decoded concepts correspond to single sparse features or to directions distributed across many features."],"forward_implications":["The probing protocol can be applied to any symbolic music generator to audit which musical concepts are linearly accessible at which layers.","The architecture-aware intervention rule predicts that cross-attention-conditioned generators can be steered safely at all layers, while decoder-only models with a shared prompt-music residual stream should be steered at a single targeted layer.","The tuned lens and vocabulary-mass analysis give a way to detect basis changes in mixed text-music tokenizers, not just in pure language models.","The two-orientation steering protocol provides a measurable cleanliness criterion for any activation-steering direction: high antisymmetric share plus low symmetric drift.","The patch-transfer experiment offers a causal way to locate where prompt conditioning stops influencing generated musical attributes."],"supporting_citations":[{"why":"Supplies the text2midi encoder-decoder model with REMI+ tokenization that is one of the two studied systems.","marker":"[11]"},{"why":"Supplies the MIDI-LLM decoder-only model, a general-purpose language model extended with MIDI tokens, the other studied system.","marker":"[12]"},{"why":"Introduces linear classifier probes, the method used to read musical concepts from activations.","marker":"[1]"},{"why":"Provides the control-task design and the argument for restricting probes to logistic regression.","marker":"[2]"},{"why":"Introduces the logit lens used to read intermediate predictions in both models.","marker":"[3]"},{"why":"Introduces the tuned lens used to test whether the sharp logit-lens transition reflects a basis change.","marker":"[4]"},{"why":"Introduces activation steering via difference-in-means vectors, which the paper adapts.","marker":"[5]"},{"why":"Supplies the difference-in-means steering recipe for music generation models and the preference for all-layer injection.","marker":"[9]"},{"why":"Supplies the SynTheory dataset used for controlled probing of music-theory concepts in MIDI form.","marker":"[22]"},{"why":"Supplies the key-estimation method used to label estimated key on full generations.","marker":"[30]"}],"fun_headline_variants":["Text-to-MIDI notes emerge gradually or in a sharp late flip","Probing reveals architecture decides when MIDI notes form","Two text-to-MIDI models, two planning styles: gradual vs. flip","Steering music: bidirectional control of register and polyphony","Interpretability toolkit: trace and control musical concepts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the observed difference in steering stability between the two models comes from their conditioning architecture (cross-attention memory versus a shared prompt-music residual stream), but only one model of each architecture is studied, so tokenization, model size, or training data could also explain it.","fun_headline_variants_meta":{"raw":{"variants":["Text-to-MIDI notes emerge gradually or in a sharp late flip","Probing reveals architecture decides when MIDI notes form","Two text-to-MIDI models, two planning styles: gradual vs. flip","Steering music: bidirectional control of register and polyphony","Interpretability toolkit: trace and control musical concepts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000679,"raw_usage":{"total_tokens":3098,"prompt_tokens":970,"completion_tokens":2128,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":586,"completion_tokens_details":{"reasoning_tokens":2042}},"tokens_in":586,"tokens_out":2128,"duration_ms":14933,"temperature":1.0,"reasoning_tokens":2042,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T04:10:18.603994+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Apply the same single-layer and all-layer steering sweep to another cross-attention encoder-decoder text-to-MIDI model and another decoder-only model with a shared prompt-music residual stream; if the cross-attention model fails to tolerate all-layer steering, or the decoder-only model tolerates it without disruptive accumulation, the architecture-attribution claim is refuted.","supporting_citations":[{"cited_title":"Text2midi: Generating symbolic music from captions","cited_arxiv_id":null,"evidence_quote":"Supplies the text2midi encoder-decoder model with REMI+ tokenization that is one of the two studied systems."},{"cited_title":"Designing and interpreting probes with control tasks","cited_arxiv_id":null,"evidence_quote":"Provides the control-task design and the argument for restricting probes to logistic regression."},{"cited_title":"Interpreting GPT: the logit lens","cited_arxiv_id":null,"evidence_quote":"Introduces the logit lens used to read intermediate predictions in both models."},{"cited_title":"Oxford University Press, 1990","cited_arxiv_id":null,"evidence_quote":"Supplies the key-estimation method used to label estimated key on full generations."}],"review_version":1}