{"id":"9c072c52-594d-42af-9c2d-a2399f579c99","arxiv_id":"2608.10314","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Structured versus prose rendering of identical theory text did not change LLM-built programs as predicted; a 16-block randomized trial with exact randomization inference found the preregistered hypothesis unsupported.","lead":"A preregistered randomized trial tested whether presenting the same theoretical content as structured instructions versus prose changes the programs large language models generate from it. The result was null: the format did not produce the predicted differences, with one interaction effect appearing in one analysis pipeline but not the other.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The null verdict rests on an unmeasured sensitivity floor: no positive control was run, and the sparse response tensors (S2) leave the instrument's false-negative rate unknown.","rationale":"The paper is unusually careful: preregistered with a sealed analysis before the assignment randomness existed, exact randomization inference, full data release, and explicitly bounded claim language. The reader's CONDITIONAL verdict is appropriate. The stress-test focus is the same weakest point the reader identified: the null conclusion depends on an unmeasured sensitivity floor. I do not see a mathematical inconsistency in the analysis itself. The transform-dependence of the single H2 endpoint is disclosed and correctly prevents support under the registered rule. The H2 validity-gate failure is an attempt counter rather than a data-integrity check; the paper's handling is credible and documented. The concern is not that the conclusion is false but that the absence claim is made with an uncalibrated instrument. A positive control would settle that concern; until then CONDITIONAL is the right verdict, and no adjustment is needed.","tokens_in":49139,"tokens_out":3439,"duration_ms":38707,"concrete_test":"Run a registered positive-control stage on the same 16-block design, same two pinned snapshots, same five cards, same frozen probes, and same endpoints, replacing the renderer contrast with two prompts guaranteed to differ in executable content: one arm instructs 'set the belief output coefficient for the designated H1 rows to +20 and the continuation output to -20', and the other arm instructs the opposite signs. If this control fails to clear the registered M at least 0.15 and G at least 0.80 floors, or at least fails to yield a positive M with exact one-sided p at most 0.05, the instrument cannot be assumed sensitive enough for the null verdict.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is a null: renderer format did not produce the predicted geometry. For that claim to carry evidential weight within the apparatus, the apparatus must be able to detect a renderer effect of the registered size if one existed. Section 4.9 states this premise is unverified: 'No positive control was registered and none was run, so there is no measured floor on what this instrument could have found.' Supplement S2 makes the risk concrete: all 44 coordinate-by-model-by-stage medians and interquartile ranges are exactly zero, and between 0.6842 and 0.9608 of each coordinate's values are exactly 0.0. The generated programs mostly do not respond to the evaluator's input perturbations, so the exact tests are comparing response tensors that are largely inert. Two robustness checks remove only narrow artifacts: denominator floors are absent, and the support-floor audit does not show a single global family at zero. Neither establishes that a moderate renderer-induced shift would have produced a nonzero M or G contrast. A further sensitivity limitation is structural: criteria 23 and 24 (unique identity and reciprocal matching) fail in all four cells, so the matched-distance statistic is computed over matchings that in most blocks are neither injective nor reciprocal (Section 4.6). If this instrument has low or zero sensitivity, the observed NOT_SUPPORTED verdict and the 19/108 pass rate are compatible with an undetected effect. This does not make the reported analysis wrong, but it makes the boundary claim 'renderer format did not produce...' depend on an unmeasured quantity.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports SPECFORM, a prospective preregistered randomized experiment testing whether the presentation format of identical theory content changes the executable programs that LLMs generate. Sixteen random bits assigned 32 paired renderer slots across 16 blocks to either an intervention contract or connected prose; two pinned LLM snapshots translated five anonymous account cards, yielding 320 single-shot programs in a frozen sparse quadratic language. A deterministic evaluator measured finite-difference responses (H1, single-input; H2, interaction), with primary endpoints of cross-model matched-distance reduction M and closed-set same-account identifiability G, analyzed by exact randomization inference over 65,536 assignments under two pipelines (linear and rank). Both preregistered stages returned NOT_SUPPORTED; only 19 of 108 scientific criterion evaluations passed; identifiability was near chance in all cells; one H2 linear matched-distance contrast cleared the Bonferroni threshold but failed the registered magnitude floor under the rank pipeline. The paper concludes that renderer format did not produce the predicted uniform, family-invariant, classifiable behavioral geometry and emphasizes the methodological transparency of the design.","tokens_in":49376,"tokens_out":6845,"duration_ms":74696,"significance":"If read within its stated boundaries, the paper's main value is methodological: it provides a fully auditable randomized design for LLM theory-to-program translation, with preregistration sealed before the assignment randomness existed, exact randomization inference, a deterministic evaluator with no LLM in the loop, full data and code release, and unusually honest reporting of deviations, limitations, and robustness checks. These strengths are real and should be credited. The scientific contribution as a negative result is real but bounded: within this specific apparatus, the data do not support the predicted renderer effects. The absence of a positive control and the extreme sparsity of the response tensors mean that the result cannot, in its current form, carry the broader 'concrete boundary' claim about specification-format effects stated in the abstract and conclusion.","major_comments":[{"comment":"The central boundary claim requires the measurement apparatus to be able to detect an effect of the registered size if one existed. Supplement S2 shows that between 0.6842 and 0.9608 of each output coordinate's values are exactly zero and all 44 coordinate-by-model-by-stage medians and interquartile ranges are zero; §4.9 states 'No positive control was registered and none was run, so there is no measured floor on what this instrument could have found.' The two robustness audits in §4.9 exclude denominator floors and a global support floor, but they do not establish that a moderate renderer-induced shift would have produced a nonzero M or G contrast. The NOT_SUPPORTED verdict is therefore compatible with an insensitive measurement apparatus, and the abstract's phrase 'places a concrete boundary' overstates what an uncalibrated null can establish. I recommend adding a positive control or spiked-in effect calibration using synthetic programs with known effect sizes, or rewriting the abstract and conclusion to say that no support was found for the predicted geometry in this apparatus without asserting a general boundary on renderer-format effects.","section":"§4.9, Supplement S2 (Tables S25–S26)"},{"comment":"Criteria 23 and 24 fail in all four cells: the contract arm's cross-model account matching is, in most blocks, neither injective nor reciprocal. The paper acknowledges this as a limitation, but it is load-bearing for interpreting the primary endpoints. The matched-distance statistic M is computed over these matchings, so neither H1's null nor H2's one positive matched-distance result can be interpreted as a clean same-account distance. The paper needs a sensitivity analysis restricted to blocks with valid reciprocal and injective matchings, or an alternative matching-based distance, to show that the quantitative endpoints are not artifacts of many-to-one or non-reciprocal correspondences. Without this, the statement that one endpoint 'moved' is only as strong as the compromised matching on which it is based.","section":"§4.6, Supplement S3.2 (criteria 23–24)"},{"comment":"The main text emphasizes that the H2 linear matched-distance result 'survives' deletion of the largest block, but the re-enumeration in Supplement S3.7 shows that after that deletion the exact p is 0.00250, which no longer clears the registered Bonferroni threshold of 0.0025, and the corresponding rank-pipeline result drops below significance (p = 0.0569). The paper does disclose these numbers in the supplement, but §4.5 and §6 should state explicitly in the main text that the corrected-threshold status is lost after the deletion, so that readers do not infer robustness to the multiplicity correction. This is particularly important because the abstract and conclusion give the one moving endpoint prominent placement.","section":"§4.5/§6 vs Supplement S3.7 (Table S23)"}],"minor_comments":[{"comment":"The notation '2.746,582×10−5' uses a comma as a decimal separator, which is inconsistent with the rest of the manuscript and is likely to be misread; it should be formatted as 2.746582×10⁻⁵.","section":"Table 7, Table S15"},{"comment":"The caption includes the driver field 'joint_verdict=INVALID' without immediate explanation; since the paper's scientific verdict is not_supported, add a parenthetical noting that this is a provenance field and not the scientific conclusion, to avoid confusing readers who encounter the figure before reading Table 1.","section":"Figure 1 caption"},{"comment":"The discussion of the two pipelines' different achievable ranges is honest and useful, but it should cross-reference Supplement S3.7's leave-one-block-out result so that the sensitivity of the one moving endpoint is connected to the design-limitation discussion.","section":"§5.3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is exemplary in transparency: the preregistration, exact p-values, data release, and explicit limitation statements are all to the paper's credit. My main concern is the gap between the registered null and the abstract's 'concrete boundary' claim, which the absence of a positive control does not support. I do not suspect any misconduct; the repeated abandoned sealing episodes are disclosed and the design logic is internally consistent. The paper may be a better fit for a methods-oriented venue than for a broad claims venue, but as a cs.SE contribution it is publishable after the sensitivity/calibration issue and claim-language revisions are addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this is a real randomized experiment, not another prompt-sensitivity demo: 16 independent bits from a NIST beacon assign two renderer formats in matched blocks, two pinned snapshots translate five accounts, and a deterministic evaluator executes the resulting programs. The exact randomization inference over all 65,536 assignments, the byte-identical propositions across arms, the sealed claim language, and the full data and software release are genuine contributions. Second, the result is a thoroughly bounded null. The preregistered conjunction fails 89/108; identifiability sits at chance; the one nominally significant H2 matched-distance effect survives Bonferroni under the linear pipeline but misses the magnitude floor under rank and is not robust to leaving one block out under rank. The paper labels this correctly: not_supported, not a confirmed effect.\n\nWhat it does well: the provenance design is excellent. The analysis was sealed before the randomness existed, H1's null was attested before H2 was analyzed, and every deviation, including the three repeated direction-runs and the H2 INVALID driver field, is disclosed and separated from the scientific verdict. The negative controls are well thought out. The authors do not oversell: no mechanism claim, no model-generalization claim.\n\nSoft spots. The biggest is instrument sensitivity. The paper says in Section 4.9 that no positive control was run and there is no measured floor on what the instrument could find. The supplement makes this concrete: 68–96% of output-coordinate values are exactly zero, and medians and IQRs are zero in all 44 cells. The generated programs mostly do not respond to the perturbations, so the exact tests may be comparing largely inert tensors. A null from a device that might not detect an effect of the registered size is not evidence of absence. The stress-test note lands here. Second, criteria 23 and 24 fail in all four cells: the matched-distance statistic is computed over matchings that in most blocks are neither injective nor reciprocal. That weakens not just the null but the one positive H2 result. The paper acknowledges this, but it cannot be waved away. Third, the one absolute floor applied to two pipelines with different achievable ranges is a real design flaw, honestly reported. These are limitations, not signs of bad faith.\n\nWho this is for: anyone working on LLM formalization, prompt sensitivity, or preregistered evaluation methodology. The substantive answer—format did not change behavior—is narrow and provisional, but the methodological template is worth studying.\n\nRecommendation: send it to peer review. A serious referee can push on the sensitivity floor and the matching-structure issue, but this is exactly the kind of transparent, reproducible null that should be in the literature.","headline":"A scrupulously reported preregistered null whose main weakness is not the analysis but the unmeasured sensitivity of the instrument—worth a serious referee for the methodology alone.","tokens_in":49944,"tokens_out":2546,"would_cite":true,"duration_ms":24729,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A randomized trial found that rendering identical theory text as structured contracts instead of prose did not change the executable behavior of LLM-generated programs, returning the registered verdict NOT_SUPPORTED on both stages.","keywords":["preregistered randomized experiment","large language models","theory-to-program translation","autoformalization","renderer format","exact randomization inference","null result","specification format"],"falsifier":"Run the same apparatus on a contrast known to alter generated programs — for example, feed the five account cards with one proposition deliberately omitted or its direction sign-reversed, or substitute a deliberately different theory in place of one card — and check whether the deterministic evaluator detects it in the response geometry. If the evaluator cannot register such a known manipulation, then its failure to register a renderer effect is uninterpretable; if it can, the null verdict stands as a genuine finding about renderer format.","tokens_in":48884,"feed_emoji":"🤖","tokens_out":7984,"duration_ms":67477,"temperature":0.7,"pith_summary":"The paper asks whether the way a theory is written down changes the executable program an LLM builds from it. To answer it, the author held the proposition strings byte-identical and randomized each of 16 blocks between two renderer formats — an intervention contract with HOLD FIXED / INTERVENE / REQUIRED DIRECTION labels, and connected conditional prose — while two pinned LLM snapshots translated five anonymous theory accounts into a frozen sparse-quadratic language and a deterministic evaluator scored how the resulting programs responded to input changes. The preregistered verdict is negative: both hypothesis stages returned NOT_SUPPORTED, with only 19 of 108 criterion evaluations passing and same-account identifiability at chance everywhere. If the paper is right, presentation format alone is not a reliable lever on the executable content of LLM formalizations in this fixed apparatus, and the study supplies a fully auditable way to test such claims.","feed_headline":"Structured prompts don't change LLM-built programs","feed_subtitle":"A preregistered trial varied only how theory text was presented; both registered stages returned not supported.","key_machinery":"The carrying object is the matched renderer pair — an intervention contract whose lines carry HOLD FIXED, INTERVENE, and REQUIRED DIRECTION labels, versus connected conditional sentences of the form 'With [hold], if we [change], then [response]' — applied to byte-identical proposition strings from five engineered account cards. Two pinned LLM snapshots translate each account into a frozen sparse-quadratic program language of 823 legal monomials, and a deterministic evaluator with no LLM in the loop measures how the 11 outputs respond to registered input perturbations. The primary endpoint that carries the argument is the matched-distance reduction $M = 2(d_{\\mathrm{prose}} - d_{\\mathrm{contract}})/(d_{\\mathrm{prose}} + d_{\\mathrm{contract}})$ between same-account cross-model programs; inference is an exact one-sided randomization test enumerating all 65,536 renderer-label assignments under the Fisher sharp null that renderer label is irrelevant at every slot, with the registered finite-experiment estimand being the mean over 16 blocks of each block's two potential contrasts.","core_discovery":"Stated in the paper's own terms, the registered claim was that holding the proposition strings fixed, presenting them as explicit hold-change-direction contracts would reduce direct cross-model distance between same-account response functions and improve their relative identifiability compared with connected prose. The trial's answer is that this predicted uniform, family-invariant, classifiable behavioral geometry did not appear. The preregistered conjunction returned NOT_SUPPORTED for both stages: H1 and H2 each failed the registered support rule, and only 19 of 108 scientific criterion evaluations passed. The identifiability endpoint sat at chance in all four stage-by-pipeline cells (AUC 0.469–0.523 against a registered 0.80), so one model's two renderings of an account were never mutually recognizable. The single exception that moved is H2's matched-distance reduction, which was positive and nominally significant in both pipelines and cleared the Bonferroni-corrected threshold in the linear pipeline ($0.517$, exact $p = 0.00125$), but which the rank pipeline put an order of magnitude lower ($0.069$, below the registered $0.15$ floor); because the registered rule demands the magnitude floor in both transforms, even this endpoint did not constitute support.","pith_inferences":["The near-total silence in the output tensors (medians and interquartile ranges exactly zero in all 44 coordinate-model-stage cells, with 68–96% of coordinate values exactly zero) suggests the evaluator mostly measured programs that did not move; a positive control is the missing experiment that would make the null interpretable.","The order-of-magnitude disagreement between the linear and rank pipelines on the one moving endpoint implies that the choice of response scale is itself a decision that can flip a result from supported to not supported; future preregistrations should calibrate magnitude thresholds per transform rather than applying one floor to both.","The block-randomized, exact-inference design generalizes beyond renderer format to other presentation contrasts in prompt-based formalization, such as proposition ordering, variable naming conventions, or few-shot exemplar format."],"forward_implications":["In a fixed theory-to-program apparatus, renderer format alone should not be expected to reliably change the executable behavior LLMs produce; structured intervention contracts and connected prose yielded indistinguishable same-account programs in this trial.","Closed-set same-account identifiability sat at chance (AUC 0.469–0.523 against a registered 0.80), so the two renderings of one account were never mutually recognizable across model snapshots.","The single moving endpoint — H2's matched-distance reduction under the linear pipeline ($0.517$, exact $p = 0.0013$, surviving Bonferroni) — failed the joint support rule because the rank pipeline put the same contrast at $0.069$, below the registered $0.15$ floor; the pre-registered rule declines to call transform-dependent effects support.","The trial's exact randomization inference over all 65,536 renderer-label assignments and its time-stamped preregistration provide a reusable, fully auditable template for null-result experiments in LLM-based formalization."],"supporting_citations":[{"why":"Motivates the whole problem: computational modeling forces a theory's hidden implementation choices into the open, which is the freedom the renderer contrast is meant to narrow.","marker":"Guest and Martin, 2021"},{"why":"Establishes LLM autoformalization from informal math into a formal language, the translation capability the experiment treats as a stochastic compiler step.","marker":"Wu et al., 2022"},{"why":"Shows language models can translate English into executable formal specifications, making natural-language-to-program translation a measurable task.","marker":"Hahn et al., 2022"},{"why":"Supplies the warning that prompt-sensitivity findings can be inflated by evaluation artifacts, justifying the paper's frozen deterministic evaluator with no LLM in the loop.","marker":"Hua et al., 2025"},{"why":"Analyzes why meaning-preserving prompt templates still change model behavior, the alternative explanation the renderer contrast must confront.","marker":"Liu and Chu, 2026"},{"why":"The sharpest threat here: lexical choice alone, holding the task fixed, shifts output quality, which is why any observed contrast is treated as compatible with a purely lexical effect.","marker":"Xie et al., 2026"},{"why":"Provides the interventional logic the probe battery borrows: interventions separate structures that observational data leave equivalent.","marker":"Hauser and Bühlmann, 2012"},{"why":"Reports that explicit conditional or program-like structure can improve LLM causal reasoning, the competing direction this study's null result speaks against.","marker":"Liu et al., 2025"}],"fun_headline_variants":["Format-free: LLM programs identical across renderings","Randomized trial finds no presentation effect on LLM code","LLM theory-to-code translation immune to prompt style","Structured contracts fail to alter LLM-built programs","Preregistered null result: presentation style doesn't matter"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The instrument may have been too insensitive to detect a genuine renderer effect: no positive control was registered or run, and 68 to 96 percent of each output coordinate's values are exactly zero, so the null verdict could be a false negative of the measurement apparatus rather than a true absence of effect.","fun_headline_variants_meta":{"raw":{"variants":["Format-free: LLM programs identical across renderings","Randomized trial finds no presentation effect on LLM code","LLM theory-to-code translation immune to prompt style","Structured contracts fail to alter LLM-built programs","Preregistered null result: presentation style doesn't matter"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000597,"raw_usage":{"total_tokens":2872,"prompt_tokens":1104,"completion_tokens":1768,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":720,"completion_tokens_details":{"reasoning_tokens":1690}},"tokens_in":720,"tokens_out":1768,"duration_ms":13156,"temperature":1.0,"reasoning_tokens":1690,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:10:03.203879+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same apparatus on a contrast known to alter generated programs — for example, feed the five account cards with one proposition deliberately omitted or its direction sign-reversed, or substitute a deliberately different theory in place of one card — and check whether the deterministic evaluator detects it in the response geometry. If the evaluator cannot register such a known manipulation, then its failure to register a renderer effect is uninterpretable; if it can, the null verdict stands as a genuine finding about renderer format.","supporting_citations":[],"review_version":1}