REVIEW 3 major objections 3 minor 20 references
Does the way we write a theory change the program an LLM builds from it? A prospective randomized study of renderer format in LLM theory-to-program translation
T0 review · 3 major / 3 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read A randomized trial found that rendering identical theory text as structured contracts instead of prose did not change the executable behavior of LLM-generated programs, returning the registered verdict NOT_SUPPORTED on both stages.
desk verdict A scrupulously reported preregistered null whose main weakness is not the analysis but the unmeasured sensitivity of the instrument—worth a serious referee for the methodology alone. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is the matched renderer pair — an intervention contract whose lines carry HOLD FIXED, INTERVENE, and REQUIRED DIRECTION labels, versus connected conditional sentences of the form 'With [hold], if we [change], then [response]' — applied to byte-identical proposition strings from five engineered account cards. Two pinned LLM snapshots translate each account into a frozen sparse-quadratic program language of 823 legal monomials, and a deterministic evaluator with no LLM in the loop measures how the 11 outputs respond to registered input perturbations. The primary endpoint that carries the argument is the matched-distance reduction $M = 2(d_{\mathrm{prose}} - d_{\mathrm{contract}})/(d_{\mathrm{prose}} + d_{\mathrm{contract}})$ between same-account cross-model programs; inference is an exact one-sided randomization test enumerating all 65,536 renderer-label assignments under the Fisher sharp null that renderer label is irrelevant at every slot, with the registered finite-experiment estimand being the mean over 16 blocks of each block's two potential contrasts.
What would settle it
Run the same apparatus on a contrast known to alter generated programs — for example, feed the five account cards with one proposition deliberately omitted or its direction sign-reversed, or substitute a deliberately different theory in place of one card — and check whether the deterministic evaluator detects it in the response geometry. If the evaluator cannot register such a known manipulation, then its failure to register a renderer effect is uninterpretable; if it can, the null verdict stands as a genuine finding about renderer format.
Extended reading notes
Core claim
Stated in the paper's own terms, the registered claim was that holding the proposition strings fixed, presenting them as explicit hold-change-direction contracts would reduce direct cross-model distance between same-account response functions and improve their relative identifiability compared with connected prose. The trial's answer is that this predicted uniform, family-invariant, classifiable behavioral geometry did not appear. The preregistered conjunction returned NOT_SUPPORTED for both stages: H1 and H2 each failed the registered support rule, and only 19 of 108 scientific criterion evaluations passed. The identifiability endpoint sat at chance in all four stage-by-pipeline cells (AUC 0.469–0.523 against a registered 0.80), so one model's two renderings of an account were never mutually recognizable. The single exception that moved is H2's matched-distance reduction, which was positive and nominally significant in both pipelines and cleared the Bonferroni-corrected threshold in the linear pipeline ($0.517$, exact $p = 0.00125$), but which the rank pipeline put an order of magnitude lower ($0.069$, below the registered $0.15$ floor); because the registered rule demands the magnitude floor in both transforms, even this endpoint did not constitute support.
Load-bearing premise
The instrument may have been too insensitive to detect a genuine renderer effect: no positive control was registered or run, and 68 to 96 percent of each output coordinate's values are exactly zero, so the null verdict could be a false negative of the measurement apparatus rather than a true absence of effect.
Editorial extensions
If this is right
- In a fixed theory-to-program apparatus, renderer format alone should not be expected to reliably change the executable behavior LLMs produce; structured intervention contracts and connected prose yielded indistinguishable same-account programs in this trial.
- Closed-set same-account identifiability sat at chance (AUC 0.469–0.523 against a registered 0.80), so the two renderings of one account were never mutually recognizable across model snapshots.
- The single moving endpoint — H2's matched-distance reduction under the linear pipeline ($0.517$, exact $p = 0.0013$, surviving Bonferroni) — failed the joint support rule because the rank pipeline put the same contrast at $0.069$, below the registered $0.15$ floor; the pre-registered rule declines to call transform-dependent effects support.
- The trial's exact randomization inference over all 65,536 renderer-label assignments and its time-stamped preregistration provide a reusable, fully auditable template for null-result experiments in LLM-based formalization.
Reading between the lines
- The near-total silence in the output tensors (medians and interquartile ranges exactly zero in all 44 coordinate-model-stage cells, with 68–96% of coordinate values exactly zero) suggests the evaluator mostly measured programs that did not move; a positive control is the missing experiment that would make the null interpretable.
- The order-of-magnitude disagreement between the linear and rank pipelines on the one moving endpoint implies that the choice of response scale is itself a decision that can flip a result from supported to not supported; future preregistrations should calibrate magnitude thresholds per transform rather than applying one floor to both.
- The block-randomized, exact-inference design generalizes beyond renderer format to other presentation contrasts in prompt-based formalization, such as proposition ordering, variable naming conventions, or few-shot exemplar format.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports SPECFORM, a prospective preregistered randomized experiment testing whether the presentation format of identical theory content changes the executable programs that LLMs generate. Sixteen random bits assigned 32 paired renderer slots across 16 blocks to either an intervention contract or connected prose; two pinned LLM snapshots translated five anonymous account cards, yielding 320 single-shot programs in a frozen sparse quadratic language. A deterministic evaluator measured finite-difference responses (H1, single-input; H2, interaction), with primary endpoints of cross-model matched-distance reduction M and closed-set same-account identifiability G, analyzed by exact randomization inference over 65,536 assignments under two pipelines (linear and rank). Both preregistered stages returned NOT_SUPPORTED; only 19 of 108 scientific criterion evaluations passed; identifiability was near chance in all cells; one H2 linear matched-distance contrast cleared the Bonferroni threshold but failed the registered magnitude floor under the rank pipeline. The paper concludes that renderer format did not produce the predicted uniform, family-invariant, classifiable behavioral geometry and emphasizes the methodological transparency of the design.
Significance. If read within its stated boundaries, the paper's main value is methodological: it provides a fully auditable randomized design for LLM theory-to-program translation, with preregistration sealed before the assignment randomness existed, exact randomization inference, a deterministic evaluator with no LLM in the loop, full data and code release, and unusually honest reporting of deviations, limitations, and robustness checks. These strengths are real and should be credited. The scientific contribution as a negative result is real but bounded: within this specific apparatus, the data do not support the predicted renderer effects. The absence of a positive control and the extreme sparsity of the response tensors mean that the result cannot, in its current form, carry the broader 'concrete boundary' claim about specification-format effects stated in the abstract and conclusion.
major comments (3)
- [§4.9, Supplement S2 (Tables S25–S26)] The central boundary claim requires the measurement apparatus to be able to detect an effect of the registered size if one existed. Supplement S2 shows that between 0.6842 and 0.9608 of each output coordinate's values are exactly zero and all 44 coordinate-by-model-by-stage medians and interquartile ranges are zero; §4.9 states 'No positive control was registered and none was run, so there is no measured floor on what this instrument could have found.' The two robustness audits in §4.9 exclude denominator floors and a global support floor, but they do not establish that a moderate renderer-induced shift would have produced a nonzero M or G contrast. The NOT_SUPPORTED verdict is therefore compatible with an insensitive measurement apparatus, and the abstract's phrase 'places a concrete boundary' overstates what an uncalibrated null can establish. I recommend adding a positive control or spiked-in effect calibration using synthetic programs with known effect sizes, or rewriting the abstract and conclusion to say that no support was found for the predicted geometry in this apparatus without asserting a general boundary on renderer-format effects.
- [§4.6, Supplement S3.2 (criteria 23–24)] Criteria 23 and 24 fail in all four cells: the contract arm's cross-model account matching is, in most blocks, neither injective nor reciprocal. The paper acknowledges this as a limitation, but it is load-bearing for interpreting the primary endpoints. The matched-distance statistic M is computed over these matchings, so neither H1's null nor H2's one positive matched-distance result can be interpreted as a clean same-account distance. The paper needs a sensitivity analysis restricted to blocks with valid reciprocal and injective matchings, or an alternative matching-based distance, to show that the quantitative endpoints are not artifacts of many-to-one or non-reciprocal correspondences. Without this, the statement that one endpoint 'moved' is only as strong as the compromised matching on which it is based.
- [§4.5/§6 vs Supplement S3.7 (Table S23)] The main text emphasizes that the H2 linear matched-distance result 'survives' deletion of the largest block, but the re-enumeration in Supplement S3.7 shows that after that deletion the exact p is 0.00250, which no longer clears the registered Bonferroni threshold of 0.0025, and the corresponding rank-pipeline result drops below significance (p = 0.0569). The paper does disclose these numbers in the supplement, but §4.5 and §6 should state explicitly in the main text that the corrected-threshold status is lost after the deletion, so that readers do not infer robustness to the multiplicity correction. This is particularly important because the abstract and conclusion give the one moving endpoint prominent placement.
minor comments (3)
- [Table 7, Table S15] The notation '2.746,582×10−5' uses a comma as a decimal separator, which is inconsistent with the rest of the manuscript and is likely to be misread; it should be formatted as 2.746582×10⁻⁵.
- [Figure 1 caption] The caption includes the driver field 'joint_verdict=INVALID' without immediate explanation; since the paper's scientific verdict is not_supported, add a parenthetical noting that this is a provenance field and not the scientific conclusion, to avoid confusing readers who encounter the figure before reading Table 1.
- [§5.3] The discussion of the two pipelines' different achievable ranges is honest and useful, but it should cross-reference Supplement S3.7's leave-one-block-out result so that the sensitivity of the one moving endpoint is connected to the design-limitation discussion.
Circularity Check
No significant circularity: the registered design fixes thresholds and claim language before the randomized assignment, and the negative verdict is data-driven rather than derived from its inputs.
full rationale
The paper's central claim is a preregistered negative result: under a fixed theory-to-program apparatus, the renderer contrast did not produce the predicted uniform, family-invariant, classifiable geometry (Section 6). The endpoint values (M, G, exact randomization p-values) are computed by a deterministic evaluator from independently generated programs and were not fitted to, or defined by, the outcome. The support thresholds, analysis pipelines, and permitted claim language were sealed before the assignment randomness existed (Sections 2.1 and 3.4), so the verdict cannot be a retrofitting of inputs onto outputs. The few self-references (Panossian 2026a,b) are data and software deposits and are not load-bearing for the scientific inference. The limitations noted in Sections 4.9, 5.3, S2, and S4.2—no positive control, heavily zero response tensors, and non-injective or non-reciprocal matchings—are threats to instrument sensitivity and to the evidential weight of the null, not circular reductions: they do not make any equation or parameter equal to its own input. No derivation step reduces to its own definition, no fitted quantity is relabeled as a prediction, and no author-specific uniqueness theorem is invoked to rule out alternatives. Accordingly, no circular step can be exhibited.
Assumptions & free parameters
free parameters (5)
- M magnitude floor =
0.15
- Identifiability AUC floor =
0.80
- Conditional-alignment floor =
0.20
- One-sided p threshold =
0.05
- Rank-pipeline calibration reference =
pooled four-panel reference
assumptions (7)
- domain assumption No interference between slots (SUTVA-style condition)
- domain assumption One version of each renderer
- domain assumption Assignment bits are independent fair bits from NIST Randomness Beacon
- domain assumption The frozen sparse quadratic language and probe battery can detect meaningful behavioral differences
- domain assumption The signed response distance metric is a meaningful similarity measure
- domain assumption The two model snapshots are treated as fixed comparison cases, not a random sample
- domain assumption The rank pipeline's calibration reference is arm-blind and does not introduce circularity
Cite this review
Pith. "Pith review of Does the way we write a theory change the program an LLM builds from it? A prospective randomized study of renderer format in LLM theory-to-program translation." pith.science (2026). https://pith.science/paper/77W6PHS5
@misc{pith2026260810314,
author = {Pith},
title = {Pith review of: Does the way we write a theory change the program an LLM builds from it? A prospective randomized study of renderer format in LLM theory-to-program translation},
year = {2026},
howpublished = {\url{https://pith.science/paper/77W6PHS5}},
note = {Machine review of arXiv:2608.10314}
}
read the original abstract
A verbal theory does not run: translating it into an executable model requires choices about variables, interventions, and interactions. We tested whether the presentation of otherwise identical theoretical content systematically changes the programs produced by large language models. In a prospective preregistered randomized experiment, 16 independent assignment bits allocated 32 paired renderer slots between a structured intervention contract and connected prose. Two pinned LLM snapshots translated five anonymous theoretical accounts, yielding 320 preauthorized single-shot programs in a frozen sparse quadratic language. A deterministic evaluator measured atomic finite-difference responses (H1) and mixed interaction responses (H2). Two primary endpoints assessed cross-model matched-distance reduction and closed-set same-account identifiability, with exact randomization inference and 27 preregistered support criteria evaluated across signed-linear and magnitude-rank pipelines. Both H1 and H2 returned the registered verdict NOT_SUPPORTED; only 19 of 108 criterion evaluations passed. Same-account identifiability remained near chance (AUC 0.469-0.523, against a registered 0.80 threshold). One H2 matched-distance endpoint moved and survived multiplicity correction in the signed-linear pipeline, but the corresponding magnitude-rank result missed the registered effect-size floor, so it did not satisfy the joint support rule. Thus renderer format did not produce the uniform, family-invariant, classifiable behavioral geometry predicted in advance. The result places a concrete boundary on specification-format effects in LLM theory-to-program translation and provides a fully auditable randomized design, datasets, and software for studying executable formalization.
Figures
Reference graph
Works this paper leans on
-
[1]
The complete apparatus, runtime, documents, tests, criteria, and failure rules are written into the preregistration seal
-
[2]
The seal is timestamped before one exact futureNIST Randomness Beacon2.0 pulse, fixed 120 one-minute pulses after the observed anchor
-
[3]
The monitor may send only bodylessHEADrequests to the exact target URI
One durable assignment-attempt marker starts a six-hour availability window. The monitor may send only bodylessHEADrequests to the exact target URI. Redirects are rejected
-
[4]
The target is then fetched exactly once withGET, and its raw bytes are persisted before parsing
After the firstHEAD 200, one durable target-fetch marker is written. The target is then fetched exactly once withGET, and its raw bytes are persisted before parsing
-
[5]
The target pulse, its predecessor, and the Beacon certificate are verified against the pinned certificate chain, the RSA-4096/SHA-512 signature, the output hash, the predecessor link, and the predecessor commitment
-
[6]
The assignment digest is the SHA-256 of a fixed domain separator concatenated with the sealed preregistration digest bytes and the target pulse output-value bytes
-
[7]
The first 16 digest bits, most-significant bit first, assign the intervention contract to neutral slotUorVin each block
-
[8]
The remaining 240 disjoint bits generate opaque IDs, adapters, anonymous axis order, and execution order. A timeout, failure of the sole targetGET, or a cryptographic failure stops the study. There is no alternate pulse, seed, target retry, balancing rule, or reroll. S1.7.1 The complete complementary within-block mapping Step 7 does not assign one slot an...
Show all 20 references
-
[9]
timed out against the provider
Re-derive each phase Merkle root.Every phase terminal record lists the exact path, size and SHA-256 of every artifact of that phase alongside its root. Recompute the leaf hashes, rebuild the tree, and compare the root against the value in Table B2. A single added, missing, alt...
2026
-
[11]
Then Ddirectional = (Gcontract−G prose)− (Ucontract−U prose)
Unsigned magnitude, U.Replace every response with its absolute value. Then Ddirectional = (Gcontract−G prose)− (Ucontract−U prose). This requires a signed advantage beyond output support or magnitude
-
[12]
ThenDsignature = (Gcontract−G prose)− (Scontract−S prose)
Row-free signature,S.Find the best reassignment of complete response rows before computing identifiability. ThenDsignature = (Gcontract−G prose)− (Scontract−S prose). This requires an advantage beyond an unordered bag of response rows
-
[13]
p <0.001
Broken alignment.Independently permute complete rows on both model sides 512 times, separately for each arm within each randomized block. For each arm, that arm’s conditional alignmentis its signed G minus the mean of its own shuffledG values. Dconditional is the contract arm’...
-
[14]
Time stamp
Verify the RFC 3161 receipt and read its signed time.This is offline and contacts nothing. openssl ts -verify -data <sealed-file> -in <sealed-file>.tsr \ -CAfile <tsa-chain.pem> openssl ts -reply -in <sealed-file>.tsr -text | grep -i "Time stamp"
-
[15]
Upgrading a registered proof alters a bound artifact, which is why the originals on disk still carry only PendingAttestation
Verify the three OpenTimestamps calendar proofs.Work on copies. Upgrading a registered proof alters a bound artifact, which is why the originals on disk still carry only PendingAttestation. cp <sealed-file>.ots /tmp/check.ots ots upgrade /tmp/check.ots ots verify /tmp/check.ot...
-
[16]
Record its published time and its output value
Fetch the target pulse.Query the NIST Randomness Beacon 2.0 public interface for chain 2, pulse 1894079. Record its published time and its output value. Check the pulse signature and the predecessor link against the Beacon’s own certificate chain
-
[17]
Both values are printed in Table B1; the gap is 1 hour 54 minutes
Check the ordering.The signed time from step 2 must be strictly earlier than the pulse time from step 4. Both values are printed in Table B1; the gap is 1 hour 54 minutes. This is the comparisonverify_precommit()makes, and it is fail-closed in the code
-
[18]
Recompute the assignment digestfrom the seal digest of step 1 and the pulse output value of step 4, then compare it against the published value and against the digest of the released assignment bytes. SHA256( domain || precommit_sha256_bytes || target_outputValue_bytes ) # exp...
-
[48]
timed out against the provider
Treating axis 1 as 48 exchangeable units will understate dependence. • Axis 2 is ordered by the battery, and the order differs between panels.It is the panel’s own list of intervention cases in registered order — 3 forcalibration, 10 for each holdout panel. The battery file ca...
2026
-
[2013]
Renaud Jardri, Sandrine Duverne, Alexandra S
doi: 10.1093/brain/awt257. Renaud Jardri, Sandrine Duverne, Alexandra S. Litvinova, and Sophie Denève. Experimental evidence for circular inference in schizophrenia.Nature Communications, 8:14218, 2017. doi: 10.1038/ncomms14218. He Jia, Rhys Morris, Haoye Ye, Federica Sarro, a...
2017
-
[2026]
contract
doi: 10.18653/v1/2026.acl-long.2053. 30 Jay I. Myung and Mark A. Pitt. Optimal experimental design for model discrimination. Psychological Review, 116:499–518, 2009. doi: 10.1037/a0016104. Klaus Oberauer and Stephan Lewandowsky. Addressing the theory crisis in psychology.Psych...
2021
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.