REVIEW 3 major objections 4 minor 16 references
Trusting sovereign language models as scientific instruments: evidence from Portugal's AMALIA
T0 review · 3 major / 4 minor · reviewed 2026-08-02 · deepseek-v4-flash
Pith's one-line read Portugal's AMALIA matches much larger models on coder agreement, but only half its authority coding survives the theory's decomposed route.
desk verdict Sound, transparent audit with a credible empirical core; the 'only half is theory' claim is a hypothesis until the decomposed instrument is calibrated on the 9B model itself. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Grain calibration and its central quantity, the recovery gap Δ = F1u − F1d. The authority codebook is decomposed into seven binary yes/no clauses in detection, distinction and appraisal components; four are recombined by the deterministic rule (D∧(A1∨A2))∨A3 ⇒ authority, with no weights or tuning. The undecomposed prompt is the whole codebook in one call; the decomposed prompt answers each clause separately and integrates by the rule. If the clause route reproduces the holistic prompt's F1 on the same ground truth, the instrument measures the construct as the theory defines it; a large Δ means the holistic performance rides on surface correlates the theory does not license. Δ belongs to the
What would settle it
On a corpus of European-Portuguese authority texts engineered to contain no surface correlates (no authority nouns, no moral-outrage vocabulary, the construct expressed only through stance and appraisal), run the registered protocol: if AMALIA's decomposed F1 matches its holistic F1 there, the shortcut explanation collapses. A cheaper version: recalibrate the construct on AMALIA itself at a coarser, still theory-respecting grain and see whether the recovery gap closes.
Extended reading notes
Core claim
Central discovery: agreement and construct validity split. The recovery gap Δ = F1u − F1d measures how much of the holistic prompt's performance survives when the codebook is decomposed into theory clauses and recombined by (D∧(A1∨A2))∨A3 ⇒ authority. On 448 out-of-sample European-Portuguese texts, AMALIA's gap is open under both instruction languages (+0.358 pt-PT, +0.436 EN; lower bounds above the closure band): the theory route recovers about half, or just over a quarter, of holistic agreement. A 120B model closes the gap (+0.028) on the same corpus and instructions, and a 70B model only widens to +0.121, so the shortfall lies in the annotator model, not the corpus or translation. AMALIA
Load-bearing premise
The diagnosis that AMALIA is using surface shortcuts rests on the unproven assumption that the seven binary clauses and the rule (D∧(A1∨A2))∨A3 are a complete, faithful and executable operationalization of the authority construct at 9B scale; if the clauses are too fine-grained, over-specific, or unexecutable for reasons unrelated to construct validity, a low decomposed score does not by itself prove the holistic prompt is shortcutting.
Editorial extensions
If this is right
- Agreement with human coders does not certify construct validity: a model can match codes through surface correlates while failing the theory-defined route.
- National-model programmes should not treat home-language fluency or agreement benchmarks as evidence that a model measures a construct; the recovery-gap audit can run inexpensively at each release checkpoint on any construct with a decomposable codebook.
- The transcreation protocol — blind to ground-truth codes, gated by five criteria, with human adjudication and a back-translation audit — lets a calibrated instrument cross a language boundary while defending, rather than assuming, code transfer.
- The executable grain of a construct belongs to the construct–model pair: the calibrated grain that works at 120B scale is below what a 9B model can execute clause by clause, so calibration does not transfer across model sizes.
- AMALIA remains usable for triage, pre-annotation and monitored coding, but not for measurement that supports inference about the authority construct.
Reading between the lines
- Implicit extension: the recovery-gap audit should generalize to any construct whose codebook decomposes into atomic clauses, making it a general validity screen for LLM annotation beyond moral foundations.
- Testable conjecture: smaller models will systematically sit at coarser executable grain; prompt recalibration may restore reliability but cannot restore substantive construct validity, so sovereign-model programmes should budget for per-construct, per-version audits.
- One cause the paper cannot isolate is training composition: AMALIA's Portuguese instruction data is largely synthetic, generated by a much larger model, and a base-checkpoint comparison would test whether fluent agreement without theory-executable reasoning is learned there.
- Practical extension: the transcreation-and-audit pipeline could be reused to move calibrated instruments into other low-resource languages or registers, turning a one-off €30 audit into a routine release gate.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a pre-registered validity audit of AMALIA-9B as an instrument for coding the moral foundation of "authority" in a European-Portuguese transcreation of 748 MFRC texts. It compares an undecomposed codebook prompt with a decomposed seven-clause route integrated by the deterministic rule (D∧(A1∨A2))∨A3 (Table 1, Eq. 1). On 448 out-of-sample texts, AMALIA agrees with trained coders within roughly six F1 points of much larger open models, but the recovery gap Δ = F_u^1 − F_d^1 (Eq. 2) is +0.352 under Portuguese instructions and +0.468 under English instructions, while GPT-OSS-120B closes the gap at +0.028 on the same corpus. The authors interpret this as a failure of substantive construct validity: only about half of AMALIA's coding agreement is reproducible through the theory-defined clause route, and the remainder is attributed to surface shortcuts. They conclude that sovereignty buys operational and performance trust but not epistemic trust, and propose the recovery gap as a portable, inexpensive audit for LLM-based measurement instruments.
Significance. The design has genuine strengths: the analysis was pre-registered; the main confirmatory endpoints are computed out-of-sample on texts no model had seen; uncertainty is quantified by whole-text bootstrap resampling; parse rates are fully reported; the transcreation pipeline includes verification gates, human adjudication, and stratified back-translation audits; and a replication package is openly archived. If the core inference is accepted, the paper makes a significant contribution by demonstrating, in a carefully controlled setting, that agreement with human coders does not by itself certify construct-valid measurement, and by offering a cheap, portable diagnostic suitable for repeated audits of national and commercial models. It would also be the first validity study of a sovereign language model as a scientific instrument in its own language. The main weakness is interpretive: the step from a low decomposed F_d^1 to the conclusion that the holistic prompt uses non-theory shortcuts is not fully supported by the design, and the paper itself acknowledges the unresolved "principle or instruction" question in Section 5.
major comments (3)
- [§5, Eq. (2), Table 4] The central claim that AMALIA's large recovery gap shows shortcut-based coding rests on treating the decomposed seven-clause instrument as a valid and executable operationalization of the authority construct for this model. That premise is not established. The construct files were calibrated in English on much larger models and transferred to AMALIA without recalibration; the paper explicitly leaves open whether the low F_d^1 is "one of principle or of instruction" (§5). A low decomposed F1 could equally mean that AMALIA cannot execute the atomized clause format in pt-PT, or that the clauses lie below its executable grain, rather than that the holistic prompt bypasses the theory. The GPT-OSS-120B comparison does not resolve this: that model operated the English prompt, is far larger, and was part of the original calibration loop, so its closure only shows the construct is executable at t
- [§4.1, Table 3, Appendix E] The shortcut attribution is mainly carried by S1 and S2, which are exploratory and rest on an LLM reading panel rather than human raters. S1 relies on 46 false positives, with 63% pre-adjudication agreement; S2 clears its 50% threshold by two texts and is described in Appendix E as "confirmed but fragile." The paper does label these endpoints as exploratory, but Section 5 then uses them to conclude that "the holistic agreement rests on surface shortcuts" and to characterize what "fills that space." That inference outruns the registered evidence. These diagnostics should be presented as hypothesis-generating, and the strong causal language in the abstract and conclusion should be correspondingly qualified.
- [§4.2, §5] The informal decomposition of the gap (corpus component near zero, language effect about 0.07, annotator component about 0.25) is presented as "a rough decomposition rather than a formal estimate." It is used, however, to support the abstract's claim that "the shortfall lies in the annotator model," and it assumes an additivity that is not justified across models with very different scales, languages, and instruction-following behavior. The Llama widening from 0.054 to 0.121 could reflect model-specific difficulties with the Portuguese corpus rather than a separable "language effect," and GPT-OSS's closure shows only that one large model can execute the construct on this corpus. This decomposition should be clearly labeled exploratory and not treated as a quantitative attribution.
minor comments (4)
- [Table 3 / Appendix E] The S2 result is reported as 54% with no sample size in Table 3; the fragility margin (two texts) appears only in Appendix E. Adding N and the margin to the table or caption would improve transparency.
- [Figure 1] The reference-model points are plotted without confidence intervals, while AMALIA's confirmatory gaps are shown with intervals. Since the interpretation depends on comparing AMALIA to GPT-OSS/Llama, adding bootstrap intervals for the reference models or reporting the differences with intervals would help readers assess the comparison.
- [§3] The deduplication of the MFRC changes the positive labels by 37.6%, so a slightly fuller description of the deduplication procedure — beyond the duplicate-row percentages and the author confirmation — would improve reproducibility. The open replication package presumably contains this, but the main text would benefit from a sentence on how duplicates were identified.
- [§5] The sentence "Agreement this close to the human annotators and the reference models, built on an annotation process this far from the theory, is therefore precisely the failure mode the framework predicts" is stronger than the evidence for the "built on" part; consider softening in line with the exploratory status of S1/S2.
Circularity Check
Recovery-gap interpretation partly built into the definition of Δ and into the self-calibrated GPT-OSS reference; central AMALIA numbers remain out-of-sample and preregistered.
-
self definitional
[Abstract; §2 Eq. (2); §5]
"Formally, ∆ = F u 1 − F d 1 (2) ... The larger the value, the more the undecomposed prompt's performance rides on evidence the theory does not define as valid for the construct. ... the decomposed prompt recovers a little over half of the agreement the undecomposed prompt achieves. The other half cannot be attributed to moral foundations theory, but probably to 'shortcuts' via surface correlates."
The construct's 'theory-defined' coding is operationalized as the decomposed clause route, so 'only about half ... can be attributed to the theory' is a restatement of the ratio F_d^1/F_u^1 (0.359/0.711 ≈ 0.5) plus the label 'shortcuts' for the complement of Δ. The shortcut attribution is not independently measured; §5 admits the low F_d^1 could be 'one of principle or of instruction,' i.e., an execution artifact rather than evidence of non-theory shortcuts.
-
fitted input called prediction
[§2; §4.2]
"The final undecomposed and decomposed prompts result from grain calibration over seven cycles, against the English MFRC ground truth used as the reference [Pita, 2026]. ... On the original English corpus, the same instrument calibrated for the 'authority' construct closes the recovery divergence with the LLM GPT-OSS-120B (∆ = 0). ... GPT-OSS-120B closes the divergence, with F u 1 0.753 and F d 1 0.725, that is, ∆G en = +0.028."
GPT-OSS is the paper's control for 'the corpus is not at fault,' but it is one of the models on which the author's prior work calibrated the instrument to close Δ in English. Its near-zero Δ on a Portuguese transcreation of the same English corpus is therefore partly a fitted property inherited from calibration, not an independent test. The conclusion 'AMALIA's open divergence is therefore primarily a property of the annotator' leans on that calibration-selected reference.
full rationale
The AMALIA measurement itself is not circular: the study is preregistered, the 448 confirmatory texts were unseen, the integration rule (D∧(A1∨A2))∨A3 was fixed a priori, and all data/code are released. The circularity score is raised because the headline interpretation ('only about half ... attributed to the theory') is an interpretive restatement of the paper's own operationalization of 'theory-defined routing,' and because the GPT-OSS reference that localizes the failure to AMALIA was itself the model on which the instrument was calibrated in the author's prior work. The paper is unusually transparent about the main vulnerability ('Establishing whether this bound is one of principle or of instruction would require running the calibration loop on the 9B model itself'), which shows the shortcut attribution is not forced by the data. That admission prevents a higher score. No external benchmark or independent clause-level validation on AMALIA is provided, so the central inference retains a self-referential component, but it is not a full reduction of the result to its inputs.
Assumptions & free parameters
free parameters (3)
- Recovery-gap interpretation bands =
closed <0.05, open >=0.10
- Clause decomposition and prompt wording
- Integration rule =
(D∧(A1∨A2))∨A3
assumptions (4)
- domain assumption The seven-clause decomposition and Boolean rule faithfully operationalize the authority construct as defined by MFRC/MFT
- domain assumption Ground-truth labels transfer from English MFRC to the Portuguese transcreation
- ad hoc to paper A large recovery gap indicates shortcuts in the undecomposed prompt rather than, e.g., decomposed-prompt difficulty
- domain assumption The LLM reading panel's judgments approximate human error analysis
invented entities (1)
-
Recovery gap Δ = F_u^1 - F_d^1
independent evidence
Cite this review
Pith. "Pith review of Trusting sovereign language models as scientific instruments: evidence from Portugal's AMALIA." pith.science (2026). https://pith.science/paper/HCMSMHTL
@misc{pith2026260708731,
author = {Pith},
title = {Pith review of: Trusting sovereign language models as scientific instruments: evidence from Portugal's AMALIA},
year = {2026},
howpublished = {\url{https://pith.science/paper/HCMSMHTL}},
note = {Machine review of arXiv:2607.08731}
}
read the original abstract
National language models are becoming publicly funded epistemic infrastructure. Public ownership, linguistic specialization, and open weights create a presumption of trustworthiness. Such an instrument, built by and for a language community, looks like the natural choice for measuring what that community says and values. Whether such a model validly measures anything is untested at release. The evaluation of LLMs as measurement instruments is typically task-specific and stops at agreement with human coders. Agreement cannot distinguish an LLM instrument that measures a construct from one that reaches matching codes through surface correlates. We audit the presumption on a favourable case: AMALIA, Portugal's publicly funded 9B model, coding the moral foundation of authority in European Portuguese. The \textit{recovery gap} operationalizes the audit: decompose the codebook into its theory-defined clauses, recombine them through the theory's explicit rule, and measure how much of the original prompt's performance the stated theory reproduces. In a pre-registered, out-of-sample study on a transcreated (English to European Portuguese) corpus, AMALIA agrees with trained coders within six points of open models eight to thirteen times its size. Yet, the recovery gap shows that only about half of coding performance on authority can be attributed to the theory. A larger multilingual LLM closes the recovery gap on the same corpus, suggesting the shortfall lies in the annotator model, not the corpus or its translation. Sovereignty earns operational and performance trust; epistemic trust requires calibration -- and the audit method is inexpensive, and portable across models, languages and tasks.
Figures
Reference graph
Works this paper leans on
-
[5]
Jesse Graham, Jonathan Haidt, Sena Koleva, Matt Motyl, Ravi Iyer, Sean P
doi: 10.1073/pnas.2305016120. Jesse Graham, Jonathan Haidt, Sena Koleva, Matt Motyl, Ravi Iyer, Sean P. Wojcik, and Peter H. Ditto. Moral foundations theory: The pragmatic validity of moral pluralism. In Patricia Devine and Ashby Plant, editors,Advances in Experimental Social Psychology, volume 47, pages 55–130. Academic Press,
-
[10]
URLhttps://arxiv.org/abs/ 2601.18129. Manuel Pita. Correct codes for the wrong reasons? Validating LLMs as measurement instruments for theoretical constructs.arXiv preprint,
-
[11]
URL https://arxiv.org/ abs/2606.28574
doi: 10.48550/arXiv.2606.28574. URL https://arxiv.org/ abs/2606.28574. Abhinav Rao, Akhila Yerukola, Vishwa Shah, Katharina Reinecke, and Maarten Sap. NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguis...
-
[12]
URL https://arxiv.org/abs/2404.12464
doi: 10.18653/v1/2025.naacl-long.120. URL https://arxiv.org/abs/2404.12464. Steve Rathje, Dan-Mircea Mirea, Ilia Sucholutsky, Raja Marjieh, Claire E. Robertson, and Jay J. Van Bavel. GPT is an effective tool for multilingual psychological text analysis.Proceedings of the National Academy of Sciences, 121(34):e2308950121,
arXiv 2025
-
[13]
doi: 10.1073/pnas.2308950121. 15 Trusting sovereign language modelsA Preprint André da Fonseca Schuck, Gabriel Lino Garcia, João Renato Ribeiro Manesco, Pedro Henrique Paiola, and João Paulo Papa. Evaluating Large Language Models for Brazilian Portuguese Sentiment Analysis: A Comparative Study of Multilingual State-of-the-Art vs. Brazilian Portuguese Fine...
-
[14]
ISSN 1678-4804. doi: 10.5753/jbcs.2025.5793. URLhttps://doi.org/10.5753/jbcs.2025.5793. Afonso Simplício, Gonçalo Vinagre, Miguel Moura Ramos, Diogo Tavares, Rafael Ferreira, Giuseppe Attanasio, Duarte M. Alves, Inês Calvo, Inês Vieira, Rui Guerra, James Furtado, Beatriz Canaverde, Iago Paulo, Vasco Ramos, Diogo Glória-Silva, Miguel Faria, Marcos Treviso,...
arXiv 2025
-
[15]
ISBN 979-8-89176-387-6
Association for Computational Linguistics. ISBN 979-8-89176-387-6. URLhttps://aclanthology.org/2026.propor-1.38/. Petter Törnberg. Large language models outperform expert coders and supervised classifiers at annotating political social media messages.Social Science Computer Review, 43(6):1181–1195,
2026
-
[1955]
FotiosFitsilis, MariaKamilaki, BasilisGatos, VassilisKatsouros, andGeorgeMikros
doi: 10.1037/h0040957. FotiosFitsilis, MariaKamilaki, BasilisGatos, VassilisKatsouros, andGeorgeMikros. TheHellenicParliament’s Approach to Digital Sovereignty for the Era of Artificial Intelligence.International Journal of Parliamentary Studies, 6(1):155–167,
Show all 16 references
-
[1995]
David Eduardo Pereira, Daniela Thuaslar Simão Gomes, and Claudio E
doi: 10.1037/0003-066X.50.9.741. David Eduardo Pereira, Daniela Thuaslar Simão Gomes, and Claudio E. C. Campelo. Evaluating LLMs on Argument Mining Tasks in Brazilian Portuguese Debate Data.Journal of the Brazilian Computer Society, 31(1):1279–1299,
-
[2013]
Suramya Jadhav, Abhay Shanbhag, Amogh Thakurdesai, Ridhima Sinare, and Raviraj Joshi
doi: 10.1016/B978-0-12-407236-7.00002-4. Suramya Jadhav, Abhay Shanbhag, Amogh Thakurdesai, Ridhima Sinare, and Raviraj Joshi. On limitations of llm as annotator for low resource languages. InProceedings of the 8th International Conference on Natural Language and Speech Proces...
2025 doi
-
[2018]
Eurollm-9b: Technical report
Pedro Henrique Martins, João Alves, Patrick Fernandes, Nuno M Guerreiro, Ricardo Rei, Amin Farajian, Mateusz Klimaszewski, Duarte M Alves, José Pombal, Nicolas Boizard, et al. Eurollm-9b: Technical report. arXiv preprint arXiv:2506.04079,
-
[2022]
doi: 10.48550/arXiv.2208.05545. 16
-
[2023]
doi: 10.1037/pspp0000470
ISSN 0022-3514. doi: 10.1037/pspp0000470. URLhttps://doi.org/10.1037/pspp0000470. Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learn...
1901 doi
-
[2024]
Mohammad Atari, Jonathan Haidt, Jesse Graham, Sena Koleva, Sean T
doi: 10.1093/pnasnexus/pgae245. Mohammad Atari, Jonathan Haidt, Jesse Graham, Sena Koleva, Sean T. Stevens, and Morteza Dehghani. Morality beyond the WEIRD: How the nomological network of morality varies across cultures.Journal of Personality and Social Psychology, 125(5):1157–1188,
-
[2025]
doi: 10.5753/jbcs.2025.5824
ISSN 1678-4804. doi: 10.5753/jbcs.2025.5824. URLhttps://doi.org/10.5753/ jbcs.2025.5824. Kunat Pipatanakul and Pittawat Taveekitworachai. Typhoon-S: Minimal Open Post-Training for Sovereign Large Language Models. Arxiv preprint 10.48550/arXiv.2601.18129,
2025
-
[2026]
doi: 10.1163/26668912-bja10127
ISSN 2666-8912. doi: 10.1163/26668912-bja10127. URLhttps://doi.org/ 10.1163/26668912-bja10127. Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. ChatGPT outperforms crowd workers for text- annotation tasks.Proceedings of the National Academy of Sciences, 120(30):e2305016120,
Reviewed August 2, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.