REVIEW 4 major objections 5 minor 26 references
Methods to Assess the UK Government's Current Role as a Data Provider for AI
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper argues that UK government websites measurably improve LLM answers to citizen queries while data.gov.uk datasets are almost never recalled by the tested models.
desk verdict A clearly written technical report whose ablation half mostly works, but whose headline claim about data.gov.uk is undercut by the paper's own failed controls. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is a two-part audit method. The ablation probe uses LLM unlearning: a model is reverse-fine-tuned so that its loss on a target corpus of UK government welfare pages increases, while a Kullback-Leibler divergence term keeps its loss flat on a safe corpus of general encyclopedic text, so the model forgets the government pages without losing language ability; the comparison of error counts before and after quantifies how much those pages contributed. The leakage probe adapts a known prompting framework: the model is asked to complete a statistic in four templates, under zero-, one-, and five-shot prompting, and an instruct-tuned variant is asked the statistic directly, with correct recall treated as evidence that the statistic was in training data. The experiments use small open-weight models and manually coded evaluation of structural and knowledge errors.
What would settle it
Run the same leakage prompts against statistics that model documentation explicitly lists as training data; if recall is near zero for those known statistics, the method cannot separate 'not trained on' from 'not prompted successfully.' A simpler version is to re-ask the control statistics (central bank base rate, national population) in several phrasings with models that failed the controls: persistent failure would show the leakage test is too insensitive to support the paper's negative conclusion about the open data portal.
Extended reading notes
Core claim
The paper's central discovery is a contrast between two UK government data channels. After an unlearning-based ablation makes a model forget roughly twenty government welfare pages, knowledge errors on an 18-question citizen-query test rise by an average of 42.6% while fluency and formatting errors stay flat; the paper reads this as evidence that those web pages were doing real work in the model's answers. The effect is uneven across questions, and the paper finds a negative correlation between how much a query degrades after ablation and how often non-government websites can answer the same query, so government web data matters most where alternative online coverage is scarce. The second result comes from an information-leakage test: across five data.gov.uk datasets and three small models, almost none of the statistics are recalled, despite prompts modelled on known leakage methods, and the paper concludes that the portal's datasets are not part of the tested models' training corpora.
Load-bearing premise
The whole negative claim about data.gov.uk depends on treating a model's failure to volunteer a statistic as proof that the statistic was not in its training data, and the control questions show that some of the same models could not recall even widely known figures such as the central bank base rate.
Editorial extensions
If this is right
- If UK government websites are in LLM training corpora, then the government already influences AI behaviour through its ordinary web content, and changes to that content such as removals, rewrites, or paywalls will shift how models answer citizen queries.
- Because the ablation effect is concentrated on topics with little non-government coverage, improving government prose on benefits interactions and eligibility details is the most direct way to improve LLM performance where it currently fails.
- If data.gov.uk datasets are not being recalled, publishing statistics as spreadsheets alone is not an effective way to get numbers into AI training; the portal's current format and discoverability would need to change.
- The two methods together give any data-holding organisation a reusable way to audit whether its published data and prose are part of AI training mixtures.
Reading between the lines
- Editorial inference: the negative result for data.gov.uk may be a property of the small models and prompt formats tested, not of all LLMs or of retrieval-augmented systems, which could still use open data even if it is not memorised.
- Editorial inference: the leakage method's failure on controls suggests that 'not recalled' should not be read as 'absent from training' without a sensitivity calibration on statistics known to be in the training corpus.
- Editorial inference: a testable extension would be to convert selected data.gov.uk statistics into prose articles and check whether leakage scores rise, isolating format rather than content as the reason the portal is not feeding models.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper introduces two methods for assessing whether UK government data sources contribute to LLM training: (1) an ablation study that applies LLM 'unlearning' to remove a small set of government welfare-related websites from five small open-weight models and measures changes in knowledge errors on 18 citizen queries and one control query, and (2) an information leakage study that prompts the same models to recall statistics from five data.gov.uk datasets, with two external controls (Bank of England base rate and UK population estimate). The authors report that government websites are important data sources for LLMs, with heterogeneity across subject matters, and that data.gov.uk is not a data provider for AI. The paper is framed as a technical report, with code and data released and a non-technical ODI companion report. It also claims broader applicability of the two methods for probing opaque training corpora.
Significance. The paper addresses a genuinely important and timely question—whether and how government data enters proprietary and opaque training corpora—and the two proposed methods are creative. The release of code and data, the use of external controls in the leakage study, and the inclusion of a Google-based prevalence check for the ablation analysis are all positive elements. If the methods were rigorously validated, they could provide organisations with a practical framework for evaluating their data contributions to AI. However, the headline negative conclusion about data.gov.uk is not supported by the evidence as presented: the leakage test is uncalibrated and fails on known-positive controls for two of the three model families, and the paper's own limitations section acknowledges this. The ablation findings, while more plausible, also lack basic statistical grounding. The paper's contribution therefore currently rests on a partially unsupported central claim, though the methods themselves are potentially salvageable.
major comments (4)
- [§3.3, Table 11; KR4] The conclusion that data.gov.uk is not a data provider for AI is not supported by the leakage experiment, because the method's sensitivity is uncalibrated: in Table 11, Gemma 2 2B and Qwen 2.5 3B also fail to recall the Bank of England base rate and the UK population estimate, which are certainly in their pretraining data. The paper itself acknowledges in the limitations (§4) that 'poor performance in controls makes the results of the information leakage study potentially less robust'. A null result in this setting cannot distinguish 'not in the training corpus' from 'present but not extractable by these prompts at this model scale'. The claim also generalizes from five datasets to the whole of data.gov.uk. Please either add positive controls of similar format and obscurity to the data.gov.uk statistics, report sensitivity per model and per prompting condition, and/or weaken KR4 to a claim about the five tested datasets.
- [§2.4.2–2.4.3, Figures 2 and 3] The claim that the ablation causes a clear increase in knowledge errors (KR2 and KR3) rests on a single annotator's coding of 19 hand-picked queries with no inter-rater reliability, no error bars, and no statistical test; the reported 42.6% average increase is not accompanied by variance or any significance measure. The assertion that 'all LLMs were fairly homogeneously affected' is based on visual inspection of Figure 2. Please report per-model and per-query counts with appropriate uncertainty and statistical comparisons, and either provide a second annotator or a clearly defined coding protocol to establish reliability.
- [§2.5, Figure 4] The claim of a 'significant negative correlation' between the effect of ablation and the Google prevalence measure is unsupported: no correlation coefficient, p-value, or test is reported, and the prevalence measure is not precisely specified (search engine, query string, date, region, and the rule for the 'first 10 non-government websites'). Since KR3* is one of the paper's key findings, this analysis needs to be made reproducible and statistically quantified, or the claim should be softened to a descriptive observation.
- [§3.2 and §3.3, Table 10] The paper reports '5 out of 195 tests' in §3.3, but Table 11 shows 7 datasets × 3 models × 4 prompting conditions = 84 cells; please clarify how the 195 count is obtained and whether multiple prompt templates are being aggregated. In addition, the selection of the five data.gov.uk datasets is described as a 'random sample' without giving the sampling frame or procedure, which is needed to support any generalization to the entire portal.
minor comments (5)
- [§2.2.1 and footnotes] There are several typographical errors: 'analagous' should be 'analogous', 'espsecially' should be 'especially', and reference [19] lists 'Stablility.ai' which should be 'Stability AI'.
- [Table 11] The symbols used in Table 11 (!, %, 5, etc.) are visually ambiguous; please use explicit check/cross markers or a clear legend with textual labels, and ensure the table is readable without reference to the surrounding text.
- [§3.3] The sentence 'in 5 out of 195 tests, tested LLMs simply did not recall data points in data.gov.uk' appears to state the opposite of what the results show; presumably 'did' was intended instead of 'did not'.
- [Abstract and §4] The abstract's unqualified statement that 'data.gov.uk is not' a data provider should be tempered to reflect the limitations acknowledged in §4, since the leakage study's sensitivity is not established.
- [§2.4.1, control query] The single control query in the ablation study concerns US welfare; consider adding a UK-related control query outside the target welfare topics to better test whether the unlearning procedure affects general UK knowledge beyond the ablated websites.
Circularity Check
No significant circularity: both experiments compare externally defined targets and controls; the unsupported KR4 is a validity limitation, not a by-construction reduction.
full rationale
The derivation chain is not circular. In the ablation study (Section 2), the target websites, safe dataset, and evaluation queries are specified independently, ground truth is taken from the websites, and the unlearning protocol is adopted from Yao et al. (external work). KR2 follows from pre/post Type 2 error counts; no parameter is fitted to the outcome and no result is defined in terms of its own conclusion. The leakage study (Section 3) adapts an external prompting framework from Wang et al. and tests five data.gov.uk datasets against two external controls (BOE, POP). The negative inference for KR4 assumes that failed recall implies absence from training data; the paper itself concedes this is 'potentially less robust' because Gemma and Qwen fail the controls. That is an uncalibrated-detector validity threat, not circularity: the conclusion is not forced by construction, it is merely under-supported. Self-citations ([1], [2], [4]) are background and policy positioning only, not load-bearing for any experimental result. No uniqueness theorem, hidden ansatz, or renaming of known results appears. Therefore score 0.
Assumptions & free parameters
free parameters (3)
- Unlearning learning rate =
2e-4
- Unlearning steps =
1000
- Unlearning loss weightings =
[0.25, 0, 1]
assumptions (3)
- domain assumption Each target government website was present in CommonCrawl before April 2024 and therefore sits in the training window of Llama 3.1, Gemma 2, and Qwen 2.5, meaning the ablation actually removes data these models were trained on.
- domain assumption The manual qualitative coding framework (Table 1) is a reliable and consistent measure of structural and knowledge errors, with no inter-rater reliability check.
- domain assumption Google search prevalence, measured by how many of the first 10 non-government websites answer each query, is a valid proxy for the prevalence of a topic online.
Cite this review
Pith. "Pith review of Methods to Assess the UK Government's Current Role as a Data Provider for AI." pith.science (2026). https://pith.science/paper/BQ6PUZL4
@misc{pith2026241209632,
author = {Pith},
title = {Pith review of: Methods to Assess the UK Government's Current Role as a Data Provider for AI},
year = {2026},
howpublished = {\url{https://pith.science/paper/BQ6PUZL4}},
note = {Machine review of arXiv:2412.09632}
}
abstract
Governments typically collect and steward a vast amount of high-quality data on their citizens and institutions, and the UK government is exploring how it can better publish and provision this data to the benefit of the AI landscape. However, the compositions of generative AI training corpora remain closely guarded secrets, making the planning of data sharing initiatives difficult. To address this, we devise two methods to assess UK government data usage for the training of Large Language Models (LLMs) and 'peek behind the curtain' in order to observe the UK government's current contributions as a data provider for AI. The first method, an ablation study that utilises LLM 'unlearning', seeks to examine the importance of the information held on UK government websites for LLMs and their performance in citizen query tasks. The second method, an information leakage study, seeks to ascertain whether LLMs are aware of the information held in the datasets published on the UK government's open data initiative data$.$gov$.$uk. Our findings indicate that UK government websites are important data sources for AI (heterogenously across subject matters) while data$.$gov$.$uk is not. This paper serves as a technical report, explaining in-depth the designs, mechanics, and limitations of the above experiments. It is accompanied by a complementary non-technical report on the ODI website in which we summarise the experiments and key findings, interpret them, and build a set of actionable recommendations for the UK government to take forward as it seeks to design AI policy. While we focus on UK open government data, we believe that the methods introduced in this paper present a reproducible approach to tackle the opaqueness of AI training corpora and provide organisations a framework to evaluate and maximize their contributions to AI development.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
We Must Fix the Lack of Transparency Around the Data Used to Train Foundation Models
Jack Hardinges, Elena Simperl, and Nigel Shadbolt. We Must Fix the Lack of Transparency Around the Data Used to Train Foundation Models. Harvard Data Science Review , (Special Issue 5), May 2024. URL https://hdsr.mitpress.mit.edu/pub/xau9dza3
work page 2024
-
[2]
Building a better future with data and AI: a white paper, 2024
The Open Data Institute (ODI). Building a better future with data and AI: a white paper, 2024. URL https:// theodi.org/insights/reports/building-a-better-future-with-data-and-ai-a-white-paper/ . [Accessed 21-11-2024]
work page 2024
-
[3]
The role of government as a provider of data for AI, 2024
The Global Partnership for AI (GPAI). The role of government as a provider of data for AI, 2024. URL https://gpai.ai/projects/data-governance/theroleofgovernmentasaproviderofdataforai/ role-of-government-as-a-provider-of-data-for-AI-phase-1-full-report.pdf . [Accessed 21-11-2024]
work page 2024
-
[4]
The ODI’s input to the AI Action Plan: an AI-ready Na- tional Data Library, 2024
The Open Data Institute (ODI). The ODI’s input to the AI Action Plan: an AI-ready Na- tional Data Library, 2024. URL https://theodi.org/news-and-events/consultation-responses/ the-odis-input-to-the-ai-action-plan-an-ai-ready-national-data-library/ . [Accessed 21- 11-2024]
work page 2024
-
[5]
Application of large language model in intelligent Q&A of digital government
Shangsheng Gao, Li Gao, Qi Li, and Jianjun Xu. Application of large language model in intelligent Q&A of digital government. In Proceedings of the 2023 2nd International Conference on Networks, Communications and Infor- mation Technology, CNCIT ’23, page 24–27, New York, NY , USA, 2023. Association for Computing Machinery. ISBN 9798400700620. doi:10.1145/...
-
[6]
Understanding Black-box Predictions via Influence Functions, 2020
Pang Wei Koh and Percy Liang. Understanding Black-box Predictions via Influence Functions, 2020. URL https://arxiv.org/abs/1703.04730
arXiv 2020
-
[7]
Roger Grosse, Juhan Bae, Cem Anil, Nelson Elhage, Alex Tamkin, Amirhossein Tajdini, Benoit Steiner, Dustin Li, Esin Durmus, Ethan Perez, Evan Hubinger, Kamil˙e Lukoši¯ut˙e, Karina Nguyen, Nicholas Joseph, Sam McCandlish, Jared Kaplan, and Samuel R. Bowman. Studying Large Language Model Generalization with Influence Functions,
-
[8]
What is your data worth to gpt? llm-scale data valuation with influence functions, 2024
Sang Keun Choe, Hwijeen Ahn, Juhan Bae, Kewen Zhao, Minsoo Kang, Youngseog Chung, Adithya Pratapa, Willie Neiswanger, Emma Strubell, Teruko Mitamura, Jeff Schneider, Eduard Hovy, Roger Grosse, and Eric Xing. What is your data worth to gpt? llm-scale data valuation with influence functions, 2024. URL https: //arxiv.org/abs/2405.13954
arXiv 2024
Show all 26 references
-
[9]
Large Language Model Unlearning, 2024
Yuanshun Yao, Xiaojun Xu, and Yang Liu. Large Language Model Unlearning, 2024. URL https://arxiv. org/abs/2310.10683
2024 arXiv
-
[10]
Eraser: Jailbreaking Defense in Large Language Models via Unlearning Harmful Knowledge, 2024
Weikai Lu, Ziqian Zeng, Jianwei Wang, Zhengdong Lu, Zelin Chen, Huiping Zhuang, and Cen Chen. Eraser: Jailbreaking Defense in Large Language Models via Unlearning Harmful Knowledge, 2024. URL https: //arxiv.org/abs/2404.05880
2024 arXiv
-
[11]
Exact and Efficient Unlearning for Large Language Model-based Recommendation, 2024
Zhiyu Hu, Yang Zhang, Minghao Xiao, Wenjie Wang, Fuli Feng, and Xiangnan He. Exact and Efficient Unlearning for Large Language Model-based Recommendation, 2024. URL https://arxiv.org/abs/2404.10327
2024 arXiv
-
[12]
Rényi Divergence and Kullback-Leibler Divergence
Tim van Erven and Peter Harremoes. Rényi Divergence and Kullback-Leibler Divergence. IEEE Transactions on Information Theory, 60(7):3797–3820, July 2014. ISSN 1557-9654. doi:10.1109/tit.2014.2320500. URL http://dx.doi.org/10.1109/TIT.2014.2320500
2014
-
[13]
Perplexed: Understanding when large language models are confused, 2024
Nathan Cooper and Torsten Scholak. Perplexed: Understanding when large language models are confused, 2024. URL https://arxiv.org/abs/2404.06634
2024 arXiv
-
[14]
Dated Data: Tracing Knowledge Cutoffs in Large Language Models, 2024
Jeffrey Cheng, Marc Marone, Orion Weller, Dawn Lawrie, Daniel Khashabi, and Benjamin Van Durme. Dated Data: Tracing Knowledge Cutoffs in Large Language Models, 2024. URL https://arxiv.org/abs/2403.12958
2024 arXiv
-
[15]
The Llama 3 Herd of Models, 2024
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The Llama 3 Herd of Models, 2024. URL https://arxiv.org/abs/2407.21783
2024 arXiv
-
[16]
Gemma 2: Improving Open Language Models at a Practical Size, 2024
Gemma Team. Gemma 2: Improving Open Language Models at a Practical Size, 2024. URL https://arxiv. org/abs/2408.00118
2024 arXiv
-
[17]
Qwen2.5, 2024
QwenLM. Qwen2.5, 2024. URL https://github.com/QwenLM/Qwen2.5
2024
-
[18]
Microsoft Corporation, 2023
The New york Times Company v. Microsoft Corporation, 2023. URL https://www.courtlistener.com/ docket/68117049/the-new-york-times-company-v-microsoft-corporation/
2023
-
[19]
Getty Images (US), Inc. v. Stability AI, Inc., 2023. URL https://www.courtlistener.com/docket/ 66788385/getty-images-us-inc-v-stability-ai-inc/ . 16 Methods to assess the UK government’s current role as a data provider for AI - ODI
2023
-
[20]
Extracting Training Data from Diffusion Models, 2023
Nicholas Carlini, Jamie Hayes, Milad Nasr, Matthew Jagielski, Vikash Sehwag, Florian Tramèr, Borja Balle, Daphne Ippolito, and Eric Wallace. Extracting Training Data from Diffusion Models, 2023. URL https: //arxiv.org/abs/2301.13188
2023 arXiv
-
[21]
Truong, Simran Arora, Mantas Mazeika, Dan Hendrycks, Zinan Lin, Yu Cheng, Sanmi Koyejo, Dawn Song, and Bo Li
Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ri- tik Dutta, Rylan Schaeffer, Sang T. Truong, Simran Arora, Mantas Mazeika, Dan Hendrycks, Zinan Lin, Yu Cheng, Sanmi Koyejo, Dawn Song, and Bo Li. DecodingTrust: A Compreh...
2024 arXiv
-
[22]
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and...
2022 arXiv
-
[23]
Reinforcement Learning from Human Feedback: Progress and Challenges
John Schulman. Reinforcement Learning from Human Feedback: Progress and Challenges. https://eecs. berkeley.edu/research/colloquium/230419-2/, 2023. [Accessed 21-11-2024]
2023
-
[24]
Performance Law of Large Language Models, 2024
Chuhan Wu and Ruiming Tang. Performance Law of Large Language Models, 2024. URL https://arxiv. org/abs/2408.09895
2024 arXiv
-
[25]
A comprehensive evaluation of quantization strategies for large language models, 2024
Renren Jin, Jiangcun Du, Wuwei Huang, Wei Liu, Jian Luan, Bin Wang, and Deyi Xiong. A comprehensive evaluation of quantization strategies for large language models, 2024. URL https://arxiv.org/abs/2402. 16775. 17
2024
-
[2023]
URL https://arxiv.org/abs/2308.03296
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.