{"id":"a167576f-2050-4178-9820-95fcd922a233","arxiv_id":"2508.02271","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Danish Dynaword packages 4.8B tokens of openly licensed Danish text into a continuously versioned, test-gated corpus that improves language-model perplexity compared with Danish Gigaword.","lead":"Danish Dynaword is a large, openly licensed Danish text corpus built to keep growing through community contributions instead of being released once. It contains about 4.8 billion tokens, several times more than earlier open Danish corpora, and pilot language-model tests show perplexity gains over Danish Gigaword.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'exclusively openly licensed' claim is contradicted by Table 4 itself: retsinformation.dk and Domsdatabasen.dk (~904M tokens, ~19%) are labeled 'Copyright Law', not an open license, and OpenSubtitles' CC-0 label conflicts with Section 3 and the ethical note.","rationale":"The reader's weakest assumption correctly identifies the license classification in Table 4 as load-bearing. I agree, and the manuscript itself provides direct evidence that this assumption is insecure: two rows are labeled 'Copyright Law', which is not an open license and is not traceable under the paper's own standard, and the OpenSubtitles entry conflicts with the exclusion described in Section 3 and with the ethical note about seemingly open data containing copyrighted content. Because the 'exclusively openly licensed' claim is the foundation for the paper's main contribution, the paper should not be accepted as-is. However, this is a fixable documentation problem: the authors could provide legal verification for the 'Copyright Law' rows, clarify the OpenSubtitles provenance, and re-state the open-license fraction accordingly. The dataset itself may still be valuable, and the paper's reproducible scripts, versioned releases, and community-contribution process are genuine strengths. Therefore the correct disposition is CONDITIONAL acceptance on a license audit, matching the reader's verdict rather than changing it. If the audit shows that a substantial share of tokens is not openly licensed and cannot be reclassified, the central claim would fail and the appropriate verdict would move to REJECT.","tokens_in":12036,"tokens_out":6133,"duration_ms":71417,"concrete_test":"Run a source-by-source license audit: for each row in Table 4, open the corresponding datasheet and license file in the repository and check whether it grants reuse, resharing, and modification. Specifically: (1) retrieve the license terms for retsinformation.dk and Domsdatabasen.dk; if 'Copyright Law' does not contain an explicit open grant, relabel or reclassify those rows and compute the token share that is not openly licensed. (2) Inspect the OpenSubtitles processing script to confirm that the excluded copyrighted samples are absent from the released files and that the remaining subtitle text carries a valid CC-0 grant; if not, remove or relabel the row. (3) Recompute the 'exclusively openly licensed' statement against the corrected table. If the non-open share is nonzero, the claim must be revised to state the fraction of openly licensed tokens.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central resource claim is that Danish Dynaword is 'exclusively openly licensed' and therefore the largest openly licensed Danish corpus. The load-bearing condition is that every row in Table 4 corresponds to a license that permits resharing, reuse, and modification, and that no included text falls outside those licenses. That condition is not established, and the paper's own text undermines it. First, Table 4 lists retsinformation.dk (818.25M tokens) and Domsdatabasen.dk (86.35M tokens) with license 'Copyright Law'. 'Copyright Law' is not an open license; at best it names a legal regime, and it does not by itself grant reuse, resharing, or modification rights. These two rows alone are roughly 904M tokens, about 19% of the corpus, so the 'exclusively openly licensed' label cannot be read literally. If the intended meaning is that Danish law places these public documents outside copyright, the paper must say so and provide the legal basis; a vague label is exactly what the paper's own 'traceable license' principle warns against. Second, Section 3 states that 'copyrighted samples from OpenSubtitles (<1M tokens)' were excluded, yet Table 4 includes OpenSubtitles at 271.60M tokens labeled CC-0. The Ethical considerations section concedes that seemingly openly-licensed datasets may contain copyrighted content and describes a 'notable instance' involving OpenSubtitles. This leaves unresolved whether the CC-0 label actually applies to the subtitle text in the released version. Third, the Ethical considerations paragraph is an internal admission that the review process can miss copyrighted content, so the claim of exclusivity is stronger than the evidence supports. For the central claim to hold, Table 4's labels must be verified against source license documents, especially the 'Copyright Law' rows and OpenSubtitles.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces Dynaword, a framework for continuously maintained, openly licensed language corpora, and presents Danish Dynaword, a 4.8B-token Danish corpus assembled from legal, social-media, spoken, web, medical, encyclopedic, literary, news, and dialect sources. The authors report that Danish Dynaword contains over four times as many tokens as comparable Danish releases, is exclusively openly licensed, and has received contributions from industry, government, and research. They evaluate the corpus by continually pre-training and training from scratch a Gemma-3-1B model on Danish Dynaword and Danish Gigaword, reporting average relative perplexity improvements of 5.9% for continual pre-training and 26% for from-scratch training on the full corpus, with additional downstream EuroEval results.","tokens_in":12312,"tokens_out":6933,"duration_ms":81845,"significance":"If the licensing and reproducibility properties are as claimed, this is a valuable community resource and a useful template for future low-resource language corpora. The work is unusually concrete in addressing legal and ethical concerns around web-scale training data: the corpus is versioned, collection scripts are published, datasheets are promised for each source, and the repository includes lightweight tests. The language-modeling comparison is also well designed in several respects: held-out validation sources, external 2025 texts, a size-matched control, and publicly released training code. The perplexity improvements are consistent enough to support the qualitative conclusion that Danish Dynaword is at least as good as Danish Gigaword for language modeling. The main weakness is that the central 'exclusively openly licensed' claim is not fully documented in the manuscript, and the license table itself contains entries that do not appear to be open licenses.","major_comments":[{"comment":"The claim that Danish Dynaword is 'exclusively openly licensed' is not established by the evidence in Table 4. The rows for retsinformation.dk (818.25M tokens) and Domsdatabasen.dk (86.35M tokens) list 'Copyright Law' as the license; this is the name of a legal regime, not a license that grants resharing, reuse, or modification, and no Danish-law analysis is provided to show that these official texts can be freely redistributed. These two rows represent roughly 19% of the 4.80B-token total, so the headline claim cannot be read literally. Please replace these labels with a precise legal basis (for example, a specific public-sector-information exception) and cite the relevant legal provisions, or revise the claim. The same concern applies to the labels 'Gutenberg' and 'DanNet 1.0', which are not standard open-license identifiers and are not traceable under the paper's own licensing principle.","section":"Abstract, §2.2, Table 4"},{"comment":"There is a direct inconsistency between Section 3, which states that 'copyrighted samples from OpenSubtitles (<1M tokens)' were excluded, and Table 4, which lists OpenSubtitles as a 271.60M-token source with license CC-0. The Ethical considerations section further says that a 'notable instance' of copyrighted content occurred with the initial release of OpenSubtitles as part of Danish Gigaword. The paper must clarify exactly which OpenSubtitles content is included in Danish Dynaword, how the CC-0 status of the included subset was verified, and whether all included subtitle text is openly licensed. Without this clarification, the reader cannot determine whether the 'exclusively openly licensed' property holds.","section":"§3, Table 4, Ethical considerations"}],"minor_comments":[{"comment":"The sentence 'which is we eloborate on in the following section' contains a typo, and several other grammatical slips appear (e.g., 'It is by no mean uncommon' and 'it is therefore encourages that model developers exclude evaluation data'). These should be corrected.","section":"§4.1, Appendix B"},{"comment":"The unit 'Llama 3 tokens' is used without specifying the exact tokenizer version or counting script; please provide a pointer to the tokenizer and the script used for token counting.","section":"Table 4, Figure 2"},{"comment":"The downstream EuroEval results are mostly within one standard error of the baseline; the statement that Danish Dynaword 'yields gains on 7 out of 9 tasks' should be phrased as directional improvements rather than confirmed gains.","section":"Tables 6 and 7"},{"comment":"There is a typo in 'OCR'ed Newwspapers from NCC', and the license column would be easier to audit if every license name included a URL or a stable identifier pointing to the full license text.","section":"Table 4"},{"comment":"The mechanism for marking benchmark-containing datasets is described only briefly; please add a pointer to the repository documentation that shows how the marking is implemented and maintained.","section":"§3.2"},{"comment":"The sentence stating that the repository 'includes light-weight tests to ensure data formatting, quality, and documentation' would be more useful if it named the specific tests and the CI system that runs them.","section":"Introduction"}],"recommendation":"major_revision","confidential_remarks":"This is a useful resource and the central language-modeling comparison is reasonably careful. The main obstacle is the licensing documentation: the 'Copyright Law' labels in Table 4 and the OpenSubtitles inconsistency directly affect the paper's most prominent claim. I believe these issues are fixable if the authors have the underlying legal analysis and can document it clearly, so I would not reject. I would ask the editor to ensure that the licensing review is not simply asserted but traceable to the datasheets and to the legal basis for each row."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look. Danish Dynaword is a real resource: 4.8B tokens of Danish text, openly released, with a versioned repository, CI tests, datasheets, and a documented contribution process. That operational package—test-gated community contributions, reproducible collection scripts, traceable license notes—is the genuinely new part. The perplexity experiments are also reasonable for a resource paper: consistent gains over Danish Gigaword on held-out and contemporary 2025 data, and downstream EuroEval improvements on most tasks. I'd trust the direction of those results.\n\nThe soft spot is exactly what the stress-test note says. The central claim 'exclusively openly licensed' is not supported by the paper's own Table 4. Two large sources totaling ~904M tokens (about 19% of the corpus) are labeled 'Copyright Law,' which is a legal regime, not an open license. It may well be that Danish law places these official legal texts in the public domain, but the paper needs to say so and point to the legal basis. Right now the label conflicts with the paper's own 'traceable license' principle. And the OpenSubtitles row is internally inconsistent: Section 3 says copyrighted samples were excluded, Table 4 lists OpenSubtitles as CC-0 at 271M tokens, and the Ethical considerations paragraph mentions a 'notable instance' with OpenSubtitles without resolving it. If those labels are wrong, the 'largest openly licensed Danish corpus' headline fails.\n\nMinor points: the perplexity results are single-run, no variance, and the baseline is only Danish Gigaword; that is fine for a first comparison but makes the 26% number less precise than it looks. The paper's own limitations section already concedes bias toward legal documents and limited social media, which is honest and useful.\n\nBottom line: the resource is real and the framework is worth adopting, but the license documentation needs to be fixed and verified before the strongest claims are taken at face value. I'd send this to review, with the expectation of revisions on Table 4 and license traceability. The authors have clearly done serious work; the gaps are in the write-up, not the underlying effort.","headline":"A genuinely useful Danish dataset and a real community-maintenance framework, but the 'exclusively openly licensed' claim is contradicted by the paper's own Table 4 and needs fixing before the headline claim is credible.","tokens_in":13015,"tokens_out":2117,"would_cite":true,"duration_ms":22257,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Danish Dynaword offers an openly licensed, continuously updated Danish corpus that beats Danish Gigaword for language modeling.","keywords":["dynaword","Danish language corpus","openly licensed data","dataset licensing","continuous dataset development","language model pre-training","perplexity","community contributions"],"falsifier":"An independent audit of the license documents for the 'Copyright Law' rows and OpenSubtitles that finds one of them restricts resharing or modification would falsify the 'exclusively openly licensed' claim; alternatively, re-running the reported training regime on the released v1.2.7 data and failing to reproduce the 5.9% and 26% perplexity improvements would falsify the performance claim.","tokens_in":11828,"feed_emoji":"🇩🇰","tokens_out":5787,"duration_ms":63418,"temperature":0.7,"pith_summary":"The paper argues that pre-training corpora should be living, community-maintained resources rather than one-shot releases, and it offers a four-principle recipe for building them: traceable and open licensing, reproducible collection, documentation, and extensibility. It then presents Danish Dynaword as a working implementation of that recipe. The corpus contains roughly 4.8 billion tokens from openly licensed Danish sources, over four times the public segment of Danish Gigaword, and it grew through contributions from industry, government, and research. The paper's performance claim is that models trained on Danish Dynaword beat models trained on Danish Gigaword in language modeling perplexity, by 5.9% with continual pre-training and 26% when training from scratch. If this holds, Danish Dynaword is both the largest openly licensed Danish corpus and evidence that community-maintained data can be a practical foundation for language modeling.","feed_headline":"Danish Dynaword quadruples open Danish training data","feed_subtitle":"Openly licensed, continuously updated corpus improves Danish language modeling and keeps growing.","key_machinery":"The carrying mechanism is the dynaword approach itself: a four-principle framework (traceable and open licensing, reproducibility, documentation, and extensibility) enforced by per-source datasheets, reproducible collection scripts, versioned releases, lightweight tests for format, quality, and documentation, and a maintenance pipeline for accepting community contributions. The framework turns dataset curation from a one-shot release into a continuous process, and Danish Dynaword is the testbed demonstrating that the process produces a corpus with the claimed size and license properties while improving language modeling.","core_discovery":"The paper's central discovery is that a pre-training corpus built on the dynaword principles can be both large and legally clean: Danish Dynaword v1.2.7 contains roughly 4.8 billion tokens, more than four times the public segment of Danish Gigaword, and is assembled exclusively from openly licensed sources with traceable license documentation. As evidence that the corpus is not only bigger but better, the paper reports that 1-billion-parameter language models continually pre-trained on Danish Dynaword improve perplexity over Danish Gigaword by 5.9% on average, and models trained from scratch improve by 26%; even a size-matched Dynaword subset beats Gigaword by 2.6% in continual pre-training and 18% when training from scratch. The paper also reports gains on seven of nine Danish downstream tasks after continual pre-training. The claim is therefore that a continuously developed, openly licensed dataset can serve as a sustainable foundation for language modeling.","pith_inferences":["Editorial inference: the 'exclusively openly licensed' guarantee is only as strong as the legal review behind each source, and because license interpretations vary across jurisdictions, the guarantee should be treated as a claim about documentation and intent rather than absolute legal certainty.","Editorial inference: if the corpus grows at the pace shown in the paper's timeline, it may gradually close part of the gap with Common-Crawl-based Danish data; a natural test is whether continued growth yields continued perplexity gains or diminishing returns.","Editorial inference: the same pipeline could transfer to other low- and mid-resource languages, but the likely bottleneck will be finding enough contributing institutions and individuals, not the technical tooling.","Editorial inference: because the authors mark and exclude evaluation data, downstream users can train models with reduced risk of accidental test-set contamination, a feature that could become more valuable as benchmark overlap grows."],"forward_implications":["Danish Dynaword is currently the largest openly licensed Danish corpus, at roughly 4.8 billion tokens, more than four times the size of comparable releases.","Models trained on it achieve lower perplexity than on Danish Gigaword: 5.9% average improvement in continual pre-training and 26% when training from scratch, with gains persisting even when dataset size is matched.","Continual pre-training on Danish Dynaword improves performance on seven of nine Danish downstream tasks, a sign the corpus transfers beyond language modeling.","The dynaword framework gives other languages and domains a documented blueprint for building openly licensed, continually updated training data.","Versioned releases with a changelog make license and content changes transparent, which lowers the legal and ethical risk of downstream models."],"supporting_citations":[{"why":"Defines the three tiers of dataset openness and the legal risks that motivate the dynaword licensing requirements.","marker":"Baack et al. (2025)"},{"why":"Provides Danish Gigaword, the baseline corpus whose public segments Danish Dynaword extends and outperforms.","marker":"Derczynski et al. (2021)"},{"why":"Supplies the datasheets methodology used to document each source in the corpus.","marker":"Gebru et al. (2021)"},{"why":"Represents the Common-Crawl-based iterative release approach that dynaword compares against in Table 1.","marker":"Penedo et al. (2024)"},{"why":"Provides Common Corpus, the large openly licensed corpus whose Danish subset is a direct comparison point.","marker":"Langlais et al. (2025)"},{"why":"Provides the Danish evaluation suite used for the downstream task results.","marker":"Nielsen (2023)"},{"why":"Serves as the model of a continually developed, community-supported language resource.","marker":"Nivre et al. (2016)"},{"why":"Motivates the dataset-poisoning and review-quality concerns discussed in the limitations.","marker":"Goldblum et al. (2022)"}],"fun_headline_variants":["Danish Dynaword quadruples open Danish corpus to 4.8B tokens","Community-built Danish Dynaword: 4x data, better models","Openly licensed Danish Dynaword quadruples corpus size","Danish Dynaword: 4x tokens, 5.9% perplexity gain"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole 'exclusively openly licensed' claim rests on the license review being correct for every source in Table 4, including the rows marked 'Copyright Law' and OpenSubtitles; if any of those labels is wrong, the corpus is not exclusively open.","fun_headline_variants_meta":{"raw":{"variants":["Danish Dynaword quadruples open Danish corpus to 4.8B tokens","Community-built Danish Dynaword: 4x data, better models","Openly licensed Danish Dynaword quadruples corpus size","Danish Dynaword: 4x tokens, 5.9% perplexity gain"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001087,"raw_usage":{"total_tokens":4518,"prompt_tokens":897,"completion_tokens":3621,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":513,"completion_tokens_details":{"reasoning_tokens":3539}},"tokens_in":513,"tokens_out":3621,"duration_ms":30783,"temperature":1.0,"reasoning_tokens":3539,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T05:02:25.238930+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An independent audit of the license documents for the 'Copyright Law' rows and OpenSubtitles that finds one of them restricts resharing or modification would falsify the 'exclusively openly licensed' claim; alternatively, re-running the reported training regime on the released v1.2.7 data and failing to reproduce the 5.9% and 26% perplexity improvements would falsify the performance claim.","supporting_citations":[],"review_version":1}