{"id":"1b244b13-910c-49ac-b7ba-7184ec43f2ec","arxiv_id":"2508.10246","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A corpus study of Toki Pona finds diachronic shifts in word classes and register-based variation, indicating constructed languages undergo sociolinguistically driven change.","lead":"This study analyzes millions of Toki Pona sentences from chat and published texts to track how word usage changes over time. It finds that body-part words are increasingly used as verbs and that informal and formal registers differ, suggesting constructed languages evolve like natural ones.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unvalidated parser/heuristic pipeline may manufacture POS trends; a gold-standard annotation test is needed.","rationale":"The reader identified the unvalidated parser and heuristic scoring algorithm as the weakest assumption, and I agree. This is indeed the most load-bearing concern: every quantitative trend in the paper depends on the POS tags produced by the pipeline, and the heuristic's explicit preference for observed usage patterns creates circularity risk. The paper itself does not report validation against any gold-standard parse or tag set. My proposed check is a concrete gold-standard annotation study that would determine whether the trends are artifacts. Since the reader's verdict is already CONDITIONAL and this concern does not move it, the verdict should remain UNCHANGED.","tokens_in":570,"tokens_out":2469,"duration_ms":57633,"concrete_test":"Create a gold standard by having two fluent Toki Pona speakers independently annotate POS tags (following Table 2) on a stratified random sample of 1,000 sentences (e.g., 100 per year/corpus) drawn from the same filtered corpora. Compute the parser's precision/recall/F1 for each tag, especially TVERB, NOUN, MOD, and INTJ. Then recompute the reported proportions—luka/uta as transitive verbs, pu as noun vs. modifier, and pona/wawa as interjections—using the gold annotations (via direct substitution or correction factors) and check whether the trends in Figures 3, 4, 5, and 6 survive within 95% confidence intervals.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central empirical claims are frequency trends in POS tags (Figs. 3–7). These counts come from a custom Earley parser and a heuristic ambiguity-resolution scoring algorithm (§3.3–§3.4). No gold-standard validation is reported. The heuristic is explicitly designed to prefer interpretations that 'reflect observed usage patterns' (§3.4), so the counts are not purely data-driven: they encode the author's prior about what Toki Pona usage should look like. If the parser or heuristic is systematically biased—especially if errors correlate with word, year, or formality—then the claimed increases in luka/uta as transitive verbs, the decrease of pu as a standalone noun, and the interjection register differences could be artifacts. The specific trends are not internally inconsistent, but the pipeline's unvalidated status is the weakest load-bearing point.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper applies a corpus-based, computational pipeline to study diachronic and register variation in the constructed language Toki Pona. Using a Discord corpus (6.39M tokens) and the poki Lapo literary corpus, the authors filter with sona-toki, parse with a custom Earley-grammar implementation, resolve ambiguity with a hand-written heuristic scorer, and tag parts of speech. They report that luka and uta increasingly occur as transitive verbs in informal data (Fig. 3), that pu is decreasingly used as a standalone noun (Fig. 4), and that interjection use differs across registers (Figs. 5–6), concluding that Toki Pona changes like natural languages under sociolinguistic pressures.","tokens_in":7574,"tokens_out":4469,"duration_ms":46381,"significance":"If the reported trends are real, the study is a valuable empirical demonstration of language change in a constructed language, with a publicly available pipeline and corpora. Its strengths are the scale of the informal corpus, the explicit grammar, public code, and the comparison of two registers. However, the central quantitative claims rest on a parser and scoring heuristic that have not been validated against gold-standard annotations, and the analysis lacks inferential statistics; as a result, the findings are currently suggestive rather than established.","major_comments":[{"comment":"The load-bearing assumption is that the Earley parser and heuristic scorer produce accurate POS counts. §3.4 says the scorer 'favors specific patterns' and that this 'prioritization ... reflects observed usage patterns'; this means the counts encode the author's prior, and no gold-standard evaluation or inter-annotator agreement is reported. If parse errors correlate with word, year, or formality, every trend in Figs. 3–7 could be an artifact. Report precision/recall per tag on a stratified annotated sample, or remove/redraft the causal and general claims.","section":"§3.3–§3.4, Figs. 3–7"},{"comment":"Words were chosen because they 'showed a high degree of change' and were 'referenced often' (quote), introducing selection bias. No significance tests, confidence intervals, or raw token counts are given; percentages normalize but do not show magnitude. Because the survey window is 2020–2025 and the sample is one community, the claim of general language change is not supported. Provide counts and a regression model (e.g., year as continuous predictor) and report uncertainty.","section":"§4.1, Figs. 3–4"},{"comment":"The claim that adoption is 'not delayed' in the formal corpus rests on visual inspection of a single line plot. The figure has no error bars or fitted trajectories. Quantify the onset or slope difference between corpora (e.g., interaction of year and corpus in a model), or phrase the result as descriptive.","section":"§4.2.2, Fig. 7"},{"comment":"The conclusion extends the results to 'proper nouns may be used more as adjectives,' but the analysis only addresses pu as a modifier versus standalone noun; pu is not a proper name. The cognitive/sociolinguistic explanations (processing efficiency, identity) are plausible but not tested by the corpus data. Separate observational trends from explanatory hypotheses.","section":"§5"}],"minor_comments":[{"comment":"Axis labels do not state units or whether the plotted value is a proportion of tokens, sentences, or messages. Include raw counts or a clear denominator; error bars or confidence bands would also help.","section":"Figures 3–7"},{"comment":"The parse tree in Figure 1 appears inconsistent with the grammar in Table 1: the tree shows 'li' as part of a content-word constituent rather than as a daughter of S between the subject and predicate. Please correct the figure or clarify the tree notation.","section":"Figure 1 / Table 1"},{"comment":"Typographical error: 'T oki Pona' appears with an extra space in the Introduction and in the reference for Lang (2014).","section":"Introduction / References"},{"comment":"The heuristic scoring algorithm is described only informally and via an external repository. Include a pseudocode or formal description in the paper, since this is the step that most affects the reported counts.","section":"§3.4"},{"comment":"The sona-toki filtering configuration is said to be 'the same configuration' used by an external tool, but no specific parameters are listed. Provide the configuration or a pointer to the exact code version so the filtering step is reproducible.","section":"§3.2"},{"comment":"The phrase 'proper nouns may be used more as adjectives' should be 'proper names may be used more as modifiers,' since Toki Pona names are not a distinct noun class and the paper's own analysis concerns pu, an adjective.","section":"§5"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a modest, clearly written corpus study that reports genuinely new things about Toki Pona usage — body-part nouns sliding into transitive-verb slots, pu losing its standalone-noun use, and interjection use differing across the Discord and long-form corpora. The grammar is grounded in Lang (2014), the code and aggregated data are public, and the authors are upfront about the short time window and the written-only data. That is real credit: the work is reproducible in a way much corpus sociolinguistics is not.\n\nThe main soft spot is exactly where your stress-test note lands. The POS counts come from a custom Earley parser plus a heuristic ambiguity-resolution step that explicitly prefers parses \"reflecting observed usage patterns.\" No gold-standard evaluation of either component is reported. If the heuristic biases toward or against certain readings in ways that correlate with year or formality, the trends in Figures 3–7 could be partly manufactured. The authors do not hide the heuristic, and the grammar itself is a simple, publicly inspectable CFG, but until the pipeline is validated against a manually annotated sample, the specific magnitudes and even some directions should be treated as suggestive rather than established. I'd also like token-count baselines and some significance testing for the diachronic claims; the post hoc word selection in §4.1 makes the absence of statistics more consequential. The pu-as-proper-name explanation is interesting but speculative.\n\nThat said, the central point — that a constructed language with a small, deliberately constrained grammar still shows sociolinguistically patterned change — does not rest entirely on the parser. The examples in the paper (e.g., luka/uta used as transitive verbs, the lipu pu construction) are real and community-discussed, and the cross-corpus comparison is a sensible design. The paper knows its limits and does not overclaim.\n\nWho it is for: people working on language change, usage-based theory, and conlang communities. It is a solid subfield contribution, not a field reorientation. I'd send it to peer review asking for parser validation and a statistical pass; a gold-standard sample of a few hundred sentences should be feasible given the small lexicon and simple grammar. With that, the empirical claims would be much firmer.","headline":"A worthwhile first corpus study of Toki Pona with new empirical findings, but the unvalidated parser/heuristic means the trend lines need a gold-standard check before they carry the argument.","tokens_in":7986,"tokens_out":2851,"would_cite":false,"duration_ms":31044,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that Toki Pona, a constructed language of about 120 words, shows the same kinds of sociolinguistic change and variation as natural languages when a real community uses it.","keywords":["Toki Pona","constructed language","language change","sociolinguistic variation","corpus linguistics","part-of-speech tagging","fluid word classes","transitivity"],"falsifier":"Hand-annotate a random sample of sentences from each year of both corpora using the paper's tag set, then compare the parser's transitive-verb rates for luka and uta with the hand labels; if the parser systematically over- or under-tags these body-part verbs, the upward trend in Figure 3 will not reproduce under gold-standard annotation.","tokens_in":7274,"feed_emoji":"📈","tokens_out":7250,"duration_ms":72946,"temperature":0.7,"pith_summary":"This paper asks whether a deliberately designed language with a tiny vocabulary, Toki Pona, undergoes the same kinds of change and variation as natural languages when a real community uses it. Using a custom Earley parser and part-of-speech tagger over two corpora, an informal Discord server with millions of sentences and a formal corpus of published works, the authors track how content words shift among syntactic positions over time and across registers. They find that body-part nouns such as luka 'hand' and uta 'mouth' are increasingly used as transitive verbs in casual use, that pu is losing ground as a standalone noun in favor of lipu pu, and that interjections such as wawa 'strong' have surged in conversational text but not in formal writing. The conclusion is that even a constructed language evolves through community use, shaped by processing efficiency, conformity to existing patterns, and identity signaling.","feed_headline":"Toki Pona changes under real use","feed_subtitle":"Corpus data shows body-part nouns like luka are becoming transitive verbs in casual chat.","key_machinery":"The analysis rests on an Earley context-free parser that converts Toki Pona sentences into hierarchical phrase structures, making syntactic positions such as subject, predicate, and direct object explicit. A heuristic scoring module then resolves ambiguity, for example whether tawa is a preposition or a content verb, by preferring interpretations that match observed community usage. Finally, a part-of-speech tagger labels each content word as noun, modifier, intransitive verb, or transitive verb, and the counts are aggregated by year and corpus.","core_discovery":"The central claim is that Toki Pona, despite being designed with fluid word classes and a minimal lexicon, behaves sociolinguistically like a natural language: content words show systematic diachronic shifts in their preferred syntactic positions, and usage differs by register. Concretely, the paper argues that in informal Toki Pona from 2020 to 2024, luka and uta have become noticeably more frequent as transitive verbs, while pu 'interacting with the official Toki Pona book' has declined as a standalone noun and increasingly appears in the phrase lipu pu, mirroring the language's name-construction pattern. The formal corpus does not lag behind on these innovations, which the authors attribu","pith_inferences":["One testable extension the authors do not run: hold out a later time slice after 2025 and check whether the luka/uta transitive-verb trend continues, saturates, or reverses, since a genuine community innovation should show a logistic adoption curve rather than noise.","If the parser's heuristic scoring is nudged to favor different parses, the reported frequencies could shift; a natural check is to compare against a small hand-annotated gold standard or to perturb the heuristic weights and see how much the trends move.","The social mechanisms invoked, processing efficiency, analogy to name constructions, and identity signaling, are standard explanations in natural-language change; if they hold in a language with no native speakers, it suggests these mechanisms arise from general communication pressures rather than from particular language histories.","Because Toki Pona's lexicon has only about 120 words, a complete inventory of word classes over time is feasible, so this pipeline could be extended to map a trajectory for every content word, offering a near-exhaustive view of change in a small system."],"forward_implications":["If the reported trends are accurate, Toki Pona speakers are innovating within the existing grammar: body-part nouns luka and uta now routinely serve as transitive verbs in informal chat, a usage that is shorter than the periphrastic alternative.","pu as a standalone noun for the official book is declining, replaced by the pattern-conforming lipu pu, which suggests speakers normalize novel constructions to existing name patterns.","Register differences appear: pona is consistently more frequent as an interjection in informal chat, and wawa's rise as an interjection is confined to the informal corpus.","Innovations in the informal community are not delayed in formal written Toki Pona, pointing to a homogeneous, tightly connected speech community.","The finding generalizes: constructed languages with fluid word classes can serve as empirical evidence that sociolinguistic mechanisms operate independently of a language's historical depth."],"supporting_citations":[{"why":"Supplies the informal conversational corpus, a Discord server with over 5.97 million sentences spanning 2016 to 2025, from which all diachronic informal trends are drawn.","marker":"ma pona pi toki pona, 2025"},{"why":"Supplies poki Lapo, the formal curated corpus of published Toki Pona works used for the register comparison.","marker":"kala Asi et al., 2025"},{"why":"Provides sona-toki, the Python library used to clean, sentence-split, tokenize, and filter raw text into Toki Pona sentences before parsing.","marker":"Danielson, 2025"},{"why":"Establishes the specific n-gram configuration of sona-toki used consistently to filter both corpora.","marker":"Danielson, 2024b"},{"why":"Provides the parsing algorithm that generates all possible syntactic interpretations of a sentence, enabling the ambiguity-resolution and tagging pipeline.","marker":"Earley (1970)"},{"why":"Supplies the nearley JavaScript parsing toolkit used to implement the Toki Pona grammar and produce hierarchical phrase structures.","marker":"Chandra & Radvan, 2020"},{"why":"Defines the official Toki Pona grammar, the fluid word classes, and the sense of pu 'interacting with the official Toki Pona book' that ground the analysis.","marker":"Lang, 2014"},{"why":"Supplies the theoretical framing that mental processing and repetition drive language change, used to explain the luka/uta and pu trends.","marker":"Bybee, 2015"},{"why":"Supplies the act-of-identity principle used to explain community-level adoption of innovations in both informal and formal Toki Pona.","marker":"Labov, 2010"}],"fun_headline_variants":["Toki Pona's luka evolves into a verb in chat","Chat data shows Toki Pona's body words turning verb","Constructed Toki Pona evolves like natural tongues","Corpus analysis: Toki Pona changes with social use","Toki Pona's luka and uta go transitive in chat"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The trends stand on the unvalidated assumption that the parser-and-tagger pipeline assigns the correct part of speech to each token; because the ambiguity-resolving heuristic deliberately prefers interpretations that 'reflect observed usage patterns,' the frequency counts are partly shaped by the authors' prior expectations, and a systematic tagging bias would make the reported changes artifacts.","fun_headline_variants_meta":{"raw":{"variants":["Toki Pona's luka evolves into a verb in chat","Chat data shows Toki Pona's body words turning verb","Constructed Toki Pona evolves like natural tongues","Corpus analysis: Toki Pona changes with social use","Toki Pona's luka and uta go transitive in chat"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00043,"raw_usage":{"total_tokens":1963,"prompt_tokens":605,"completion_tokens":1358,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":349,"completion_tokens_details":{"reasoning_tokens":1271}},"tokens_in":349,"tokens_out":1358,"duration_ms":11059,"temperature":1.0,"reasoning_tokens":1271,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T20:32:57.900270+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Hand-annotate a random sample of sentences from each year of both corpora using the paper's tag set, then compare the parser's transitive-verb rates for luka and uta with the hand labels; if the parser systematically over- or under-tags these body-part verbs, the upward trend in Figure 3 will not reproduce under gold-standard annotation.","supporting_citations":[{"cited_title":"(2025, April).sona-toki [GitHub repository]","cited_arxiv_id":null,"evidence_quote":"Provides sona-toki, the Python library used to clean, sentence-split, tokenize, and filter raw text into Toki Pona sentences before parsing."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the parsing algorithm that generates all possible syntactic interpretations of a sentence, enabling the ambiguity-resolution and tagging pipeline."},{"cited_title":"(2020, June)","cited_arxiv_id":null,"evidence_quote":"Supplies the nearley JavaScript parsing toolkit used to implement the Toki Pona grammar and produce hierarchical phrase structures."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the official Toki Pona grammar, the fluid word classes, and the sense of pu 'interacting with the official Toki Pona book' that ground the analysis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the theoretical framing that mental processing and repetition drive language change, used to explain the luka/uta and pu trends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the act-of-identity principle used to explain community-level adoption of innovations in both informal and formal Toki Pona."}],"review_version":1}