{"id":"bca589b6-5d24-44b2-bc3b-0ae95dc60c1f","arxiv_id":"2408.04596","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Insertional code-switching in Chinese-English text and speech is not purely speaker-driven, since secondary-language productions are less predictable than primary-language alternatives.","lead":"This paper uses language models on Chinese-English forum posts and speech transcripts to show that insertional code-switching is not driven only by speaker ease. English insertions have lower predictability than meaning-equivalent Chinese alternatives, pointing to listener-oriented functions.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Cross-lingual perplexity comparison between English inserts and meaning-equivalent Chinese alternatives lacks explicit normalization or joint modeling.","rationale":"The reader's identified assumption (perplexity as production-difficulty proxy) is related but downstream; the more immediate load-bearing step is the cross-lingual comparison required to establish that English is not easier. Full methods would be needed to check normalization, so the verdict remains conditional rather than rejected.","tokens_in":1680,"tokens_out":281,"duration_ms":27471,"concrete_test":"Recompute the key comparison (English vs. meaning-equivalent Chinese) on the same bilingual data using bits-per-character or a single joint model; if the English items no longer show reliably lower predictability, the rejection of the speaker-driven account does not hold.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central rejection of purely speaker-driven code-switching rests on the claim that English productions exhibit even lower predictability than meaning-equivalent Chinese alternatives. Perplexity values from separate monolingual LMs are not directly comparable across languages due to differing baseline entropies, tokenization, and character densities. The abstract provides no indication of bits-per-character normalization, a single multilingual LM, or other calibration; without this, the directional claim that English is 'lower' (hence not easier) does not follow from the reported numbers.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.3","summary":"The paper uses language modeling on Chinese-English bilingual online forum posts and spontaneous speech transcripts to replicate that low primary-language (Chinese) predictability correlates with insertional switches to English. It then claims that the English productions exhibit even lower predictability than meaning-equivalent Chinese alternatives and are therefore not easier to produce, rejecting a purely speaker-driven information-theoretic account of code-switching in both text and speech.","tokens_in":1792,"tokens_out":455,"duration_ms":26624,"significance":"If the cross-lingual predictability comparison holds after proper calibration, the result would meaningfully challenge speaker-centric accounts by showing that switches do not occur at points of relative production ease. The replication of the primary-language predictability correlation plus the use of meaning-equivalent alternatives across both written and spoken modalities constitute a clear empirical contribution via new corpus measurements.","major_comments":[{"comment":"Abstract and central results: the directional claim that English productions have lower predictability than meaning-equivalent Chinese alternatives (and are therefore not easier) rests on direct comparison of perplexities from separate monolingual LMs. No bits-per-character normalization, joint multilingual modeling, or other calibration for differing baseline entropies and tokenization is described, so the inequality does not follow from the reported numbers. This comparison is load-bearing for the rejection of the speaker-driven theory.","section":"Abstract and Results"},{"comment":"Methods and analysis sections: details on LM training (architecture, data, tokenization), the procedure for constructing and aligning meaning-equivalent Chinese alternatives, and any statistical controls or significance tests on the perplexity differences are not supplied. Without these, the robustness of the key negative result cannot be evaluated.","section":"Methods"}],"minor_comments":[{"comment":"The abstract and introduction could more explicitly flag the cross-lingual comparability issue and how it is addressed (or why separate monolingual perplexities suffice).","section":"Abstract"},{"comment":"Notation for predictability/perplexity should be defined once and used consistently when switching between Chinese and English contexts.","section":null}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the thoughtful and constructive report. The two major comments identify important gaps in the presentation of our cross-lingual predictability comparison and in the methodological documentation. We address each point below and commit to a major revision that supplies the missing details and strengthens the empirical claim.","responses":[{"response":"We agree that the manuscript does not describe any cross-lingual calibration and that this omission weakens the central negative result. In the revision we will (1) report bits-per-character normalized perplexities for both languages, (2) train and evaluate a single multilingual model on the combined data to provide a joint baseline, and (3) include an explicit discussion of why the directional inequality survives these controls. If the calibrated comparison no longer supports the claim, we will qualify or retract the rejection of the purely speaker-driven account. We therefore treat this as a required methodological correction rather than a minor clarification.","revision_made":"yes","referee_comment":"[Abstract and Results] Abstract and central results: the directional claim that English productions have lower predictability than meaning-equivalent Chinese alternatives (and are therefore not easier) rests on direct comparison of perplexities from separate monolingual LMs. No bits-per-character normalization, joint multilingual modeling, or other calibration for differing baseline entropies and tokenization is described, so the inequality does not follow from the reported numbers. This comparison is load-bearing for the rejection of the speaker-driven theory."},{"response":"We acknowledge that the current Methods section is insufficiently detailed. The revised manuscript will add: (a) full LM specifications (architecture, training corpora, tokenizers, and hyper-parameters), (b) the exact procedure used to generate and align meaning-equivalent Chinese alternatives (including any translation or parallel-corpus resources), and (c) statistical tests (paired Wilcoxon signed-rank tests with effect sizes and confidence intervals) on the perplexity differences. These additions will make the negative result reproducible and allow readers to assess its robustness directly.","revision_made":"yes","referee_comment":"[Methods] Methods and analysis sections: details on LM training (architecture, data, tokenization), the procedure for constructing and aligning meaning-equivalent Chinese alternatives, and any statistical controls or significance tests on the perplexity differences are not supplied. Without these, the robustness of the key negative result cannot be evaluated."}],"tokens_in":1330,"tokens_out":506,"duration_ms":16868,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper replicates that insertional switches to English happen at low-predictability points in Chinese, using both written forum data and speech transcripts. That part lines up with earlier work and gives a clean empirical check. The new move is the direct comparison showing English tokens are even less predictable than meaning-equivalent Chinese alternatives, which is meant to rule out a purely speaker-driven account. That comparison is the part that does not hold up. Perplexity scores from separate monolingual models cannot be ranked across languages without normalization for baseline entropy, tokenization, or character density; the abstract gives no sign that any such calibration was done. Without it, you cannot conclude the English productions are harder or easier to produce. The rest of the logic is straightforward, but the load-bearing step rests on an unadjusted cross-lingual metric. This work is aimed at people already following computational models of bilingual production. A reader in that niche would find the replication useful and the proposed test worth discussing, even if the numbers need reworking. The paper shows honest engagement with the literature and a clear experimental contrast, so it is worth sending out for review once the authors supply the missing calibration or switch to a single multilingual model.","headline":"The central claim fails because separate monolingual LMs give incomparable perplexities across Chinese and English, so the rejection of speaker-driven code-switching does not follow.","tokens_in":2235,"tokens_out":317,"would_cite":false,"duration_ms":24819,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[{"relation":"unclear","rs_module":"IndisputableMonolith/Cost/FunctionalEquation.lean","rs_theorem":"washburn_uniqueness_aczel","paper_passage":"We then demonstrate that the predictability of the English productions is even lower than that of meaning-equivalent Chinese alternatives"},{"relation":"unclear","rs_module":"IndisputableMonolith/Foundation/RealityFromDistinction.lean","rs_theorem":"reality_from_one_distinction","paper_passage":"surprisal(wi) = −log2P(wi|wi−1, ..., wi−t)"}],"headline":"Cross-lingual surprisal analysis of code-switching has no structural overlap with RS cost or forcing machinery","alignment":"orthogonal","rationale":"The paper's core apparatus (5-gram/LLM surprisal on CS1 vs. monolingual baselines, normalized ratios, logistic regression on frequency/length/POS) operates entirely within information-theoretic linguistics. No reciprocal cost J, φ-ladder, 8-tick periodicity, or distinction-forcing appears. Domain is orthogonal to RS theorems such as reality_from_one_distinction, washburn_uniqueness_aczel, or alexander_duality_circle_linking.","tokens_in":52872,"confidence":"high","tokens_out":295,"duration_ms":8433,"cache_read_input_tokens":38528,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Code-switching to English occurs even when English is less predictable than equivalent Chinese phrases, rejecting purely speaker-driven accounts.","keywords":["code-switching","language modeling","predictability","bilingual speech","Chinese-English","information theory","speaker design"],"falsifier":"A dataset of Chinese-English code-switches in which the English insertions show higher predictability than their meaning-equivalent Chinese alternatives would contradict the central observation.","tokens_in":2590,"feed_emoji":"🗣","tokens_out":605,"duration_ms":15479,"temperature":0.7,"pith_summary":"The paper examines whether insertional code-switching is driven only by the need to make speech or writing easier for the bilingual speaker. It replicates the finding that switches from Chinese to English tend to occur where Chinese text or speech is hard to predict. However, when the actual English insertions are compared to what a Chinese continuation with the same meaning would have been, the English turns out to be even harder to predict. This pattern appears in both online forum writing and spontaneous speech transcripts, implying that speakers are not switching languages merely to lower their own production cost.","feed_headline":"English code-switches less predictable than Chinese alternatives","feed_subtitle":"This pattern in writing and speech rejects the claim that speakers switch languages only to ease their own production burden.","key_machinery":"Language-model perplexity on primary-language text, used to quantify predictability at potential switch points and to compare the actual secondary-language insertion against a meaning-matched primary-language continuation.","core_discovery":"Low predictability in the primary language (Chinese) correlates with switches to the secondary language (English), yet the English material actually produced has lower predictability than meaning-equivalent Chinese alternatives; therefore the switches do not reduce production difficulty and the purely speaker-driven account of code-switching is rejected for both written and spoken data.","pith_inferences":["Similar comparisons could be run on other language pairs to test whether the pattern is specific to Chinese-English bilinguals.","Listener comprehension studies could check whether the harder-to-predict English insertions actually improve understanding or signal emphasis.","The results leave open the possibility that code-switching frequency changes with audience expectations or conversational setting."],"forward_implications":["Code-switching in both writing and speech serves communicative goals beyond reducing speaker effort.","Speakers may insert the secondary language to direct listener attention rather than to simplify their own output.","The same rejection of a purely speaker-driven account holds for both online forum posts and spontaneous speech transcripts.","Information-theoretic models of speaker design must incorporate listener-oriented or social functions to explain observed switching patterns."],"fun_headline_variants":["Low Chinese predictability leads to English switches but English less predictable","English code switches have lower predictability than equivalent Chinese","Code switching data rejects purely speaker driven account","Both written and spoken data challenge information theoretic speaker design"],"cache_read_input_tokens":64,"weakest_assumption_plain":"Language-model perplexity on the primary language serves as a valid stand-in for the real-time production difficulty a bilingual speaker faces when choosing which language to use next.","fun_headline_variants_meta":{"raw":{"variants":["Low Chinese predictability leads to English switches but English less predictable","English code switches have lower predictability than equivalent Chinese","Code switching data rejects purely speaker driven account","Both written and spoken data challenge information theoretic speaker design"]},"model":"grok-4.3","cost_usd":0.00653,"raw_usage":{"total_tokens":3032,"prompt_tokens":625,"num_sources_used":0,"completion_tokens":52,"cost_in_usd_ticks":65299500,"prompt_tokens_details":{"text_tokens":625,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2355,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":625,"tokens_out":52,"duration_ms":20779,"temperature":1.0,"reasoning_tokens":2355,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-23T21:58:11.436558+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A dataset of Chinese-English code-switches in which the English insertions show higher predictability than their meaning-equivalent Chinese alternatives would contradict the central observation.","supporting_citations":[],"review_version":1}