{"id":"f59e283b-a3f2-46ef-a35c-7e22e22748cb","arxiv_id":"2608.10109","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new 6,800-post Persian-English code-mixed corpus with LLM-generated Universal Dependencies POS tags, human-validated on a sample, underpins the first cross-platform analysis of Persian-English code-mixing.","lead":"PERCEPT is a new dataset of 6,800 Persian-English code-mixed posts from X, Instagram, and Digikala, with part-of-speech tags for code-mixed words generated by an LLM and validated by human annotators. The dataset enables the first large-scale linguistic analysis of Persian-English code-mixing across platforms, and it is publicly released.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 4 statistics are computed directly from LLM labels whose measured token recall is 86.6% and POS macro F1 is 76.8%; without error propagation, the paper's quantitative conclusions (noun predominance, triggering ordering, topic ranking) are not established.","rationale":"The reader's weakest assumption identifies exactly the same load-bearing issue: Section 4 treats automatically generated labels as ground truth despite the measured error rates reported in Sections 3.1 and 3.2. I agree that this is the most important threat to the paper's central claim. The corpus resource itself is still genuinely useful, and the human validation provides some support, but the quantitative findings cannot be taken at face value without error propagation. The proposed test—recomputing all Section 4 statistics on the validated sample with bias correction—would settle the concern directly. If the conclusions survive the correction, the paper is in good shape; if not, the empirical claims should be conditional. The reader's CONDITIONAL verdict is appropriate; I see no reason to change it. A secondary concern about keyword-based tweet selection inflating the noun share is real but more limited in scope; the error-propagation test would partially expose it as well, because the added tweets are part of the full corpus.","tokens_in":28563,"tokens_out":8305,"duration_ms":76062,"concrete_test":"Using the 150-text, 268-code-mixed-word gold sample (Section 3.1), compute the confusion matrix between LLM labels and gold labels for both token detection and POS tagging. Apply inverse-propensity weighting or matrix-inversion correction to the full-corpus counts, then recompute the POS distribution (Table 2), the length-controlled triggering rates (Table 4), and the per-topic code-mixing means (Figure 2). Bootstrap the gold sample to obtain confidence intervals. If the corrected noun share drops below 50% on any platform, or if the Digikala-vs-X triggering ordering flips, or if the Industry & Commerce / Brand & Business ranking in Figure 2 changes, the Section 4 conclusions as stated should be revised or downgraded to hypotheses.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim includes the first cross-platform empirical analysis of Persian–English code-mixing. Those analyses (Section 4) are computed on the full corpus from labels generated by Gemini 3.5 Flash, but Section 3.1 reports that the LLM misses 13.4% of gold code-mixed tokens (recall 86.6%) and achieves a macro F1 of only 76.8% for POS tagging; topic labels have macro F1 of 65.6% (Section 3.2). These error rates are not propagated into any of the Section 4 statistics. If the missed tokens or the mislabeled tokens are non-random (e.g., verbs in transliterated compound constructions are harder to detect, or rare POS classes are systematically collapsed into NOUN), the reported 60% noun share, the platform-specific POS orderings, the length-controlled triggering rates, and the topic–mixing ranking in Figure 2 could all be biased. The paper provides no confidence intervals, no sensitivity analysis, and no re-analysis on the 150-text validated subset. Because the abstract and introduction present the noun predominance and the Digikala triggering effect as findings, the quantitative claims are load-bearing; they are currently unsupported by the published validation numbers.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces PERCEPT, a corpus of 6,800 Persian–English code-mixed texts from X, Instagram, and Digikala, with Universal Dependencies POS tags for code-mixed words and document-level topic labels generated by Gemini 3.5 Flash and validated by human annotators on a sample. The paper then presents cross-platform analyses of POS distributions, the position of code-mixed words within utterances, a triggering effect measured by co-occurrence of multiple code-mixed words per utterance, and variation in mixing degree across topics. The authors claim that PERCEPT is the first publicly available large-scale POS-annotated Persian–English code-mixed corpus and that their analyses constitute the first cross-platform empirical study of Persian–English code-mixing.","tokens_in":28810,"tokens_out":7413,"duration_ms":62979,"significance":"The dataset fills a clear gap: no existing Persian–English code-mixed corpus provides UD POS annotations, and the resource could support POS tagging, language identification, and other downstream tasks. The release of the corpus, prompts, and analysis code is a concrete contribution. The annotation framework is straightforward, but the human validation protocol (two native annotators, adjudication, explicit instructions without AI assistance) is sound as far as it goes. The quantitative linguistic findings are interesting, but they are only as reliable as the underlying automatic labels; the paper does not currently establish that reliability for the full-corpus statistics.","major_comments":[{"comment":"The corpus-level statistics in Section 4 are direct aggregates of the LLM-generated labels, but the evaluation errors reported in Section 3.1 are never propagated into them. The LLM's code-mixed token recall is 86.6% (F1 90.8%), its POS macro F1 is 76.8%, and its topic macro F1 is 65.6%. The paper offers no confidence intervals, no sensitivity analysis, and no re-analysis on the 150-text validated subset. Because the abstract and introduction present noun predominance, positional consistency, and the Digikala triggering effect as findings, these are load-bearing claims. For instance, a 13.4% false-negative rate for token identification could shift the Table 2 percentages if missed tokens are non-random with respect to POS or platform, and the Table 4 triggering rates are computed entirely from the same token identifications.","section":"§4 (Tables 2–4, Figure 2) and §3.1–3.2"},{"comment":"The exclusion criteria for code-mixed status are defined by prompt instructions to the LLM, but the prompt contradicts itself: Figure 7 instructs the model to exclude transliterated proper nouns, including company and brand names, yet the same example output tags 'Asus' as PROPN, and Table 2 reports 35.0% PROPN for Digikala, which the text attributes to exactly the brand names the instructions say to exclude. The paper does not report any validation of the exclusion boundary, and the human evaluation does not measure errors in the exclusion decision itself. Since the boundary between NOUN and PROPN directly affects the headline noun-preponderance finding and the platform comparisons, the current validation does not establish that the corpus definition was applied consistently.","section":"§3.1, Appendix A (Figures 7–8), Table 2"},{"comment":"The human validation sample is too small to support the corpus-level claims for the categories that matter. It contains 150 texts and 268 code-mixed words, which is about 2.2% of the texts and roughly 2.5% of the 10,785 code-mixed words in Table 2. No confidence intervals are reported for precision, recall, or F1, and the sample is not stratified by POS category or by lexical type. The macro F1 of 76.8% implies that rare POS tags perform substantially worse than the 97.0% overall accuracy, but no per-category breakdown is given. Given that the paper's platform-specific findings hinge on secondary POS categories (VERB in X, ADJ in Instagram, PROPN in Digikala) and on topic-conditioned rates in Figure 2, the absence of per-category validation makes it difficult to distinguish real patterns from annotation bias.","section":"§3.1"}],"minor_comments":[{"comment":"The reported pipeline F1 of 88.1% appears to be the product of token F1 (90.8%) and POS accuracy (97.0%); using the POS macro F1 of 76.8% instead would yield roughly 69.7%, so the paper should state explicitly which metric is used and why.","section":"§3.1"},{"comment":"The chi-square test of position by platform treats each code-mixed word as an independent observation, but multiple words from the same utterance are not independent; a test that accounts for within-utterance clustering would be more appropriate.","section":"§4.2"},{"comment":"The 'length-controlled' triggering rate is described only verbally; the paper should report the length-group definitions, the group-specific rates, and the denominators for the numbers in Table 4.","section":"§4.3"},{"comment":"The retry policy for 'no code-mixed words' responses is applied only to the keyword-collected X tweets, not to the other subsets; the potential effect of this asymmetric policy on the platform comparisons should be discussed.","section":"§3.1"},{"comment":"The white cells for missing topic–platform combinations are not accompanied by a minimum-cell-size threshold, so the reader cannot tell whether some means are based on very few utterances.","section":"Figure 2"}],"recommendation":"major_revision","confidential_remarks":"The paper's main value is the dataset and the validation protocol; the linguistic analysis sections are presented with more confidence than the error rates support. I would recommend requiring the authors to reframe the Section 4 findings as exploratory or to provide a rigorous error-propagation/sensitivity analysis on the validated sample before publication. The internal inconsistency in the Digikala prompt example (Asus as PROPN despite the instruction to exclude brand names) should be fixed in the public repository as well as the paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this paper's corpus is real and useful; its linguistic claims are under-supported by its own validation numbers.\n\nThe contribution is PERCEPT, the first publicly available large-scale POS-annotated Persian-English code-mixed corpus (6,800 texts from X, Instagram, and Digikala), with UD tags for code-mixed words and topic labels. That is a genuine gap filler. The corpus is released, the annotation prompts are in the appendix, and the human validation protocol is reasonable: two native annotators with adjudication, and the numbers they report (token F1 90.8%, POS macro F1 76.8%) are honest. The observation that nouns dominate and that POS distributions vary across platforms is a useful baseline for a language pair that had none.\n\nThe soft spot is exactly what the stress-test note flags. The headline analyses in Section 4 (noun percentages, positional consistency, triggering ordering, topic ranking) are computed on the full corpus from Gemini labels, not from the gold sample. The known error rates are not propagated. Token recall is 86.6%: the LLM misses 13.4% of code-mixed tokens. POS macro F1 is 76.8%. Topic macro F1 is 65.6%. If the misses or mislabels are systematic (transliterated compound verbs, rare categories), then the 60% noun share and the Digikala triggering effect could shift. The paper gives no confidence intervals, no sensitivity analysis, and no re-analysis on the validated subset. That is a load-bearing gap for the empirical claims, though not for the corpus itself.\n\nA second, smaller issue: the additional 1,173 X posts were gathered by searching for the 71 most frequent code-mixed words. That inevitably inflates the noun share and affects any corpus-level POS distribution. The authors acknowledge the keyword step but don't discuss its selection effect.\n\nThe citation pattern looks fine; the related work on code-mixed corpora and POS tagging is current, and the Persian resources are covered.\n\nWho benefits: anyone building or evaluating POS taggers or code-mixed NLP for Persian. The resource deserves referee time. The right path is conditional acceptance: keep the corpus, fix the analysis by adding error bounds, re-running key statistics on the gold sample, and discussing the keyword sampling bias. If the authors do that, this is a standard reference.","headline":"A genuinely useful new Persian-English code-mixed corpus built with LLM annotation; the corpus is solid, but the paper's quantitative linguistic analyses overstate their support given known annotation error rates.","tokens_in":29327,"tokens_out":1886,"would_cite":true,"duration_ms":18928,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The PERCEPT corpus is the first large-scale Persian-English code-mixed dataset with Universal Dependencies POS tags for code-mixed words.","keywords":["code-mixing","Persian-English","part-of-speech tagging","Universal Dependencies","LLM annotation","social media corpus","code-switching","topic detection"],"falsifier":"Take a random sample of roughly 500 PERCEPT texts, have native-speaker annotators label code-mixed tokens and POS tags following the same instructions, and compare the resulting noun proportions, triggering rates, and platform differences with the LLM-based numbers; if, for example, the Digikala noun share drops toward 40% or the triggering gap narrows below the reported margins, the corpus-level analyses are not robust.","tokens_in":28337,"feed_emoji":"💬","tokens_out":9144,"duration_ms":71816,"temperature":0.7,"pith_summary":"This paper introduces PERCEPT, a corpus of 6,800 Persian-English code-mixed posts from X, Instagram, and Digikala, with Universal Dependencies part-of-speech tags for every English or English-transliterated word judged to be code-mixed. It is offered as the first publicly available large-scale resource of its kind for Persian. The paper also presents an annotation pipeline in which a large language model assigns POS tags and topic labels, with human validation on a 150-post sample. Using PERCEPT, the authors report that code-mixed words are overwhelmingly nouns, that their sentence positions are distributed almost identically across platforms, and that the tendency for code-mixing to recur within one utterance is strongest in Digikala reviews.","feed_headline":"Persian-English code-mixing gets its first POS-tagged corpus","feed_subtitle":"6,800 posts from X, Instagram, and Digikala show nouns dominate and e-commerce triggers more mixing.","key_machinery":"The central object is the PERCEPT corpus: 6,800 anonymized Persian posts containing 5,660 code-mixed tokens from X, 1,635 from Instagram, and 3,490 from Digikala, each code-mixed token tagged with one of 17 Universal Dependencies POS tags. The mechanism that carries the argument is the LLM-assisted annotation pipeline, which uses platform-specific prompts to identify English and English-transliterated code-mixed words, exclude established loanwords, proper nouns, and platform-specific interface terms, and assign POS tags with the model's temperature set to zero. Human validation on 150 texts (268 code-mixed words) reports a token-level F1 of 90.8%, a POS accuracy of 97.0% on correctly identified tokens, and a POS macro F1 of 76.8%, with an overall pipeline F1 of 88.1%.","core_discovery":"The central claim is that PERCEPT is the first large-scale, publicly available Persian-English code-mixed corpus whose code-mixed words carry Universal Dependencies POS tags, covering both English-script borrowings and English words transliterated into Persian script. The paper argues that this resource, built from 6,800 posts across X, Instagram, and Digikala, supports the first multi-platform empirical analysis of Persian-English code-mixing. The analyses assert that nouns are the dominant POS category for code-mixed words, that positional distributions are consistent across platforms (roughly one-third each in initial, medial, and final positions, with a chi-square test yielding no significant platform association), that the triggering effect—multiple code-mixed words within one utterance—is stronger in Digikala (42.0% length-controlled) than in X (27.1%) or Instagram (10.2%), and that Industry & Commerce and Brand & Business topics show the highest code-mixing density.","pith_inferences":["If the reported human-LLM agreement generalizes, it suggests that LLM-assisted annotation can produce linguistic corpora for low-resource code-mixed languages at scale, but the 2.2% validation sample leaves the error bars on the corpus-level statistics unknown.","The large PROPN share on Digikala (35.0%) is partly a product of prompt design: the annotation instructions exclude established loanwords and platform-specific terms from tagging, so cross-platform differences in PROPN may be amplified by the exclusion lists rather than by underlying language use.","The triggering-effect operationalization (whether an utterance contains more than one code-mixed word) is a coarse density measure; a more direct test of the triggering hypothesis would model inter-code-mixed-word distances or use a regression that controls for topic and utterance length.","Because texts labeled as offensive were removed from the source datasets, PERCEPT may underrepresent certain topics and styles, which could bias the reported topic-code-mixing correlations."],"forward_implications":["Persian NLP systems can use PERCEPT to train and evaluate POS taggers, language identifiers, and syntactic parsers on code-mixed input, a setting where such tools are currently weak.","The prompt-based annotation framework, being tied to the language-independent UD tagset, can be adapted to other low-resource code-mixed language pairs by translating the prompts and adjusting the exclusion lists.","The finding that nouns dominate code-mixing and that e-commerce platform vocabulary inflates the PROPN category can guide the design of code-mixed lexicons and text-normalization tools.","The stronger triggering effect in Digikala implies that e-commerce reviews are a particularly concentrated site of code-mixing, which matters for sentiment analysis and product-summary systems.","Because positional distributions are consistent across platforms, a single positional model may suffice for predicting where code-mixing appears in Persian social media text across genres."],"supporting_citations":[{"why":"Collected 3,640 Persian-English code-mixed tweets for sentiment analysis; the existing Persian resource that lacks POS annotations for code-mixed words.","marker":"Sabri et al. (2021b)"},{"why":"Introduced PinLID for Persian-English code-mixed language identification but excluded English-transliterated words; marks the gap PERCEPT fills.","marker":"Ghafouri et al. (2025)"},{"why":"Defines the Universal Dependencies tagset and annotation guidelines used for the POS labels.","marker":"De Marneffe et al. (2021)"},{"why":"Documents the Gemini 3.5 Flash model used for automatic POS tagging and topic detection in the annotation pipeline.","marker":"Google, 2026"},{"why":"Source dataset for the Digikala comments included in PERCEPT.","marker":"Asli et al. (2020)"},{"why":"Source dataset for X tweets included in PERCEPT.","marker":"Golazizian et al. (2020)"},{"why":"Source dataset for Instagram comments included in PERCEPT.","marker":"Heidari and Shamsinejad (2020)"}],"fun_headline_variants":["First POS-tagged Persian-English code-mixed corpus","6,800 posts reveal nouns dominate Persian-English code-mixing","E-commerce sparks more Persian-English code-mixing than X or Instagram","New Persian-English code-mixing corpus tags every word's POS"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The quantitative conclusions treat the automatically generated POS and topic labels as correct for all 6,800 texts, but only 150 texts (268 code-mixed words) were checked by humans, so systematic annotation errors could change the reported distributions.","fun_headline_variants_meta":{"raw":{"variants":["First POS-tagged Persian-English code-mixed corpus","6,800 posts reveal nouns dominate Persian-English code-mixing","E-commerce sparks more Persian-English code-mixing than X or Instagram","New Persian-English code-mixing corpus tags every word's POS"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001065,"raw_usage":{"total_tokens":4490,"prompt_tokens":995,"completion_tokens":3495,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":611,"completion_tokens_details":{"reasoning_tokens":3425}},"tokens_in":611,"tokens_out":3495,"duration_ms":26667,"temperature":1.0,"reasoning_tokens":3425,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:10:11.684823+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of roughly 500 PERCEPT texts, have native-speaker annotators label code-mixed tokens and POS tags following the same instructions, and compare the resulting noun proportions, triggering rates, and platform differences with the LLM-based numbers; if, for example, the Digikala noun share drops toward 40% or the triggering gap narrows below the reported margins, the corpus-level analyses are not robust.","supporting_citations":[],"review_version":1}