{"id":"c3518994-75e0-4a2a-8b6c-cc5a91a6eeca","arxiv_id":"2506.13513","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"The paper introduces an automated RAG-based pipeline that generates and filters 7,555 Korean neutral-toxic sentence pairs, plus 539 English pairs, for training detoxification models.","lead":"K/DA is a pipeline that automatically creates pairs of neutral and offensive Korean sentences for training language models that remove toxic language. It uses slang retrieved from online communities and LLM-based filtering, and the paper reports gains in implicit offensiveness and pair consistency versus existing Korean datasets.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The dataset-quality claim rests on GPT-4 Turbo as generator, filter, and evaluator; the human checks are too small and too low-agreement to break the circularity, so independent human replication is needed.","rationale":"Reader's weakest assumption is exactly the one I find most load-bearing: GPT-4's dual role as filter and evaluator. Section 5.1's Table 3 is the quantitative backbone of the abstract's 'greater implicit offensiveness' claim, and every score in that table is a G-Eval output from GPT-4 Turbo using criteria (Tables 17-18) that were developed to match the GPT-4-coined taxonomy of Section 3. There is no independent human rating of the same comparison: Table 4 compares only K/DA vs K-OMG on 50 samples and even there the instructions were not identical (fluency criterion differed), and no significance tests are reported. The low Fleiss kappa (0.17-0.23) in Appendix I.1 further weakens the claim that GPT-4's scores reflect a stable human notion of implicit offensiveness; 86-90% agreement with GPT-4 is reported, but with such low inter-annotator agreement the human 'gold' is itself not well-defined. The detoxification claim also depends on G-Eval on 100 test sentences and disappears on BEEP, so the headline should at least be qualified. The proposed concrete test—independent human ratings of a stratified sample of all compared datasets with the same rubrics—directly resolves whether the GPT-4 circularity is benign. If it passes, conditional acceptance is justified; if not, the dataset-quality claim is unverified. I therefore keep the reader's CONDITIONAL verdict.","tokens_in":24291,"tokens_out":4048,"duration_ms":37947,"concrete_test":"Run a preregistered human rating study: sample 200 pairs from each dataset in Table 3 (Ours, K-OMG, BEEP, KODOLI, translated CADD), have at least 5 fresh native Korean annotators per item rate implicit offensiveness and consistency using the same rubrics as in Tables 17-18, and test whether Ours ranks highest in mean implicit offensiveness and consistency using a paired bootstrap or mixed-effects model. Also compute per-item human-GPT-4 agreement. If the human ranking reproduces Table 3, the circularity concern is resolved; if not, the dataset-comparison claim should be downgraded.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing assumption is that GPT-4 Turbo's judgments of implicit offensiveness and pair consistency are a valid proxy for human perception. This assumption enters three places: Section 3 uses GPT-4 to coin the 'trend-aligned slang' taxonomy; Section 4.2 uses GPT-4 as the filter that retains only outputs it deems context-preserving and implicitly offensive; Section 5.1 uses G-Eval with GPT-4 Turbo to compare Ours against K-OMG, BEEP, KODOLI, and translated CADD (Table 3). Because the same model defines, selects, and scores the construct, the headline result 'greater implicit offensiveness and pair consistency' can be a self-consistency artifact rather than a property of the data. The human checks are too small to break the loop: 15 annotators, 50 samples for dataset quality (Table 4), 45 sentences for detoxification preference, and inter-annotator agreement is only Fleiss kappa 0.17-0.23; even if GPT-4 matches the majority vote at 90%, individual annotators disagree substantially. An additional discrepancy: Table 5 shows the trained detoxification model is not better than the Vanilla LM on BEEP (Overall O. 1.580 vs 1.481; Implicit O. 1.506 vs 1.393), so the abstract's 'high-performing detoxification model' claim is only in-distribution or near-distribution. Both issues are addressable, but as published the central claim is not independently verified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces K/DA, a two-stage automated pipeline for generating Korean neutral-toxic paired data for detoxification training. Stage one uses retrieval-augmented generation with a vector database of 92,953 crawled online-community comments to produce toxic variants of neutral sentences while incorporating recently emerging slang. Stage two filters generations with LLM prompts for pair consistency and implicit offensiveness, and the authors define the new construct 'trend-aligned slang' based on GPT-4 Turbo's categorization of implicitly offensive comments. The released dataset contains about 7.5K Korean pairs and 539 English pairs. The authors evaluate dataset quality using G-Eval with GPT-4 Turbo, compare against K-OMG, BEEP, KODOLI, and translated CADD, and train an instruction-tuned Ko-LLaMA3-Luxia-8B detoxification model. They report higher implicit offensiveness and pair consistency than existing Korean datasets, demonstrate cross-lingual and cross-model applicability, and show improved detoxification performance when tested on their own dataset and, partially, on KOLD.","tokens_in":24569,"tokens_out":4578,"duration_ms":43602,"significance":"If the reported results are accepted, the pipeline offers a genuinely useful alternative to expensive human annotation for building paired detoxification data, with a mechanism (RAG from live communities) for staying current with slang. The paper contributes a new Korean paired dataset, an openly released pipeline and code, cross-lingual and open-model replication experiments, and a careful ethical release statement. These are concrete strengths that make the work reproducible and testable by others. However, the central empirical claims rest on GPT-4 Turbo acting simultaneously as category definer, generator, filter, and evaluator, with only minimal human validation. Because the same model defines, selects, and scores the target constructs, the headline advantages in implicit offensiveness and consistency could be self-consistency artifacts rather than properties of the data. The limited human checks (50 dataset-quality samples, 45 preference judgments, Fleiss kappa 0.17-0.23) do not break this loop. The practical claim of a 'high-performing detoxification model' is also only reliably supported in-distribution, with no advantage on BEEP.","major_comments":[{"comment":"The evaluation of dataset quality is circular. GPT-4 Turbo is used to define the taxonomy of trend-aligned slang (Section 3, Figure 1), to generate toxic variants (Section 4.1), to filter for pair consistency and implicit offensiveness (Section 4.2), and to score the final dataset with G-Eval (Section 5.1, Table 3). The same model is then used to compare Ours against K-OMG, BEEP, KODOLI, and translated CADD. The human-check in Appendix I is too small to break this loop: dataset quality is scored on 50 samples (Table 4), and inter-annotator agreement on the filtering tasks is only Fleiss kappa 0.17-0.23, with the majority-vote agreement to GPT-4 Turbo rising to 97%/94% but individual agreement at 86%/90%. With 15 annotators disagreeing at this level, the reported 'greater implicit offensiveness and pair consistency' may reflect GPT-4 Turbo's internal consistency rather than a property that human users would reliably recognize. I recommend adding an independent evaluation with a different large language model (or a larger, more carefully designed human study with per-annotator disagreement reported) and, at minimum, tempering the claims in Table 3 accordingly.","section":"Sections 3, 4.2, 5.1, and Appendix I"},{"comment":"The claim in the abstract that K/DA 'enables effective training of a high-performing detoxification model' is not supported by the out-of-domain results. On BEEP, the model trained on K/DA has higher Overall O. (1.580) and Implicit O. (1.506) than the Vanilla LM (1.481 and 1.393, respectively), meaning it performs worse on that transfer set. On KOLD, the improvements over Vanilla LM (Overall O. 1.606 vs 1.741; Implicit O. 1.566 vs 1.682) are within the reported standard errors and are not accompanied by any significance test. The only clear and robust improvement is on the in-distribution test set, Ours. The paper should state this limitation explicitly in the abstract and conclusions, or provide statistical significance testing and a more careful characterization of where the model does and does not help.","section":"Section 5.3, Table 5"},{"comment":"The human evaluation of dataset quality uses only 50 randomly sampled pairs, which is a very small basis for concluding that K/DA has 'greatest implicit offensiveness' beyond the GPT-4 evaluation. In addition, the comparison with K-OMG is approximate, as acknowledged: the evaluation instructions differ (e.g., fluency instructions are not identical), and the K-OMG column for implicit offensiveness is missing because K-OMG was not scored on that dimension. The paper reports Cronbach's alpha in the table but no confidence intervals or effect sizes for the 50-sample comparison. This human check is too weak to independently validate the Table 3 G-Eval conclusions, and I would like to see either a larger human sample or a downgrade of the claim to 'suggestive evidence'.","section":"Table 4 and Appendix I.2"},{"comment":"The definition of the paper's key construct, 'trend-aligned slang', rests on a GPT-4 Turbo categorization of 1,000 comments with no human agreement check. The paper reports that 64% of implicitly offensive comments fall into categories (2) and (3) (community-specific slang and detection-evading profanity), and this figure motivates the entire pipeline design. If this categorization is not stable across human raters, the central construct may be model-specific. Given that the paper already conducts human evaluations for other tasks, I recommend adding a small human-annotation study of the 1,000 comments (or a representative sample) to confirm the proportions and the three-way taxonomy.","section":"Section 3, Figure 1"}],"minor_comments":[{"comment":"The sentence 'The tendency for overall offensiveness to be the lowest, while implicit offensiveness remains the highest, indicates that the dataset has been appropriately constructed' is contradicted by Table 3, where BEEP has a lower Overall O. (2.300) than Ours (2.719). The claim should be revised to exclude BEEP or to note that the pattern holds only among the paired/generated datasets.","section":"Section 5.1, paragraph 3"},{"comment":"The caption says 'The numbers in parentheses represent the Cronbach's α' but the table displays values in square brackets. Also, the Overall O. and Implicit O. columns for Ours both show 4.196, which may be a coincidence but deserves a brief note or verification.","section":"Table 4 caption"},{"comment":"There is a typo: '7,555 Korean netural-toxic paired' should be 'neutral-toxic'. Please fix.","section":"Appendix J"},{"comment":"The English generalization results are based on only 539 total pairs, and the 500-pair evaluation covers nearly the entire set. It would be helpful to state whether these are the same 539 pairs or a separate held-out set, and to include a brief note on the small size when claiming 'applicability to other languages'.","section":"Section 5.2, Appendix G"}],"recommendation":"major_revision","confidential_remarks":"The paper's engineering contributions and released assets are valuable, and the automated pipeline is a useful direction for the community. The main reviewer concern is methodological: the central claims are supported primarily by the same model used as evaluator, generator, and filter, and the human validation described in the appendix is too limited to break the circularity. The authors should be encouraged to add an independent evaluation and to revise the abstract's overbroad 'high-performing detoxification model' claim. This is a major-revision rather than a rejection because the issues are addressable with additional experiments and stricter claim phrasing."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"K/DA is a genuinely useful pipeline: multi-n RAG plus LLM filtering for pair consistency and implicit offensiveness is a new combination, and the released dataset (7,555 Korean pairs plus 539 English) and code make the contribution concrete. The cross-model experiments with open-source models are a good addition, and the authors are transparent about several limitations.\n\nThe soft spot is the evaluation loop. GPT-4 Turbo defines the categories, generates the toxic variants, filters them, and then scores the final dataset with G-Eval. That makes the reported 'higher implicit offensiveness and consistency' partly self-confirmatory. The human checks do not break the loop: 15 annotators, 50 samples for dataset quality, 45 sentences for preference, and Fleiss kappa of 0.17-0.23 indicate substantial individual disagreement, so the 86-90% agreement with GPT-4 is not as reassuring as it looks. The stress-test note gets this right.\n\nThe second issue is the abstract's 'high-performing detoxification model.' Training on K/DA helps on the in-distribution test set and on KOLD, but on BEEP the fine-tuned model is worse than the vanilla LM (Overall O. 1.580 vs 1.481; Implicit O. 1.506 vs 1.393). The authors acknowledge this, and the explanation is plausible, but the claim as stated overreaches.\n\nThese are fixable problems, not fatal ones. I would send this to peer review with a request for independent human evaluation on a larger sample, significance tests, release of the retrieval database, and careful rewording of the cross-domain claims. The paper is a solid practical contribution and the authors are engaged with the right questions. It deserves a serious referee.","headline":"A useful, honestly-reported pipeline for building paired Korean detoxification data, but the headline quality claims are still largely self-confirmatory because the same model generates, filters, and scores the data.","tokens_in":25171,"tokens_out":2860,"would_cite":true,"duration_ms":29033,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An automated pipeline can build Korean detoxification training data that beats existing datasets on implicit offensiveness and pair consistency.","keywords":["language detoxification","implicit offensiveness","Korean offensive language","retrieval-augmented generation","paired data generation","LLM-based filtering","trend-aligned slang","instruction tuning"],"falsifier":"A human study with a larger and more diverse annotator pool that asked whether GPT-4's filter decisions match the majority label better than chance, or a blind test where a detoxification model trained on human-annotated paired data outperforms the K/DA-trained model on a held-out set of naturally occurring offensive comments.","tokens_in":24065,"feed_emoji":"🛡️","tokens_out":4905,"duration_ms":40123,"temperature":0.7,"pith_summary":"The paper introduces K/DA, an automated pipeline that builds paired neutral-toxic Korean sentences for training language detoxification models. The pipeline retrieves current slang from online communities and uses an LLM to filter generations for pair consistency and implicit offensiveness, including newly defined 'trend-aligned slang.' The authors report that the resulting dataset of about 7.5K pairs scores higher on implicit offensiveness than existing Korean datasets, and that a model fine-tuned on it detoxifies better than models trained on prior datasets. This matters because it removes the need for expensive human annotation and keeps training data current with rapidly evolving offensive language.","feed_headline":"Pipeline auto-builds Korean detox pairs that beat existing sets","feed_subtitle":"Retrieving fresh slang and LLM filtering removes the need for hand-labeled training data.","key_machinery":"The load-bearing mechanism is a two-stage pipeline called K/DA. Stage one, slang retrieval, uses retrieval-augmented generation with multiple retrieval counts ($n \\in \\{0,3,5,7,9\\}$) to pull comments from a Sentence-BERT-embedded corpus of Korean online communities, prompting an LLM to rewrite a neutral sentence with that slang while keeping the meaning. Stage two, generation filtering, uses two GPT-4-based prompts: a pair-consistency filter that rejects responses, paraphrases, and context shifts, and an implicit-offensiveness filter that accepts only outputs with trend-aligned slang or disguised profanity. The definition of trend-aligned slang, which splits implicit offensiveness into three subcategories, is what the filters operationalize.","core_discovery":"The central claim is that an automated, retriever-based generation pipeline can produce training data for detoxification that is both contextually aligned and implicitly offensive, two properties that prior Korean datasets achieve only partially. K/DA generates toxic versions of neutral sentences by injecting slang retrieved from a vector database of 92,953 online comments, then keeps only candidates that pass two LLM-based filters: one for meaning preservation and one for implicit offensiveness. The paper defines 'trend-aligned slang' as community-specific slurs and disguised profanity variants, and reports that these make up most implicitly offensive comments in real Korean online data. Evaluated by GPT-4, the resulting dataset shows the highest implicit offensiveness and pair consistency among the compared Korean datasets; human evaluation on a 50-pair sample and preference judgments on 45 detoxified sentences also favor the K/DA-trained model.","pith_inferences":["If the GPT-4 judgments are as reliable as the 86–90% agreement reported, a similar pipeline could be used to bootstrap paired detoxification data for other low-resource languages by swapping the vector database and prompt language.","The low human inter-annotator agreement (Fleiss Kappa 0.17–0.23) suggests that 'implicit offensiveness' is far from a settled category; K/DA's operational definition could serve as a more consistent labeling standard, though it inherits whatever bias GPT-4 has.","The pipeline could be pointed at targeted communities or demographic groups to generate adversarial examples that stress-test hate-speech detectors, not just detoxifiers."],"forward_implications":["Instruction-tuned models trained on K/DA pairs lower overall and implicit offensiveness on in-distribution and KOLD test sets, with statistically significant gains over models trained on K-OMG or translated CADD.","Re-running the pipeline on freshly scraped comments lets a dataset track new slang without human annotation, directly addressing the stale-dataset problem in detoxification.","The same pipeline applied to English data yields the highest implicit offensiveness among compared datasets, suggesting the method is language-agnostic.","Open-source LLMs such as Trillion-7B and Gemma2-9B can drive generation and filtering, so the pipeline does not depend on a proprietary model."],"supporting_citations":[{"why":"Supplies the retrieval-augmented generation method used to inject community slang into neutral sentences.","marker":"Lewis et al., 2020"},{"why":"Provides Sentence-BERT embeddings that form the retrieval index over 92,953 online comments.","marker":"Reimers and Gurevych, 2019"},{"why":"Defines G-Eval, the LLM-based scoring protocol used to measure offensiveness, consistency, and fluency.","marker":"Liu et al., 2023"},{"why":"Gives the prior definition of implicit offensiveness that the paper extends into trend-aligned slang.","marker":"Wiegand et al., 2021"},{"why":"Provides K-OMG, the main LLM-generated Korean dataset baseline for comparison.","marker":"Shin et al., 2023"},{"why":"Offers ToxiGen, a machine-generated implicit-hate dataset used as an English comparison for implicit offensiveness.","marker":"Hartvigsen et al., 2022"},{"why":"Supplies BEEP, a Korean news-comment toxicity dataset used as a generalization test set.","marker":"Moon et al., 2020"},{"why":"Supplies KOLD, a Korean offensive language dataset used as a test set for detoxification generalization.","marker":"Jeong et al., 2022"}],"fun_headline_variants":["Slang-savvy pipeline auto-builds Korean detox pairs","LLM-filtered slang yields implicit-offense Korean detox data","K/DA: Retriever + LLM = better Korean detox training pairs","Automated Korean detox: trend-aligned slang, implicit offense","Fresh slang + LLM filters: automated Korean detox pair factory"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole pipeline relies on GPT-4 Turbo being a reliable judge of what is consistent and implicitly offensive, so if its judgments are systematically different from human perception, the reported quality improvements would not carry over to real users.","fun_headline_variants_meta":{"raw":{"variants":["Slang-savvy pipeline auto-builds Korean detox pairs","LLM-filtered slang yields implicit-offense Korean detox data","K/DA: Retriever + LLM = better Korean detox training pairs","Automated Korean detox: trend-aligned slang, implicit offense","Fresh slang + LLM filters: automated Korean detox pair factory"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000239,"raw_usage":{"total_tokens":1471,"prompt_tokens":857,"completion_tokens":614,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":473,"completion_tokens_details":{"reasoning_tokens":524}},"tokens_in":473,"tokens_out":614,"duration_ms":6578,"temperature":1.0,"reasoning_tokens":524,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:59:28.588000+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A human study with a larger and more diverse annotator pool that asked whether GPT-4's filter decisions match the majority label better than chance, or a blind test where a detoxification model trained on human-annotated paired data outperforms the K/DA-trained model on a held-out set of naturally occurring offensive comments.","supporting_citations":[],"review_version":2}