Pith. sign in

REVIEW 5 cited by

HC3 Plus: A Semantic-Invariant Human ChatGPT Comparison Corpus

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.02731 v4 pith:WMOO6QJL submitted 2023-09-06 cs.CL cs.AI

classification cs.CLcs.AI
keywords taskssemantic-invariantchatgptdetectingtextchallengingfine-tuninginstruction
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

ChatGPT has garnered significant interest due to its impressive performance; however, there is growing concern about its potential risks, particularly in the detection of AI-generated content (AIGC), which is often challenging for untrained individuals to identify. Current datasets used for detecting ChatGPT-generated text primarily focus on question-answering tasks, often overlooking tasks with semantic-invariant properties, such as summarization, translation, and paraphrasing. In this paper, we demonstrate that detecting model-generated text in semantic-invariant tasks is more challenging. To address this gap, we introduce a more extensive and comprehensive dataset that incorporates a wider range of tasks than previous work, including those with semantic-invariant properties. In addition, instruction fine-tuning has demonstrated superior performance across various tasks. In this paper, we explore the use of instruction fine-tuning models for detecting text generated by ChatGPT.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 6 citations worldwide. Full citation record

  1. Telescope: Improving Zero Shot Detection of LLM Generated Content By Measuring Token Repetition Probability

    cs.CL 2026-07 accept novelty 7.0 of 10

    Telescope Perplexity, the average negative log probability a reference LM assigns to each token immediately after seeing it, yields strong zero-shot LLM-text detection by probing an early-training aversion to repetition.

  2. MAGA-Bench: Machine-Augment-Generated Text via Alignment Detection Benchmark

    cs.CL 2026-01 conditional novelty 6.0 of 10

    Adding human-alignment augmentation (roleplaying, BPO, self-refine, RLDF) to machine-generated text both fools existing detectors and improves the generalization of detectors fine-tuned on it.

  3. The Arabic AI Fingerprint: Stylometric Analysis and Detection of Large Language Models Text

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Arabic text written by LLMs carries detectable stylometric signatures, and fine-tuned XLM-RoBERTa detectors reach near-perfect F1 on academic abstracts but degrade on social media.

  4. Stylometry recognizes human and LLM-generated texts in short samples

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Stylometric features and tree-based classifiers separate human-written Wikipedia summaries from LLM-generated texts with high cross-validated accuracy on a new seven-class benchmark, though performance drops on other ...

  5. StylOch at PAN: Gradient-Boosted Trees with Frequency-Based Stylometric Features

    cs.CL 2025-07 conditional novelty 3.0 of 10

    A non-neural stylometric detector using LightGBM on spaCy-derived features reached a final mean score of 0.897 on the PAN 2025 task, below the 0.922 TF-IDF SVM baseline.

Pith tools