Pith. sign in

REVIEW 3 cited by

Detecting AI-Generated Sentences in Human-AI Collaborative Hybrid Texts: Challenges, Strategies, and Insights

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.03506 v4 pith:IYJDQ2ZA submitted 2024-03-06 cs.CL cs.AI

classification cs.CLcs.AI
keywords hybridtextssegmentstextai-generatedauthorshipsentencesdetecting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This study explores the challenge of sentence-level AI-generated text detection within human-AI collaborative hybrid texts. Existing studies of AI-generated text detection for hybrid texts often rely on synthetic datasets. These typically involve hybrid texts with a limited number of boundaries. We contend that studies of detecting AI-generated content within hybrid texts should cover different types of hybrid texts generated in realistic settings to better inform real-world applications. Therefore, our study utilizes the CoAuthor dataset, which includes diverse, realistic hybrid texts generated through the collaboration between human writers and an intelligent writing system in multi-turn interactions. We adopt a two-step, segmentation-based pipeline: (i) detect segments within a given hybrid text where each segment contains sentences of consistent authorship, and (ii) classify the authorship of each identified segment. Our empirical findings highlight (1) detecting AI-generated sentences in hybrid texts is overall a challenging task because (1.1) human writers' selecting and even editing AI-generated sentences based on personal preferences adds difficulty in identifying the authorship of segments; (1.2) the frequent change of authorship between neighboring sentences within the hybrid text creates difficulties for segment detectors in identifying authorship-consistent segments; (1.3) the short length of text segments within hybrid texts provides limited stylistic cues for reliable authorship determination; (2) before embarking on the detection process, it is beneficial to assess the average length of segments within the hybrid text. This assessment aids in deciding whether (2.1) to employ a text segmentation-based strategy for hybrid texts with longer segments, or (2.2) to adopt a direct sentence-by-sentence classification strategy for those with shorter segments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Segmenting Human-LLM Co-authored Text via Change Point Detection

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    A weighted change point detection algorithm locates human- and LLM-written segments in mixed text and attains minimax-optimal localization under heterogeneous sentence scores.

  2. Discovering Coordinated Processes From Social Online Networks

    cs.SI 2025-06 conditional novelty 6.0 of 10

    Retweet event logs mined into stochastic Petri nets show structural and timing differences between coordinated bot and ordinary user behavior.

  3. Do people rely on ChatGPT more than their peers to detect deepfake news?

    econ.GN 2026-08 conditional novelty 5.0 of 10

    In a lab deepfake-detection task, students shifted more toward ChatGPT's advice than toward peers' advice (weight-of-advice 0.59 vs 0.33), though in 2025 sessions they trusted linguistic experts slightly more than ChatGPT.

Pith tools