{"id":"d205a678-3ac8-4e11-854b-e55e22ffa481","arxiv_id":"2603.17061","paper_version":2,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"A content-controlled, on-device smartphone protocol yields usable prosodic features at scale (N=560, 9,877 clips), with strong sex classification but weak concurrent affect prediction.","lead":"Researchers built a smartphone protocol that has people read fixed positive, neutral, or negative sentences, extracts acoustic features on-device with OpenSMILE, and immediately deletes the raw audio. It lets labs collect large-scale prosody data while controlling word content and reducing privacy risk from identifiable voice recordings.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Unverifiable lexical fidelity is the softest point for the content-control claim, but does not undercut feasibility or the speaker-signal evidence.","rationale":"The reader correctly isolates lexical fidelity as the weakest assumption and still recommends ACCEPT. That judgment is sound: the strongest empirical claims (scale feasibility, usable on-device features after QC, good compliance, speaker-informative signal via sex classification and non-degenerate diagnostics) do not require perfect verbatim compliance—only that the pipeline yields analyzable, speaker-stable prosodic summaries under realistic conditions. Affect prediction is weak and honestly reported; residual privacy risks of features are discussed. No internal contradiction or overclaim appears. The fidelity gap is real for the “content-controlled” framing but is already transparent in the Limitations and does not justify moving off ACCEPT. The proposed ASR pilot would tighten the claim for future work without altering the present methodological contribution.","tokens_in":10553,"tokens_out":513,"duration_ms":23000,"concrete_test":"Run a small pilot (n≈50) with the same protocol plus a lightweight on-device ASR keyword/phrase matcher (as suggested in §4.2) that scores word-overlap or exact-sentence detection without storing audio or transcripts. Report the fraction of recordings exceeding a high-match threshold (e.g., ≥80 % word overlap). If that fraction is substantially below ~85 %, the content-control claim needs stronger qualification; if high, the concern is largely mitigated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The protocol’s central value is content-controlled collection that standardizes lexical content (including valence) so prosody can be isolated from semantics (§1, §2.1). That isolation requires participants to read the displayed sentences essentially verbatim. QC in §3.2 only drops clips with low voicing probability, few/short voiced segments, or non-positive HNR; raw audio is deleted immediately, so lexical fidelity cannot be checked post hoc (§4.2). Occasional paraphrasing, skipping, or disfluency would reintroduce semantic variance and weaken the very confound the design claims to remove. The paper acknowledges this but supplies no quantitative bound on the deviation rate. Feasibility, compliance (~68 % initiation, 97 % completion once started), acoustic diagnostics, and ~92 % blocked-CV sex accuracy remain well supported; the content-control claim itself is the least secure link.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces and evaluates a smartphone protocol for collecting prosodic speech data in everyday life that standardizes lexical content (including prompt valence) via scripted read-aloud sentences, extracts eGeMAPS and ComParE features on-device with OpenSMILE, immediately deletes raw audio, and transmits only feature vectors. Deployed in a quota-matched German panel (final N=560; 9,877 recordings after QC), it reports compliance (67.8% initiation; 96.8% completion once started), acoustic diagnostics and mixed-effects condition contrasts, and diagnostic prediction tasks under participant-blocked CV: ~92% balanced accuracy for speaker sex and modest performance for concurrent single-item valence/arousal. The authors position the protocol as a field-ready, privacy-preserving baseline for content-controlled prosody sampling.","tokens_in":10821,"tokens_out":648,"duration_ms":5770,"significance":"If the feasibility and signal-characterization claims hold, the work supplies a practical, reproducible template that jointly addresses the prosody–semantics confound and raw-audio privacy barriers that currently limit large-scale in-the-wild prosody research. Strengths include the large quota sample, explicit QC rules, FDR-corrected mixed-effects contrasts, participant-blocked RF CV, and planned release of analysis scripts plus the Android pipeline. The sex-classification positive control and non-degenerate acoustic diagnostics after QC give concrete evidence that usable speaker-informative features can be obtained under realistic smartphone conditions. The residual privacy discussion and the framing of the protocol as a within-participant baseline for unconstrained speech are useful for the field.","major_comments":[{"comment":"§2.1 and §4.2: The central content-control claim (standardizing lexical content, including valence, so that prosody can be isolated from semantics) rests on participants reading the displayed sentences essentially verbatim. Because raw audio is deleted immediately, lexical fidelity cannot be verified post hoc; QC in §3.2 only flags low voicing probability, few/short voiced segments, or non-positive HNR. Occasional paraphrasing, skipping, or disfluency would reintroduce semantic variance. The paper acknowledges this limitation but supplies no quantitative bound or sensitivity analysis on the deviation rate. A concrete estimate (e.g., pilot ASR keyword-match rates, or a small retained-audio validation subset) would substantially strengthen the content-control claim without changing the privacy design.","section":null},{"comment":"§3.4: Affect prediction is reported as modest (eGeMAPS R² ≈ 0.03–0.04 for arousal/valence; ComParE similar) with no condition differences. While the authors correctly treat these tasks as diagnostic rather than primary claims, the abstract and introduction still frame the protocol as enabling prosodic analysis of affective states. The manuscript would be clearer if it more sharply separated the strong feasibility/speaker-signal evidence from the weak affect-signal evidence, and if it quantified how much of the modest performance is attributable to single-item EMA reliability versus scripted-read limitations.","section":null}],"minor_comments":[],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a methods paper that packages known pieces—valence-balanced read-aloud sentences, OpenSMILE eGeMAPS/ComParE on device, immediate raw-audio deletion, EMA prompting—into one Android protocol and then runs it at real scale (N=560 after QC, 9,877 clips). That packaging and the empirical characterization are what is new. Prior work had unconstrained wild speech, raw audio, or lab scripts; this puts the controlled, privacy-first version in the field and reports what you get.\n\nWhat it does well is concrete. Compliance is usable (68% initiation, 97% completion once started). QC rules are explicit (voicing, voiced-segment rate/length, HNR>0). Acoustic diagnostics look non-degenerate. Mixed-effects condition contrasts with FDR correction show only modest prompt-set shifts and solid speaker ICCs. The sex classifier under participant-blocked CV (~92% balanced accuracy) is a clean positive control that the features carry speaker-informative signal. Affect prediction is weak and they say so plainly—no overclaim. Scripts and aggregated outputs are promised; raw features stay private for good reason.\n\nThe soft spot the stress-test flags is real but bounded: content control assumes people read the sentences essentially verbatim, and with audio deleted you cannot check lexical fidelity—only that something voiced and harmonic was present. The paper owns this in the limitations and sketches ASR keyword checks as future work. That weakens the strongest version of the “prosody isolated from semantics” claim; it does not undercut feasibility, compliance, or the speaker-signal evidence. Other limits (single-item affect, engineered features vs embeddings, heterogeneous phone mics) are ordinary for this setting and proportionately discussed.\n\nMath is light and appropriate (LME, RF with blocked CV). Citations track the right lab-to-wild and privacy literature. No circularity: labels are external (sex, concurrent affect).\n\nWho it is for: people building digital phenotyping, affective computing, or speech-in-the-wild pipelines who need a reproducible privacy-first baseline module. Not a theory paper. I would send it to peer review; a serious referee can tighten the fidelity discussion and the future-work checks without rewriting the contribution. Worth engaging if you collect voice in the wild under GDPR-like constraints.","headline":"Field-ready methods package that actually ships: content control + on-device OpenSMILE + delete-raw, evaluated at useful scale; feasibility and speaker signal hold, lexical fidelity is the softest claim.","tokens_in":11432,"tokens_out":583,"would_cite":true,"duration_ms":10275,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A content-controlled smartphone protocol can collect privacy-preserving prosodic features at scale by having people read fixed sentences while processing audio only on the device.","keywords":["prosody","speech data collection","smartphones","privacy","on-device processing","ecological momentary assessment","acoustic features"],"falsifier":"Re-run the protocol while temporarily retaining raw audio for a held-out subset; if a large share of clips that pass the voicing and HNR filters are paraphrased, silent, or otherwise non-compliant, or if participant-blocked sex classification falls far below the reported ~92% balanced accuracy on a new matched sample, the claim of reliable content-controlled, speaker-informative signal fails.","tokens_in":11476,"feed_emoji":"📱","tokens_out":818,"duration_ms":25150,"temperature":0.7,"pith_summary":"Everyday speech research is hard because what people say confounds how they say it, raw audio is privacy-sensitive, and recording tasks often get skipped. This paper introduces a smartphone protocol that shows participants fixed, valence-balanced sentences to read aloud, extracts standard acoustic features on the phone, deletes the raw recording immediately, and transmits only the derived features. In a large field deployment (560 participants, 9,877 retained recordings) compliance was good and the resulting features passed basic acoustic quality checks. Diagnostic models recovered speaker sex with about 92% balanced accuracy under participant-blocked evaluation, while prediction of momentary valence and arousal was only modest. The result is a practical, field-ready way to gather controlled prosody without storing identifiable voice audio.","feed_headline":"Phone app collects speech prosody without storing audio","feed_subtitle":"Fixed sentences and on-device features enable private, large-scale field studies of how people speak.","key_machinery":"The content-controlled, privacy-first smartphone protocol: at each evening prompt, participants read three valence-matched sentences (positive, neutral, negative); OpenSMILE extracts eGeMAPS and ComParE features on-device; the raw WAV is deleted at once; only feature vectors leave the phone. It standardizes lexical content (including valence) while capturing delivery variation and removes raw audio from the research pipeline.","core_discovery":"The authors claim that a content-controlled, privacy-first smartphone protocol—scripted read-aloud sentences of controlled lexical valence, on-device extraction of standard acoustic feature sets, immediate deletion of raw audio, and transmission of features only—is feasible in everyday life at scale. Deployed with 560 participants and 9,877 retained recordings, it produced good compliance, analyzable prosodic summaries with substantial speaker-level stability, strong sex classification under blocked cross-validation, and weaker prediction of concurrent self-reported affect.","pith_inferences":["Lightweight on-device keyword checks (without storing audio or text) could close the unverifiable lexical-fidelity gap the authors note.","The same design can serve as a private reference track for calibrating passive continuous-audio embeddings collected in parallel.","Pairing the scripted module with brief free-speech or acted probes in the same session would quantify how much affective signal content control removes.","Because engineered features can still be re-identifying, future deployments may need feature-level anonymization as inference models improve."],"forward_implications":["Scripted read-aloud modules can be added as semantic baselines inside larger in-the-wild speech studies.","Prosody panels become practical under strict privacy rules because raw audio never leaves the device.","Speaker-stable metrics from the protocol can calibrate analyses of unconstrained daily speech.","On-device features carry enough speaker signal for basic demographic recovery under blocked evaluation.","Affect prediction stays limited under fixed scripts, so the method is better used as a control baseline than as a primary emotion sensor."],"fun_headline_variants":["Smartphone protocol captures prosody from scripted sentences, deletes audio","Privacy-first app extracts on-device speech features at field scale","Content-controlled read-alouds yield private prosody data from 560 users","On-device features from 9877 recordings enable sex and affect prediction","Scripted sentences standardize content while logging everyday prosody privately"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The protocol assumes people actually read the displayed sentences as written, which cannot be verified later because the raw audio is deleted immediately.","fun_headline_variants_meta":{"raw":{"variants":["Smartphone protocol captures prosody from scripted sentences, deletes audio","Privacy-first app extracts on-device speech features at field scale","Content-controlled read-alouds yield private prosody data from 560 users","On-device features from 9877 recordings enable sex and affect prediction","Scripted sentences standardize content while logging everyday prosody privately"]},"model":"grok-4.5","effort":"low","cost_usd":0.00303,"raw_usage":{"total_tokens":1047,"prompt_tokens":717,"num_sources_used":0,"completion_tokens":80,"cost_in_usd_ticks":30300000,"prompt_tokens_details":{"text_tokens":717,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":250,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":717,"tokens_out":80,"duration_ms":3179,"temperature":1.0,"reasoning_tokens":250,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-13T23:21:43.509339+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run the protocol while temporarily retaining raw audio for a held-out subset; if a large share of clips that pass the voicing and HNR filters are paraphrased, silent, or otherwise non-compliant, or if participant-blocked sex classification falls far below the reported ~92% balanced accuracy on a new matched sample, the claim of reliable content-controlled, speaker-informative signal fails.","supporting_citations":[],"review_version":1}