{"id":"208e266a-3d86-40c0-b784-d035e4be0626","arxiv_id":"2502.01402","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Presents an open-source podcast annotation tool and a small annotated dataset (7 episodes, 1,960 utterances, 300 check-worthy claims) for claim detection and stance classification in English, Norwegian, and German.","lead":"An open-source tool that lets annotators fact-check podcasts while listening to them, along with a small multilingual dataset of podcast utterances labeled for claim detection and stance. It matters because most fact-checking resources are text-only, and this is one of the first attempts to produce labeled podcast data with fine-grained claim types.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claim of competitive utility rests on unvalidated gold labels and a tiny test set; reported F1 gaps have no confidence intervals.","rationale":"The paper's contribution is a tool plus a dataset, and the reader's conditional verdict already requires data release, inter-annotator agreement, and confidence intervals. My stress-test agrees that the most fragile point is the empirical demonstration of utility: the gold labels are filtered through unanimous agreement among three crowdworkers with no reported reliability, and the test set is far too small for the F1 differences in Table 2 to be meaningful. The tool itself appears to be a genuine engineering contribution with open-source code, so there is no fundamental objection to the resource. The missing dataset link is an additional verifiability problem that reinforces the conditional verdict. I would keep the conditional acceptance but make the conditions explicit: report IAA and retention rate, publish the dataset at a persistent URL, and provide bootstrap confidence intervals or significance tests for the model comparison. If those conditions are not met, the comparative claims should be treated as unverified. The reader's weakest assumption identified the same label-quality and small-test-set concern, so no verdict change is needed.","tokens_in":901,"tokens_out":3072,"duration_ms":67727,"concrete_test":"Obtain the raw per-annotator judgments for all three annotators and the final released dataset, then (1) compute Fleiss' kappa on the original triple annotations and report the retention rate after unanimous filtering; and (2) bootstrap 10,000 resamples over the 24 true claim-detection instances and the 50 stance instances to obtain 95% confidence intervals for each F1 in Table 2 and for the XLM-R versus GPT-4 differences. If kappa is below 0.6, the retention rate materially changes the label distribution, or the confidence intervals for the key F1 differences include zero, the comparative claims in Section 4 should be withdrawn or heavily qualified. Provide a persistent dataset URL in the same release; without it the released-resource claim cannot be tested.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the released annotations support claim detection and stance classification that are competitive with GPT-4. Two linked premises are load-bearing: the gold labels are reliable, and the Table 2 comparisons are statistically stable. Section 2.2 says each task was assigned to three annotators and \"unanimous labels retained to account for subjectivity,\" but no inter-annotator agreement, retention rate, or adjudication procedure is reported. If many non-unanimous judgments were discarded, the retained 1,960 utterances and the 50-instance stance test set are biased toward easy cases, making the gold labels unrepresentative. Table 3 shows only 24 check-worthy test instances, 18 refutes, and 32 supports. The key F1 differences (0.45 vs. 0.57 for true claims; 0.79 vs. 0.56 for supports) come from double-digit counts, so without confidence intervals or significance tests the claim that fine-tuned XLM-RoBERTa is competitive with GPT-4 is unsupported. The manuscript also says the annotated dataset is released but gives only a tool repository URL, not a dataset URL, so the released-resource claim is not yet verifiable. Section 2.2 defers annotation details to the self-cited master's thesis [1], which does not substitute for reporting agreement metrics in the paper.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an open-source web tool for transcribing and annotating podcasts, with simultaneous audio playback, and describes a multilingual (English, Norwegian, German) dataset built from 531 episodes of 38 podcasts, of which 7 episodes were annotated for check-worthy claims, claim types, fact-checking motivations, and stance/verification. The authors fine-tune XLM-RoBERTa-Large on the annotated utterances for binary claim detection and stance classification, and compare against GPT-4 in few-shot settings. The central claims are that the tool is the first designed for simultaneous audio playback and transcription annotation for podcasts, and that the released annotations allow small multilingual models to be competitive with GPT-4.","tokens_in":5777,"tokens_out":4500,"duration_ms":37063,"significance":"If the claims hold, the contribution is a useful new resource: an open tool for spoken-document annotation, a released multilingual podcast transcript corpus, and fine-grained claim/verification annotations that are currently scarce. The paper's strengths are that the tool is released as open source, the transcript collection is large (531 episodes, three languages), and the proposed annotation taxonomy is concrete and task-oriented. However, the experimental support for the comparative model claim is thin: the test sets are very small, no significance testing is reported, and the quality of the gold labels is not documented beyond a unanimous-agreement rule. The resource may still be valuable for the community, but the current evidence as presented is not yet sufficient to establish the stated competitive-performance conclusion.","major_comments":[{"comment":"The gold-label quality is not established. The paper states that each task was assigned to three annotators and 'unanimous labels retained to account for subjectivity,' but it does not report inter-annotator agreement, the number or proportion of labels discarded, or any adjudication procedure for non-unanimous cases. This is load-bearing because if a large fraction of non-unanimous judgments was discarded, the retained 1,960 utterances and the 176-instance test split would be biased toward easy instances, and the subsequent fine-tuning results in Table 2 would overstate model utility. Please report per-task agreement metrics (e.g., Fleiss' kappa or Krippendorff's alpha), retention rates, and an explicit adjudication policy, or otherwise justify the filtering rule.","section":"§2.2"},{"comment":"The comparison between fine-tuned XLM-RoBERTa and GPT-4 is statistically unsupported. Table 3 shows that the test set contains only 24 check-worthy utterance labels, 18 stance 'True/Supports' instances, and 32 'False/Refutes' instances, yet Table 2 reports F1 differences such as 0.45 vs. 0.57 for true claims and 0.79 vs. 0.56 for supports. With double-digit denominators, these differences are compatible with sampling noise. No confidence intervals, bootstrap estimates, or significance tests are provided. Please add uncertainty quantification (e.g., bootstrap CIs) and, where appropriate, a paired test such as McNemar's test, or explicitly reframe the results as exploratory with no comparative claim.","section":"§4, Table 2 and Table 3"},{"comment":"The claimed release of the annotated dataset is not verifiable from the manuscript. Section 1 gives only the tool repository URL (https://github.com/factiverse/factcheck-podcasts) and says 'we release transcripts for 531 episodes... alongside an annotated dataset,' but no dataset URL, archival link, license, or download procedure is provided. The abstract's phrase 'sample annotations' also leaves unclear whether the full annotations are released. Please specify the data repository, exact contents, licensing, and version, and ensure the URLs are resolvable.","section":"§1 (dataset release)"}],"minor_comments":[{"comment":"The manuscript repeatedly defers annotation details to the self-cited master's thesis [1]; the thesis cannot substitute for reporting agreement and filtering statistics in the paper itself, and the self-citation should be supported by the primary data in the paper.","section":"§2.2"},{"comment":"The column headers 'True' and 'False' in Table 3 are ambiguous because §4 defines stance labels as 'Supports' and 'Refutes.' Please unify the terminology and state clearly whether True=Supports and False=Refutes.","section":"§4, Table 3"},{"comment":"The transcription evaluation lacks details: it is not stated on what corpus, language, or audio conditions the Whisper error rates were computed, whether the transcripts were compared with human reference transcripts, or how the 'prompted' condition differs from the standard one. Please provide this context or remove the table if it is not comparable.","section":"§4, Table 4"},{"comment":"The fine-tuning setup is underspecified: no training hyperparameters, number of runs, seeds, early-stopping criteria, or label imbalance handling are reported, which makes it difficult to reproduce the XLM-RoBERTa results in Table 2.","section":"§4.1"},{"comment":"The sentence 'the pipeline supports over 90 languages (excluding co-reference resolution)' is presented without supporting evidence; the paper only evaluates Whisper on what appears to be a small sample, and no multilingual evaluation of the full pipeline is reported. Please either provide such evidence or soften the claim.","section":"§1"},{"comment":"The dataset analysis is based on only 7 annotated episodes, a limitation the paper acknowledges only in passing ('We plan larger-scale annotations in the future'); this scale should also be reflected in the abstract and conclusion, where the resource is described as an 'annotated dataset specifically for end-to-end fact-checking.'","section":"§3"}],"recommendation":"major_revision","confidential_remarks":"The paper is a resource/application paper for a venue that likely values the tool and dataset. The main risk is that the headline comparison with GPT-4 rests on a tiny test set and undocumented label quality; both are fixable in revision. If the authors can supply IAA/retention statistics, confidence intervals, and a working dataset URL, the contribution becomes solid. I would not reject on the basis of the current small-scale experiments, but I would not accept until the release and statistical support are confirmed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I\\'ll be blunt: this is a resource paper, and judged as that, it is basically sound. The genuinely new thing is a purpose-built podcast annotation tool that plays audio and lets the annotator correct transcripts, label check-worthy spans, and assign fine-grained claim types and fact-checking motivations in one interface. That combination does not exist in the open-source ecosystem, and the annotation scheme is more detailed than ClaimBuster or debate-focused check-worthiness datasets. The authors also transcribe 531 episodes across three languages and release the tool code. For the fact-checking community, that is a real contribution, especially since the Spotify Podcast Dataset is no longer public.\n\nThe paper does not oversell the scale. Section 3 is explicit that these are 7 episodes and that larger annotation is planned. But the experimental section overreaches. Table 2 reports F1 gaps (0.45 vs 0.57 for true claims, 0.79 vs 0.56 for supports) on a test set of 24 check-worthy, 18 refute, and 32 support instances. There are no confidence intervals or significance tests. Those differences are within the noise for counts that small. The claim that fine-tuned XLM-RoBERTa is \\\"competitive with GPT-4\\\" is therefore not supported by the data. It is a plausible preliminary observation, but the paper should either add intervals or soften the wording.\n\nThe more load-bearing soft spot is the gold-label reliability. Section 2.2 says each task was assigned to three annotators and unanimous labels were retained, but it reports no agreement statistics and no retention rate. If unanimity discarded many judgments, the surviving labels are biased toward easy cases, which would weaken both the dataset and the model comparison. The authors point to a master\\'s thesis for details, but a resource paper should include that information on the page.\n\nAlso, the abstract says \\\"we release the annotated podcast transcripts and sample annotations,\\\" but the only URL is the tool repository. No dataset link appears. That is a concrete missing artifact for a paper whose contribution is a dataset.\n\nNone of this is fatal. The tool and the annotation scheme stand on their own, and the experiments are explicitly preliminary. The missing agreement numbers and dataset URL are fixable in revision. I\\'d send this to review. The right outcome is a conditional acceptance that requires releasing the data, reporting inter-annotator agreement, and either adding confidence intervals or clearly labeling the comparisons as illustrative.","headline":"A genuinely useful podcast annotation tool and a small but well-structured dataset, with a model comparison that is too underpowered to support its headline claim.","tokens_in":6242,"tokens_out":2979,"would_cite":false,"duration_ms":28185,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims to build the first podcast annotation tool that plays audio while showing the transcript, and that the released multilingual labels let a fine-tuned XLM-RoBERTa-Large compete with GPT-4 on claim and stance detection.","keywords":["podcasts","fact-checking","data annotation","claim detection","stance classification","multilingual datasets","crowdsourcing","spoken content"],"falsifier":"Re-annotate a held-out sample of podcast utterances with more than three independent workers, measure inter-annotator agreement, and count how many labels the unanimous rule discards; then rerun the fine-tuned XLM-RoBERTa-Large and the GPT-4 prompts on a larger test set built from those labels. If the stance-detection advantage (0.79 versus 0.56 F1) shrinks or flips, or if a large share of labels is discarded, the dataset is too weak to support the comparative conclusions.","tokens_in":5289,"feed_emoji":"🎙️","tokens_out":11346,"duration_ms":84136,"temperature":0.7,"pith_summary":"The paper tries to establish that podcast fact-checking can be supported end to end by a lightweight open-source pipeline: transcribe the audio, assign speaker labels, resolve mentions, and let crowdworkers mark claims while listening to the recording. It claims this is the first tool designed for podcast annotation that plays audio and shows the transcript at the same time, and it releases transcripts for 531 episodes in English, Norwegian, and German plus fine-grained claim annotations for selected episodes from seven podcasts. The paper then argues the annotations are useful by fine-tuning XLM-RoBERTa-Large and showing it competes with GPT-4 in few-shot settings on claim detection and stance classification. A sympathetic reader would care because spoken fact-checking needs both the annotation interface and the labeled data, which have been missing since the large podcast corpus used by earlier work became inaccessible.","feed_headline":"Open tool transcribes, annotates, and fact-checks podcasts in one pass","feed_subtitle":"Fine-grained labels from 531 episodes in three languages let small models match GPT-4 on stance.","key_machinery":"The load-bearing mechanism is the annotation interface itself: a web application that plays the podcast audio while displaying a Whisper transcription with word-level timestamps, speaker labels from a diarization model, sentence splits, and coreference-resolved utterances, so annotators can label claims in real time. Around it sits a crowdsourcing pipeline that assigns each utterance to three crowdworkers and keeps only unanimous labels, and an evaluation pipeline that fine-tunes XLM-RoBERTa-Large on the resulting binary claim-detection and stance-classification tasks. The interface is what makes the claimed first-of-its-kind contribution possible.","core_discovery":"The central discovery, on the paper's own terms, is that simultaneous audio-plus-transcription annotation changes what can be collected from podcasts: annotators can correct transcription errors, mark check-worthy claims and claim spans, choose claim types and fact-checking motivations, and assign verdicts while the audio plays, producing fine-grained labels that generic text tools cannot capture. With those labels, the paper reports that a fine-tuned XLM-RoBERTa-Large reaches a weighted F1 of 0.85 for claim detection and 0.67 for stance detection, while GPT-4 reaches 0.86 and 0.55, respectively; GPT-4 is better at identifying true check-worthy claims (0.57 versus 0.45), while XLM-RoBERTa is better at identifying supporting evidence (0.79 versus 0.56). The paper reads this as evidence that smaller multilingual models are competitive for podcast fact-checking and that the released data can support both tasks.","pith_inferences":["The unanimous-only labeling rule probably makes the gold standard conservative: ambiguous utterances are discarded, so the reported F1 values are likely measured on the easier instances and may overstate how the models would perform on borderline real-world claims.","Because numerical claims make up 62.6 percent of the check-worthy labels, a natural next step the paper does not pursue is a dedicated numerical-claim verifier that checks statistics against external sources, and the dataset would support training such a system.","A testable extension is to re-annotate a sample with more workers and measure inter-annotator agreement; if agreement is low, the three-worker unanimous rule would need to be replaced by majority voting or adjudication.","The same simultaneous audio-and-transcription design could be applied to live audio fact-checking by feeding a real-time transcript into the annotation interface, connecting this offline dataset work to streaming fact-checking systems the paper cites."],"forward_implications":["If the tool is adopted, podcast fact-checking no longer requires a separate transcription step: annotators can correct automatic speech recognition errors and mark claims in one pass, making the process faster and cheaper than listening and annotating separately.","The released transcripts and annotations give the community a multilingual resource for claim detection and stance classification in English, Norwegian, and German, filling the gap left by the closed large podcast corpus.","Fine-grained labels for claim types and motivations allow downstream systems to prioritize numerical claims and public-interest claims, which together dominate the check-worthy annotations.","The competitive performance of XLM-RoBERTa-Large suggests that specialized smaller models can serve as a lightweight alternative to large API-based language models for podcast fact-checking.","The claim-type taxonomy and the annotation schema can be reused for other long-form audio domains, such as lectures, interviews, and meeting recordings."],"supporting_citations":[{"why":"It supplies the detailed annotation-process description that the paper points to for the crowdsourcing and labeling rules.","marker":"[1]"},{"why":"It is the large podcast corpus whose inaccessibility motivates the need for a new open podcast dataset.","marker":"[2]"},{"why":"It is an earlier check-worthy claim dataset from political debates that lacks the fine-grained claim-type and verification annotations this paper adds.","marker":"[3]"},{"why":"It is audio-based check-worthy claim detection work whose scope is political speech rather than podcasts.","marker":"[4]"},{"why":"It provides the coreference resolution that turns pronouns into self-contained utterances ready for fact-checking.","marker":"[6]"},{"why":"It provides speaker diarization used to assign speaker labels to transcribed audio segments.","marker":"[7]"},{"why":"It supplies the automatic speech recognition model that produces the transcripts with word-level timestamps and punctuation.","marker":"[8]"},{"why":"It is a fact-checking text editor that is not designed for data annotation, used to position the new tool's contribution.","marker":"[10]"},{"why":"It supplies the prompting strategy used for the GPT-4 few-shot comparison in the experiments.","marker":"[11]"},{"why":"It is a live fact-checking system for audio streams that does not support annotation, used to position the new tool's contribution.","marker":"[13]"}],"fun_headline_variants":["Real-time podcast annotation tool rivals GPT-4 on fact-checking","Tool transcribes and annotates podcasts live, enabling claim checks","Multilingual podcast dataset helps small model match GPT-4 on claims","Annotate claims as you listen: new tool for podcast fact-checking","Small model beats GPT-4 on stance, trails on check-worthy claims"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire label set rests on the assumption that unanimous agreement among just three crowdworkers, with no reported inter-annotator agreement and no count of discarded non-unanimous labels, produces gold-standard annotations strong enough to support the model comparisons.","fun_headline_variants_meta":{"raw":{"variants":["Real-time podcast annotation tool rivals GPT-4 on fact-checking","Tool transcribes and annotates podcasts live, enabling claim checks","Multilingual podcast dataset helps small model match GPT-4 on claims","Annotate claims as you listen: new tool for podcast fact-checking","Small model beats GPT-4 on stance, trails on check-worthy claims"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000919,"raw_usage":{"total_tokens":3909,"prompt_tokens":876,"completion_tokens":3033,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":492,"completion_tokens_details":{"reasoning_tokens":2939}},"tokens_in":492,"tokens_out":3033,"duration_ms":20370,"temperature":1.0,"reasoning_tokens":2939,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T15:25:25.983391+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-annotate a held-out sample of podcast utterances with more than three independent workers, measure inter-annotator agreement, and count how many labels the unanimous rule discards; then rerun the fine-tuned XLM-RoBERTa-Large and the GPT-4 prompts on a larger test set built from those labels. If the stance-detection advantage (0.79 versus 0.56 F1) shrinks or flips, or if a large share of labels is discarded, the dataset is too weak to support the comparative conclusions.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the detailed annotation-process description that the paper points to for the crowdsourcing and labeling rules."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It is the large podcast corpus whose inaccessibility motivates the need for a new open podcast dataset."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It is audio-based check-worthy claim detection work whose scope is political speech rather than podcasts."},{"cited_title":"F-coref: Fast, Accurate and Easy to Use Coreference Resolution","cited_arxiv_id":"2209.04280","evidence_quote":"It provides the coreference resolution that turns pronouns into self-contained utterances ready for fact-checking."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It is a fact-checking text editor that is not designed for data annotation, used to position the new tool's contribution."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the prompting strategy used for the GPT-4 few-shot comparison in the experiments."},{"cited_title":"LiveFC: A System for Live Fact-Checking of Audio Streams","cited_arxiv_id":"2408.07448","evidence_quote":"It is a live fact-checking system for audio streams that does not support annotation, used to position the new tool's contribution."}],"review_version":1}