{"id":"e237f5bf-c446-42fb-9901-72e769bf169e","arxiv_id":"2506.04848","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"MockConf provides manual span- and word-level alignments for 7 hours of Czech-centered simultaneous interpreting, together with an annotation tool and baseline alignment results.","lead":"This paper introduces MockConf, a 7-hour dataset of student simultaneous interpreting in five European languages, with manual word- and span-level alignments, plus a web annotation tool and automatic alignment baselines. It is a resource for studying how interpreters compress, reformulate, and sometimes distort speech in real time.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The corpus-level analyses and baselines rest on single-annotator gold labels, yet the only inter-annotator check shows target label Kappa 0.25 and exact link agreement 15-30%; without broader double annotation, the quantitative claims are annotator-dependent.","rationale":"The paper is a resource paper, and its strongest claim is that a usable, manually aligned interpreting corpus now exists. The reader's verdict is CONDITIONAL, citing transcript fidelity and gold-standard reliability. I agree with the second premise as the more load-bearing one: even with perfect transcripts, the alignment and span-label layers are what distinguish MockConf from prior corpora, and those layers show fair-to-moderate inter-annotator agreement on the one recording where agreement was measured. The absence of per-annotator control in the corpus analysis means the reported label distributions and baseline scores could be artifacts of annotator style; Figure 2 documents substantial style differences. This is not a claim of fraud or sloppiness: the authors disclose the relevant numbers and discuss the difficulty of the task. Rather, it is a correctness risk for every downstream use of the resource. The concrete test would settle whether the quantitative claims survive reannotation. Since the reader already arrived at CONDITIONAL for essentially this reason, my pass does not change the verdict; it sharpens the condition by pointing to a specific, inexpensive analysis of the released double-annotated recording and a stratified reannotation protocol.","tokens_in":20367,"tokens_out":4473,"duration_ms":53544,"concrete_test":"Using the released MockConf data, recompute the main corpus statistics (Table 5 label proportions, Figure 3 length ratios, relay vs. direct differences) separately for Annotator 2 and Annotator 3 on the double-annotated recording. If the two annotators differ by more than 10 relative percentage points on any major label, single-annotator labels are too unstable to support the paper's quantitative analyses. Independently, recruit a second annotator for a stratified sample covering at least one recording per source-target direction and per primary annotator, and compute Kappa and exact-link agreement; if target label Kappa remains near 0.25 and exact link agreement below 30%, the gold standard is too unreliable for the benchmarking use case.","verdict_should_be":"UNCHANGED","load_bearing_attack":"MockConf's central value is as a manually aligned reference corpus for analysis and evaluation. That value requires the gold annotations to be stable across annotators. Section 3.1 reports the only reliability check: one cs-en recording annotated twice. Cohen's Kappa is 0.56/0.57 for source/target segmentation, 0.41/0.25 for labels, and exact alignment-link agreement is between 14.85% and 30.46% (Table 4). Figure 2 further shows large annotator-specific span-length styles, with Annotator 4 producing nearly twice the span length of Annotators 3 and 5, and Annotator 2 producing fewer links. Because each recording is labeled by a single annotator and the dev/test label statistics (Table 5, Figure 3, relay vs. direct comparisons) are aggregated over those single annotations, the reported corpus-level findings and the Table 7 baseline scores may reflect annotator style rather than interpreting behavior. No analysis in the paper controls for annotator identity, and the one double-annotated recording is not stratified across languages, directions, or the five annotators. The transcript-revision process (Section 2.1) is a secondary but related risk: non-Czech sides were checked only by native Czech speakers with self-reported proficiency, so ASR errors in those languages could propagate into the alignments. The paper is transparent about these limitations, but the limitations directly bound the central claim of a usable, manually aligned corpus.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents MockConf, a new publicly released corpus of simultaneous interpreting collected from student mock conferences. The corpus comprises roughly 7 hours of recordings in Czech, English, French, German, and Spanish, with transcripts revised from WhisperX output and manual annotations at both span and word levels. A seven-label taxonomy (Translation, Paraphrase, Summarization, Generalization, Factual/Uninformative Addition, Replacement) is applied to spans, and word links are marked as sure or possible. The authors also release InterAlign, a web-based annotation tool designed for long, non-sentence-aligned inputs. They report descriptive analyses of span lengths, label distributions, relay versus direct interpreting, and multi-track interpreting, and they propose evaluation metrics and baselines for automatic span and word alignment, comparing a BERTAlign-based pipeline (with optional sub-segmentation and label classification) against random baselines and against inter-annotator agreement.","tokens_in":20549,"tokens_out":3726,"duration_ms":45058,"significance":"If the annotations are reliable, MockConf fills a genuine gap: it is, to my knowledge, the first publicly released simultaneous-interpreting corpus with manual span-level labels and word-level alignment links, and the InterAlign tool addresses a real practical need. The paper is transparent in reporting inter-annotator agreement, including low agreement figures, and it honestly includes random baselines. The resource and tool will likely be useful for corpus-based interpreting studies, for benchmarking alignment methods, and for educational monitoring. However, the significance is currently bounded by concerns about annotation stability and transcript fidelity, which the paper itself documents but does not resolve; the quantitative corpus analyses and baseline comparisons inherit those uncertainties.","major_comments":[{"comment":"The inter-annotator reliability evidence is thin and partly discouraging: only one cs-en recording was double-annotated, target-side label agreement is Cohen's Kappa 0.25, and exact alignment-link agreement ranges from 14.85% to 30.46% depending on the reference annotator. Because every other recording in the corpus is annotated by a single annotator, the corpus-level analyses in Sections 3.2-3.4 and the baseline scores in Table 7 aggregate annotations that may be strongly influenced by individual annotator style (as Figure 2 itself shows for span lengths and link counts). The paper should either provide a quantitative control for annotator identity (for example, mixed-effects models or per-annotator breakdowns of the key statistics), or double-annotate a stratified sample of recordings and show that the reported trends and baseline rankings are stable across annotators. Without this, the central claim that MockConf is a usable reference corpus for evaluating automatic alignment is not fully supported.","section":"Section 3.1, Table 4"},{"comment":"The label classifier is trained on 80% of the development set ('we use devset for training where we take 80% of devset for actual training and 20% as held-out data for the evaluation during the training'), yet Table 7 reports results for the full development set. If the development-set rows for BA+sub+lab are computed on data that include the training portion, the label-match scores are in-sample estimates and overstate performance. This should be clarified: either the devset rows must be restricted to the 20% held-out portion, or the evaluation should use cross-validation on the development set. The same issue may affect the development-set label F1 and accuracy numbers reported in Table 7.","section":"Section 4.1 and Section 4.4, Table 7"},{"comment":"The transcriptions are produced by WhisperX and then manually revised by native Czech speakers with self-reported proficiency in English, French, German, and Spanish. Since every alignment and downstream analysis is performed on these transcripts, undetected ASR errors in the non-Czech sides will propagate into the span labels, word links, and all corpus statistics. The paper should report the revisers' actual language proficiency or, more usefully, validate transcript fidelity on a sample with a native-speaker check per language and direction, and state the resulting error rates. This is particularly important for the French, German, and Spanish recordings, where the revision pool may be small and self-reported proficiency is not externally verified.","section":"Section 2.1"}],"minor_comments":[{"comment":"The label 'Summariaztion' is a typo and should read 'Summarization'.","section":"Table 2"},{"comment":"The caption says 'ISO-632-2'; the correct standard is ISO 639-2.","section":"Table 1 caption"},{"comment":"The phrase 'e.i.' should be 'i.e.' in 'fully anonymized, e.i. they do not contain'.","section":"Ethics Statement"},{"comment":"The random baseline is said not to use Reformulation or Replacement labels, which means its label-match scores are not a fully comparable lower bound; this should be stated where the baseline is introduced, not only in the Limitations section.","section":"Section 4.3 and Limitations"},{"comment":"The claim that the span-length distribution 'seems to be uniform' would be more convincing with error bars or a statistical test, especially given the small corpus size and the annotator variability documented in Figure 2.","section":"Section 3.2, Figure 3"}],"recommendation":"major_revision","confidential_remarks":"The dataset and tool are potentially valuable contributions, and the authors are transparent about limitations. The main risk is that the single-annotator gold standard, with low measured agreement, may undermine the quantitative claims unless additional reliability evidence or annotator-identity controls are provided. The development-set leakage issue, if confirmed, is a straightforward but important correction. I would support publication after these points are addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"MockConf is a genuine new resource: a Czech-centered simultaneous interpreting corpus with manual span- and word-level alignments (including relay tracks), plus the InterAlign tool and honest baselines. That combination doesn't exist in the prior corpora (EPIC, EPTIC, ESIC, NAIST-SIC-aligned), so the paper fills a real gap.\n\nWhat it does well: the corpus and tool are released, the annotation guidelines are detailed (Appendix D), and the baselines include a random baseline, which is rare and refreshing. The authors are transparent about their own limitations, and the descriptive analyses are framed as observations, not definitive findings. The citation pattern looks fine, and the evaluation is anchored to human annotations rather than circular.\n\nThe soft spots, in order of importance. First, the gold standard is single-annotator for every recording except one, and that one double-annotated recording shows modest segmentation agreement but low label agreement on the target side (Kappa 0.25) and exact link agreement between 15-30%. That makes the corpus-level statistics in Section 3 and the baseline scores in Table 7 potentially annotator-dependent. The authors acknowledge this and discuss annotator style differences, but they don't control for annotator identity in any analysis, and they don't provide confidence intervals on the descriptive claims. Second, transcript revision for non-Czech languages was done by Czech natives with self-reported proficiency, so ASR errors could leak into the alignments; this is a real but secondary risk. Third, the dev set covers only the cs->xx direction, so the label classifier is trained without seeing other directions.\n\nNone of these are fatal for a resource paper, but they bound what can be concluded from the quantitative analyses. The paper would be stronger with a second double-annotated recording, per-language revision checks, and a note in the abstract that the 'usable' claim is contingent on the single-annotator reliability.\n\nThe stress-test note is fair; it's the same concern I'd raise. For researchers working on interpreting, alignment annotation, or speech translation evaluation, this is a useful artifact worth engaging with. I'd send it to serious review. The referee should verify the release actually contains the data and tool, and push for the authors to be more explicit about the reliability limits in the abstract.","headline":"MockConf is a genuinely new, honestly reported interpreting corpus with manual span/word alignments, but gold-standard reliability rests on single-annotator labels with low agreement—still worth serious review.","tokens_in":21226,"tokens_out":2810,"would_cite":true,"duration_ms":31762,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces MockConf, a seven-hour corpus of simultaneous interpreting in five European languages, manually aligned at span and word level with a divergence taxonomy, plus a web annotation tool and automatic alignment baselines.","keywords":["simultaneous interpreting","span alignment","word alignment","annotation tool","divergence taxonomy","MockConf","Czech-English","speech translation corpus"],"falsifier":"Re-transcribing a sample of the non-Czech recordings with a different speech recognizer plus native-speaker revision, then comparing word error rates and re-running the alignment, would test transcript fidelity directly; if alternate transcripts shift span boundaries or word links significantly, the gold labels are not trustworthy. Independently, double-annotating additional recordings and checking whether inter-annotator agreement rises above the reported target-side label Kappa of 0.25 would test the reliability of the single-annotator gold standard.","tokens_in":20054,"feed_emoji":"🎙️","tokens_out":6047,"duration_ms":63730,"temperature":0.7,"pith_summary":"This paper is trying to establish that simultaneous interpreting can be fruitfully studied and automatically benchmarked with a parallel corpus that is manually aligned at two levels: spans of text labeled by how the interpreter transformed the source, and word links marked sure or possible. It introduces MockConf, roughly seven hours of student mock-conference recordings in Czech, English, French, German, and Spanish, transcribed and annotated with a seven-label taxonomy of translation, paraphrase, summarization, generalization, addition, and replacement. The authors argue that existing parallel corpora and alignment tools fail for interpreting because the speech has no clean sentence boundaries and contains long-range divergences such as shortening and simplification. A sympathetic reader cares because a verified resource of this kind would let researchers measure interpreting behavior, monitor trainee interpreters, and test whether automatic alignment can reproduce human judgments on a hard, realistic task.","feed_headline":"7-hour corpus maps interpreting at word and span level","feed_subtitle":"Student mock-conference recordings in five languages, with a free alignment tool and baseline system.","key_machinery":"The central objects are the span-label taxonomy and the alignment pipeline. The taxonomy, adapted from Barik's classification of translation departures, defines seven labels — translation, paraphrase, summarization, generalization, factual addition, uninformative addition, and replacement — applied to maximal-length corresponding spans between the source and interpreted transcripts. The machinery of the alignment method is a three-stage pipeline: BERTAlign produces n-m coarse span alignments without relying on sentence boundaries; punctuation-matched splits refine spans into sub-segments while SimAlign with XLM-R embeddings generates word links inside each span pair; and a small neural classifier using LaBSE similarity and span lengths assigns labels. InterAlign, the released annotation tool, is itself part of the machinery, since it is what makes word- and span-level annotation on long, unsegmented inputs feasible.","core_discovery":"The central claim is that MockConf provides a reusable reference for the alignment and analysis of simultaneous interpreting, with a span-label taxonomy adapted from translation departures in interpreting studies and word links that mark sure versus context-dependent correspondences. The paper reports quantitative properties of the corpus: translation spans cover roughly half of tokens, paraphrase about a fifth, and summarization spans are shorter on the target side (ratio ~0.6), while relay interpreting shows a higher proportion of translations and fewer additions than direct interpreting. On the evaluation side, the paper claims that a three-stage baseline system (coarse BERTAlign, punctuation-driven sub-segmentation with XLM-R word alignment, and LaBSE feature labeling) beats random baselines and in some inter-annotator comparisons rivals a single human annotator for segmentation, but that exact match with manual links remains low, and standard statistical word alignment (SimAlign) performs poorly on this data.","pith_inferences":["A natural next step is to train a supervised divergence classifier on the aligned spans directly from speech or transcript features, instead of relying on punctuation-based sub-segmentation, which the paper itself flags as unreliable.","The sure/possible word-link distinction could support studies of ear-voice span and cognitive load by revealing which target words depend on context beyond a one-to-one translation.","The corpus's design could transfer to consecutive interpreting or medical settings to test whether the same divergence categories recur outside conference-style mock scenarios.","As the authors collected additional recordings without full consent, a plausible roadmap is to use MockConf's test split as a held-out benchmark while extending training data from the broader pool."],"forward_implications":["Researchers can use MockConf to quantify interpreting strategies, for example the finding that summarization spans shrink to roughly 0.6 times the source length while translation spans stay near 1.0.","The InterAlign tool makes it practical to align long parallel speech transcripts without any pre-existing sentence segmentation.","The proposed metrics (segmentation F1, relaxed and exact span match, word AER, token-level label accuracy) give the community a shared way to compare future alignment systems on interpreting data.","The reported baselines show that standard word alignment (SimAlign) performs much worse on interpreting than on written MT, indicating that dedicated models are needed.","The low inter-annotator agreement (target-label Kappa 0.25; exact link match 14.85–30.46 percent) defines an upper bound on how well any automatic system can be expected to match an individual annotator."],"supporting_citations":[{"why":"Supplies the automatic transcription (WhisperX) that produced the raw transcripts later manually revised for all recordings.","marker":"(Bain et al., 2023)"},{"why":"Provides the taxonomy of omissions, additions, and replacements that the span labels are modeled on.","marker":"(Barik, 1994)"},{"why":"Defines the two-step coarse-to-fine alignment approach that the baseline system adapts to MockConf.","marker":"(Zhao et al., 2024)"},{"why":"Provides SimAlign, used both as the word-alignment baseline and as the engine for word links in sub-segmentation.","marker":"(Jalili Sabet et al., 2020)"},{"why":"Supplies BERTAlign, the sentence alignment tool used for the coarse span-level alignment stage.","marker":"(Liu and Zhu, 2023)"},{"why":"Provides LaBSE sentence embeddings used as features for the span label classifier.","marker":"(Feng et al., 2022)"},{"why":"Provides the XLM-R model used to compute contextual word embeddings for sub-segmentation word links.","marker":"(Conneau et al., 2020)"},{"why":"Documents a large-scale English-Japanese interpreting corpus with manual annotation, a precedent the label categories build on.","marker":"(Doi et al., 2021)"}],"fun_headline_variants":["MockConf: 7-hr corpus with word and span alignments","Student interpreting data: 7 hours, 5 languages, aligned","InterAlign tool plus MockConf corpus for interpreting","Word and span alignment benchmark from mock conferences","New corpus and tool track interpreting divergences"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire dataset rests on the transcripts being faithful: non-Czech speech was transcribed automatically by WhisperX and then only revised by native Czech speakers with self-reported proficiency in each foreign language, so any transcription errors these revisers missed are inherited by every span and word alignment in the corpus.","fun_headline_variants_meta":{"raw":{"variants":["MockConf: 7-hr corpus with word and span alignments","Student interpreting data: 7 hours, 5 languages, aligned","InterAlign tool plus MockConf corpus for interpreting","Word and span alignment benchmark from mock conferences","New corpus and tool track interpreting divergences"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000294,"raw_usage":{"total_tokens":1699,"prompt_tokens":923,"completion_tokens":776,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":539,"completion_tokens_details":{"reasoning_tokens":698}},"tokens_in":539,"tokens_out":776,"duration_ms":8716,"temperature":1.0,"reasoning_tokens":698,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:33:21.542042+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-transcribing a sample of the non-Czech recordings with a different speech recognizer plus native-speaker revision, then comparing word error rates and re-running the alignment, would test transcript fidelity directly; if alternate transcripts shift span boundaries or word links significantly, the gold labels are not trustworthy. Independently, double-annotating additional recordings and checking whether inter-annotator agreement rises above the reported target-side label Kappa of 0.25 would test the reliability of the single-annotator gold standard.","supporting_citations":[],"review_version":1}