{"id":"95ed8da6-a683-4691-9346-79854ce9235c","arxiv_id":"2412.18495","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of 110 SimulST papers shows most systems rely on unrealistic human pre-segmented audio and inconsistent terminology, and it offers a taxonomy and recommendations to fix both.","lead":"This paper surveys 110 papers on simultaneous speech translation and finds that most research trains and tests on human-pre-segmented audio, not continuous speech streams. It proposes a standardized terminology and a taxonomy of system components, and recommends that the field move toward more realistic, unbounded audio inputs.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline percentages rest on a non-reproducible manual coding; Appendix A's tallies are internally inconsistent (110 vs 111), so the 81.8%/91.8% central claim needs an independent re-coding before it can be taken as established.","rationale":"The reader's weakest assumption was that the manual categorization is accurate and reproducible; my independent reading of Appendix A confirms that this is the most load-bearing point. The textual claim in §4 is only as strong as the appendix tallies, and those tallies contain concrete signs of instability: category counts sum to 111 rather than 110, at least one citation is ambiguous or duplicated, and one entry is explicitly a boundary case. The paper does not provide a coding rubric or agreement measure, so a reader cannot tell whether the 81.8% figure would survive an independent re-coding. This does not mean the central message is wrong; the general direction is plausible and the taxonomy is useful. But the conditional verdict is appropriate: the statistical backbone should be verified. I also note that §5 contains empty citation placeholders ('recent advances in the field ( )'), which is an editorial defect but not central to the main claim. Since the reader already assigned CONDITIONAL, I recommend keeping that verdict rather than moving to ACCEPT or REJECT.","tokens_in":34686,"tokens_out":8023,"duration_ms":67744,"concrete_test":"Have two independent annotators re-code all 110 papers from Appendix A using a written rubric that defines bounded vs unbounded speech, gold vs automatic pre-segmentation, and what counts as explicit acknowledgment of the gold-segmentation assumption; compute Cohen's kappa; then re-estimate the §4 percentages with bootstrap confidence intervals. If kappa < 0.8, or the bounded proportion's 95% CI moves by more than about 3 percentage points from 81.8%, the central claim is not stable enough to support the field-wide conclusion. Also fix the 90+20+1=111 tally and resolve the duplicated 'Polák et al. (2023)' entries before re-running the statistics.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is empirical: §4 reports that 81.8% of papers rely on pre-segmented audio, 97.7% of those use gold segmentation, and 91.8% do not explicitly acknowledge the gold-segmentation assumption. These figures are ratios over the authors' assignments in Appendix A, but those assignments are judgment calls. For example, A.1.1 classifies Ma et al. (2021) as gold pre-segmentation with a footnote saying 'Unbounded speech theoretically possible but not tested', and identical citations such as 'Polák et al. (2023)' appear twice in the same list, leaving it unclear whether two distinct papers or one paper are being counted. More concretely, Appendix A states 90 bounded + 20 unbounded + 1 undefined = 111 papers, although the paper says 110; A.4 sums to 110 under different duplicate handling. No coding rubric, inter-annotator agreement, or sensitivity analysis is provided. If only a few boundary papers are reclassified, the percentages in §4 shift. Since the field-level conclusion of unrealistic evaluation depends on those exact percentages, the classification must be independently reproducible before the claim is accepted.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper surveys 110 papers on simultaneous speech-to-text translation (SimulST), proposes a six-step process model and a standardized terminology/taxonomy for the task, and then uses the survey to argue that the field has largely evaluated systems on pre-segmented gold audio rather than on realistic unbounded speech streams. The headline empirical claims in §4 are that 81.8% of papers rely on pre-segmented audio, 97.7% of those use gold segmentation, and 91.8% do not explicitly acknowledge the gold-segmentation assumption. The paper also reports widespread terminological confusion among \"simultaneous\", \"streaming\", \"online\", and \"real-time\", and it closes with recommendations on automatic segmentation, latency reporting, evaluation frameworks for unbounded speech, context handling, output visualization, and user-centric evaluation.","tokens_in":34903,"tokens_out":4697,"duration_ms":45032,"significance":"If the survey statistics are correct, the paper would provide an important field-level correction: SimulST research would largely be optimizing under an unrealistic input assumption, and reported results would not transfer to continuous speech streams. The conceptual contributions—the six-step process decomposition, the taxonomy in Figure 2, and the terminology in Table 1—are genuinely useful and are likely to be adopted regardless of the exact percentages. The paper is also transparent in providing the full categorized list in Appendix A, which is a strength for reproducibility. However, the empirical claims are load-bearing for the paper's central message, and, as detailed below, the appendix as printed does not yet allow an independent reader to verify the headline percentages.","major_comments":[{"comment":"The headline statistics in §4 are not reproducible from Appendix A as printed. The appendix reports 90 bounded + 20 unbounded + 1 undefined = 111 papers, although the text states 110; the entry \"Polák et al. (2023)\" appears twice in the same category list while the citation does not distinguish between the two distinct 2023 Polák et al. papers; and the category in A.4 is \"Papers Mentioning Automatic Segmentation\", not \"papers that explicitly acknowledge gold pre-segmentation\", so the 91.8% figure in §4 cannot be checked against the appendix. The authors should provide a machine-readable one-row-per-paper table, a precise definition of each variable used in the percentages, and a coding rubric.","section":"§4 and Appendix A"},{"comment":"The classification involves judgment calls that are not subjected to any sensitivity analysis. For example, Ma et al. (2021) is listed under gold pre-segmentation with the footnote \"Unbounded speech theoretically possible but not tested\", and many papers do not state their input conditions explicitly. Since the central claim is that \"up to 81.8%\" of papers rely on pre-segmented audio and \"97.7%\" of those use gold segmentation, the authors should report inter-annotator agreement on a sample and show how the percentages shift when boundary cases such as Ma et al. (2021) are reclassified.","section":"A.1.1"},{"comment":"The sample is restricted to open-access English papers retrieved from Semantic Scholar with no flow diagram. The queries in Table 2 each returned between 69 and 265 papers, and the final set of 110 papers is obtained after unspecified deduplication and filtering. This selection could bias the field-level percentages. The authors should report the number of unique papers per query, the exclusion counts, and the exact search date, or temper the conclusions from \"the field\" to \"the open-access English literature surveyed\".","section":"Appendix A, selection criteria"},{"comment":"The claim that \"over 65% of the papers mix and match these terms\" is not supported by the appendix. No per-paper term-usage data are provided, and Figure 3 is referenced without showing the underlying computation or a table of which terms each paper uses. The authors should include the term-usage data or explicitly define the numerator and denominator behind the 65% figure.","section":"§4, terminological chaos"}],"minor_comments":[{"comment":"The first paragraph of §5 contains two empty parentheticals, \"( )\", where references to recent advances and future directions are missing; these need to be filled in.","section":"§5, first paragraph"},{"comment":"The label \"Simultaneous (Automatic) Segmentation\" is confusing because it can be read as \"simultaneous segmentation\" rather than \"automatic segmentation performed in real time\"; consider aligning the label with the taxonomy terms from Table 1 and Figure 2.","section":"A.1.2"},{"comment":"Two distinct 2023 papers by Polák et al. are cited in the appendix as identical \"Polák et al. (2023)\" entries; the authors should disambiguate them, for example as Polák et al. (2023a) and Polák et al. (2023b), so that the counts in Appendix A can be verified.","section":"References"},{"comment":"The paper would benefit from a table mapping each Appendix A category name to the corresponding taxonomy node in Figure 2; the current labels such as \"Bounded Speech\" and \"Unbounded Speech\" are close to but not identical with the terminology proposed in Table 1.","section":"Figure 2 and Appendix A"}],"recommendation":"major_revision","confidential_remarks":"The survey covers a well-scoped and timely topic, and the authors are established contributors to the SimulST field. The main risk is not the conceptual framework but the reproducibility of the empirical percentages. I would recommend requiring release of the coding data and an independent or inter-annotator re-coding before acceptance; the paper's central message depends on those numbers."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear Colleague,\n\nI've read the Papi et al. survey on how 'real' SimulST systems are. The paper makes a strong, plausible case that the field has been evaluating on gold pre-segmented audio and that terminology is a mess. The taxonomy in Figure 2 and the 6-step process decomposition are genuinely useful contributions, and the quantitative claims — 81.8% of papers using pre-segmented audio, 97.7% gold, 91.8% not acknowledging it — are the kind of field-level statement that could reorient evaluation practice.\n\nWhat is actually new here is the systematic coding of 110 papers. That coding is also the soft spot. Appendix A gives no rubric, no inter-annotator agreement, and no sensitivity analysis. It also contains concrete internal problems: the bounded/unbounded/undefined counts sum to 111 rather than 110, and 'Polák et al. (2023)' appears twice in the same list, so it's unclear whether two distinct papers are being counted. Section 5 has at least two empty citations. These are fixable, but they undermine the reproducibility of the exact percentages. A modest misclassification of a few boundary papers would shift the 81.8% and 91.8% numbers.\n\nI don't think these problems are fatal. The directional claim that most SimulST work uses gold segmentation and rarely says so is consistent with my own reading of the literature, and the authors are not hiding their data — the appendix lists every paper. The sample restriction to open-access English papers is a limitation but a defensible one for a survey. The taxonomy itself is the authors' construction, but the statistics are not derived from it in a circular way; they are observations about the papers.\n\nThe paper deserves peer review. The topic is important, the framework is useful, and the recommendations (use at least automatic pre-segmentation, be explicit about input conditions, report computationally unaware latency) are sound. What I'd ask of the authors is simple: release a proper coding protocol, settle the duplicate-counting and the 110/111 discrepancy, fill the empty citations, and report some sensitivity or agreement numbers. If they do that, this becomes a standard reference for SimulST evaluation; as it stands, treat the headline percentages as strong estimates rather than established fact.\n\nFor your group: worth reading if you work on speech translation evaluation; not a must-cite until the coding is cleaned up, but I'd probably cite it anyway for the taxonomy. I'd send it to review.","headline":"Useful survey, plausible central claim, but the headline percentages need a reproducible coding before they can be trusted.","tokens_in":35443,"tokens_out":3486,"would_cite":true,"duration_ms":29131,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that most simultaneous speech-to-text translation research evaluates on human-pre-segmented audio and rarely acknowledges this, so reported results may not transfer to real-world continuous speech.","keywords":["simultaneous speech translation","speech-to-text translation","literature review","audio segmentation","streaming speech","terminology standardization","latency evaluation","unbounded speech"],"falsifier":"Have two independent annotators re-classify the same 110 papers using the taxonomy from Appendix A and measure inter-annotator agreement; if agreement is low (e.g., Cohen's kappa below 0.6) or the re-derived percentages differ materially from the reported 81.8% and 91.8%, the paper's key statistics would not survive.","tokens_in":34501,"feed_emoji":"🎙️","tokens_out":7421,"duration_ms":58650,"temperature":0.7,"pith_summary":"This paper argues that the field of simultaneous speech-to-text translation (SimulST) has drifted from its stated goal of translating a continuous, unsegmented audio stream. The authors surveyed 110 papers and found that up to 81.8% of them operate only on pre-segmented audio, and 97.7% of those rely on gold (human-made) segment boundaries that do not exist in real-time use. They also found that over 65% of papers use \"streaming,\" \"online,\" or \"real-time\" interchangeably with \"simultaneous\" without defining these terms. To address these problems, the paper proposes a standardized taxonomy and a six-step formalization of the SimulST process, along with concrete recommendations for evaluation and reporting. If the survey's statistics are correct, much of the reported progress in SimulST may simply not carry over to the real-world conditions the task is meant for.","feed_headline":"81.8% of simultaneous translation papers test on pre-cut audio","feed_subtitle":"A 110-paper survey suggests real-world continuous speech could break most reported SimulST results.","key_machinery":"The central mechanism is a taxonomy built on three dichotomies: input (bounded vs unbounded speech), architecture (direct vs cascade), and output strategy (incremental vs re-translation), which the authors apply by hand to classify all 110 surveyed papers. Supporting this is a six-step decomposition of the SimulST process, from audio acquisition through buffer updating, hypothesis generation, buffer trimming, and output presentation. The taxonomy does the work of converting a loosely defined body of literature into countable categories, which is what makes the headline percentages (81.8%, 97.7%, 91.8%) possible at all.","core_discovery":"The paper's central claim is that the SimulST community has been evaluating its models under unrealistic input conditions while rarely acknowledging it. Its key quantitative finding is that, among the 110 reviewed papers, up to 81.8% rely on pre-segmented audio, with 97.7% of those using gold (human) segmentation, and 91.8% of all papers do not explicitly state that they assume gold pre-segmented speech. Only 20 papers address unbounded speech at all, and only two explore replacing gold with automatic segmentation in the bounded scenario. The paper also documents terminological chaos, with over 65% of papers mixing at least one of \"streaming,\" \"online,\" or \"real-time\" with \"simultaneous\" without clear definitions. As a remedy, it formalizes SimulST as a six-step process and introduces a taxonomy of system components (input type, architecture, output strategy) to make the field's assumptions visible.","pith_inferences":["A direct test the paper leaves implicit would be to run the same SimulST model on continuous audio and on gold-segmented audio; a large quality drop on the continuous stream would confirm that the field's bounded-input focus is the main gap.","The taxonomy could double as a reporting checklist for future papers, and venues could require authors to state input type explicitly; this is a policy implication the paper hints at but does not spell out.","Because the terminology problem is inherited from neighboring fields (ASR's \"streaming\" and MT's \"online\"), a durable fix likely requires coordinated glossary efforts across communities rather than a single survey's definitions.","The 97.7% gold-segmentation statistic also implies that segmentation algorithms have been evaluated only indirectly, so a dedicated segmentation-aware benchmark could turn the survey's critique into a shared task."],"forward_implications":["SimulST results obtained on gold-segmented benchmarks should not be treated as predictive of performance on unbounded, continuous audio streams.","The dominant evaluation toolkit, SimulEval, needs extensions or a successor that can score systems on streams without relying on pre-segmented inputs.","Researchers working with bounded speech inputs should adopt automatic pre-segmentation instead of gold segmentation to more closely approximate real conditions.","Adopting a unified terminology and explicitly stating the type of speech input in every paper would make results across the field comparable.","Human-centered evaluation of output visualization and of the quality-latency trade-off is needed before automatic metric improvements can be trusted to improve user experience."],"supporting_citations":[{"why":"Formalized the SimulST task as taking a continuous, unsegmented audio stream as input, which is the benchmark the paper holds the field to.","marker":"Fügen et al. (2007)"},{"why":"Provides the widely used definition of SimulST as concurrent translation, which the paper's taxonomy adopts.","marker":"Ren et al. (2020)"},{"why":"SimulEval, the dominant evaluation toolkit used by 61 of the 110 papers, which the paper argues is built around gold pre-segmented inputs.","marker":"Ma et al. (2020a)"},{"why":"Supplies the direct-vs-cascade architecture distinction that organizes one axis of the taxonomy.","marker":"Sperber and Paulik (2020)"},{"why":"One of only two papers in the corpus that test automatic instead of gold segmentation in the bounded scenario, the counterexample that makes the gold-segmentation statistic meaningful.","marker":"Kolss et al. (2008)"},{"why":"The other paper testing automatic segmentation, reinforcing that the exception exists but is rare.","marker":"Shimizu et al. (2013)"},{"why":"A recent survey that itself assumes human-segmented audio, evidence that the narrow focus persists even in meta-reviews.","marker":"Liu et al. (2024)"}],"fun_headline_variants":["81.8% of SimulST papers test on pre-cut audio","Most SimulST models never see continuous speech","Survey: 82% of SimulST research ignores real-world audio","Only 20 of 110 SimulST papers handle unbounded speech","Terminology chaos: 65% of SimulST papers misuse 'real-time'"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's percentages are produced by the authors' manual classification of the 110 papers, so the central claim stands or falls with whether those classifications are accurate and reproducible.","fun_headline_variants_meta":{"raw":{"variants":["81.8% of SimulST papers test on pre-cut audio","Most SimulST models never see continuous speech","Survey: 82% of SimulST research ignores real-world audio","Only 20 of 110 SimulST papers handle unbounded speech","Terminology chaos: 65% of SimulST papers misuse 'real-time'"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000317,"raw_usage":{"total_tokens":1775,"prompt_tokens":909,"completion_tokens":866,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":525,"completion_tokens_details":{"reasoning_tokens":772}},"tokens_in":525,"tokens_out":866,"duration_ms":7131,"temperature":1.0,"reasoning_tokens":772,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T04:40:53.154876+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have two independent annotators re-classify the same 110 papers using the taxonomy from Appendix A and measure inter-annotator agreement; if agreement is low (e.g., Cohen's kappa below 0.6) or the re-derived percentages differ materially from the reported 81.8% and 91.8%, the paper's key statistics would not survive.","supporting_citations":[],"review_version":1}