{"id":"d88e9e79-d34e-43fd-8cd2-b744a1fcc24a","arxiv_id":"2504.17366","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"LiveLongBench, a new spoken-text long-context benchmark from live streams, shows that current long-context models and compression methods degrade on redundant speech and that hybrid compression works best.","lead":"This paper builds a bilingual benchmark for long-context understanding from live-stream shopping transcripts and shows that current LLMs and compression methods struggle with redundant spoken text. It also reports that combining existing compression methods, such as MInference plus LLMLingua, can beat every single method.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Undefined 'Score' metric underlies all overall rankings, DEA efficiency scores, and the hybrid-baseline claim; exact match alone cannot carry them, so core empirical results are not auditable.","rationale":"The reader's weakest_assumption (no written-text control) is a real confound for causal claims about redundancy, and it is worth fixing with a cleaned or rewritten control corpus. I do not dispute that concern. However, the most load-bearing issue for the central empirical claims is the undefined 'Score' metric. Every overall ranking, task-average column, human-vs-model comparison, and DEA efficiency output in the paper is computed from Score, yet Section 4.1 gives no formula, rubric, or scorer. Exact Match is well-defined but saturates near zero on reasoning tasks (human EM 4.8/8.3 vs Score 41.0/65.8), so it cannot independently support the same conclusions. This means the headline finding 'no single method consistently outperforms others' and the proposed hybrid baseline's superiority are not auditable from the manuscript alone. The concern is fixable: the authors likely have a precise scoring rule, and the released repository may define it. I therefore keep the reader's CONDITIONAL verdict; no change is needed, but I would make the Score definition an explicit acceptance condition. The concrete check is to reproduce all Score-based columns under one stated rule and verify the DEA ranking.","tokens_in":17126,"tokens_out":6118,"duration_ms":63055,"concrete_test":"Request or reconstruct the exact Score definition (e.g., token-level F1 against gold answers, LLM-as-judge Likert rubric, or human partial-credit scale) and recompute the Avg./Overall columns of Tables 2, 5, and 6 from the released predictions and gold labels. If any reported Score cannot be reproduced under a single stated rule, or if the DEA input changes materially, the ranking and 'no single method' conclusions are unverified. As a secondary check, rerun the headline comparisons using only Exact Match; if the conclusions survive, the undefined-Score concern is less damaging.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.1 defines the primary reporting metric only as 'a complementary metric, Score, which offers a softer and more fine-grained assessment by capturing partial correctness and enabling a continuous measure of model performance across tasks.' No formula, rubric, scoring procedure, or annotation protocol is given anywhere in the paper or appendix. Score is then used for every 'Avg.' and 'Overall' column in Tables 2, 5, and 6, for the human-vs-model comparisons, and as the output variable in the DEA efficiency analysis that identifies the claimed optimal compression combination. The Exact Match columns cannot substitute: for the reasoning tasks, human Exact Match is 4.8% and 8.3%, while the matching Score values are 41.0 and 65.8, so the two metrics carry different content, and only Score produces interpretable signal on those tasks. Because Score is undefined, a reader cannot verify the central empirical claims: that Gemini-1.5-pro outperforms other LLMs overall, that no single method consistently outperforms others, and that MInference+Lingua-4x or KIVI+MInference+Lingua-2x is optimal. This is not a stylistic nit; the benchmark's conclusions, including the DEA recommendation, are numerically driven by this unreported metric.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces LiveLongBench, a bilingual (Chinese-English) long-context benchmark constructed from e-commerce live-stream transcripts with an average sequence length of roughly 97K tokens. It defines nine tasks grouped into three categories—retrieval-dependent, reasoning-dependent, and hybrid—and evaluates both closed and open LLMs as well as KV-cache compression methods (KIVI, MInference, LLMLingua) used individually and in combination. A DEA-based analysis is used to recommend a performance-memory trade-off. The central claims are that current LLMs, even with long context windows, perform substantially below humans on redundant spoken inputs; that no single compression method consistently outperforms the others; and that hybrid compression combinations such as MInference+LLMLingua-4x and KIVI+MInference+LLMLingua-2x provide the best balance of performance and memory use.","tokens_in":17343,"tokens_out":4873,"duration_ms":46463,"significance":"If the reported results are auditable, LiveLongBench fills a genuine and practical gap: long-context evaluation has focused on written documents, whereas live-stream transcripts with repetition, fillers, and topic drift are an important real-world domain. The paper's strengths include a detailed dataset-construction pipeline (Whisper transcription with manual proofreading, WER 0.53%, category and length statistics in Table 3), a documented six-step human-annotation quality-control protocol, an adapted Needle-in-a-Haystack stress test on spoken-style backgrounds, and a released code and benchmark repository. The hybrid-compression framing and the DEA analysis address a real deployment question. However, the significance is currently conditional: the primary evaluation metric is never defined, the attribution of model failures to redundancy lacks a matched written-text control, and the reliability of the human gold labels is not quantified.","major_comments":[{"comment":"The metric 'Score' is never defined. Section 4.1 states only that Score 'offers a softer and more fine-grained assessment by capturing partial correctness and enabling a continuous measure of model performance across tasks,' but no formula, rubric, per-task scoring procedure, or annotation protocol appears in the paper or appendix. Score is used for every Overall and Avg. column in Tables 2, 5, and 6, for the human-vs-model comparisons, and as the DEA output variable in Figure 4. The Exact Match columns cannot substitute: on reasoning tasks, human Exact Match is 4.8% and 8.3% while the corresponding Score values are 41.0 and 65.8, so the two metrics carry different content. Because Score is undefined, a reader cannot verify the core empirical claims that Gemini-1.5-pro outperforms other LLMs overall, that no single method consistently outperforms others, and that MInference+Lingua-4x or KIVI+MInference+Lingua-2x is optimal. Please provide a full definition, the scoring instructions, and the scoring code.","section":"§4.1, Tables 2, 5, 6; Figures 3–4"},{"comment":"The conclusion that current methods 'perform poorly on highly redundant inputs' is not supported by the experimental design. LiveLongBench contains only spoken live-stream transcripts; there is no matched written-text corpus and no de-redundified or cleaned version of the same transcripts. The observed performance gaps could therefore be caused by task difficulty, answer format, input length, or domain-specific vocabulary rather than by redundancy specifically. A control condition—for example, running the same questions on a version of the same transcripts with repetition and fillers removed, or on written e-commerce text of matched length and topic—is needed to attribute the degradation to redundancy.","section":"Abstract; §4.1"},{"comment":"The paper announces 'semantic multi-span' as a novel task type—an advanced form of multi-span reasoning over semantically distributed spans—but no task in Table 3 corresponds to it and Section 3.3 does not define how it is operationalized or scored. If semantic multi-span is a contribution, it needs an explicit task definition and dataset statistics; if it is intended to be covered by the 'Multiple Document QA' or 'Price Comparison' tasks, that mapping should be stated directly.","section":"§1; §3.3; Table 3"},{"comment":"No inter-annotator agreement statistic is reported for the human gold labels, and no variance or repeated-run statistics are reported for model Scores. Because the human scores are the reference point for every model comparison in Tables 2, 5, and 6, a reliability measure (e.g., Cohen's kappa or per-task agreement) is needed to establish that the labels are stable enough to support the benchmark's conclusions. Without such a measure, it is difficult to know how much of the reported human advantage over models is due to annotation noise.","section":"Appendix A.1"}],"minor_comments":[{"comment":"The benchmark name is inconsistently rendered as 'LiveLongBench,' 'LongLiveBench,' and 'LifelongBench'; please standardize it throughout.","section":"§3.2; §A.2; Table 3 caption"},{"comment":"The table reports infinite audio durations ('∞') for Whisper and Paraformer-zh; please clarify what this means, presumably that these systems can process arbitrarily long segments without a fixed length limit.","section":"Table 4"},{"comment":"The word cloud in Figure 5 is not rendered as readable text in the submitted manuscript; please replace it with a legible figure.","section":"Figure 5"},{"comment":"The reference list contains two entries for 'Leave no document behind' (Wang et al., 2024a and 2024b) with the same title; please disambiguate or merge them.","section":"References"},{"comment":"The annotation-cost calculation ('five full-time students over two days... total cost... around 400 RMB') is difficult to reconcile with a monthly salary of 800 RMB per student; please clarify the computation.","section":"Appendix A.1"}],"recommendation":"major_revision","confidential_remarks":"The undefined Score metric is the main technical blocker and is fixable in revision, but it is a hard condition: without a formal definition of Score, the empirical rankings, the human-vs-model comparison, and the DEA-based recommendation cannot be audited. The absence of a matched written or de-redundified control also weakens the paper's central conceptual claim, and I would treat that control as a required addition rather than an optional experiment. The dataset itself and the adaptation of NIAH to spoken-style transcripts are potentially valuable contributions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a genuinely useful resource — a bilingual long-context benchmark built from real live-stream transcripts, with retrieval, reasoning, and hybrid task categories and a decent amount of dataset documentation. The gap it targets is real; existing long-context benchmarks are written-text-centric. The authors also evaluate a sensible set of LLMs and compression baselines, and the observation that compression can denoise redundant speech is interesting.\n\nThe problems are concentrated in the evaluation section. The primary metric, \"Score\", is never defined. The paper says it captures partial correctness and gives a continuous measure, but no formula, rubric, or scoring protocol appears anywhere in the main text or appendix. That is not a stylistic nit: every Avg./Overall column, the human-vs-model comparisons, and the DEA efficiency analysis all rest on Score. Exact Match is reported alongside, but on the reasoning tasks human EM is 4.8% and 8.3% while the corresponding Scores are 41.0 and 65.8 — so Score carries the actual signal, and a reader cannot reproduce or interpret it. The DEA \"optimal combination\" conclusion is therefore not auditable.\n\nA second, softer problem: the paper attributes model failures to spoken redundancy, but there is no control — no matched written transcript or cleaned version of the same content. Task difficulty or input length could confound the comparison. That is a fixable design gap, and the claim is presented as a finding rather than a hypothesis.\n\nMinor issues: no error bars or repeated runs, no inter-annotator agreement (the appendix describes a six-step QC process but gives no IAA number), and the annotation cost math looks off (800 RMB/month per student over two days should not sum to ~400 RMB — likely a typo, but confusing).\n\nWhat is genuinely good: the dataset construction details (ASR WER 0.53% after proofreading, category distribution, length statistics), the task types including the semantic multi-span idea (a modest extension but reasonable), and the attempt to benchmark compression methods on real spoken text. The benchmark itself deserves a serious referee, but the empirical claims as written need the Score metric defined and a control analysis before they can be trusted.","headline":"LiveLongBench fills a real gap with a useful spoken-text benchmark, but the undefined 'Score' metric makes the headline empirical results unauditable.","tokens_in":17898,"tokens_out":1838,"would_cite":false,"duration_ms":17754,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces LiveLongBench, a bilingual benchmark of live-stream transcripts averaging 97K tokens, and uses it to claim that no current LLM or compression method reliably understands highly redundant spoken text.","keywords":["long-context understanding","spoken language","live-stream e-commerce","benchmark","KV cache compression","redundancy","retrieval tasks","hybrid tasks"],"falsifier":"Take a sample of LiveLongBench transcripts, produce two versions, the original noisy transcript and a cleaned rewrite that removes filler words, repetitions, and topic drift while keeping all factual content, then run the same models on both under identical settings. If scores on the cleaned versions do not rise substantially, or if a single method starts to dominate both, the paper's central claims about redundancy and the absence of a best method would be undercut.","tokens_in":16906,"feed_emoji":"🎙️","tokens_out":6989,"duration_ms":62336,"temperature":0.7,"pith_summary":"This paper builds LiveLongBench, the first benchmark built from live-stream transcripts rather than written documents, with Chinese and English examples averaging about 97,000 tokens. It defines nine tasks across retrieval, reasoning, and hybrid categories, and evaluates eight LLMs plus several KV-cache compression methods on them. The central claim is that current systems are strongly task-specific and degrade on highly redundant spoken inputs, with no single method winning across tasks. The paper also argues that combining compression methods, for example MInference with LLMLingua, improves both accuracy and memory use, and that aggressive compression can act as denoising rather than information loss. A sympathetic reader would care because it targets a real deployment setting, live-stream e-commerce, and claims that the way LLMs are normally benchmarked misses what actually goes wrong there.","feed_headline":"No single method masters redundant live-stream speech","feed_subtitle":"A bilingual 97K-token benchmark finds retrieval hardest and hybrid compression best for spoken inputs.","key_machinery":"The central object is the benchmark itself plus its evaluation protocol. Transcripts come from Douyin e-commerce live streams, transcribed with Whisper and manually proofread, with only light filtering so filler words and repetitions survive. Tasks are categorized by information layout, including single-span retrieval, multi-span, semantic multi-span, and global spans, and grouped into retrieval-dependent, reasoning-dependent, and hybrid categories. Compression methods are grouped into token pruning, attention sparsification, and KV-cache quantization, and their combinations are ranked by Data Envelopment Analysis, which treats memory as input and score as output. The semantic multi-span task type is introduced as a novel extension, requiring models to integrate conceptually related but dispersed segments.","core_discovery":"On the paper's own terms, the central discovery is that long-context spoken text is a distinct failure regime: even a 1M-token model trails human annotators overall, and retrieval tasks are the hardest, not the long-reasoning ones. On the compression side, single methods show task-specific preferences; quantization preserves retrieval accuracy, while token pruning helps reasoning by removing noise, and no single method dominates. The paper's proposed hybrid baseline, MInference with LLMLingua 4x, reaches the best overall performance, and the three-way combination of KIVI 4-bit, MInference, and LLMLingua 2x gives the best performance-per-memory trade-off according to Data Envelopment Analysis. A further claim is that compression in high-redundancy contexts can improve accuracy, with LLMLingua 4x beating LLMLingua 2x.","pith_inferences":["If redundancy is the causal driver, then a matched control, the same transcripts after filler-word removal or written-style rewriting, should reproduce the performance gap; the paper does not run that control, and such an experiment would separate redundancy from raw length or task difficulty.","The benchmark's transcripts come from one platform and genre, e-commerce live streams, so the claim that it represents spoken texts generally is an extrapolation; applying the same pipeline to lectures, meetings, or news broadcasts would show how far the findings carry.","The DEA efficiency ranking depends on the specific models, context windows, and memory measurements used; retraining the ranking on larger or newer models could shift which combination is optimal, while the qualitative finding that hybrids beat singles might survive.","Because the dataset keeps ASR errors very low after proofreading, it could double as a controlled testbed for compression robustness; one could inject synthetic disfluencies at varying rates to map exactly how redundancy hurts retrieval."],"forward_implications":["If the benchmark is accepted as representative, current leaderboard results on written long-context benchmarks overstate readiness for deployed conversational AI, since retrieval from redundant speech is systematically worse.","The finding that compression can improve accuracy implies that redundancy filtering is a legitimate inference-time technique, not just a cost-saving one, and that 4x pruning may be preferable to 2x in noisy input.","Hybrid compression combinations should be treated as a design space, and the paper's DEA ranking provides a principled way to choose among them under memory constraints.","Domain-specific fine-tuning helps hybrid tasks but hurts reasoning, so specialization is a trade-off rather than a free lunch.","The benchmark can serve as a testbed for studying the 'lost in the middle' effect in spoken rather than written inputs."],"supporting_citations":[{"why":"Provides the long-context benchmark design and the comparison in Table 1 that LiveLongBench extends with spoken-language characteristics.","marker":"(Bai et al., 2023)"},{"why":"Supplies the task taxonomy of retrieval, reasoning, and hybrid tasks that the benchmark's nine tasks are built around.","marker":"(Wang et al., 2024a)"},{"why":"Supplies the single-span, multi-span, and global span categorization that motivates the new semantic multi-span task type.","marker":"(Kwan et al., 2023)"},{"why":"LLMLingua is the token-pruning method whose 4x compression yields the best single-method reasoning and is a key component of the proposed hybrid baseline.","marker":"(Pan et al., 2024)"},{"why":"MInference is the attention-sparsification method that, combined with LLMLingua, forms the best overall hybrid configuration.","marker":"(Jiang et al., 2024)"},{"why":"KIVI is the KV-cache quantization method whose information retention explains retrieval success and which anchors the memory-efficient three-way combination.","marker":"(Liu et al., 2024b)"},{"why":"Provides eCeLLM-M, the domain fine-tuned model used to test whether e-commerce specialization helps or hurts long spoken inputs.","marker":"(Peng et al., 2024)"},{"why":"Provides the Needle-in-a-Haystack passkey template adapted to live-stream background text for the retrieval stress test.","marker":"(Mohtashami and Jaggi, 2023)"}],"fun_headline_variants":["Retrieval is the real wall for long spoken text","Hybrid compression beats single tricks on live-stream speech","Redundant speech defeats even 1M-token models","New benchmark exposes spoken long-text failure mode","Best overall: prune noise, keep facts, compress 4x"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the poor results come from spoken redundancy specifically, yet the evaluation never compares against a cleaned or written version of the same transcripts, so length, task difficulty, or answer format could drive the gap instead.","fun_headline_variants_meta":{"raw":{"variants":["Retrieval is the real wall for long spoken text","Hybrid compression beats single tricks on live-stream speech","Redundant speech defeats even 1M-token models","New benchmark exposes spoken long-text failure mode","Best overall: prune noise, keep facts, compress 4x"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000535,"raw_usage":{"total_tokens":2573,"prompt_tokens":945,"completion_tokens":1628,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":561,"completion_tokens_details":{"reasoning_tokens":1549}},"tokens_in":561,"tokens_out":1628,"duration_ms":12032,"temperature":1.0,"reasoning_tokens":1549,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:41:35.878785+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a sample of LiveLongBench transcripts, produce two versions, the original noisy transcript and a cleaned rewrite that removes filler words, repetitions, and topic drift while keeping all factual content, then run the same models on both under identical settings. If scores on the cleaned versions do not rise substantially, or if a single method starts to dominate both, the paper's central claims about redundancy and the absence of a best method would be undercut.","supporting_citations":[],"review_version":1}