{"id":"ae425a40-2632-4ca4-9a20-30e1ea9bf3c0","arxiv_id":"2505.19125","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"RTime-QA is a video-question benchmark where models choose between temporally opposite descriptions of the same event, and current AI models score far below humans.","lead":"This paper introduces RTime-QA, a benchmark of 822 human-checked video questions that test whether AI models understand brief, atomic events such as folding versus unfolding a chair. The best model scores 34.6 on the strict metric versus 97.3 for humans, and a companion training set raises that to 65.9, though the gain may be inflated by shared data sources.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"RTime-IT effectiveness claim is not yet established: the 34.6-to-65.9 strict-ACC gain on RTime-QA comes from training and test sets built from the same RTime source and the same multiple-choice template, so it could be template/formula overfitting rather than general temporal understanding.","rationale":"The reader's weakest assumption and my analysis converge on the same point: the RTime-IT effectiveness claim in Table 5 is not yet validated as a measure of general temporal understanding. The benchmark claim itself—822 human-annotated paired questions with strict-ACC, hard for current LMMs, requiring multiple frames—appears solid and is supported by consistent evidence: human performance near ceiling (97.3 strict-ACC), mostly below-random zero-shot strict-ACC, monotonic frame-count scaling (Table 3), and coherent model rankings. The instruction-tuning claim, by contrast, rests on a single in-distribution result with strong shared template/source factors, so it is the least secure part of the central contribution. The required fix is feasible and concrete: evaluate the RTime-IT-tuned model on external temporal benchmarks and add a control fine-tuning condition. I therefore keep CONDITIONAL rather than ACCEPT or REJECT. I do not see an internally inconsistent step, and the paper is not claiming anything impossible; the gap is in missing controls. My one concrete test would settle the dispute by comparing transfer behavior across independent datasets and a position-shuffled control.","tokens_in":10499,"tokens_out":2687,"duration_ms":23311,"concrete_test":"Fine-tune Qwen2-VL-7B under the identical recipe (same 14,096 samples and 6 epochs) on (a) RTime-IT with answer options shuffled so that position carries no signal, and (b) a matched-size temporal instruction set from an independent source such as TemporalBench or VITATECS; then evaluate both variants on RTime-QA and on an external temporal benchmark (e.g., TemporalBench, VITATECS, or Perception Test). If the RTime-IT model gains >20 strict-ACC on RTime-QA but underperforms the independent-source model on the external benchmark, template/distribution overfitting is confirmed; if the gain transfers, the effectiveness claim is supported.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central evidence for RTime-IT improving temporal understanding is Table 5: fine-tuning Qwen2-VL on RTime-IT raises RTime-QA strict-ACC from 34.6 to 65.9. This Table is the load-bearing support for the claim that RTime-IT 'effectively enhance[s] LMMs' capacity in temporal understanding' (Abstract and Section 3.4/4.2). The risk is that the gain reflects overfitting to the training/eval distribution rather than a genuine, transferable improvement. RTime-IT and RTime-QA are both derived from RTime (Du et al., 2024); both use the identical template ('Which sentence accurately describes the video? A/B Answer:'); and both use temporally-contrastive sentence pairs of similar style and length. Training on 14,096 such samples for 6 epochs can exploit template priors, answer-position bias, and annotation-style regularities that transfer directly to the 822 RTime-QA test items but would not generalize to other temporal benchmarks. No external temporal benchmark (e.g., VITATECS, TemporalBench, Perception Test) is reported after fine-tuning, and no control condition (e.g., fine-tuning on an equally sized independent temporal instruction set, or on RTime-IT with shuffled answer options) is provided. The paper does state it excludes RTime-IT videos that overlap RTime-QA, but that only removes video-level leakage; it does not remove the shared template, shared source distribution, or shared annotation process, which are exactly the dimensions most likely to be memorized.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents RTime-QA, a multiple-choice video-language benchmark focused on atomic temporal event understanding. The benchmark is constructed from the authors' prior RTime dataset: 822 questions, each pairing a video with a correct description and a temporally opposite distractor, with all annotations human-verified. The paper also introduces RTime-IT, a 14,096-sample instruction-tuning dataset with the same structure, and evaluates eight LMMs plus five vision-language alignment models. The main empirical results are that state-of-the-art Qwen2-VL achieves only 34.6 strict-ACC (vs. 97.3 for humans), that increasing the number of sampled frames improves Qwen2-VL, and that fine-tuning Qwen2-VL on RTime-IT raises strict-ACC from 34.6 to 65.9. The paper concludes that RTime-QA is a challenging temporal benchmark and that RTime-IT effectively improves LMM temporal understanding.","tokens_in":10806,"tokens_out":7460,"duration_ms":57123,"significance":"The benchmark addresses a genuine gap: many existing video QA benchmarks can be solved by image-only models, and the design here forces a model to distinguish temporally opposite descriptions. The evidence that RTime-QA is challenging is reasonably strong—image-centric LLaVA1.5 is near chance, the best LMMs are far below human performance, and performance scales with the number of frames. The release of the dataset and the paired-question Strict-ACC protocol are useful contributions. However, the claim that RTime-IT improves general temporal understanding is not yet supported, because the instruction set and the evaluation set are drawn from the same source distribution and share the same template; this is a load-bearing weakness that requires additional external evaluation and control experiments.","major_comments":[{"comment":"The claim that RTime-IT 'effectively enhance[s] LMMs' capacity in temporal understanding' is not established because RTime-IT and RTime-QA are constructed from the same source dataset (RTime) using the same annotation process and the same question template ('Which sentence accurately describes ...? A ... B ... Answer:'). Excluding overlapping videos removes video-level leakage but does not remove template-level, annotation-style, or source-distribution leakage. Training for 6 epochs on 14,096 such samples could memorize these regularities, including answer-position biases, sentence-length/style cues, and the specific temporal contrast patterns, which then transfer directly to the 822 RTime-QA items. To support the transfer claim, the authors should evaluate the fine-tuned model on independent temporal benchmarks (e.g., VITATECS, TemporalBench, Perception Test, MVBench, Next-QA) and include a control condition, such as fine-tuning on RTime-IT with shuffled answer options or on a similarly sized instruction set with a different template. Without such controls, the 34.6-to-65.9 improvement cannot be attributed to general temporal understanding rather than overfitting to the benchmark distribution.","section":"Section 3.4, Table 5"},{"comment":"The Strict-ACC metric and the 'below random' results need a position-bias analysis. Because each quadruple yields two questions with opposite correct letters, a model that always picks option A will obtain around 50% ACC and near 0% Strict-ACC, explaining the very low Strict-ACC numbers for models like LLaVA1.5 (3.9 vs. random 26.8). The paper does not report whether the correct answer is balanced across A and B in the benchmark, nor whether models exhibit a systematic letter preference. Since Strict-ACC is the central metric used to establish that RTime-QA is challenging, the authors should either rotate option order in the benchmark or report accuracy under permuted options, and show that the gap to random persists.","section":"Section 4.1, Table 2"},{"comment":"The paper asserts that the benchmark is 'carefully-curated' and that annotators excluded quadruples where the event could be inferred from a single static image, but it provides no inter-annotator agreement statistics and no direct verification that the final 822 questions are not solvable from spatial appearance alone. The single-frame and few-frame results for Qwen2-VL in Table 3 are suggestive, but a dedicated single-frame or image-only baseline on the final benchmark (e.g., evaluating a strong image-based model or human annotators with one frame) would directly validate the central 'temporal-only' design claim. Reporting these statistics and controls would also make the quality-control process reproducible.","section":"Section 3.2"}],"minor_comments":[{"comment":"The number 65.9 appears both as the zero-shot ACC of Qwen2-VL in Table 2 and as the fine-tuned Strict-ACC in Table 5, which is confusing; please disambiguate the two uses in the abstract and Section 4.2.","section":"Section 4.2, Table 5"},{"comment":"The paper should clarify the relation between 822 questions and the underlying quadruples (V, ¯V, T, ¯T): if each quadruple yields two questions, then there are 411 quadruples, and Strict-ACC is computed per quadruple rather than per question. Table 1 should report exact counts instead of '0.8K'.","section":"Section 3.2, Table 1"},{"comment":"The evaluation protocol is underspecified: the paper does not state the number of frames used for each model (other than Qwen2-VL in Table 3), the exact prompt template, the decoding parameters, or how many runs were averaged. These details are necessary for reproducibility of Tables 2–5.","section":"Section 4.1"},{"comment":"The human evaluation protocol is not described: the paper does not specify the number of human annotators, their instructions, or whether they viewed the same video and option format as the models. This information is needed to interpret the 97.3 human strict-ACC figure.","section":"Table 2"},{"comment":"For RTime-IT, the paper only states that videos do not overlap with RTime-QA; it does not say whether RTime-IT videos were also filtered against WebVid, VideoChatGPT, or other public training corpora, which is relevant to potential leakage when fine-tuning a model like Qwen2-VL.","section":"Section 3.4"},{"comment":"There are several typos and formatting issues, including 'Evalutation' in the Section 4.1 heading, 'epoches' in Section 4.2, and inconsistent spacing in model names such as 'LLaV A' and 'Vide-oLLaV A'. The footnote 'Work in progress' should also be removed for a formal submission.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The core dataset may be a useful contribution, and the paper's zero-shot results provide credible evidence that RTime-QA is challenging. My main concern is the over-claiming of RTime-IT's effectiveness without external evaluation or control experiments; this can be addressed within the scope of a revision, so I recommend major revision rather than rejection. I would also suggest the authors consider positioning the paper primarily as a benchmark paper and de-emphasizing the instruction-tuning claim until it is substantiated."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know about RTime-QA. First, the benchmark itself is a genuine step forward: 822 human-annotated multiple-choice questions where correct and incorrect answers differ only in temporal semantics (e.g., \"unfolds\" vs \"folds\" a chair). The paired-question strict-ACC metric—requiring the model to answer correctly for both the video and its temporal negative—is a good idea, and it exposes a real weakness: image-only LLaVA1.5 sits near chance, while Qwen2-VL gets only 34.6 strict-ACC versus 97.3 for humans. The frame-count ablation (Table 3) shows performance climbs with more frames, which existing benchmarks largely don't show. That's evidence the dataset actually tests temporal understanding.\n\nThe soft spot is the instruction-tuning claim. RTime-IT and RTime-QA are both built from the same source—the authors' earlier RTime dataset—using the same question template (\"Which sentence accurately describes the events happened in the video? A/B Answer:\"). Fine-tuning Qwen2-VL on 14,096 such samples for 6 epochs improves strict-ACC from 34.6 to 65.9. That gain could reflect the model learning the template, the answer-position distribution, and the annotation style, rather than acquiring generalizable temporal understanding. The paper excludes overlapping videos between RTime-IT and RTime-QA, but that only rules out video-level memorization; it doesn't rule out distribution-level overfitting. There's no evaluation on an external temporal benchmark (e.g., VITATECS, TemporalBench, Perception Test) and no control condition (e.g., an equally sized independent instruction set, or scrambled answer options). Without those, the \"effectively enhance LMMs' capacity\" claim in the abstract is not yet supported.\n\nI'd also note the paper is explicitly \"work in progress,\" which shows: the related work section is thin, and the claim that public video datasets lack temporal emphasis is supported more by anecdote than measurement. But these are minor.\n\nBottom line: the benchmark and the strict-ACC protocol deserve serious attention from anyone working on video-language evaluation. The instruction-tuning story needs a redo with external validation before it's credible. If you're refereeing this, I'd send it back for major revision with a clear request: fine-tune on RTime-IT and test on out-of-distribution temporal benchmarks, and include a control that uses the same template but shuffled answers. That would separate genuine temporal learning from template memorization. The benchmark itself is solid enough to warrant that revision rather than a rejection.","headline":"The benchmark part is worth your time; the instruction-tuning claim needs external validation before you trust the 65.9 number.","tokens_in":11391,"tokens_out":2795,"would_cite":true,"duration_ms":23777,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"RTime-QA measures whether video models grasp atomic temporal events, and finds they mostly do not.","keywords":["atomic temporal event understanding","temporal negative samples","video question answering","large multimodal models","instruction tuning","benchmark","strict accuracy","temporal order"],"falsifier":"Fine-tune a model on RTime-IT and evaluate it on a held-out temporal-negative QA set built from a different video source with independent annotations and paraphrased options; if strict accuracy collapses back toward the zero-shot level, the reported improvement is mostly template or distribution overfitting, not general temporal understanding.","tokens_in":10276,"feed_emoji":"🎬","tokens_out":7926,"duration_ms":76221,"temperature":0.7,"pith_summary":"This paper argues that existing video-language benchmarks can be answered from static appearances, so they do not measure whether large multimodal models actually track temporal order. To fix this, the authors build RTime-QA, 822 human-annotated multiple-choice questions where each video of an atomic temporal event is paired with a correct description and a temporally opposite distractor that is visually similar. On this benchmark the best models score only 34.6–38.7 strict accuracy against 97.3 for humans, while image-only models fall near random. The paper also contributes RTime-IT, a 14,096-sample instruction-tuning set built with the same temporal-negative design, and reports that fine-tuning Qwen2-VL on it raises strict accuracy to 65.9. The central claim is that atomic temporal events are a distinct, currently unsolved capability that can be isolated and improved with paired temporal negatives.","feed_headline":"Best video model gets atomic temporal events right only 34.6%","feed_subtitle":"A new 822-question benchmark with temporal negatives shows the best model trailing humans by 63 points; fine-tuning cuts the gap.","key_machinery":"The load-bearing device is the temporal negative pair $(T, \\bar{T})$: two concise captions that differ only in temporal semantics while describing nearly identical spatial scenes. Each benchmark item is built from a quadruple $(V, \\bar{V}, T, \\bar{T})$ where $V$ and $\\bar{V}$ are visually similar videos with opposite temporal events, and the strict-accuracy metric only credits a model that gets both directions right. This design removes the static-image shortcut that the paper says lets image-only models ace older video benchmarks. The atomicity requirement—brief, single-event captions averaging about six words—forces the model to commit to the specific temporal progression in the clip.","core_discovery":"RTime-QA's organizing unit is the triplet $(V, T, \\bar{T})$: a video $V$ depicting an atomic temporal event—a brief event whose identity is fixed by temporal progression rather than static appearance—a correct short description $T$, and a temporally negative description $\\bar{T}$ that shares the same objects and static appearance but describes the opposite event, such as folding versus unfolding a chair or upward versus downward motion. Each question is a forced choice between $T$ and $\\bar{T}$, and the strict-accuracy metric requires a model to answer correctly for both $V$ and $\\bar{V}$, so always choosing one description is counted as failure. The paper reports that most evaluated large multimodal models score below random choice on strict accuracy, with Qwen2-VL at 34.6 and Qwen2.5-VL at 38.7 against 97.3 for humans, and that training on RTime-IT improves Qwen2-VL's strict accuracy to 65.9 and accuracy from 65.9 to 77.9. This is presented as evidence that current models lack robust atomic temporal event understanding and that instruction tuning explicitly structured around temporal negatives can substantially improve it.","pith_inferences":["Because RTime's original captions are long descriptions, RTime-QA is essentially testing how well models perform with minimal temporal contrast; we would expect performance on the same videos to rise if the model is also given the longer caption as context, which the paper does not test.","The large drop from accuracy to strict accuracy suggests many models pick whichever caption is more plausible regardless of the video; an independent diagnostic could report the share of same-answer errors to quantify this shortcut directly.","If RTime-IT's benefit transfers, it should appear on other temporal video benchmarks and in retrieval accuracy on the source corpus; that transfer test is absent from the paper.","The paper's own footnote calls the project work in progress, so the 65.9 figure is a first snapshot rather than a settled endpoint."],"forward_implications":["Modern video large multimodal models are far from human-level on atomic temporal events, with the best reported models scoring 34.6–38.7 strict accuracy versus 97.3 for humans.","Models trained only on images perform worst, supporting the benchmark's claim that single-frame shortcuts are not enough here.","Frame count matters: Qwen2-VL improves from 5.1 with 2 frames to 34.6 with 32 frames, unlike older benchmarks where extra frames gave little gain.","Instruction tuning on temporal-negative pairs is a promising lever: Qwen2-VL's strict accuracy nearly doubles after training on RTime-IT.","Publicly available video training data may lack enough temporal emphasis, since models trained with private video data lead the table."],"supporting_citations":[{"why":"Supplies the RTime video corpus with temporal negative pairs that RTime-QA and RTime-IT are filtered and annotated from.","marker":"Du et al., 2024"},{"why":"Qwen2-VL is the strongest zero-shot model in the paper's focus and the model fine-tuned on RTime-IT, so it carries the main results.","marker":"Wang et al., 2024a"},{"why":"LLaVA1.5 is the image-only baseline whose low strict accuracy demonstrates the benchmark's temporal, non-static demand.","marker":"Liu et al., 2024b"},{"why":"TemporalBench is the prior fine-grained temporal benchmark that the paper contrasts as lacking atomic temporal negatives.","marker":"Cai et al., 2024"},{"why":"VideoMME is the prior comprehensive video benchmark the paper contrasts for the same limitation.","marker":"Fu et al., 2024"},{"why":"FreeVA's strong performance on old benchmarks motivates the claim that existing video QA can be solved without temporal understanding.","marker":"Wu, 2024"},{"why":"EgoSchema supplies the observation that more frames do not help on existing benchmarks, which RTime-QA's frame scaling contrasts with.","marker":"Mangalam et al., 2023"}],"fun_headline_variants":["RTime-QA: Best model only 34.6% on atomic temporal understanding","New benchmark exposes video models' atomic time blind spot","Fine-tuning on RTime-IT boosts temporal accuracy to 65.9%","Humans beat top LMM by 63 points on atomic temporal events","Video models struggle with atomic events: new benchmark RTime-QA"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The RTime-IT improvement is claimed as a gain in general temporal understanding, but the only test used to demonstrate it, RTime-QA, is built from the same source dataset and the same annotation template as RTime-IT, so some or all of the gain could be matching the benchmark's distribution rather than learning temporal semantics generally.","fun_headline_variants_meta":{"raw":{"variants":["RTime-QA: Best model only 34.6% on atomic temporal understanding","New benchmark exposes video models' atomic time blind spot","Fine-tuning on RTime-IT boosts temporal accuracy to 65.9%","Humans beat top LMM by 63 points on atomic temporal events","Video models struggle with atomic events: new benchmark RTime-QA"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000762,"raw_usage":{"total_tokens":3425,"prompt_tokens":1032,"completion_tokens":2393,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":648,"completion_tokens_details":{"reasoning_tokens":2297}},"tokens_in":648,"tokens_out":2393,"duration_ms":11714,"temperature":1.0,"reasoning_tokens":2297,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T14:19:44.644500+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Fine-tune a model on RTime-IT and evaluate it on a held-out temporal-negative QA set built from a different video source with independent annotations and paraphrased options; if strict accuracy collapses back toward the zero-shot level, the reported improvement is mostly template or distribution overfitting, not general temporal understanding.","supporting_citations":[],"review_version":1}