{"id":"9e850e5d-06f0-46db-a5cd-5bfff6693df5","arxiv_id":"2411.15106","paper_version":2,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A comprehensive survey organizes video action understanding into three temporal scopes: recognizing full actions, predicting ongoing actions, and forecasting unseen future actions.","lead":"This paper surveys over 1,200 research works on teaching computers to understand human actions in video, grouping them by how much of an action is already visible. It offers a map of recognition, prediction, and forecasting tasks for researchers choosing problems and methods.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The temporal taxonomy misfiles text-to-video generation: Section 6.2's forecasting scope includes methods not conditioned on any observed action, which undermines the claimed complete organizational scheme.","rationale":"The reader's weakest assumption focused on the qualitative selection of landmark papers, which is a valid concern about representativeness. My stress-test identifies a more direct, internal threat to the central claim: the taxonomy itself appears to misplace a major task category. Because the survey's value is explicitly organizational, a task whose placement contradicts the paper's own definitions undermines the claimed completeness and coherence. This is not an external-consensus disagreement but an internal consistency problem: Section 1.1 defines forecasting as requiring observed actions, while Section 6.2 includes unconditional and text-conditioned generation. If the concrete test confirms this, the paper should be accepted only conditionally, for example by reframing video generation as a separate scope or by narrowing the forecasting definition to action-conditioned generation. The survey is otherwise comprehensive and well written; the concern is specific and testable rather than a general skepticism about qualitative survey methodology.","tokens_in":56027,"tokens_out":2904,"duration_ms":30925,"concrete_test":"Audit every work cited in Section 6.2 (especially Section 6.2.2 'Text-conditioned generation') and record whether the method's input includes an observed video of an action. Concretely: for each paper, read its abstract (or the survey's own description) and classify the conditioning input as (a) observed action video, (b) text only, (c) image only, (d) other. If more than roughly 10% of the cited generation papers are text-only or image-only conditioned, the forecasting scope is not faithful to the definition in Section 1.1, and the taxonomy needs revision or a qualifying explanation.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim is that the three temporal scopes (recognition, prediction, forecasting) provide a complete organizational scheme for action understanding (Section 1.1). Section 1.1 defines forecasting as using 'the currently observed action(s) to reason about future actions not yet observed.' Section 6.2, however, places 'Video Generation' under forecasting and explicitly includes text-to-video (T2V) and image-to-video (I2V) generation. T2V methods condition on a textual prompt, not on an observed action; I2V methods condition on a single static image, which need not depict an action in progress. The section's own examples (e.g., 'A vibrant underwater scene of a scuba diver exploring a shipwreck') have no observed-action input. This is an internal inconsistency: the forecasting definition requires an observed ongoing action, yet a substantial portion of the section's cited works generate future or arbitrary content from non-action inputs. The taxonomy therefore does not cleanly partition the tasks it claims to organize, and the 'holistic' coverage claim is weakened because a major task family (generative video modeling) is placed under a scope that does not match its inputs. This is a load-bearing issue for the central claim, not merely a presentational quibble: if the scopes are not well-defined by their own criteria, the survey's organizing contribution is less secure than the text asserts.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey proposes a taxonomy of video action understanding organized by three temporal scopes: recognition of fully observed actions, prediction from partially observed ongoing actions, and forecasting of subsequent unobserved actions. Within this structure it reviews video modeling approaches, datasets, and a wide range of tasks, and it concludes with suggested research directions. The paper explicitly claims that prior surveys focus on specific aspects and that a holistic survey of action understanding is missing, which this work aims to fill using 1,284 references and a bibliometric trend analysis.","tokens_in":56300,"tokens_out":5141,"duration_ms":50374,"significance":"The survey's main strength is breadth: it assembles a very large literature across uni- and multimodal action understanding, provides useful dataset tables and task-by-task overviews, and offers a clear conceptual organization through the temporal-scope distinction. The paper does not rely on the authors' own methods for its central structure, so circularity risk is low. If the boundary issues in the taxonomy are resolved, this would be a valuable reference for the community. However, the completeness claim is currently stronger than the evidence: the forecasting scope contains tasks that do not satisfy its own definition, and the paper's selection and bibliometric methodology is not reproducible.","major_comments":[{"comment":"The forecasting scope is defined in Section 1.1 as using 'the currently observed action(s) to reason about future actions not yet observed,' but Section 6.2 places Video Generation under this scope and explicitly includes text-to-video generation, whose input is a textual prompt rather than an observed action, as well as unconditional generative models. The example in Figure 15 ('A vibrant underwater scene of a scuba diver exploring a shipwreck') illustrates that the conditioning input need not contain any observed action at all. This is an internal inconsistency: a substantial family of tasks placed in the forecasting scope does not meet the scope's defining condition. Since the central claim is that the three temporal scopes provide a complete organizational scheme, the forecasting definition should either be tightened to cover only action-conditioned future synthesis, with text-to-video and similar generation treated separately, or the completeness claim should be qualified accordingly.","section":"Section 6.2 / Section 1.1"},{"comment":"The paper asserts holistic coverage ('this survey fills this void', Section 1.1) and reports 1,284 cited papers, but the selection protocol is not described. Figure 1 states that landmark papers are selected 'by their relevance to the period's trends,' and the Figure 17 counts are 'approximated from all works citing influential papers with >=300 citations' per the footnote, using Google Scholar. Without a reproducible search strategy, inclusion and exclusion criteria, and an estimate of the approximation error, the completeness claim is difficult to verify independently. Please add a methodology paragraph or supplementary protocol describing how the literature was searched and screened, and qualify the bibliometric numbers as indicative rather than exact.","section":"Section 7 / Figure 17"}],"minor_comments":[{"comment":"Two subfigures are both labeled '(d)': the Visual Abductive Reasoning panel and the Video Alignment panel, which makes the subsequent (e) and (f) labels inconsistent. Please renumber the panels.","section":"Figure 13"},{"comment":"The sentence 'We overview of benchmarks in three groups' is grammatically incomplete; it should read 'We overview benchmarks in three groups' or similar.","section":"Section 3.3"},{"comment":"The column header 'Video adaptations 1' contains a dangling superscript '1' that is not explained anywhere in the text or table caption.","section":"Table 3"},{"comment":"In the list of state-based tasks, the second entry says 'defining start-end times (action progress prediction)', which repeats the first entry; this should presumably refer to Event Boundary Detection.","section":"Section 5.3.2"},{"comment":"The dataset list labels entry 89 as 'EK-101 (Damen et al 2022)', but the text and Table 2 consistently refer to this benchmark as EPIC-KITCHENS-100 (EK-100).","section":"Figure 3"}],"recommendation":"major_revision","confidential_remarks":"To the editor: this is a broad and potentially highly cited survey, and I do not see a need for rejection. The main issue is that a core organizational category is internally inconsistent with its own definition, and the completeness claim is backed by a non-reproducible selection process. Both are fixable within the manuscript's scope, but they should be addressed before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe paper you'd want to know about: a 1,284-paper survey of video action understanding organized by three temporal scopes—recognition of complete actions, prediction from partial observations, and forecasting of future actions. The genuinely new thing is the temporal framing and the breadth. Table 1's comparison of prior surveys by scope is a real service, and the dataset table and task-by-task organization make it a usable reference. The authors do a good job of connecting overlapping tasks (TAL, STAD, VRC, anticipation, anomaly detection), and the future-directions section is thoughtful, not filler. Their own prior work appears in method lists, but nothing in the taxonomy depends on it.\n\nThe main soft spot is one the stress-test flagged, and it holds up on reading. Section 1.1 defines forecasting as using \"the currently observed action(s) to reason about future actions not yet observed.\" Section 6.2 then places video generation under forecasting and explicitly includes text-to-video and image-to-video generation. T2V conditions on a prompt, not on an observed action; the section's own examples (scuba diver exploring a shipwreck) show this. So the taxonomy does not cleanly partition everything it claims to organize. This is not fatal—the survey's value does not rest entirely on the taxonomy being a perfect partition—but it is a real inconsistency in the paper's central organizational claim. A revision should either narrow Section 6.2 to generation conditioned on observed video, or add a fourth scope (synthesis) and say plainly that the three-scope scheme covers understanding tasks while generation is adjacent.\n\nThe secondary weakness is the selection method: landmark papers are chosen qualitatively and the bibliometric trend counts are approximated from Google Scholar. That limits the weight of the \"comprehensive\" claim, but for a survey of this size it is an acceptable choice if stated as a limitation, which it mostly is.\n\nWho gets value: anyone entering the field or needing a map of task boundaries, and researchers writing related-work sections. It deserves a serious referee; the right referee would send it back with a request to fix Section 6.2 rather than reject. I'd cite it as a reference, and I'd bring it to the reading group to argue about the taxonomy.","headline":"A comprehensive, genuinely useful survey whose promised three-scope temporal taxonomy does not cleanly contain text-to-video generation; fix that section and it is a solid reference.","tokens_in":56776,"tokens_out":2549,"would_cite":true,"duration_ms":26400,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey argues that the field of video action understanding is holistically organized by three temporal scopes—recognizing observed actions, predicting ongoing ones, and forecasting unseen ones—and that no prior survey has covered all…","keywords":["action understanding","action recognition","action prediction","action anticipation","video datasets","multimodal video understanding","video generation","survey taxonomy"],"falsifier":"A citation audit of the field would settle the completeness claim: take a representative sample of recent action-understanding publications and check whether each one can be assigned to one of the three temporal scopes with high agreement among independent coders, and whether the bibliometric counts behind the survey's research-trend figure (approximated from all works citing influential papers with at least 300 citations on Google Scholar) match exact citation data. If a substantial research line fits none of the three scopes, or if the approximated counts misstate the growth of major lines, the taxonomy and the claimed gap it fills would need revision.","tokens_in":55841,"feed_emoji":"🕒","tokens_out":6201,"duration_ms":54698,"temperature":0.7,"pith_summary":"This survey sets out to establish that the field of video action understanding, which has grown into dozens of task-specific research lines, is best understood as three temporal scopes: recognition of actions observed in full, prediction from partial observations of ongoing actions, and forecasting of actions not yet seen. The authors argue that prior surveys each cover one slice of the field—specific tasks, modalities, or scopes—and that no critical overview spanning all three has existed; this survey of 1,284 works is presented as filling that gap. The three-way division matters because each scope concentrates its own modeling problems: representing long-range dependencies for recognition, handling procedural proximity and partial evidence for prediction, and managing error accumulation and fixed anticipation intervals for forecasting. If the taxonomy holds, researchers and practitioners get a unified map of the field and of where the open problems concentrate.","feed_headline":"Three time scopes organize the whole field of video actions","feed_subtitle":"A 1,284-paper survey claims recognition, prediction, and forecasting give the field its missing holistic map.","key_machinery":"The organizing device is the temporal-scope taxonomy, a partitioning of tasks by how much of the action sequence the model can access: full observation (recognition), an observable prefix of an ongoing action (prediction), and only the current action while reasoning about a future, unobserved one (forecasting). Figure 2 formalizes this with a timeline in which an action of duration $\\tau_1$ is only partially observable ($\\tau_{1,\\rho} < \\tau_1$), a transition period $0 \\leq \\tau_{1\\to2}$ separates actions, and the next action has duration $\\tau_2$. The taxonomy is what carries the survey's argument: it is the grid into which the 1,284 surveyed works, the dataset tables, and the per-section challenges are placed, and it is the basis for the claim that no prior survey covers the field holistically.","core_discovery":"The central claim, stated in the survey's taxonomy section, is that the field of video action understanding can be holistically organized by where a model sits on the action timeline: recognition tasks use the full observation of an action of duration $\\tau_1$, prediction tasks use only an observable prefix $\\tau_{1,\\rho}$ of an ongoing action, and forecasting tasks use the currently observed action to reason about a subsequent unobserved action after a transition interval. Around this division the survey arranges the field's main task families—temporal localization, spatiotemporal detection, repetition counting, and language-based recognition under the first scope; early action prediction, frame prediction, state-change tasks, and anomaly detection under the second; action anticipation and video generation under the third—along with the modeling approaches, datasets, and benchmarks each family relies on. The paper also uses the three-scope division to expose challenges specific to each temporal position, such as procedural proximity between similar actions under partial observation, the modality gap in video-language models, and the absence of standardized evaluation for generated video. The discovery, if accepted, is that action understanding is not a loose cluster of tasks but a field with a coherent temporal backbone.","pith_inferences":["If the temporal taxonomy is adopted as a field-wide map, an implicit prediction follows: models trained for tasks within one scope should transfer more readily to other tasks in the same scope than across scopes—a testable hypothesis the paper does not itself run.","The taxonomy's 'time as a stepping stone' framing suggests a natural extension beyond the paper: the same three-scope division could organize audio-only action understanding or robot policy learning, where partial observation and forecasting are equally central.","Because the paper's landmark selection is qualitative, its comprehensiveness claim could be independently stress-tested with an automated citation-network analysis that checks whether the three scopes cover the field's citation-dense research lines without residue."],"forward_implications":["If the taxonomy is right, prior surveys are complementary slices of a single whole, and the field's history from early template matching to video-language foundation models reads as one continuous timeline rather than disconnected task communities.","Researchers entering action understanding can locate any task—temporal action localization, early action prediction, action anticipation, video generation—by its temporal scope and inherit the challenges and solution families the survey attaches to that scope.","The three scopes highlight shared weaknesses: real-time and multi-person settings are under-addressed in prediction and forecasting alike, and video generation lacks standardized benchmarks that test physical plausibility and prompt alignment.","Datasets and modeling approaches surveyed under each scope expose where benchmarks are missing, notably high-resolution video for frame prediction and unified evaluation protocols for generated video."],"supporting_citations":[{"why":"Supplies the early action-recognition survey against which the recognition scope's coverage is positioned.","marker":"Poppe (2010)"},{"why":"Anchors the survey of learned-feature action recognition that the paper extends toward multimodal tasks.","marker":"Herath et al (2017)"},{"why":"Provides the prior overview of predictive models that defines the prediction scope's baseline coverage.","marker":"Rasouli (2020)"},{"why":"Surveys action anticipation in egocentric video, the main precedent for the forecasting scope.","marker":"Rodin et al (2021)"},{"why":"Overview of short- and long-term action anticipation methods that the forecasting chapter builds on.","marker":"Zhong et al (2023b)"},{"why":"Comprehensive review of language-enabled action understanding that frames the multimodal recognition discussion.","marker":"Madan et al (2024)"},{"why":"Prior survey of attention-based video approaches used to delimit the spatiotemporal modeling review.","marker":"Selva et al (2023)"}],"fun_headline_variants":["Video actions: one timeline, three scopes","Recognition, prediction, forecasting: the action trifecta","Three temporal scopes unify video action research","Time splits action understanding into three tasks","A survey maps video action via time horizons"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's claim of holistic coverage rests on the assumption that the landmark papers it selects and the tasks it includes are truly representative of the field; the selection criteria are qualitative, since landmark papers are chosen by their relevance to the period's trends, so any bias in that selection would leave the claimed completeness unestablished.","fun_headline_variants_meta":{"raw":{"variants":["Video actions: one timeline, three scopes","Recognition, prediction, forecasting: the action trifecta","Three temporal scopes unify video action research","Time splits action understanding into three tasks","A survey maps video action via time horizons"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000332,"raw_usage":{"total_tokens":1835,"prompt_tokens":923,"completion_tokens":912,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":539,"completion_tokens_details":{"reasoning_tokens":841}},"tokens_in":539,"tokens_out":912,"duration_ms":7366,"temperature":1.0,"reasoning_tokens":841,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:28:28.566668+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A citation audit of the field would settle the completeness claim: take a representative sample of recent action-understanding publications and check whether each one can be assigned to one of the three temporal scopes with high agreement among independent coders, and whether the bibliometric counts behind the survey's research-trend figure (approximated from all works citing influential papers with at least 300 citations on Google Scholar) match exact citation data. If a substantial research line fits none of the three scopes, or if the approximated counts misstate the growth of major lines, the taxonomy and the claimed gap it fills would need revision.","supporting_citations":[],"review_version":1}