{"id":"dce62e9f-a1a7-4b36-b833-bd518873aace","arxiv_id":"2412.18688","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of long video generation that groups methods into frame-by-frame, planned-segment, and all-at-once approaches, but is weakened by inconsistent data tables.","lead":"This paper surveys recent efforts to generate long videos, grouping methods into frame-by-frame, planned-segment, and all-at-once approaches. It is intended as a reference for researchers, but its summary tables contain several internal inconsistencies.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'comprehensive foundation' claim depends on accurate catalogs, but Tables 1, 7, and 8 contain repeated misclassifications and provenance errors; until these are verified against primary sources, the central claim is not secure.","rationale":"I read the paper in good faith: it is a broad survey with a plausible high-level structure, a useful timeline, and coverage spanning 2021-2025. The divide-and-conquer emphasis and the distinction among autoregressive, divide-and-conquer, and implicit paradigms are reasonable organizing ideas. The paper also deserves credit for attempting to catalog datasets and metrics rather than only methods. However, the central claim is explicitly that the survey 'would serve as a comprehensive foundation' for the field. A foundation must be trustworthy at the level of individual table entries and category labels. The specific errors identified above - StyleGAN-V and DIGAN listed under autoregressive approaches despite their own descriptions, duplicate and conflicting references in the benchmark comparison, and misattributed commercial models - directly undermine that trust. These are correctable, and they do not invalidate the entire survey, which is why I do not recommend REJECT. They are also serious enough that ACCEPT would be premature while such entries remain. The reader's CONDITIONAL verdict is therefore appropriate and unchanged. I partially agree with the reader's weakest-assumption analysis: the reader emphasizes sampling representativeness and taxonomy exhaustiveness, whereas I find the more decisive issue to be internal accuracy and provenance of the catalogs that implement that taxonomy. Both concerns bear on the same central claim, but mine is more concretely checkable and more directly tied to the paper's evidentiary core.","tokens_in":30736,"tokens_out":6349,"duration_ms":59344,"concrete_test":"Re-extract every row of Tables 1, 7, and 8 from the primary sources cited in those rows, recording for Table 1 whether the cited paper itself describes the method as autoregressive, and for Tables 7 and 8 whether the reported FVD/FID/IS/CLIPSIM/VBench values match the source table. At minimum, check StyleGAN-V and DIGAN self-descriptions, the duplicate Phenaki/ModelScope/Gen-2/Gen-3 entries, and the Gen-2 and Gen-4 Alpha reference attributions. If mismatches affect more than roughly 10% of sampled rows, the 'comprehensive foundation' claim should remain CONDITIONAL; if the errors are isolated typos across otherwise accurate tables, ACCEPT becomes appropriate.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The survey's central value claim (Abstract, Section 1.1) is that it provides a reliable 'comprehensive foundation' for long-video generation. That claim rests on the accuracy of its taxonomy and benchmark tables, and several load-bearing inconsistencies are present. First, Table 1 and Section 3.1 classify StyleGAN-V [85] and DIGAN [40] as 'Auto Regressive Approaches', but Section 2.1.2 and the cited papers describe them as continuous/implicit GAN generators; this makes the three-paradigm partition in Section 3 appear overlapping rather than cleanly exhaustive. Second, Table 7, labeled 'Comprehensive benchmark comparison', duplicates references (Phenaki [15]/[63], ModelScope [177] and ModelScopeT2V [177], VideoDirectorGPT [11]/[166], Video LDM [179]/Video Diffusion [173]) and includes unlabeled or otherwise unverifiable entries, so a reader cannot determine which numbers come from which source. Third, Table 8 cites reference [188] for both Gen-2 and Gen-3, and the references attribute Gen-4 Alpha [17] and Gen-2 [13] to Midjourney even though the running text and public sources identify RunwayML. These are not stylistic quibbles: a reference survey must let readers trust table entries and category labels. Until the tables are re-verified, the central claim of serving as a reliable foundation is not supported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey reviews the emerging area of long video generation, organizing the literature into three generation paradigms (autoregressive, divide-and-conquer, and implicit latent-space generation) and covering backbone architectures, tokenization strategies, input control mechanisms, datasets, and evaluation metrics. The authors state in the abstract and in Section 1.1 that the survey 'would serve as a comprehensive foundation' for the field, and they position their contribution against two earlier surveys [18,19] by focusing on divide-and-conquer methods, agent-based frameworks, and short-to-long video transitions. The paper catalogs over 190 papers from 2021 to 2025, with timeline figures and several comparison tables.","tokens_in":31020,"tokens_out":4035,"duration_ms":35912,"significance":"If the cataloging is accurate, the survey would be a useful entry point for researchers: it gathers recent work (2023-2025) that earlier surveys do not cover in depth, and its three-paradigm taxonomy, particularly the divide-and-conquer subsection, is a reasonable high-level organization of the field. The paper also draws attention to and compares datasets and metrics, which are often treated separately in the literature. However, the central value of a survey of this kind is reliability of its factual claims; the inconsistencies in the taxonomy and in the benchmark tables described below mean that the 'comprehensive foundation' claim is not currently supported by the manuscript as written. The paper contains no derivations or experiments, so its correctness hangs entirely on the accuracy of its reporting.","major_comments":[{"comment":"Table 1, which is labeled 'Auto Regressive Approaches,' lists StyleGAN-V [85] and DIGAN [40] as autoregressive methods. This contradicts the paper's own description in Section 2.1.2, where both are presented as continuous/implicit GAN generators, and it also conflicts with the definition of implicit video generation in Section 3.3, which explicitly excludes extrapolation (autoregressive) and interpolation (divide-and-conquer). As a result, the three-paradigm partition in Section 3 is not clean or exhaustive, and the reader cannot determine which papers belong to which paradigm. Please reclassify these models or justify their placement in the autoregressive category.","section":"Section 3.1, Table 1"},{"comment":"Table 7, titled 'Comprehensive benchmark comparison,' mixes FVD numbers from different evaluation protocols (UCF-101, BAIR) with FID and CLIPSIM numbers from MSR-VTT, and the footnote does not resolve which cell comes from which benchmark. The table also contains duplicate entries: Phenaki appears as both [15] and [63], ModelScope and ModelScopeT2V both cite [177], VideoDirectorGPT appears as both [11] and [166], and Video LDM [179] is the same work as Video Diffusion [173]. Without per-row benchmark provenance and deduplication, these numbers cannot be verified or compared. Please provide a source for every reported value and remove the duplicate rows.","section":"Table 7"},{"comment":"Table 8 reports per-dimension VBench scores that are largely in the 80-99 range (e.g., LaVie: Subject Consistency 91.41, Background Consistency 97.47, Motion 96.38) but lists Overall scores around 26-28 for the same models (e.g., LaVie Overall 26.41). Since the caption states that the table evaluates models across 12 VBench dimensions, the Overall score should be derivable from those dimensions; the reported values appear to be from a different, unlabeled benchmark. In addition, Gen-2 is cited to [188], which is the RunwayML Gen-3 Alpha page, and references [13] and [17] attribute Gen-2 and Gen-4 Alpha to 'Midjourney Team,' contradicting the running text in Section 1, which identifies both as RunwayML models. These provenance errors prevent readers from trusting the table.","section":"Table 8"},{"comment":"The text states that ARLON [84] achieves '128× compression (8× spatial + 6× temporal downsampling),' but 8 multiplied by 6 is 48, not 128. Either the compression factor or the downsampling factors are incorrect. Please verify this claim against the ARLON paper and correct the numbers.","section":"Section 3.1, ARLON description"}],"minor_comments":[{"comment":"The survey organization paragraph contains placeholder cross-references 'Section datasets' and 'Section metrics' that do not correspond to any section, and Section 1.2 uses bracket-style references such as '[3.2]' and '[5.1]' that do not match the manuscript's section numbering.","section":"Section 1.4"},{"comment":"The reference list contains duplicate entries for the same works, including Phenaki (appearing as [15] and [63]), MEVG (as [101] and [104]), VideoDirectorGPT (as [11] and [166]), and Video Diffusion/Video LDM (as [173] and [179]). These should be consolidated.","section":"References"},{"comment":"The sentence 'Here are examples from the MSR-VTT dataset...' appears twice verbatim, once in the main text and once immediately after the figure caption for Figure 17.","section":"Section 6.2"},{"comment":"The sentence about Cosmos cites reference [84], which is the ARLON paper; the Cosmos World Foundation Model is separately cited as [113] later in the same section. The earlier citation appears to be a mismatch.","section":"Section 3.3"},{"comment":"Reference [51], an anomaly-detection paper by one of the authors, is cited in the autoencoder background section. The connection to video generation is not explained, and it is not used to support any of the survey's central claims; consider removing it or providing a clearer rationale.","section":"Section 2.2.1"},{"comment":"The introduction to Section 7 contains stray LaTeX commands: 'textit Video Quality Metrics' and 'textit Semantics Quality Metrics'. Additionally, Table 5 lists 'Pandas 70m' while the reference [141] is titled 'Panda-70M'; please make the names consistent.","section":"Section 7"},{"comment":"The caption of Figure 3 claims that most papers on long video generation were published in 2023-2025, but no data source or counting methodology is provided for the histogram, so the claim cannot be verified.","section":"Figure 3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a literature survey, so the bar for acceptance should be accuracy and completeness of reporting. The taxonomy inconsistencies and table provenance errors are fixable through re-verification against primary sources, but they are load-bearing for the paper's stated contribution of being a 'comprehensive foundation.' I would recommend asking the authors to re-check every entry in Tables 1, 7, and 8 against the cited papers, to correct the ARLON compression arithmetic, and to resolve the duplicate references. The self-citation [51] is peripheral and should be removed or justified. Given the fast-moving nature of the field, the survey's coverage through May 2025 is a strength, but the manuscript needs a careful revision pass before it can serve as a reliable reference."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague, this survey is a mixed bag. The good: it pulls together 190+ recent papers on long video generation, organizes them into a three-way taxonomy (autoregressive, divide-and-conquer, implicit), and offers a handy catalog of datasets and metrics. For someone entering the field, it is a reasonable starting point. The divide-and-conquer thread, especially the LLM-as-director and agent-based variants, is timely, and the authors clearly spent time on the recent literature.\n\nThe problems are in the details, and they are not cosmetic. The benchmark tables are unreliable. In Table 7, metrics come from different benchmarks (UCF-101 FVD, MSR-VTT FID/CLIPSIM) without clear labeling; several entries duplicate references or cite the wrong paper. Table 8 has overall VBench scores around 26 next to per-dimension scores in the 80–99 range, which is arithmetically impossible: with most dimensions that high, the weighted overall cannot collapse to 26. That table also cites the Gen-3 runway reference for Gen-2, and the references attribute Gen-4 Alpha and Gen-2 to Midjourney even though the text correctly identifies them as RunwayML. Table 1 misclassifies StyleGAN-V and DIGAN as autoregressive; they are continuous GAN generators, as the paper itself says in Section 2.1.2. There is also a duplicated block of text between Sections 3.2.1 and 3.2.2, suggesting the manuscript needed a careful edit.\n\nDo these flaws sink it? Not entirely. The narrative is coherent, and the taxonomy is a reasonable way to frame the field even if it is mostly a re-labeling of existing categories. The survey could genuinely help a newcomer. But the central promise of being a comprehensive, trustworthy map is undercut by the table errors. A reader cannot use these tables without going back to primary sources, which diminishes the survey's value as a reference.\n\nMy take: it deserves a serious referee, but only if the authors fix the tables and reference list. As it stands, I hesitate to cite it for any specific number. For the reading group, it might spark a useful discussion about how to evaluate the reliability of surveys, but I would not put it high on the list.\n\nRecommendation: send it to peer review with the expectation of substantial revision. The topic is important, the coverage is current, and the taxonomy is a reasonable skeleton. But every number should be verified against primary sources, the references corrected (Gen-2, Gen-4, Phenaki duplicates), and Table 1 fixed. If the authors do that, this could become the default survey for the field; right now it is not there yet.","headline":"A broad, useful survey that is let down by unreliable benchmark tables and attribution errors; worth revising, not rejecting.","tokens_in":31521,"tokens_out":4105,"would_cite":false,"duration_ms":61631,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey claims that long video generation is best understood as three paradigms—autoregressive, divide-and-conquer, and implicit latent-space synthesis—and argues that divide-and-conquer, especially LLM-guided planning, is the key to…","keywords":["long video generation","text-to-video synthesis","divide-and-conquer paradigm","LLM-guided generation","diffusion models","autoregressive generation","video generation survey","video evaluation metrics"],"falsifier":"A systematic literature search for long video generation papers published between 2021 and mid-2025 that are absent from this survey and do not fit any of the three paradigms (for example, direct world-simulator models that generate video without planning, sequential prediction, or compressed latent-space synthesis) would refute the claim that the taxonomy is exhaustive. A reproducible annotation study measuring the fraction of relevant papers that independent annotators cleanly assign to one of the three paradigms would likewise test whether the partition is well-defined.","tokens_in":30537,"feed_emoji":"🎬","tokens_out":3760,"duration_ms":34368,"temperature":0.7,"pith_summary":"The paper is a survey aiming to serve as the comprehensive foundation for long video generation research. It organizes the field into three generation paradigms and dives deepest into divide-and-conquer, where an LLM plans scenes and a separate generator fills frames. The authors argue this approach addresses scalability and narrative control better than pure autoregressive generation. The survey also catalogs datasets, metrics, and open problems, positioning itself as the missing focused review of a fast-growing field.","feed_headline":"Survey maps long video generation into three paradigms","feed_subtitle":"Autoregressive, divide-and-conquer, and latent-space methods compared, with a deep dive into LLM-guided planning.","key_machinery":"The central organizing mechanism is the divide-and-conquer paradigm: generate keyframes or short clips from prompts, often planned by an LLM, then interpolate or stitch them into a continuous long video. The LLM-as-director pattern is the load-bearing instance, in which an LLM produces a narrative blueprint (scene descriptions, layouts, bounding boxes, actions) that a separate video diffusion module executes. This is contrasted with autoregressive prediction, which conditions each frame on previous frames, and implicit generation, which synthesizes the whole video from a compressed spatiotemporal latent representation.","core_discovery":"On its own terms, the paper establishes that long video generation can be systematically mapped into three paradigms: autoregressive frame prediction, divide-and-conquer keyframe-and-interpolation, and implicit synthesis from a compressed latent space. Within divide-and-conquer, it identifies three sub-patterns: LLM-as-director, multi-stage or agent-based frameworks, and transition or compositional stitching. The paper further claims that this taxonomy, together with its catalog of datasets and evaluation metrics, provides a comprehensive and current foundation that earlier surveys lacked, particularly in its detailed treatment of divide-and-conquer.","pith_inferences":["A consequence the paper leaves implicit: if the three-paradigm taxonomy is accepted, the implicit category is likely to absorb most future foundation models, while divide-and-conquer will dominate controllable and story-driven applications; the two may eventually converge.","The paper's emphasis on divide-and-conquer is a bet on the value of explicit planning; one could test whether planning-based methods actually beat end-to-end latent synthesis on long-range coherence when compute budgets are held equal.","A practical extension: use the survey's taxonomy to build a benchmark that samples methods from each paradigm and evaluates them on identical prompts and durations; the results would either validate the taxonomy's utility or reveal overlaps that blur its boundaries."],"forward_implications":["If the survey's map is correct, future long video models will increasingly separate a planning stage from a frame-generation stage, since divide-and-conquer supports parallel keyframe generation and finer narrative control.","The identified scarcity of large-scale, richly captioned video datasets becomes a concrete bottleneck; building datasets that combine scale with spatial and temporal detail would directly accelerate progress.","Metrics that rely on manual feedback, such as FETV, VBench, and MiraBench, are a scalability bottleneck, making fully automated semantic and temporal evaluation a clear research target.","The divide-and-conquer sub-taxonomy defines a modular design space—planner, generator, transition module—that is testable by comparing how well different combinations reconstruct the same storylines.","The survey's open challenges, including audio alignment and physical dynamics, point to next steps that are largely orthogonal to the core generation paradigm and can be pursued on top of any of the three approaches."],"supporting_citations":[{"why":"The prior long-video-generation survey whose divide-and-conquer treatment is described as lacking, motivating this paper's gap claim.","marker":"[18]"},{"why":"The broader video-generation survey that treats long video as one topic among many, establishing the need for a focused review.","marker":"[19]"},{"why":"Sora is presented as the current state of the art and the motivating example of the one-minute length limit.","marker":"[1]"},{"why":"VideoDirectorGPT is the canonical LLM-as-director divide-and-conquer system whose planning and layout execution anchor the taxonomy.","marker":"[11]"},{"why":"Gen-L-Video demonstrates the early short-clip stitching approach that the paper extends into the transition sub-paradigm.","marker":"[16]"},{"why":"SEINE frames transitions as masked diffusion, a key reference for the compositional/transition category of divide-and-conquer.","marker":"[25]"},{"why":"Mora is the multi-agent divide-and-conquer example that defines the agent-based sub-paradigm.","marker":"[93]"},{"why":"Free-Bloom supplies the zero-shot, training-free LLM-director pattern that the paper contrasts with training-based planning models.","marker":"[21]"}],"fun_headline_variants":["Three paradigms for long video: a comprehensive survey","Divide-and-conquer emerges as key for long video generation","Long video generation: autoregressive, latent, and divide-and-conquer","Survey: planning and consistency are unsolved for long video generation","Video is worth a thousand images: survey on long video generation trends"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's comprehensiveness rests on the assumption that snowball sampling of 190+ articles from a manually selected list of venues captures a representative and complete picture of the field, and that the three-way partition of generation paradigms is exhaustive.","fun_headline_variants_meta":{"raw":{"variants":["Three paradigms for long video: a comprehensive survey","Divide-and-conquer emerges as key for long video generation","Long video generation: autoregressive, latent, and divide-and-conquer","Survey: planning and consistency are unsolved for long video generation","Video is worth a thousand images: survey on long video generation trends"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000404,"raw_usage":{"total_tokens":2059,"prompt_tokens":854,"completion_tokens":1205,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":470,"completion_tokens_details":{"reasoning_tokens":1120}},"tokens_in":470,"tokens_out":1205,"duration_ms":12703,"temperature":1.0,"reasoning_tokens":1120,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T04:33:40.944384+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A systematic literature search for long video generation papers published between 2021 and mid-2025 that are absent from this survey and do not fit any of the three paradigms (for example, direct world-simulator models that generate video without planning, sequential prediction, or compressed latent-space synthesis) would refute the claim that the taxonomy is exhaustive. A reproducible annotation study measuring the fraction of relevant papers that independent annotators cleanly assign to one of the three paradigms would likewise test whether the partition is well-defined.","supporting_citations":[],"review_version":1}