{"id":"338d3022-86ea-4228-bf78-5c2f82e411c0","arxiv_id":"2508.12644","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"DyCrowd is a video-based framework for spatiotemporally consistent 3D crowd reconstruction that uses group-guided motion optimization, a VAE motion prior, and a new virtual benchmark, VirtualCrowd.","lead":"This paper introduces DyCrowd, a method that reconstructs the poses, positions, and shapes of hundreds of people from a single large-scene video, rather than from separate still images. It adds group-guided optimization to handle occlusions and a virtual benchmark called VirtualCrowd for evaluation.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"SOTA claim rests on synthetic-only evaluation; real-video evidence is needed before the central claim can be assessed.","rationale":"The reader's UNVERDICTED verdict is appropriate. The paper's strongest claim is an empirical SOTA claim, and the only evaluation benchmark named in the abstract is VirtualCrowd, a synthetic dataset contributed by the authors. The load-bearing weak point is therefore external validation: without real large-scene video results, the SOTA claim could reflect train/test distribution matching on a self-created benchmark rather than genuine reconstruction ability on real crowds. The reader's weakest assumption concerned the group-guided occlusion mechanism itself—that visible group members with similar motion can guide occluded individuals. My concern is related but distinct: even if the mechanism is internally well designed, its success is demonstrated only on synthetic data, so the empirical central claim is not yet established. I also note that the full text provided is illegible and contains a mismatched arXiv identifier, which precludes independent verification of formulas, ablations, and comparison tables. A real-data evaluation with occlusion-stratified metrics would settle the concern. No change to the reader's UNVERDICTED verdict is warranted: the evidence is insufficient to accept or reject the central claim.","tokens_in":14269,"tokens_out":4630,"duration_ms":49138,"concrete_test":"Obtain or collect a real-world large-scene video with independent 3D ground truth (for example, a synchronized multi-view camera rig or motion-capture floor with at least 50 pedestrians and natural occlusions), run DyCrowd and the same baselines under an identical protocol, and compare per-person pose, position, and shape errors broken down by occlusion status (visible, partially occluded, and fully occluded). If the reported SOTA margin does not persist on real video, or if fully occluded individuals and fully occluded groups are systematically excluded from the metrics, the central claim is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that DyCrowd is the first framework for spatio-temporally consistent 3D reconstruction of hundreds of people from a large-scene video and achieves state-of-the-art performance—is not yet supported by the evidence visible in the abstract. The only evaluation artifact named is VirtualCrowd, a synthetic benchmark created by the authors, and the abstract explicitly motivates it by saying there is no existing well-annotated large-scene video dataset. If the experiments are limited to VirtualCrowd, the SOTA comparison is performed on a dataset whose occlusion statistics, camera parameters, and human motion and shape distributions are generated by the same group that designed the method. In that setting, the core occlusion-recovery mechanism (group-guided optimization plus a VAE motion prior) can appear strong because the prior and test distribution are matched, not because it reliably reconstructs real occluded pedestrians. Additionally, the supplied full text is encoding-corrupted and unreadable, so the equations, ablations, and comparison tables cannot be checked; the inserted line 'arXiv:2508.12645v5 [cs.IR] 18 Jan 2026' does not match the paper's claimed identifier, which is a provenance concern. These issues do not prove the method is wrong, but they make the SOTA claim unverifiable from the available material.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces DyCrowd, a framework for reconstructing the 3D poses, positions, and shapes of hundreds of people over time from a single large-scene video. The proposed method combines coarse-to-fine group-guided motion optimization with a VAE-based human motion prior and an Asynchronous Motion Consistency (AMC) loss, targeting long-term dynamic occlusions in large scenes. The authors also present VirtualCrowd, a synthetic benchmark dataset for evaluating dynamic crowd reconstruction. The abstract claims state-of-the-art performance, but the submitted full text is encoding-corrupted and largely unreadable, so the methodology, equations, experimental setup, and quantitative comparisons cannot be verified from the available material.","tokens_in":14523,"tokens_out":3109,"duration_ms":34390,"significance":"The task addressed by DyCrowd is timely and practically relevant for city surveillance and crowd analysis, and the goal of spatio-temporally consistent reconstruction of hundreds of individuals from a single video is a genuine gap in the current literature. The contribution of a benchmark dataset, even a synthetic one, is useful if it is carefully validated and released. However, the significance of the work cannot currently be assessed: the central state-of-the-art claim rests on the abstract alone, no quantitative metrics are reported in readable form, and the only named evaluation artifact is a synthetic dataset constructed by the same authors. If the method and dataset are made fully readable and are validated against real-world large-scene data, the contribution could be substantial; at present it remains unsubstantiated.","major_comments":[{"comment":"The central claim of state-of-the-art performance is not supported by any readable quantitative result. The abstract provides no metrics, ablations, or error bars, and the full text is encoding-corrupted, making the comparison tables and experimental sections inaccessible. The authors should provide a readable manuscript with concrete numbers on the VirtualCrowd benchmark and, ideally, on at least one real-world large-scene video with annotations or proxy metrics.","section":"Abstract and Full Text"},{"comment":"The only evaluation benchmark named in the paper, VirtualCrowd, is introduced by the authors, and the abstract justifies its creation by the absence of existing well-annotated large-scene video datasets. If the synthetic data are generated using the same group-motion assumptions and occlusion patterns that DyCrowd explicitly encodes, then the reported state-of-the-art result would be at least partly self-confirming. The paper should describe the data generation process in detail, provide statistics on occlusion rates and motion diversity, and include a transfer experiment to real video to demonstrate that the method generalizes beyond the synthetic distribution.","section":"VirtualCrowd evaluation and circularity"},{"comment":"The core mechanism assumes that visible, unoccluded members of a group with similar motion segments can reliably guide the reconstruction of occluded members. The abstract does not address failure cases where an occluded individual has a unique motion not shared by any visible group member, or where an entire group becomes occluded simultaneously. Without an analysis of these failure modes or an ablation that varies occlusion severity and group size, the claim of 'robust and plausible motion recovery' remains unsupported.","section":"Group-guided occlusion mechanism"},{"comment":"The full text contains the line 'arXiv:2508.12645v5  [cs.IR]  18 Jan 2026', which does not match the claimed identifier '2508.12644' nor the subject area cs.CV. This mismatch raises a provenance concern: the authors should confirm that the correct version of the paper was uploaded, and the inserted line should be removed from the camera-ready version.","section":"Provenance of the submitted file"},{"comment":"The submitted full text is not readable due to encoding corruption; the equations, including the definitions of the AMC loss and the VAE motion prior, along with all tables and algorithm pseudocode, are unreadable. This prevents the verification of every technical contribution and every experimental result, and it must be fixed before any further assessment can occur.","section":"Readability of the full text"}],"minor_comments":[{"comment":"The abstract promises that 'code and dataset will be available for research purposes' but gives no details on licenses or access conditions; please specify the intended release terms.","section":"Abstract"},{"comment":"The acronym 'AMC' is used without a definition in the abstract; please add a brief clarification when the loss is first mentioned.","section":"Abstract"},{"comment":"Once the full text is readable, please ensure that all tables report error bars or statistical significance, and that ablations isolate the contributions of the group-guided optimization, the VAE prior, and the AMC loss.","section":"General"}],"recommendation":"uncertain","confidential_remarks":"The manuscript as submitted is not assessable: the full text is corrupted, the only evaluation benchmark is author-created, and the stated arXiv identifier mismatches the inserted header line. I recommend that the editor return the manuscript to the authors for a resubmission with a readable text and a more complete evaluation, rather than attempting a substantive review of the current artifact."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know. First, the actual idea is a genuine extension of prior static-image crowd reconstruction to the video setting, with a group-guided optimization scheme and a VAE motion prior aimed at occlusion recovery. That is a sensible, worthwhile direction. Second, the central SOTA claim is currently unsupported by anything we can inspect: the full text I was given is encoding-corrupted to the point of illegibility, and the only evaluation artifact named is VirtualCrowd, a synthetic benchmark the authors built because no well-annotated real large-scene video dataset exists. So the headline result is unverified as it stands.\n\nWhat is actually new: the task framing itself (video input, hundreds of people, consistent 3D pose/position/shape over time) is not in the static-image papers they cite. The AMC loss and the segment-level group-guided optimization are plausible mechanisms, and the decision to release code and dataset is good practice. If the method works on real video, this would be a useful advance for surveillance and crowd analysis, though not a paradigm shift.\n\nThe soft spots, in order of seriousness. The evaluation concern is the big one. Testing on your own synthetic dataset, generated with the same assumptions you encode in your method, can look great without transferring to real occlusions. The stress-test note is right that the abstract provides no external benchmark or real-video result. A second, smaller concern is the lack of any reported numbers in the abstract—no metrics, no ablations, no error bars—which makes even the synthetic comparison hard to weigh. And there is a provenance oddity: a garbled line in the full text references a different arXiv identifier and cs.IR category, which is probably a rendering artifact but should be cleaned up.\n\nI want to be clear that none of this disproves the method. The group-guided occlusion-recovery heuristic is reasonable; the VAE prior is standard. I just cannot verify any of the load-bearing claims from the material I have. If the authors produce a readable manuscript with results on a real large-scene video (or at least a synthetic benchmark with independent baselines and clear ablations), this paper deserves a serious referee. As it stands, I would tell the editor: send it to review, but the reviewers should be asked to focus on whether the evaluation actually validates the occlusion-recovery claim. That is a better use of referee time than desk-rejecting a plausible idea, and a better outcome than accepting based on a synthetic SOTA.","headline":"A plausible video-based crowd reconstruction idea with a genuinely new task framing, but the SOTA claim currently rests on an authors-made synthetic benchmark and the supplied text is unreadable, so the evidence cannot yet be checked.","tokens_in":15004,"tokens_out":2335,"would_cite":false,"duration_ms":22790,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single large-scene video can yield consistent 3D poses, positions, and shapes for hundreds of people over time.","keywords":["dynamic crowd reconstruction","large-scene video","3D human pose estimation","temporal consistency","occlusion reasoning","group-guided optimization","motion prior","VirtualCrowd"],"falsifier":"Take a video where a person walks behind a long wall while everyone else in view either stands still or moves differently. If DyCrowd reconstructs the hidden walker's motion as following the standing crowd rather than continuing along the observed pre- and post-occlusion path, the group-guidance premise is falsified; a multi-camera ground-truth capture of the same scene would settle it.","tokens_in":14120,"feed_emoji":"🎥","tokens_out":5239,"duration_ms":49831,"temperature":0.7,"pith_summary":"The paper introduces DyCrowd, a framework that takes one large-scene video and reconstructs the 3D pose, ground position, and body shape of hundreds of individuals frame by frame while keeping each person's motion consistent over time. This matters because existing methods reconstruct crowds from static images, so they have no temporal consistency and are vulnerable to occlusions. DyCrowd's central idea is that crowd motion is collective: people with similar motion segments are grouped, and clearly visible members of a group are used to guide the recovery of members who are occluded. The paper also contributes VirtualCrowd, a virtual benchmark dataset for evaluating large-scene dynamic crowd reconstruction, and reports state-of-the-art results on it.","feed_headline":"One video in, 3D poses and shapes for hundreds of people over time","feed_subtitle":"DyCrowd tracks each person consistently and lets visible crowd members guide reconstruction of hidden people.","key_machinery":"The load-bearing mechanism is segment-level group-guided optimization. Motion sequences are divided into segments, individuals with similar motion segments are clustered, and each cluster's motion is optimized together in a coarse-to-fine manner; within a cluster, visible unoccluded segments act as the temporal reference for reconstructing occluded segments. The VAE-based human motion prior constrains optimizations to plausible body motions, while the AMC loss enforces consistency between group members despite asynchrony and rhythmic variation. This turns occlusion from a per-person missing-data problem into a group inference problem.","core_discovery":"The central claim is that DyCrowd is the first framework to reconstruct hundreds of individuals' poses, positions and shapes with spatio-temporal consistency from a single large-scene video. The argument rests on a coarse-to-fine, group-guided motion optimization: motion sequences are split into segments, similar segments are grouped, and the group's motions are optimized jointly so that reliable, unoccluded segments transfer their temporal information to occluded ones. A VAE-based human motion prior keeps reconstructed motions plausible, and the Asynchronous Motion Consistency (AMC) loss makes the transfer robust to people moving out of sync or at different rhythms. The paper reports that this method achieves state-of-the-art performance on the large-scene dynamic crowd reconstruction task, evaluated on its new VirtualCrowd benchmark.","pith_inferences":["The method's advantage should grow with crowd density, since dense crowds provide more similar visible segments to guide each occluded person, until everyone is simultaneously hidden.","Group-guided temporal transfer of the same kind could apply to other correlated multi-object scenes, such as vehicle traffic or animal herds, wherever motion is shared enough to form reliable groups.","A concrete stress test is to move the VAE motion prior outside its training distribution: unusual gaits or choreographed actions may remain unrecoverable even when visible peers move similarly, because the prior will pull reconstructions back to familiar motions."],"forward_implications":["Large-scene video becomes a viable input for crowd reconstruction, giving temporal consistency that static-image methods cannot offer.","Long-term occlusions can be resolved by borrowing movement patterns from similarly behaving people who are visible, rather than relying only on the occluded person's own observed past.","Reconstruction quality should scale with crowd regularity: a stream of pedestrians with similar gaits should recover occluded members better than a crowd of independent, dissimilar actors.","The VirtualCrowd benchmark provides a quantitative way to compare methods on temporal consistency and occlusion recovery in large scenes."],"supporting_citations":[],"fun_headline_variants":["DyCrowd: 3D pose and shape for hundreds from one video","First dynamic 3D crowd reconstruction from large-scene video","3D crowd reconstruction from video, occlusion-robust","Temporal-consistent 3D poses for hundreds from one video","DyCrowd: first spatio-temporally consistent crowd reconstruction"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method bets that people with similar motion can be grouped and that seeing some group members clearly is enough to reconstruct the hidden members; if a hidden person's motion resembles no one visible, or the entire group is occluded at once, the temporal guidance has nothing to draw on.","fun_headline_variants_meta":{"raw":{"variants":["DyCrowd: 3D pose and shape for hundreds from one video","First dynamic 3D crowd reconstruction from large-scene video","3D crowd reconstruction from video, occlusion-robust","Temporal-consistent 3D poses for hundreds from one video","DyCrowd: first spatio-temporally consistent crowd reconstruction"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001078,"raw_usage":{"total_tokens":4534,"prompt_tokens":991,"completion_tokens":3543,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":607,"completion_tokens_details":{"reasoning_tokens":3451}},"tokens_in":607,"tokens_out":3543,"duration_ms":25786,"temperature":1.0,"reasoning_tokens":3451,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:19:55.038005+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a video where a person walks behind a long wall while everyone else in view either stands still or moves differently. If DyCrowd reconstructs the hidden walker's motion as following the standing crowd rather than continuing along the observed pre- and post-occlusion path, the group-guidance premise is falsified; a multi-camera ground-truth capture of the same scene would settle it.","supporting_citations":[],"review_version":2}