{"id":"fd3052ca-9e34-4102-8a6f-00b6fec7a80f","arxiv_id":"2411.18314","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"A vague proposal for a CNN plus optical flow tracking method with unverifiable performance claims.","lead":"This paper sketches a CNN-based video tracker that combines detection with optical flow and online updates. It claims better success and failure rates than SIFT and Transformer baselines, but gives no data, code, or evaluation details to back the claim.","discovery_kind":"incremental","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed performance advantage is unsupported because Section III provides no quantitative results, and its favorable figures are described as YOLOv3's results rather than the proposed CNN+Flow algorithm's results.","rationale":"The reader's weakest assumption is that Figures 2-5 come from real, fair experiments. I agree that this is the core problem, but my concern is slightly more specific: even treating the figures as genuine, the accompanying text in Section III attributes the key favorable comparisons in Figures 4 and 5 to YOLOv3, not to the proposed algorithm. This internal mismatch means the figures, as described, cannot support the claimed success-rate advantage of the proposed method. The absence of dataset, metrics, protocol, and code compounds the problem: there is no way to check whether the comparison was fair or whether the proposed architecture was actually evaluated. The paper's equations are standard R-CNN/R-FCN/YOLOv3 formulations, so there is no parameter-free or formally verified component that could independently support the headline result. Because the central claim is entirely dependent on evidential support that is neither present nor reproducible, the REJECT verdict is appropriate. I see no reason to change the reader's determination.","tokens_in":7275,"tokens_out":3532,"duration_ms":32247,"concrete_test":"Ask the authors to release the raw per-sequence tracking results, dataset names (e.g., OTB or VOT), evaluation metric definitions, and the exact model variant used to generate Figures 2-5. Then independently rerun the comparison on the same data. Specifically, determine whether Figures 4 and 5 were produced by the proposed CNN+Flow+online-update model or by YOLOv3; if they were produced by YOLOv3, recompute the success and failure rates using the actual proposed model. If the data or code is not supplied, the performance claim remains unverified and the REJECT verdict stands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires evidence that the proposed CNN-based tracker (CNN plus optical flow and online updating, Section II) achieves higher tracking success rates and lower failure rates than mainstream trackers under rapid motion, partial occlusion, and complex backgrounds. That evidence is absent. Section III reports no dataset names, no metric definitions, no success/failure numbers, no baseline versions, no evaluation protocol, and no code. Figures 2 through 5 are referenced, but their quantitative content is never given in the text. More importantly, the prose attributes the favorable comparisons in Figures 4 and 5 to YOLOv3: 'Figure 4's detailed comparison highlights YOLOv3's significant advantage' and 'Figure 5 ... shows YOLOv3's notable advantages.' The paper's method, however, is a joint detection-and-tracking network with a flow branch and an inter-frame regression loss, not YOLOv3. If the figures actually show YOLOv3 versus SIFT and Transformer, they do not test the proposed algorithm. If they are supposed to show the proposed algorithm, the text does not state that. Either way, the claimed success-rate and failure-rate advantage is not established. The method itself is also underspecified: Figure 1 is not described layer by layer, training data and hyperparameters are absent, and equations (1)-(7) are generic CNN/R-FCN/YOLOv3 formulas with no derivation specific to the proposed architecture. This is not a disagreement with the field's consensus; it is a missing evidential link between the described architecture and the asserted result.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a real-time video target tracking algorithm that combines CNN-based detection with optical flow correlation features and online model updating, targeting scenarios with rapid motion, partial occlusion, and complex backgrounds. Section II describes a joint detection-and-tracking architecture with classification, single-frame regression, and inter-frame regression losses, presented as an extension of the R-FCN and YOLO families. Section III claims experimental verification through Figures 2–5, reporting higher tracking success rates, lower failure rates, and real-time speeds (25–30 FPS on GPU) against SIFT, YOLOv3/YOLOv5, and Transformer baselines. The manuscript contains no quantitative results, no dataset description, no evaluation metrics, no baseline versions, and no code, so the central performance claim is not supported by any checkable evidence.","tokens_in":7565,"tokens_out":3132,"duration_ms":29430,"significance":"If the claimed results were substantiated, the contribution would be an incremental engineering integration of known components (a flow branch added to an R-FCN/YOLO-style detector with an extra inter-frame regression loss). No parameter-free derivations, machine-checked proofs, reproducible code, or falsifiable quantitative predictions are provided; the paper's only empirical content is a set of bare figures with textual assertions. At present the manuscript cannot be independently evaluated or reproduced, and the central claim of improved tracking success and lower failure rates is unverified.","major_comments":[{"comment":"The central performance claim is made entirely on the basis of Figures 2–5, but the text provides no dataset names, no metric definitions, no numerical results, no error bars, no baseline versions, and no evaluation protocol. Worse, the prose in Section III explicitly attributes the favorable comparisons in Figures 4 and 5 to YOLOv3: “Figure 4's detailed comparison highlights YOLOv3's significant advantage” and “Figure 5's in-depth comparison shows YOLOv3's notable advantages.” This does not test the proposed CNN+Flow algorithm, and if the figures are instead meant to show the proposed algorithm, the text does not say so. The claimed advantage in tracking success rate and failure rate is therefore not established.","section":"Section III, Figures 2–5"},{"comment":"Equations (2)–(6) are the standard R-CNN/R-FCN coordinate parameterization for bounding-box prediction, and Equation (7) is the generic YOLOv3 decoding formula for tx, ty, tw, th. No derivation is given that connects these formulas to the proposed joint detection-and-tracking architecture, the Flow branch, the offset prediction module, or the online updating mechanism. Figure 1 is never described layer by layer, and hyperparameters, the loss weights λ1 and λ2, training data, optimizer settings, and implementation details are absent. The method is therefore underspecified and not reproducible as written.","section":"Section II, Equations (1)–(7)"},{"comment":"The speed claim of 25–30 FPS on GPU and 15 FPS on CPU cannot be assessed because no hardware model, video resolution, code version, or measurement methodology is reported. The statement that “YOLOv5 excels in speed” is presented in passing, but no YOLOv5 comparison is shown in any figure or table, and no quantitative speed values are given for the proposed algorithm and the baselines. The absence of a baseline table or a per-sequence breakdown makes the speed and robustness claims untestable.","section":"Section III, FPS discussion"},{"comment":"There is an unresolved inconsistency in the role of YOLOv3. Section II introduces the framework as an extension of R-FCN with three parallel branches and an inter-frame loss, while Section III discusses YOLOv3 as the subject of comparisons and claims. If YOLOv3 is intended to be the detection backbone of the proposed algorithm, that identification is never stated; if it is a separate baseline, the text does not say how its results are relevant to evaluating the proposed architecture. Either way, the paper fails to attribute the experimental curves to the method being proposed.","section":"Section II versus Section III"}],"minor_comments":[{"comment":"Equation (1) is garbled by typesetting: the input dimension appears as “321”, the summation bound is inconsistent with the notation “x1, x2, x3”, and the variables Wi and bi are introduced without a clear indexing convention.","section":"Equation (1)"},{"comment":"The figure captions are uninformative: they provide no axis labels, no units, no legend descriptions, and no explanation of what each curve represents. It is impossible to determine from the manuscript what quantity is being plotted or which method each line corresponds to.","section":"Figures 2–5"},{"comment":"The term “correlation features” is used repeatedly, but it is never formally defined or distinguished from standard CNN feature maps. It is unclear whether this refers to cross-correlation of frame features, optical flow, or a different mechanism.","section":"Section II.A, “correlation features”"},{"comment":"Several references are cited in contexts that do not match the corresponding entries; for example, the citation of [1] for reinforcement-learning decision-making, [8] for CNN-LSTM weather forecasting, and [21] for multi-UAV navigation do not connect to the specific claims they are attached to. The reference list also contains many entries that are unrelated to the technical content of the paper.","section":"References"}],"recommendation":"reject","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this should not be published in current form, and I would not send it to review. The idea—a CNN-based tracker with an optical-flow branch and an inter-frame regression loss—is plausible, and real-time tracking under occlusion and rapid motion is a genuine problem. But the paper provides no evidence that the proposed algorithm works. The abstract promises higher success rates and lower failure rates, and Section III is supposed to demonstrate this. Instead, the text gives no dataset names, no metric definitions, no numerical results, no error bars, no baseline versions, and no evaluation protocol. Figures 2 and 3 compare with SIFT but contain only images, not numbers. Worse, Figures 4 and 5 are explicitly described as showing YOLOv3's advantages over SIFT and Transformer, not the proposed method. That is a load-bearing flaw: the reader cannot tell if the authors ran any experiments at all.\n\nTo give credit where it is due: the authors correctly identify known limitations of traditional trackers and list sensible ingredients like depthwise separable convolutions, feature pyramids, and online model updates. The inter-frame loss is a minor variation of standard R-CNN regression losses. But none of this is new, and the equations are generic, garbled, and not tied to any specific architecture. There is no training data, no hyperparameter tuning details, and no code. The reference list is padded with loosely related arXiv papers, several self-cited, which does not help.\n\nThe internal inconsistency between the abstract's real-time claim and the conclusion's statement that computational complexity is the primary limitation is another soft spot. The paper reads less like a completed study and more like an assembled outline with placeholder figures. A serious reader gets no reproducible result and no reason to trust the central claim.\n\nI would desk-reject this. It does not deserve referee time in its present state. If the authors have real experiments, they need to report them properly; until then, this is an assertion with missing evidence.","headline":"The paper's central performance claim is entirely unsupported: the results section reports no numbers for the proposed algorithm, and two of the favorable figures are actually attributed to YOLOv3.","tokens_in":8041,"tokens_out":1967,"would_cite":false,"duration_ms":19221,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A CNN tracker with an optical-flow branch claims higher success in rapid-motion and occlusion scenes.","keywords":["convolutional neural networks","real-time video tracking","optical flow","target detection and tracking","online learning","inter-frame regression","video surveillance","intelligent transportation"],"falsifier":"Run the described algorithm and the claimed baselines (SIFT-based tracking, YOLOv3, YOLOv5, and a Transformer-based tracker) on a standardized video-tracking benchmark with annotated sequences covering rapid motion, partial occlusion, and background clutter, and measure success rate, recall, and FPS under the same protocol. If the plotted advantages cannot be reproduced, or the method does not beat the baselines on those metrics, the central claim is settled against the paper.","tokens_in":7097,"feed_emoji":"🎯","tokens_out":7922,"duration_ms":62653,"temperature":0.7,"pith_summary":"This paper proposes a real-time video target tracking algorithm that couples a convolutional neural network detector with an optical-flow branch, an online model update mechanism, and a multi-task loss that predicts both single-frame boxes and inter-frame motion. The authors claim that this design tracks targets more successfully and fails less often than several mainstream tracking algorithms when targets move rapidly, are partially occluded, or sit in cluttered backgrounds. They report 25-30 FPS on GPU, which they argue is fast enough for live video surveillance and intelligent transportation. The paper's contribution is the specific integration of detection and tracking, not a new CNN architecture.","feed_headline":"CNN tracker claims fewer failures in fast-motion, occluded scenes","feed_subtitle":"Method fuses CNN features with optical flow and online model updates to hit 25-30 FPS on GPU.","key_machinery":"The load-bearing mechanism is a joint detection-and-tracking network with two streams: a CNN feature branch for appearance and a Flow branch that processes optical flow between frames to supply temporal context. The Flow branch predicts the target's next-frame position and feeds an offset prediction module that refines the current detection. Training uses a multi-task loss under the R-FCN framework that combines a classification loss, a single-frame regression loss, and the added inter-frame regression loss for motion offsets; detection boxes follow YOLOv3's anchor-based scheme. Real-time speed comes from depthwise separable convolutions, bottleneck layers, cross-layer feature sharing, feature pyramids, and GPU-accelerated parallel frame processing.","core_discovery":"The central claim is that combining CNN-based detection with inter-frame optical flow in a joint detection-tracking framework yields a tracker that is both accurate and real-time in challenging conditions. On the paper's account, the Flow branch encodes motion between consecutive frames and predicts the target's next position; an offset prediction module then refines the detected box using that motion feature map. The training objective extends the R-FCN multi-task loss with an inter-frame regression loss, and the whole pipeline is made fast with depthwise separable convolutions, bottleneck layers, feature pyramids, and GPU parallel processing. Experimental figures are said to show higher tracking success rates and recall than a SIFT-based tracker, and better time efficiency and loss convergence than YOLOv3 and a Transformer-based model, with an average of 25 FPS in dense multi-target scenes and 30 FPS on GPU.","pith_inferences":["The paper does not report ablations, so the most direct next test would be to remove the Flow branch, the inter-frame loss, or the online update one at a time to see which component actually drives the claimed improvement.","Because the experimental evidence is presented only as figures with no dataset or metric values, a reader cannot yet compare this method against published numbers; re-running the same design on a public tracking benchmark with standardized evaluation would settle the comparison.","The optical-flow prior could plausibly be extended to long-term tracking, where targets disappear for many frames and a motion model alone cannot reacquire them, but the paper does not address full occlusion.","If the speed claim transfers to lighter CNN backbones, the approach might run on embedded devices, though that is not demonstrated in the paper."],"forward_implications":["If the reported results hold, video surveillance systems can track targets through partial occlusion and rapid motion at frame rates suitable for live monitoring.","The inter-frame regression loss gives the tracker a temporal consistency cue that single-frame detectors lack, which is what the paper credits for higher recall in complex scenes.","The method reaches a practical real-time operating point (25-30 FPS on GPU) without the computational overhead the paper attributes to Transformer-based trackers.","The same joint detection-tracking design could be carried over to intelligent transportation tasks such as vehicle and pedestrian tracking, where the paper expects its main applications."],"supporting_citations":[{"why":"It supplies the R-FCN multi-task framework whose classification and single-frame regression losses the paper extends with an inter-frame regression branch.","marker":"[19]"},{"why":"It provides the YOLOv3 anchor-box prediction formulas the algorithm uses to turn grid-cell offsets into bounding boxes.","marker":"[21]"},{"why":"It motivates depthwise separable convolutions and bottleneck layers, the design choices the paper credits for reduced parameters and faster forward propagation.","marker":"[15]"},{"why":"It motivates cross-layer connections and feature pyramids, which the paper uses for feature sharing and reuse to cut redundant computation.","marker":"[16]"},{"why":"It is cited for the offset prediction module trained by minimizing offset loss, the component that refines detected boxes using the optical-flow feature map.","marker":"[13-14]"},{"why":"It is cited as the basis for adding an inter-frame regression loss that trains the motion prediction branch across consecutive frames.","marker":"[20]"}],"fun_headline_variants":["CNN tracker adapts in real time to beat occlusion and clutter","Real-time CNN tracker boosts success under occlusion and fast motion","CNN tracker at 30 FPS with online learning for changing targets","CNN detection plus optical flow yields robust real-time tracker"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the figures in Section III are real experimental outputs from a fair comparison; the paper gives no dataset, no quantitative numbers, and no evaluation protocol, so if those curves do not reflect actual measurements, the central performance claim collapses.","fun_headline_variants_meta":{"raw":{"variants":["CNN tracker adapts in real time to beat occlusion and clutter","Real-time CNN tracker boosts success under occlusion and fast motion","CNN tracker at 30 FPS with online learning for changing targets","CNN detection plus optical flow yields robust real-time tracker"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000903,"raw_usage":{"total_tokens":3874,"prompt_tokens":923,"completion_tokens":2951,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":539,"completion_tokens_details":{"reasoning_tokens":2882}},"tokens_in":539,"tokens_out":2951,"duration_ms":18633,"temperature":1.0,"reasoning_tokens":2882,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:17:55.980727+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the described algorithm and the claimed baselines (SIFT-based tracking, YOLOv3, YOLOv5, and a Transformer-based tracker) on a standardized video-tracking benchmark with annotated sequences covering rapid motion, partial occlusion, and background clutter, and measure success rate, recall, and FPS under the same protocol. If the plotted advantages cannot be reproduced, or the method does not beat the baselines on those metrics, the central claim is settled against the paper.","supporting_citations":[{"cited_title":"Dynamic Fraud Detection: Integrating Reinforcement Learning into Graph Neural Networks[J]","cited_arxiv_id":null,"evidence_quote":"It supplies the R-FCN multi-task framework whose classification and single-frame regression losses the paper extends with an inter-frame regression branch."},{"cited_title":"Research on Move-to-Escape Enhanced Dung Beetle Optimization and Its Applications[J]","cited_arxiv_id":null,"evidence_quote":"It motivates depthwise separable convolutions and bottleneck layers, the design choices the paper credits for reduced parameters and faster forward propagation."},{"cited_title":"Fine-grained imbalanced leukocyte classification with global-local attention transformer[J]","cited_arxiv_id":null,"evidence_quote":"It motivates cross-layer connections and feature pyramids, which the paper uses for feature sharing and reuse to cut redundant computation."},{"cited_title":"A memorizing and generalizing framework for lifelong person re-identification[J]","cited_arxiv_id":null,"evidence_quote":"It is cited as the basis for adding an inter-frame regression loss that trains the motion prediction branch across consecutive frames."}],"review_version":1}