{"id":"ec24b856-ebd8-4a18-af26-27cbc39e7032","arxiv_id":"2412.09991","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey that taxonomizes deep-learning visual trackers across RGB, thermal, LiDAR, and four multi-modal combinations, with benchmark tables and future directions.","lead":"This paper reviews hundreds of visual object tracking methods across seven data modalities, from ordinary RGB video to thermal, LiDAR, and combined sensor inputs. It organizes them into pipeline families and benchmark tables, making it a useful map rather than a new tracker or experiment.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Benchmark-table fidelity is the load-bearing risk: the MLSSNet dataset row in Table 12 contradicts Section 5.2 and mixes method with dataset, so the survey's reference value needs a primary-source audit.","rationale":"Read in good faith: the paper is a survey, not a new algorithm. Its strongest claim is to be the first comprehensive modality-organized review including LiDAR-based, RGB-LiDAR, and RGB-Language tracking; that claim requires accurate coverage and trustworthy tables. The reader's weakest assumption is table fidelity, and I agree with that identification. The internal MLSSNet mismatch is objective and verifiable from the text alone: Section 5.2 says 500 videos/228k frames, while Table 12 says 430 videos/200k boxes. It cannot both be right. The presence of method names in the dataset table (MLSSNet, MMNet, ECO-MM) strengthens the suspicion of misattributed or copy-pasted statistics. Table 11's three-number cells contradict the stated Precision/AUC legend, which is another transcription-level problem. None of these issues invalidate the organizational thesis or the proposed taxonomies; the conditional verdict remains appropriate rather than a rejection. The proposed concrete test is feasible: check the primary source for the MLSSNet dataset and LSOTB-TIR, then re-audit the affected tables. No other concern, such as the novelty of the survey or taxonomy overlap, is as directly load-bearing, because the tables are the part readers will actually reuse as reference data.","tokens_in":61277,"tokens_out":4156,"duration_ms":46159,"concrete_test":"Independently fetch reference [293] (MLSSNet, TMM 2020) and reference [423] (LSOTB-TIR) and determine which paper actually introduces the 500-sequence/228k-frame TIR benchmark. If [293] does not, replace the Table 12 MLSSNet row with the correct benchmark citation and statistics. Then audit all Table 12 rows that cite method papers (MLSSNet, MMNet, ECO-MM) and all three-number cells in Table 11 against their primary sources, and publish a corrections table before the survey is used as a reference.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central contribution is to be a reliable modality-organized map of VOT, with Tables 1 through 12 and Section 5 as the reference data. That contribution fails if table entries are not faithful transcriptions. There is at least one concrete internal contradiction: Section 5.2 describes \"MLSSNet\" as a large-scale TIR dataset with 500 video sequences and 228k frames, while Table 12 reports \"MLSSNet [293] 2020, 430 videos, 200k boxes, 20 classes, 15.5s.\" Only one figure can be the dataset statistic. The same table also lists MMNet and ECO-MM as Thermal datasets, even though Table 4 and Section 3.2 introduce those references as tracker methods; either the dataset designations are wrong or the dataset column is mixing method names with benchmark names. Separately, Table 11 promises Precision/AUC but several TNL2K cells contain three numbers (for example, \"0.27/0.34/0.25\"), so the metric legend does not match the data. These are transcription-level defects, not interpretive disagreements, and they directly undermine the claimed reference utility.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is a survey of visual object tracking (VOT) organized by data modality. It reviews three single-modal families (RGB, thermal infrared, LiDAR) and four multi-modal combinations (RGB-Depth, RGB-Thermal, RGB-LiDAR, RGB-Language). For RGB trackers it proposes a taxonomy of four deep-learning paradigms (discriminative correlation filters, Siamese trackers, instance classification/detection, and one-stream transformers); for TIR and LiDAR it provides modality-specific taxonomies; and for multi-modal methods it applies an early/middle/late fusion categorization. The survey compiles comparison results on many benchmarks in Tables 1-11 and dataset statistics in Table 12, and closes with ten short discussions of future directions such as parameter-efficient transfer learning, online learning, and multi-modal tracking. The paper claims to be the first review covering LiDAR-based, RGB-LiDAR, and RGB-Language VOT methods.","tokens_in":61627,"tokens_out":8243,"duration_ms":77893,"significance":"If the benchmark tables and dataset statistics are reliable, this survey would be a useful entry point for researchers, especially for the less-covered LiDAR, RGB-LiDAR, and RGB-Language areas. The four-RGB-paradigm taxonomy and the early/middle/late fusion categorization are clear organizational principles, and the breadth of coverage (300+ papers, seven modality families) is a genuine strength. Because the survey's main contribution is reference value rather than new methods or derivations, the fidelity of the tables is load-bearing: a reader consulting this paper will use the tables to compare trackers and to select datasets. The paper ships no code or proofs, but as a survey that is not expected; its value rests on accurate transcription of the cited literature.","major_comments":[{"comment":"The dataset statistics for MLSSNet are internally contradictory: Section 5.2 states that MLSSNet [293] is a large-scale TIR dataset with 500 video sequences and 228k frames, while Table 12 reports 430 videos, 200k boxes, 20 classes, and an average duration of 15.5s for the same reference. Since the survey's reference value depends on faithful transcription of primary sources, the authors must verify the original paper and make the text and table agree.","section":"Section 5.2 and Table 12 (Thermal block)"},{"comment":"References [293] (MLSSNet), [294] (MMNet), and [30] (ECO-MM) are presented as trackers with benchmark results in Table 4 and Section 3.2, but the same references are listed as Thermal datasets in Table 12 and described as datasets in Section 5.2. If these papers indeed introduce both a tracker and a dataset, the survey should explicitly say so and clearly separate the two roles; as it stands, a reader cannot tell whether the rows in Table 12 are datasets, methods, or both, which undermines the dataset table.","section":"Table 12 and Section 5.2 vs. Section 3.2 and Table 4"},{"comment":"The caption of Table 11 states that results are evaluated by Precision/AUC, but several TNL2K cells contain three values (e.g., Feng et al. 0.27/0.34/0.25, Wang et al. 0.06/0.11/0.11 and 0.42/0.50/0.42, VLTTT 0.53/0.53). The legend must be expanded to explain what the third number represents, or the cells must be corrected, because the table is not interpretable as presented.","section":"Table 11"}],"minor_comments":[{"comment":"The sentence claiming the survey covers \"eight multiple modalities\" should read \"four multiple modalities,\" since Section 4 reviews exactly four multi-modal combinations (RGB-Depth, RGB-Thermal, RGB-LiDAR, RGB-Language).","section":"Section 2.1"},{"comment":"LaSOT appears in the RGB block with year 2019 and in the RGB-La block with year 2018; the year should be made consistent, and the double listing (the same dataset in two modality groups) should be explicitly justified.","section":"Table 12"},{"comment":"The last column header appears as \"FPSSR(%)\", which seems to merge the FPS and SR(%) columns; please split the header into separate columns for FPS and SR(%) or correct the label.","section":"Table 1"},{"comment":"The citation for VLTTT appears as \"VLTT T[411]\" in the text but as \"V LTT T[41]\" in Table 11; the reference number should be consistent (the reference list entry is [41]).","section":"Section 4.4"},{"comment":"The phrase \"most of them are shotted at night\" contains a typo; \"shotted\" should be \"shot.\"","section":"Section 5.2"},{"comment":"The claim of being the first review to cover LiDAR-based, RGB-LiDAR, and RGB-Language VOT is plausible but should be substantiated by a more explicit comparison with the related surveys listed in Section 2, since the current discussion does not fully rule out partial coverage in prior works.","section":"Section 1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript fits the scope of a survey-oriented computer vision journal and the modality-based organization is a reasonable contribution. My recommendation is driven entirely by the benchmark-table fidelity issues: the MLSSNet contradiction, the method/dataset conflation, and the uninterpretable TNL2K cells are load-bearing for a survey whose primary value is as a reference. These problems appear fixable locally, but they require the authors to audit Tables 1-12 against their cited sources, not merely to correct the three cited examples. I would encourage the editor to ask for that audit before resubmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a wide-ranging survey of VOT across seven modalities, and it earns its keep as an organizational contribution. The four-paradigm split for RGB trackers (DCF, Siamese, ICD, OST) is a sensible update over older surveys that stop at DCF and Siamese, and the early/middle/late fusion framing for multi-modal methods is clear enough to be useful. The claimed first coverage of LiDAR, RGB-LiDAR, and RGB-Language tracking appears legitimate—I do not know of another survey that bundles these. The development-lineage figures and the sheer number of entries in the method tables will save a newcomer a lot of digging.\n\nThe soft spots are real but concentrated in the reference tables, which is a problem because the paper's value is as a reference. The stress-test note is right: Table 12 lists MLSSNet as 430 videos and 200k boxes, while Section 5.2 says 500 sequences and 228k frames; only one of those can be the dataset stat. Worse, the same table lists MMNet and ECO-MM as Thermal datasets, when Sections 3.2 and Table 4 treat both as trackers. That is a method/dataset mix-up, not a style preference. Table 11 also has three-number cells under a two-metric legend for TNL2K. There is also a stray phrase in Section 2.1 about 'eight multiple modalities' when you only cover four. These are transcription- and copy-paste-level defects, but they are exactly the load-bearing surface for a survey whose selling point is being a reliable map.\n\nOn substance, the taxonomy holds up on reading. I did not find a misclassification that would break the four-paradigm or fusion-strategy claims. The paper is a review, so there is no new method or measurement to verify, and that is fine given the title.\n\nWho is this for? Practitioners wanting a quick orientation across modalities, and newcomers who need one entry point. It deserves a serious referee, but the referee should require a primary-source audit of Tables 1–12 before publication. I would cite it after that, and I would probably bring it to a reading group as a starting point for discussion.","headline":"A genuinely useful modality-organized survey whose reference tables need a primary-source audit before the paper can be trusted as a citation.","tokens_in":62006,"tokens_out":2311,"would_cite":true,"duration_ms":23894,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A survey maps visual object tracking across seven data modalities, from RGB to LiDAR and language.","keywords":["visual object tracking","multi-modal tracking","RGB-thermal tracking","LiDAR tracking","RGB-Language tracking","deep learning","benchmarks","survey"],"falsifier":"Pick any row in Tables 1 through 11, locate the cited paper's reported score, and check whether the numbers match; one confirmed mismatch, such as the MLSSNet entry described in Section 5.2 as a 500-sequence dataset versus the 430 videos listed in Table 12, would show that the reference value of the tables depends on verification against primary sources.","tokens_in":61092,"feed_emoji":"🎯","tokens_out":4833,"duration_ms":51463,"temperature":0.7,"pith_summary":"This review aims to give a complete map of visual object tracking (VOT) methods organized by data modality rather than by algorithm family alone. It covers three single-modality tracks — RGB video, thermal infrared, and LiDAR point clouds — and four multi-modal combinations — RGB-Depth, RGB-Thermal, RGB-LiDAR, and RGB-Language. The authors claim this is the first survey to include LiDAR-based, RGB-LiDAR, and RGB-Language tracking, and they support the map with abstracted pipeline schemas, benchmark comparisons, and dataset statistics. A reader who wants to know which tracking paradigms exist, which methods inherit them, and what numbers they achieve on standard benchmarks would use this as a reference.","feed_headline":"One survey maps visual object tracking across seven data modalities","feed_subtitle":"RGB, thermal, and LiDAR trackers sorted into shared paradigms, plus fusion stages, benchmarks, and the newest RGB-Language work.","key_machinery":"The central organizing device is a modality-by-modality taxonomy with abstracted pipeline diagrams and named schemas. For RGB trackers the four schemas are DCF, Siamese, ICD, and OST; for multi-modal trackers the three fusion strategies are early, middle, and late fusion. These schemas do the argumentative work: they turn a list of several hundred methods into a small set of inheritance relations, and they let the survey transfer paradigm knowledge across modalities — for example, TIR trackers building on DCF and Siamese, and LiDAR trackers borrowing the Siamese schema before moving to motion-modeling and one-stream Transformers.","core_discovery":"On its own terms, the paper's central claim is that VOT is best understood from the perspective of data modalities, and that a survey organized that way can be both comprehensive and novel. For single-modality RGB tracking, it abstracts four paradigms: discriminative correlation filters (online-trained filters convolved with search features), Siamese trackers (a shared network matching a template to a search region), instance classification/detection (a network specialized to one target instance), and one-stream Transformers (a single Transformer that jointly extracts features and relates template to search). Thermal trackers are shown to inherit the DCF and Siamese schemas, and LiDAR trackers are shown to follow Siamese, motion-modeling, and one-stream Transformer designs. For multi-modal tracking, the organizing distinction is fusion stage: early fusion at the input, middle fusion at the feature level, or late fusion at the result level. The survey concludes that these taxonomies, together with its benchmark tables, constitute the first systematic reference for the newly emerged LiDAR-based, RGB-LiDAR, and RGB-Language tracking directions.","pith_inferences":["If the modality-first organization is right, the missing next piece is a cross-modal evaluation protocol that reuses the same target categories and metrics across RGB, TIR, and LiDAR, so that paradigm-transfer claims can be tested quantitatively rather than by inspection.","The fusion-stage taxonomy suggests a testable conjecture: middle fusion will keep dominating RGB-Thermal tracking because it offers learnable cross-modal parameters without requiring the strict input alignment that early fusion demands; a meta-analysis of the table entries could check whether late-fusion methods ever surpass middle-fusion ones at similar speed.","The inclusion of RGB-Language tracking implies the field may treat natural-language descriptions as a first-class query channel alongside boxes and point clouds, which would connect VOT to open-vocabulary and referring-expression benchmarks beyond those listed.","Because the survey records FPS alongside accuracy, a reader could use its tables to test whether the OST paradigm's accuracy gains come at a speed cost, a question the paper raises but does not resolve."],"forward_implications":["A newcomer can identify the paradigm of any RGB tracker by matching its pipeline to one of four schemas rather than reading each paper in full.","Multi-modal trackers can be classified by fusion stage, which predicts whether the method requires aligned inputs, learns cross-modal feature interactions, or fuses final predictions.","The benchmark tables provide a single place to compare trackers across LaSOT, TrackingNet, GOT-10k, VOT, KITTI, nuScenes, Waymo, PTB, DepthTrack, RGBT234, LasHeR, and TNL2K, with numbers transcribed from the cited papers.","Identifying OST as the emerging RGB paradigm points to one-stream Transformers as the schema most likely to be transferred to TIR and LiDAR tracking."],"supporting_citations":[{"why":"Previous multi-modal survey covering only RGB-Depth and RGB-Thermal; defines the gap this survey fills.","marker":"[51]"},{"why":"Earlier multicue tracking survey that this paper extends with LiDAR, RGB-LiDAR, and RGB-Language modalities.","marker":"[52]"},{"why":"Recent RGB-only survey whose generative/discriminative taxonomy omits the ICD and one-stream Transformer paradigms.","marker":"[46]"},{"why":"KCF defines the discriminative correlation filter paradigm that anchors the first RGB schema and TIR inheritors.","marker":"[2]"},{"why":"SiamFC establishes the fully-convolutional Siamese matching paradigm inherited by many RGB, TIR, and LiDAR trackers.","marker":"[16]"},{"why":"MDNet is the canonical instance classification/detection tracker that anchors the ICD schema.","marker":"[25]"},{"why":"MixFormer is cited as a leading one-stream Transformer that defines the newest RGB schema.","marker":"[29]"},{"why":"P2B is the baseline LiDAR tracker that most later 3D tracking methods improve upon.","marker":"[33]"},{"why":"LaSOT is the large-scale long-term RGB benchmark whose numbers fill the main comparison tables.","marker":"[18]"},{"why":"KITTI supplies the shared benchmark for LiDAR-based and RGB-LiDAR tracking comparisons.","marker":"[340]"}],"fun_headline_variants":["A map of object tracking across seven data modalities","Tracking any target, in any sensor, in one survey","From RGB to LiDAR: the complete tracking survey","Seven modalities, four paradigms, one review","Object tracking across every data type, surveyed"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark numbers and dataset statistics in the tables are faithful copies of the cited papers, and the modality and fusion taxonomies assign every method to exactly one correct box.","fun_headline_variants_meta":{"raw":{"variants":["A map of object tracking across seven data modalities","Tracking any target, in any sensor, in one survey","From RGB to LiDAR: the complete tracking survey","Seven modalities, four paradigms, one review","Object tracking across every data type, surveyed"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000487,"raw_usage":{"total_tokens":2412,"prompt_tokens":969,"completion_tokens":1443,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":585,"completion_tokens_details":{"reasoning_tokens":1371}},"tokens_in":585,"tokens_out":1443,"duration_ms":9553,"temperature":1.0,"reasoning_tokens":1371,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T16:27:47.539443+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Pick any row in Tables 1 through 11, locate the cited paper's reported score, and check whether the numbers match; one confirmed mismatch, such as the MLSSNet entry described in Section 5.2 as a 500-sequence dataset versus the 430 videos listed in Table 12, would show that the reference value of the tables depends on verification against primary sources.","supporting_citations":[{"cited_title":"Multi-modal Visual Tracking: Review and Experimental Comparison","cited_arxiv_id":"2012.04176","evidence_quote":"Previous multi-modal survey covering only RGB-Depth and RGB-Thermal; defines the gap this survey fills."}],"review_version":1}