REVIEW 25 cited by
MOT20: A benchmark for multi object tracking in crowded scenes
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Standardized benchmarks are crucial for the majority of computer vision applications. Although leaderboards and ranking tables should not be over-claimed, benchmarks often provide the most objective measure of performance and are therefore important guides for research. The benchmark for Multiple Object Tracking, MOTChallenge, was launched with the goal to establish a standardized evaluation of multiple object tracking methods. The challenge focuses on multiple people tracking, since pedestrians are well studied in the tracking community, and precise tracking and detection has high practical relevance. Since the first release, MOT15, MOT16, and MOT17 have tremendously contributed to the community by introducing a clean dataset and precise framework to benchmark multi-object trackers. In this paper, we present our MOT20benchmark, consisting of 8 new sequences depicting very crowded challenging scenes. The benchmark was presented first at the 4thBMTT MOT Challenge Workshop at the Computer Vision and Pattern Recognition Conference (CVPR) 2019, and gives to chance to evaluate state-of-the-art methods for multiple object tracking when handling extremely crowded scenarios.
Forward citations
Cited by 25 Pith papers
-
Higher-Order Cell Tracking Transformer
An edge-centric Transformer with line-to-line geometric attention achieves SOTA cell lineage tracking without pretrained image encoders and fine-tunes far more efficiently than node-embedding baselines.
-
WaspMOT: A Benchmark for Long-Term Multi-Object Tracking of Trichogramma Wasps
A new insect-tracking benchmark reveals that five standard MOT methods suffer severe identity fragmentation on 8-minute sequences even with oracle detections, with simple spatial stitching recovering significant gains.
-
Polycepta: Object-Centric Appearance Estimation for Multi-Object Tracking
Polycepta recursively estimates per-object appearance states so visual cues improve over time, reducing identity switches and lifting tracking-by-detection performance at real-time speed.
-
COVTrack++: Learning Open-Vocabulary Multi-Object Tracking from Continuous Videos via a Synergistic Paradigm
Continuous TAO annotations plus multi-cue fusion, hierarchical aggregation, and temporal confidence propagation raise novel TETA to 35.4%/30.5% on TAO val/test.
-
Video Individual Counting and Tracking from Moving Drones: A Benchmark and Methods
GD3A and DVTrack, driven by optimal-transport descriptor matching with an adaptive dustbin score, set state-of-the-art results on a new moving-drone dense-crowd counting and tracking benchmark.
-
Generative Semantic Multi-Object Tracking: A Large-Scale Benchmark and an MLLM-Driven Reasoning Framework
A new LLM-generated dataset and an MLLM-based tracker claim state-of-the-art semantic multi-object tracking, but the evaluation protocol masks missed objects and ID switches.
-
MVTrajecter: Multi-View Pedestrian Tracking with Trajectory Motion Cost and Trajectory Appearance Cost
An end-to-end multi-view pedestrian tracker that aggregates motion and appearance costs over K past timestamps, outperforming prior methods on GMVD, Wildtrack, and MultiviewX, though with a validation-protocol concern...
-
To New Beginnings: A Survey of Unified Perception in Autonomous Vehicle Software
A taxonomy that classifies unified perception methods in autonomous driving into Early, Late, and Full Unified Perception based on task integration, tracking formulation, and representation flow.
-
GRASPTrack: Geometry-Reasoned Association via Segmentation and Projection for Multi-Object Tracking
A depth-aware MOT tracker using mask-guided 3D point clouds, voxelized 3D IoU association, adaptive Kalman noise, and 3D motion consistency surpasses prior TBD methods on MOT17, MOT20, and DanceTrack.
-
RoundaboutHD: High-Resolution Real-World Urban Environment Benchmark for Multi-Camera Vehicle Tracking
A high-resolution, real-world roundabout dataset for multi-camera vehicle tracking, with 512 identities, four non-overlapping 4K cameras, and baselines across four tasks.
-
USVTrack: USV-Based 4D Radar-Camera Tracking Dataset for Autonomous Driving in Inland Waterways
A new USV dataset combining 4D radar, camera, GPS, and IMU data for tracking boats, ships, and vessels on inland waterways, together with a radar-camera matching method that consistently improves two-stage trackers.
-
Progressive Scaling Visual Object Tracking
A progressive scaling training strategy with small-teacher distillation and masked-input alignment improves tracking accuracy and powers a new 12-dataset benchmark.
-
Learning better representations for crowded pedestrians in offboard LiDAR-camera 3D tracking-by-detection
A BEVFusion-based offboard tracker with density-aware loss weighting, nearest-neighbor relationship targets, and high-resolution sparse features doubles MOTA on a new crowded-pedestrian benchmark (PCP-MV) from 0.172 to 0.353.
-
Motion Estimation for Multi-Object Tracking using KalmanNet with Semantic-Independent Encoding
SIKNet, a learned Kalman filter with a semantic-independent encoder, improves bounding-box motion prediction in multi-object tracking over the classic Kalman filter and prior KalmanNet variants.
-
MeMoSORT: Memory-Assisted Filtering and Motion-Adaptive Association Metric for Multi-Person Tracking
MeMoSORT reports state-of-the-art multi-person tracking HOTA of 67.9% on DanceTrack and 82.1% on SportsMOT by combining a memory-compensated Kalman filter with a motion-adaptive IoU association metric.
-
Head Anchor Enhanced Detection and Association for Crowded Pedestrian Tracking
FocusTrack, a YOLOX-based tracker using head keypoints and detector features, matches but does not beat simple baselines on MOT17 and MOT20, and its 3D trajectory completion underperforms 2D interpolation.
-
CrowdTrack: A Benchmark for Difficult Multiple Pedestrian Tracking in Real Scenarios
CrowdTrack is a dense, first-person-view pedestrian tracking benchmark that exposes large performance drops in existing multi-object trackers.
-
Lightweight Multi-Frame Integration for Robust YOLO Object Detection in Videos
Multi-frame early fusion with single-frame supervision improves YOLOv7-tiny detection on MOT20Det and a new BOAT360 fisheye dataset.
-
YOLOv8-SMOT: An Efficient and Robust Framework for Real-Time Small Object Tracking via Slice-Assisted Training and Adaptive Association
A YOLOv8 detector trained on overlapping slices plus an OC-SORT tracker with EMA motion direction and expanded IoU distance penalty achieves 55.205 SO-HOTA on the SMOT4SB public test set.
-
Glance-MCMT: A General MCMT Framework with Glance Initialization and Progressive Association
Glance-MCMT combines BoT-SORT single-camera tracks, a short glance phase to seed global IDs, and progressive cross-view association, reaching 51.34 HOTA on AI City 2025 validation data.
-
LazyVLM: Neuro-Symbolic Approach to Video Analytics
LazyVLM decomposes multi-frame video queries into vector-search entity matching, SQL-style relationship lookup, and lightweight VLM refinement, but provides no experimental evaluation of its claims.
-
Incremental Optimal Assignment for Real-Time Crowd Tracking
An incremental assignment solver with warm-started dual potentials claims 1.1–6.5× speedups over Hungarian on synthetic block-sparse crowd matrices, with the headline 3.7–6.5× range not matching its own data.
-
A Framework for Multi-View Multiple Object Tracking using Single-View Multi-Object Trackers on Fish Data
A YOLOv8-ByteTrack pipeline plus stereo triangulation can produce 3D fish tracks for some underwater video pairs, but the claimed multi-view accuracy improvement is not demonstrated.
-
A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects
A comprehensive survey of video scene parsing that organizes methods, datasets, metrics, and benchmark results across VSS, VIS, VPS, VTS, and OVVS.
-
Trajectory Prediction in Dynamic Object Tracking: A Critical Study
A survey of dynamic object tracking and trajectory prediction that identifies gaps and proposes a conceptual feedback-loop integration, but presents no formal model or experimental validation.
Discussion (0). Continue with ORCID to comment.