REVIEW 37 cited by
BoT-SORT: Robust Associations Multi-Pedestrian Tracking
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The goal of multi-object tracking (MOT) is detecting and tracking all the objects in a scene, while keeping a unique identifier for each object. In this paper, we present a new robust state-of-the-art tracker, which can combine the advantages of motion and appearance information, along with camera-motion compensation, and a more accurate Kalman filter state vector. Our new trackers BoT-SORT, and BoT-SORT-ReID rank first in the datasets of MOTChallenge [29, 11] on both MOT17 and MOT20 test sets, in terms of all the main MOT metrics: MOTA, IDF1, and HOTA. For MOT17: 80.5 MOTA, 80.2 IDF1, and 65.0 HOTA are achieved. The source code and the pre-trained models are available at https://github.com/NirAharon/BOT-SORT
Forward citations
Cited by 37 Pith papers
-
Sonic Stage: Auto-Generating Interactive Spatial Soundscapes to Facilitate Dialogue Video Comprehension for Blind Viewers
A system that spatializes dialogue, adds diegetic sound, and offers tap-to-hear descriptions improved blind viewers' understanding of character position, movement, actions, and visual details in dialogue-only scenes.
-
Training-Free Off-Screen Player Imputation for Broadcast-Based Spatial Football Analytics
Role-anchored centroid voting, a training-free online imputer, roughly halves hidden-zone pitch-control error from ignoring off-screen players and cuts control-share error to 28–48% of the ignore baseline across three...
-
Unsupervised Detection of Entry and Exit Regions from Vehicle Trajectories for Camera-Agnostic Turning Movement Counts
Clustering start and end points of vehicle trajectories yields persistent entry/exit polygons that classify turning movements at ~3% median error across 25 uncalibrated cameras.
-
WaspMOT: A Benchmark for Long-Term Multi-Object Tracking of Trichogramma Wasps
A new insect-tracking benchmark reveals that five standard MOT methods suffer severe identity fragmentation on 8-minute sequences even with oracle detections, with simple spatial stitching recovering significant gains.
-
Resonance-enhanced integrated acousto-optic beam steering
A TFLN ring-resonator-enhanced acousto-optic beam steerer reaches 26% efficiency and 18° FOV and supports FMCW LiDAR via electro-optic resonance locking.
-
Video Individual Counting and Tracking from Moving Drones: A Benchmark and Methods
GD3A and DVTrack, driven by optimal-transport descriptor matching with an adaptive dustbin score, set state-of-the-art results on a new moving-drone dense-crowd counting and tracking benchmark.
-
Learning Association via Track-Detection Matching for Multi-Object Tracking
TDLP uses a link-prediction head to match tracks to detections, beating heuristic and metric-learning trackers on several MOT benchmarks while underperforming on MOT17.
-
DeepSea MOT: A benchmark dataset for multi-object tracking on deep-sea video
DeepSea MOT, a public four-video benchmark with manually corrected ground truth, enables the first standardized HOTA comparisons of multi-object detectors and trackers on deep-sea ROV footage.
-
MVTrajecter: Multi-View Pedestrian Tracking with Trajectory Motion Cost and Trajectory Appearance Cost
An end-to-end multi-view pedestrian tracker that aggregates motion and appearance costs over K past timestamps, outperforming prior methods on GMVD, Wildtrack, and MultiviewX, though with a validation-protocol concern...
-
VIFSS: View-Invariant and Figure Skating-Specific Pose Representation Learning for Temporal Action Segmentation
Claims a view-invariant pose representation method for figure skating jump segmentation reaches over 92% F1@50, but the submitted text is a different manuscript, leaving the claim unverified.
-
GRASPTrack: Geometry-Reasoned Association via Segmentation and Projection for Multi-Object Tracking
A depth-aware MOT tracker using mask-guided 3D point clouds, voxelized 3D IoU association, adaptive Kalman noise, and 3D motion consistency surpasses prior TBD methods on MOT17, MOT20, and DanceTrack.
-
TrackOR: Towards Personalized Intelligent Operating Rooms Through Robust Tracking
TrackOR keeps persistent staff identities in the operating room by matching 3D geometric signatures, reporting +11% association accuracy over its strongest baseline and enabling per-person workflow trajectories.
-
Multi-tracklet Tracking for Generic Targets with Adaptive Detection Clustering
A tracklet-based multi-hypothesis tracker with adaptive detection clustering achieves competitive MOTA and IDF1 on GMOT-40 without category-specific knowledge.
-
LAVA: Language Driven Scalable and Versatile Traffic Video Analytics
A natural-language video analytics system combining bandit-based sampling, open-vocabulary detection, and trajectory linking reports higher query accuracy than closed-world baselines on a new 18-predicate traffic benchmark.
-
RoundaboutHD: High-Resolution Real-World Urban Environment Benchmark for Multi-Camera Vehicle Tracking
A high-resolution, real-world roundabout dataset for multi-camera vehicle tracking, with 512 identities, four non-overlapping 4K cameras, and baselines across four tasks.
-
USVTrack: USV-Based 4D Radar-Camera Tracking Dataset for Autonomous Driving in Inland Waterways
A new USV dataset combining 4D radar, camera, GPS, and IMU data for tracking boats, ships, and vessels on inland waterways, together with a radar-camera matching method that consistently improves two-stage trackers.
-
Multiple Object Tracking in Video SAR: A Benchmark and Tracking Baseline
A new public video SAR multi-object tracking benchmark, plus a DETR-based tracker with line feature enhancement and motion-aware association, reports state-of-the-art results on that benchmark.
-
MOSE: A Novel Orchestration Framework for Stateful Microservice Migration at the Edge
A new edge-computing framework orchestrates stateful container migration, preserving live connections and meeting user-chosen KPI targets with up to 77 percent lower downtime.
-
Out of the Past: An AI-Enabled Pipeline for Traffic Simulation from Noisy, Multimodal Detector Data and Stakeholder Feedback
An AI pipeline (computer vision, quadratic optimization, and LLM-generated constraints) builds a Strongsville traffic simulation from noisy camera and loop detector data, but the simulation's accuracy is only checked ...
-
Event-RGB Adaptive Tracking for Nighttime Highway Perception
JEAT jointly associates RGB and event detections with NIS-adapted measurement noise, raising MOTA on unlit nighttime highways from 46% (RGB) / 69% (event) to 77% on a new CARLA dataset.
-
CLIFE: Camera-LiDAR Fusion Framework for Edge-Deployable Roadside VRU Perception
An edge-deployed camera–LiDAR late-fusion system with targetless online calibration achieves real-time VRU tracking on a single Jetson, but its robustness claims are only partially supported by the experiments.
-
Efficient Perception in Automotive Detection and Tracking Using Neuromorphic Computing
Transfer-learned SpikeYOLO achieves mAP 0.937/0.771 and HOTA 0.701/0.445 on KITTI and BDD100K for two-class automotive detection and tracking, competitive with conventional deep networks.
-
A Task-Driven Evaluation of UAV Detection and Tracking under Synthetic Fog
Fog degrades UAV detection and tracking mainly via missed detections; fog-inclusive training is more robust than test-time dehazing, and restoration quality does not proportionally improve downstream perception.
-
From Camera-Based Sensing to Reasoning: A Comprehensive Review Toward Proactive Vulnerable Road User Safety
A survey organizing recent camera-based AI methods for vulnerable road user safety into four interlocking visual tasks and four open deployment challenges.
-
Motion Estimation for Multi-Object Tracking using KalmanNet with Semantic-Independent Encoding
SIKNet, a learned Kalman filter with a semantic-independent encoder, improves bounding-box motion prediction in multi-object tracking over the classic Kalman filter and prior KalmanNet variants.
-
SoccerNet 2025 Challenges Results
The SoccerNet 2025 challenges benchmarked four football video understanding tasks, and top submissions beat the organizer baselines by large margins.
-
MeMoSORT: Memory-Assisted Filtering and Motion-Adaptive Association Metric for Multi-Person Tracking
MeMoSORT reports state-of-the-art multi-person tracking HOTA of 67.9% on DanceTrack and 82.1% on SportsMOT by combining a memory-compensated Kalman filter with a motion-adaptive IoU association metric.
-
Head Anchor Enhanced Detection and Association for Crowded Pedestrian Tracking
FocusTrack, a YOLOX-based tracker using head keypoints and detector features, matches but does not beat simple baselines on MOT17 and MOT20, and its 3D trajectory completion underperforms 2D interpolation.
-
CrowdTrack: A Benchmark for Difficult Multiple Pedestrian Tracking in Real Scenarios
CrowdTrack is a dense, first-person-view pedestrian tracking benchmark that exposes large performance drops in existing multi-object trackers.
-
SportMamba: Adaptive Non-Linear Multi-Object Tracking with State Space Models for Team Sports
A Mamba-plus-attention motion predictor with a height-adaptive IoU matching metric achieves state-of-the-art HOTA on SportsMOT and strong zero-shot results on VIP-HTD.
-
Depth-Aware Scoring and Hierarchical Alignment for Multiple Object Tracking
A training-free MOT framework that adds zero-shot depth histograms and a hierarchical box/mask alignment score to association, with mixed state-of-the-art results.
-
Glance-MCMT: A General MCMT Framework with Glance Initialization and Progressive Association
Glance-MCMT combines BoT-SORT single-camera tracks, a short glance phase to seed global IDs, and progressive cross-view association, reaching 51.34 HOTA on AI City 2025 validation data.
-
VisionGuard: Synergistic Framework for Helmet Violation Detection
A two-module post-processing framework (tracking-based relabeling plus virtual bounding box injection) reports +1.6% to +3.1% mAP@50 on helmet violation detection, but only on a self-annotated test set with fitted con...
-
YASMOT: Yet another stereo image multi-object tracker
A lightweight detection-only multi-object tracker uses Gaussian distance plus Hungarian matching to track objects over time, link stereo views, and combine ensemble detector output.
-
Robust Video-Based Pothole Detection and Area Estimation for Intelligent Vehicles with Depth Map and Kalman Smoothing
A pothole detection and area-estimation pipeline combining a modified YOLOv8 detector, monocular depth, and Kalman smoothing, with area estimates validated only for internal consistency, not against ground truth.
-
A Novel Tuning Method for Real-time Multiple-Object Tracking Utilizing Thermal Sensor with Complexity Motion Pattern
A simple SORT-based tracker with manually tuned hyperparameters won the PBVS TP-MOT thermal pedestrian tracking challenge, beating ReID and diffusion-based trackers.
- NOVA: Navigation via Object-Centric Visual Autonomy for High-Speed Target Tracking in Unstructured GPS-Denied Environments
Discussion (0). Continue with ORCID to comment.