{"id":"b11eb3f2-d40f-4ed8-9d1d-d1f4cd914a5b","arxiv_id":"2505.04793","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"DetReIDX is a new multi-session aerial-ground person re-identification dataset spanning 5.8 to 120 meters altitude, and the paper reports that current SOTA detection and ReID methods degrade heavily on it.","lead":"This paper introduces DetReIDX, a new aerial-ground dataset for person re-identification with over 13 million claimed bounding boxes, two clothing-change sessions, and drone altitudes up to 120 meters. It reports that state-of-the-art detection and ReID models degrade sharply on this data, positioning the set as a stress test for real-world UAV surveillance.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"ReID 'catastrophic degradation' claim lacks same-model baselines on existing datasets; Table IX reports only DetReIDX numbers, so the central degradation claim is unverified.","rationale":"The reader's verdict is CONDITIONAL and lists both missing baselines and annotation quality as issues. I focus on the missing ReID baselines because it attacks the central empirical claim directly. The dataset's stated purpose is to show that SOTA methods 'catastrophically degrade' relative to existing benchmarks; if that comparison is absent, the headline quantitative claim (abstract: 'over 70% in Rank-1 ReID') is not established by the paper. This is not a consensus disagreement but a missing-control issue: the same models are not measured on the control condition. Annotation quality is also important, but the paper at least describes a cross-verification procedure; while no inter-annotator agreement is reported, that is a quality-reporting gap rather than a demonstrated contradiction. The missing baseline is a logical gap: the text explicitly asserts 'relatively good performance on existing ground-level datasets' without providing those data. A single table of same-model results on Market-1501/DukeMTMC would settle it. Detection degradation is internally supported (Table VII), so the concern is specific to the ReID half of the central claim. I do not think this overturns the dataset's potential usefulness; it makes the paper's evidence incomplete, which matches the CONDITIONAL verdict. I therefore recommend UNCHANGED.","tokens_in":13707,"tokens_out":8240,"duration_ms":80282,"concrete_test":"Run PersonViT, SeCap, and CLIP-ReID on Market-1501 and DukeMTMC-ReID (and, if feasible, AG-ReID.v2) using the same training/inference code and evaluation protocol as in Section IV-B, and report R1/mAP alongside the DetReIDX results. If these models achieve high R1 (e.g., >80%) on the existing benchmarks but low R1 on DetReIDX A2A/A2G/G2A, the catastrophic-degradation claim is supported. If they also score low on existing benchmarks, the paper must either provide corrected baselines or weaken the claim to 'DetReIDX is a hard benchmark' rather than 'models catastrophically degrade on it.' This check also requires the authors to release the evaluation code; without it, the comparison cannot be reproduced.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that DetReIDX is a stress test exposing catastrophic failure of SOTA ReID methods. The detection part is internally supported: Table VII shows YOLOv8 AP50 dropping from 0.914 in D1->D1 to 0.137 in D1->D3, an 81% relative drop. The ReID part is not supported. Section IV-B asserts that all models perform 'relatively good on the existing ground-level datasets' and then degrade on DetReIDX, but Table IX gives only absolute mAP/R1 values on DetReIDX (e.g., PersonViT A2A R1=8.8, A2G R1=19.6, G2A R1=51.9; CLIP-ReID G2A R1=58.1). No same-model results on Market-1501, DukeMTMC-ReID, AG-ReID.v2, or any prior benchmark are reported. Without those baselines, the word 'degrade' is unsupported: low absolute performance could reflect weak training protocols, poor hyperparameters, or the models' general unsuitability, not a DetReIDX-specific stress. The abstract's 'over 70% Rank-1 loss' therefore cannot be verified. This missing comparison is the load-bearing empirical evidence for the dataset's stated purpose, and it is absent from the manuscript.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces DetReIDX, a large-scale aerial-ground multi-session person dataset with 509 identities, indoor ground references, and UAV captures from 18 viewpoints at altitudes between 5.8 m and 120 m, annotated for detection, tracking, ReID, soft biometrics, and action recognition. It reports benchmarks on three detectors and three ReID models, arguing that performance degrades sharply on DetReIDX, especially at long distances and in cross-view settings, and concludes that DetReIDX should serve as a stress-test benchmark for UAV-based person recognition.","tokens_in":13898,"tokens_out":4912,"duration_ms":46574,"significance":"If the dataset is as described and released, it fills a real gap: existing aerial-ground benchmarks have narrower altitude ranges, lack per-subject clothing changes across sessions, and provide fewer annotation types. The two-session protocol with clothing variation, the 120 m altitude range, and the multi-task annotations are valuable assets. The authors also state that the dataset and evaluation protocols will be publicly released, which supports community use and reproducibility. The empirical claims currently outrun the evidence: the headline bounding-box count is inconsistent across sections, the ReID degradation claim is unsupported by same-model baseline numbers on existing datasets, and annotation quality is not quantified. These issues are fixable, but they prevent the paper from supporting its strongest claims as written.","major_comments":[{"comment":"The headline scale claim is internally inconsistent: the abstract and Section III say \"over 13 million bounding boxes\", Table I lists 12.6M, and Table V sums to 11,797,199 (5,095,539 + 2,483,836 + 4,217,824). Since this is the paper's primary quantitative claim, the authors must reconcile the numbers and report a single official count for the released annotations.","section":"Abstract and Sections III-A, III-F; Tables I and V"},{"comment":"The ReID degradation claim is not supported by the reported experiments. Table IX gives only absolute mAP/R1/R5/R10 values on DetReIDX; no same-model results on Market-1501, DukeMTMC-ReID, AG-ReID.v2, G2APS, or any prior benchmark appear anywhere in the paper. The sentence in Section IV-B that models perform \"relatively good on the existing ground-level datasets\" is therefore an assertion, not a result. Moreover, Table X's altitude breakdown shows at most a 45.7% relative Rank-1 drop (A2A D1 to D3), so the abstract's \"over 70% Rank-1 loss\" is not derivable from the presented data. The authors should add baseline tables using identical training and evaluation protocols or revise the claims.","section":"Section IV-B and Tables IX-X"},{"comment":"Annotation quality is load-bearing and unquantified. The text says annotations were \"manually done by a set of volunteers, using the CVAT tool and cross-verified by peers,\" but no inter-annotator agreement, label-error rate, or verification protocol is reported. Since the ReID evaluation depends on consistent PID labels across sessions and the detection evaluation includes ROIs of 8x8 pixels, the absence of quality statistics undermines the benchmark numbers.","section":"Section III-D"},{"comment":"Split statistics and training details are inconsistent or incomplete. Section IV-A states a 70-20-10 split, but Table V gives Train/Val/Test videos of 120/56/109 out of 285, i.e., 42%/20%/38%. Section IV-B states a 70%-30% PID-disjoint split, but the reported 267 train and 67 test identities correspond to roughly 80%/20% of the 334 outdoor-reobserved identities, not 70/30 of the 509 total. In addition, no optimizer, learning rate, batch size, input resolution, number of epochs, or number of independent runs is reported for any model, and the detection results come from a single split, so the reader cannot assess variance or rule out undertraining as an explanation for low scores.","section":"Sections IV-A and IV-B, Tables V and VIII"}],"minor_comments":[{"comment":"The manuscript contains numerous typos and spacing inconsistencies ('real-wordlsettings', 'reserch', 'Futher Research', 'UA V'); a careful copyedit is needed.","section":"Throughout"},{"comment":"Reference [16] appears to be mismatched: the CSM dataset is cited to a paper on movie popularity prediction, which is not a person ReID dataset reference.","section":"References [16]"},{"comment":"Figure 10 highlights a 'critical distance (70 meters)', but the distance bins are defined as D1<20m, D2=20-50m, and D3>50m; the relation between 70m and the bins should be clarified.","section":"Figure 10 and Section IV-A"},{"comment":"Table I lists the DetReIDX height range as 5-120m, while the text and Table IV give 5.8-120m; the ranges should be made consistent.","section":"Table I and Table IV"}],"recommendation":"major_revision","confidential_remarks":"The dataset resource appears genuinely useful, and the authors have made a serious collection effort. My main concern is that several headline claims (bounding-box count, ReID degradation, split ratios) are not yet supported by the manuscript's own tables, and the lack of annotation-QA statistics is a risk for downstream benchmark use. I would support a major revision that adds the missing baseline experiments and quality metrics rather than a rejection, because the dataset itself could be a solid contribution to the community."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nDetReIDX is worth knowing about. It's a new aerial-ground person dataset that genuinely combines things no existing public benchmark does: drone altitudes up to 120 m, two capture sessions per subject with clothing changes, indoor ground references, and multi-task annotations including 16 soft biometric labels. If the data is clean, it fills a real gap for long-term ReID stress testing. The collection protocol is thoughtful and the scale is nontrivial.\n\nThe detection experiments are internally consistent: YOLOv8 AP50 drops from 0.914 in D1→D1 to 0.137 in D1→D3, and DDOD/Grid-RCNN collapse even harder. That supports the \"stress test\" idea for detection. The ReID story is weaker. Table IX reports absolute mAP/R1 on DetReIDX only. The paper asserts models do \"relatively good\" on ground benchmarks and then degrade, but no same-model baseline on Market-1501, Duke, or AG-ReID is reported. Without that, \"over 70% Rank-1 loss\" is unverifiable — low absolute numbers could just mean weak training or poor adaptation, not dataset-specific difficulty. That needs fixing.\n\nOther issues: the bounding box count is inconsistent across the abstract (13M+), Table I (12.6M), and Table V (11.8M). That's sloppy for a dataset paper. No inter-annotator agreement or label-stability statistics, which matters given tiny ROIs (sub-10px). No code release mentioned. Single-split detection results without error bars. The demographic skew (68% Indian, 58% male, 89% aged 18–24) is shown in a figure but not discussed as a limitation.\n\nThe central dataset claim holds up. I'd send this to a serious referee, but the ReID baselines and count consistency need to be addressed before publication. The dataset itself, if downloadable and genuinely as described, could become a standard stress-test benchmark for drone-based ReID. The paper needs revision, but it deserves the review.","headline":"A genuinely new aerial-ground dataset with real stress-test potential, undermined by missing ReID baselines and inconsistent box counts.","tokens_in":14561,"tokens_out":1707,"would_cite":true,"duration_ms":17452,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces DetReIDX, a large aerial-ground person dataset with more than 13 million bounding boxes, and reports that state-of-the-art detectors and ReID models lose up to 80% detection accuracy and over 70% Rank-1 accuracy on it.","keywords":["person re-identification","UAV surveillance","aerial-ground dataset","cross-view recognition","long-term ReID","clothing variation","soft biometrics","stress-test benchmark"],"falsifier":"Re-annotate a random subset of DetReIDX from scratch with independent annotators and compare identity assignments and bounding boxes to the released labels. If cross-session identity agreement is materially incomplete, or if many boxes under 10 pixels are misplaced, the published AP50 and Rank-1 numbers are not a clean measure of model failure. A second check: train and evaluate a ReID model on a cleanly re-annotated subset; if Rank-1 rises far above the published values, annotation noise, not the real-world variability DetReIDX claims to capture, explains the collapse.","tokens_in":13454,"feed_emoji":"🚁","tokens_out":9353,"duration_ms":86728,"temperature":0.7,"pith_summary":"DetReIDX is a person dataset built from drone and ground recordings of 509 volunteers across seven campuses on three continents, with more than 13 million bounding boxes captured at altitudes between 5.8 and 120 meters. The paper's central claim is that this is the first large-scale aerial-ground benchmark that combines two-session clothing changes with extreme range and viewpoint variation, and that current state-of-the-art detectors and ReID models fail on it: up to an 80% drop in detection accuracy and more than 70% Rank-1 loss. Existing benchmarks, the authors argue, either stay at close range, keep clothing fixed, or lack cross-view pairing, so they cannot reveal how much models depend on color, silhouette, and near-ground viewpoints. If the claim is right, DetReIDX gives the field a standardized stress test for long-term, cross-view person recognition under realistic drone conditions.","feed_headline":"Drone stress test drops ReID Rank-1 below 9%","feed_subtitle":"Spanning 5.8 to 120 meters altitude and two-session clothing changes, it exposes appearance-based shortcuts.","key_machinery":"The load-bearing object is the collection protocol rather than any single model. Each identity receives an indoor reference set (mugshots from three angles plus a gait video) and two outdoor drone sessions on different days with different clothing, captured from 18 fixed viewpoints spanning pitch angles of 30, 60, and 90 degrees and horizontal distances from 10 to 120 meters. This produces a controlled axis of degradation: bounding boxes shrink from more than 1000 pixels tall indoors to under 10 pixels at the farthest viewpoints, while the two-session design removes clothing as a reliable cue. The dataset defines three ReID evaluation modes (aerial-to-aerial, aerial-to-ground, ground-to-aerial) and distance bins D1 (<20m), D2 (20–50m), and D3 (>50m), which are the controlled variables through which the paper measures model collapse.","core_discovery":"The central discovery is empirical: when state-of-the-art models are evaluated on DetReIDX, they catastrophically degrade. On detection, training on short-range views and testing on long-range views drops YOLOv8's AP50 from 91.4% to 13.7%, while DDOD and Grid-RCNN fall below 1%; extrapolating to unseen 90-degree top-down views also drops AP50 by between 22% and 35% relative. On ReID, all three tested models score below 9% Rank-1 in aerial-to-aerial cross-session matching, and performance keeps falling as drone distance grows from D1 to D3. The authors interpret this as evidence that current models rely on appearance cues—clothing, color, texture, full-body silhouettes—that DetReIDX systematically removes through clothing changes across sessions and through sub-10-pixel top-down targets.","pith_inferences":["The paper does not decompose the failure by cause; a natural extension is to compare same-view, same-day retrieval against cross-view, cross-day retrieval within the aerial-to-aerial setting, which would isolate the contribution of clothing change alone.","The near-zero D3-to-D1 transfer implies a falsifiable prediction about scale augmentation: if long-range-trained detectors are given scale-invariant pretraining, their short-range AP50 should rise, and DetReIDX's distance bins provide the exact protocol to test this.","Because soft-biometric labels cover height, body volume, and other geometry, one could test whether attributes predicted from indoor images transfer to aerial images; a positive result would support the paper's call for geometry-aware representations.","A stress-test dataset can also be turned into a training signal: selecting hard cross-session, cross-view pairs and fine-tuning on them should improve real-world ReID if DetReIDX captures the true failure modes."],"forward_implications":["Current detectors are viewpoint-locked: training on 30°/60° pitches and testing on an unseen 90° top-down view lowers AP50 by 22–35% relative.","Current ReID models cannot bridge cross-session clothing changes in aerial settings, with Rank-1 below 9% in aerial-to-aerial matching.","Long-range drone views alone do not teach transferable pedestrian features: training on D3 and testing on D1 gives near-zero AP50 for all tested detectors.","The soft-biometric annotations make DetReIDX a testbed for geometry-aware or attribute-based ReID, which the paper argues is the direction needed for robustness."],"supporting_citations":[{"why":"the standard ground-level ReID benchmark whose close-range fixed-camera assumptions DetReIDX challenges","marker":"[1]"},{"why":"the cloth-changing ReID benchmark DetReIDX extends toward long-term clothing variation","marker":"[5]"},{"why":"the drone person dataset whose soft-biometric annotation scheme DetReIDX adapts","marker":"[6]"},{"why":"the large aerial human-understanding dataset used for comparison on altitude and task coverage","marker":"[7]"},{"why":"the closest aerial-ground ReID dataset that DetReIDX claims to surpass in altitude range and clothing change","marker":"[14]"},{"why":"cited for the G2APS aerial-ground dataset and used as the SeCap ReID baseline","marker":"[15]"},{"why":"the one-stage detector whose long-range transfer collapses to 13.7% AP50","marker":"[18]"},{"why":"the dense detector that falls below 1% AP50 in D1-to-D3 transfer","marker":"[19]"},{"why":"the region-based detector baseline used for viewpoint and distance generalization tests","marker":"[20]"},{"why":"the transformer ReID baseline that scores below 9% Rank-1 in aerial-to-aerial cross-session matching","marker":"[21]"}],"fun_headline_variants":["Aerial ReID stress test: Rank-1 crashes below 9%","Long-range drone views crush detector AP50 to 13.7%","Clothing swaps across sessions wreck ReID accuracy","From 5.8 to 120 meters, ReID models fall apart","New UAV dataset: 13M boxes, 3 continents, ReID fails"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire benchmark rests on the manual annotations being correct and consistent across sessions, especially the identity labels staying stable between indoor and outdoor captures and between Session 1 and Session 2, and the paper reports no inter-annotator agreement or label-quality statistics, so noisy labels could by themselves create part of the measured degradation.","fun_headline_variants_meta":{"raw":{"variants":["Aerial ReID stress test: Rank-1 crashes below 9%","Long-range drone views crush detector AP50 to 13.7%","Clothing swaps across sessions wreck ReID accuracy","From 5.8 to 120 meters, ReID models fall apart","New UAV dataset: 13M boxes, 3 continents, ReID fails"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000726,"raw_usage":{"total_tokens":3302,"prompt_tokens":1042,"completion_tokens":2260,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":658,"completion_tokens_details":{"reasoning_tokens":2165}},"tokens_in":658,"tokens_out":2260,"duration_ms":13927,"temperature":1.0,"reasoning_tokens":2165,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:20:42.212549+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-annotate a random subset of DetReIDX from scratch with independent annotators and compare identity assignments and bounding boxes to the released labels. If cross-session identity agreement is materially incomplete, or if many boxes under 10 pixels are misplaced, the published AP50 and Rank-1 numbers are not a clean measure of model failure. A second check: train and evaluate a ReID model on a cleanly re-annotated subset; if Rank-1 rises far above the published values, annotation noise, not the real-world variability DetReIDX claims to capture, explains the collapse.","supporting_citations":[{"cited_title":"Long-term cloth-changing person re-identification,","cited_arxiv_id":null,"evidence_quote":"the cloth-changing ReID benchmark DetReIDX extends toward long-term clothing variation"},{"cited_title":"The p- destre: A fully annotated dataset for pedestrian detection, tracking, and short/long-term re-identification from aerial devices,","cited_arxiv_id":null,"evidence_quote":"the drone person dataset whose soft-biometric annotation scheme DetReIDX adapts"},{"cited_title":"Uav-human: A large benchmark for human behavior understanding with unmanned aerial vehicles,","cited_arxiv_id":null,"evidence_quote":"the large aerial human-understanding dataset used for comparison on altitude and task coverage"},{"cited_title":"Ag-reid. v2: Bridging aerial and ground views for person re-identification,","cited_arxiv_id":null,"evidence_quote":"the closest aerial-ground ReID dataset that DetReIDX claims to surpass in altitude range and clothing change"},{"cited_title":"SeCap: Self-Calibrating and Adaptive Prompts for Cross-view Person Re-Identification in Aerial-Ground Networks","cited_arxiv_id":"2503.06965","evidence_quote":"cited for the G2APS aerial-ground dataset and used as the SeCap ReID baseline"},{"cited_title":"What is yolov8?","cited_arxiv_id":null,"evidence_quote":"the one-stage detector whose long-range transfer collapses to 13.7% AP50"},{"cited_title":"Disentangle your dense object detector,","cited_arxiv_id":null,"evidence_quote":"the dense detector that falls below 1% AP50 in D1-to-D3 transfer"},{"cited_title":"Grid r-cnn,","cited_arxiv_id":null,"evidence_quote":"the region-based detector baseline used for viewpoint and distance generalization tests"},{"cited_title":"Personvit: large-scale self-supervised vision transformer for person re-identification,","cited_arxiv_id":null,"evidence_quote":"the transformer ReID baseline that scores below 9% Rank-1 in aerial-to-aerial cross-session matching"}],"review_version":1}