{"id":"b9654c8b-fa33-46a3-a580-c41d238ea455","arxiv_id":"2501.14070","paper_version":2,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"The BRIAR dataset is extended with BGC3 and BGC4, adding group activities, winter weather, a mock city, and more subjects for long-range whole-body biometrics.","lead":"This paper describes two new collections (BGC3 and BGC4) added to the BRIAR dataset, a restricted-access benchmark for identifying people at long distances and from elevated cameras, totaling over 475,000 images and 3,450 hours of video of 1,760 subjects. It documents collection, curation, annotation, and evaluation protocols for whole-body biometrics.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Identity labels in long-range group/mock-city tracks are verified only at track endpoints; mid-track identity switches could silently contaminate the BTS benchmark, so the dataset's core utility rests on unmeasured annotation fidelity.","rationale":"The reader accepted with high confidence and identified automated annotation accuracy and sparse manual verification as the weakest assumption. I agree that this is the load-bearing point, but it deserves to be made concrete: the specific failure mode is mid-track identity switch in group and mock-city videos, which the paper's endpoint-only verification cannot catch. The manuscript explicitly discloses sparse QA, so this is not a hidden flaw; however, it is an unquantified risk to the evaluation protocol. A dataset paper can be accepted without perfect labels, but the 'comprehensive resource' claim and the BTS benchmark protocol imply identity label correctness. Since the paper includes automated models and sparse manual QA but no measured annotation error rate, a conditional acceptance requiring a sample-based QA report is proportionate. This does not impugn the collection effort; it pins the central claim to a checkable quantity. If the test shows errors are negligible, the original ACCEPT stands; if not, the dataset release should include error estimates or annotation corrections.","tokens_in":10714,"tokens_out":4404,"duration_ms":42232,"concrete_test":"Select a stratified random sample of 100 tracks from the new group-backpack, mock-city, and 500-720 m single-subject videos. For each track, have independent annotators label the visible subject ID on every 5th frame (using full-body appearance, clothing set, and entry/exit logs), blind to the automated/XML labels. Report (a) the fraction of tracks with at least one frame-level identity disagreement, and (b) the frame-level identity error rate. If either is non-negligible (e.g., >1% of tracks), the paper should quantify the resulting noise in BTS evaluation or release corrected labels before claiming the dataset is a clean benchmark; if both are essentially zero, the current sparse QA is vindicated.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim is that BRIAR BGC1-4 is a comprehensive, usable whole-body biometric resource. That usability depends on the XML identity labels being correct, especially in the new group (Section III-C 'Group Backpack') and BGC4 mock-city (Section III-D) videos, where multiple subjects occlude and re-appear. The curation pipeline (Section V-A.4) manually verifies only the first and last subject of each collection day, and the annotation stage (Section V-B) verifies only the first and last frame of each automated track. Neither step can detect a mid-track identity switch produced by BoT-SORT or an incorrect Re-ID association by DG-Net++. In mock-city videos, individual activities are not timestamped—only entry/exit times are recorded—so there is no independent record to disambiguate which subject a track should follow after an occlusion or group crossing. At 500-720 m with atmospheric turbulence and low resolution, automated whole-body recognition is known to be error-prone. If even a small percentage of tracks switch identities or are assigned to the wrong subject, the BTS probe/gallery protocol (Section V-C) is contaminated, and any recognition scores produced on the benchmark become unreliable. The paper nowhere reports a measured annotation error rate, so the load-bearing assumption—that sparse endpoint verification suffices—is untested.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports the extension of the BRIAR dataset with two new collections, BGC3 and BGC4, and describes their composition, collection methodology, curation pipeline, annotation procedure, and evaluation protocol. The new data comprises over 125,000 additional images and 83,000 additional videos across new field activities (including group backpack scenarios), a winter-weather site, and an indoor mock-city environment (Hogan's Alley), bringing the cumulative dataset to more than 475,000 images and 3,450 hours of video of 1,760 subjects (1,173 full subjects plus 587 distractors). The paper does not present any experimental results; it is a dataset description and resource announcement.","tokens_in":10974,"tokens_out":5226,"duration_ms":47106,"significance":"If the dataset is as described, it is a substantial and valuable resource for long-range, elevated-view, and whole-body biometric recognition, exceeding the subject count and distance coverage of prior public datasets in this niche. The paper's strengths include the detailed documentation of collection infrastructure, the transparent description of the automated annotation pipeline, the explicit ethical oversight and consent process, and the careful definition of the BRS/BTS partition and evaluation protocol. The dataset has already been used by the program's own evaluations (referenced in [4]), and it fills a clear gap in the community. However, the paper does not quantify the accuracy of the identity labels that underpin the dataset's benchmark utility, and this is a significant gap for a resource intended to support rigorous evaluation.","major_comments":[{"comment":"The identity-label verification in the annotation pipeline is limited to manual checking of only the first and last frame of each automated track (Section V.B), and the curation QA covers only the first and last subject of each collection day (Section V.A.4). The paper reports no measured accuracy of the automated chain (YOLOv5, MeshTransformer, DARK, DG-Net++, BoT-SORT) and no end-to-end validation of the final identity labels. Because the BTS probe/gallery protocol (Section V.C) relies on the correctness of per-track subject identity, an undetected mid-track identity switch would directly contaminate benchmark results. This is a load-bearing assumption for the dataset's core claim of being a usable benchmark. Please provide a validation study, such as a random sample of tracks manually verified frame-by-frame, or a comparison of the automated labels against an independent annotation source, or at minimum an explicit analysis of the expected error rate and its potential impact on evaluation conclusions.","section":"V.B and V.A.4"},{"comment":"In the Hogan's Alley mock-city scenario, individual subject activities are not timestamped; only entry and exit times are recorded. This makes it especially difficult to disambiguate identity after occlusions, group crossings, or subjects leaving and re-entering the scene. Combined with the sparse endpoint-only verification of automated tracks, this is a concrete source of possible identity-label errors in the new data. The paper should state this as a known limitation and describe any mitigation (e.g., additional manual checks for group videos, or flags for low-confidence tracks) so that users can properly interpret results obtained on these videos.","section":"III.D"}],"minor_comments":[{"comment":"The term 'UA Vs' appears with a stray space in the abstract and in several other places; please standardize to 'UAVs' or 'UAS'.","section":"Abstract and Section I"},{"comment":"In the Contributions paragraph, the sentence 'It is the first dataset is of its kind' contains an extra 'is' and should read 'It is the first dataset of its kind'.","section":"Section I"},{"comment":"The row for 'BRIAR BGC 1-4 (this paper) [7]' cites reference [7], which is the original BRIAR dataset paper; this should cite the present work. Also, '1000m' should be formatted as '1,000 m' for consistency with the rest of the text.","section":"Table I"},{"comment":"The dataset summary reports 'over 475,000 images' and 'up to 1,000 m', but Figures 10 and 11 show BGC3/BGC4 distance maxima of 500 m and 720 m. Please clarify explicitly that the 1,000-m figure and the 475,000-image total include data from BGC1/2, so that the contribution of the new collections is unambiguous.","section":"Section IV"}],"recommendation":"major_revision","confidential_remarks":"The dataset itself is clearly valuable and the paper is a useful technical resource, but the absence of even a basic annotation-accuracy estimate makes the central usability claim shaky. I would like to see either a small validation study or a clearly worded limitations section addressing the risk of identity-label errors, especially for the group and mock-city videos. The paper is otherwise well-organized and the ethical documentation is exemplary."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this is a good dataset paper and should go to review. BGC3 and BGC4 are genuinely new collection events—new subjects, locations, group scenarios, a mock city, winter weather, and longer-range rooftop cameras. The paper doesn't introduce recognition methods, but that's not its job. It describes the data, the curation pipeline, annotations, and the BTS protocol clearly and honestly.\n\nThe strongest part is the curation discussion. The paper walks through timestamp validation, video cutting, XML metadata generation, QA, and the automated annotation chain (YOLOv5, MeshTransformer, DARK, DG-Net++, BoT-SORT). It is refreshingly specific about what was done and what was checked. The demographic stats and distance/elevation/yaw distributions are useful. The BRS/BTS split being subject-disjoint and demographic-balanced is sensible.\n\nThe soft spot is exactly what the stress-test note says: annotation fidelity is unmeasured. Manual verification only checks the first and last frame of each track, and daily QA only checks the first and last subject. In the new group-backpack and mock-city videos, where subjects occlude and cross paths, a mid-track identity switch would silently contaminate the labels. The paper does disclose this—it says 'only the first and last frame of a track were used for verification'—but it never quantifies the error rate. That is a real limitation for a benchmark whose whole value depends on identity labels being right. It is not a fatal flaw: the paper is a resource, not a results paper, and the field can assess annotation quality through use. But a future revision should either provide a measured annotation error rate or at least a more detailed discussion of likely failure modes and their impact on the BTS protocol.\n\nWho is this for? Anyone working on long-range, whole-body biometrics, person re-ID, or turbulence mitigation. It is already cited by over 100 papers, so it clearly fills a need. I would accept it with minor revisions; ask for a more explicit treatment of annotation error and its consequences for benchmark scores. The dataset itself is access-restricted, which is stated clearly and is expected for this kind of data.","headline":"Solid, honest dataset resource; the annotation-QA limitation is real but disclosed, and the paper deserves peer review.","tokens_in":11537,"tokens_out":2194,"would_cite":true,"duration_ms":19849,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper extends the BRIAR dataset with collections 3 and 4, making it the largest resource for whole-body biometric recognition at extreme distances and altitudes.","keywords":["biometric recognition","whole-body","long-range","altitude","dataset","curation","annotation","surveillance"],"falsifier":"An independent audit that inspects every frame of a random sample of group and mock-city videos, or that re-runs tracking with a different algorithm and compares label switches, would reveal the rate of identity mislabelling; a nontrivial switch rate in the released annotations would contradict the dataset's claim to be a reliable ground-truth resource.","tokens_in":10558,"feed_emoji":"📹","tokens_out":8335,"duration_ms":65457,"temperature":0.7,"pith_summary":"The paper reports an extension of the BRIAR dataset, a large-scale resource for whole-body biometric recognition under extreme distances, high altitudes, and real-world conditions. With the addition of Government Collections 3 and 4, the dataset now contains over 475,000 images and 3,450 hours of video of 1,760 subjects (1,173 full subjects and 587 distractors), captured at ranges up to 1,000 meters and view angles up to 50 degrees. The new collections introduce group activities, winter weather, and a mock-city setting, and the paper describes the collection, curation, and annotation methods that turn raw footage into a benchmark. The authors argue that this makes the dataset the largest of its kind and a foundation for developing recognition algorithms that work in operational surveillance scenarios.","feed_headline":"BRIAR biometric dataset grows to 475,000 images and 3,450 hours","feed_subtitle":"New collections add winter scenes, group activities, and a mock city for whole-body recognition.","key_machinery":"The load-bearing mechanism is the multi-stage curation and annotation pipeline that converts raw multi-sensor footage into a usable benchmark. It includes automated timestamp validation to catch scheduling conflicts, activity-based video segmentation, XML metadata generation linking subjects, sensors, weather, and atmospheric measurements, and a chain of automated models (YOLOv5 for whole-body detection, MeshTransformer and DARK for pose and mesh estimation, DG-Net++ for re-identification, and BoT-SORT for tracking). Manual annotators verify only the first and last frame of each track and the first and last subject of each day, and their corrections are merged with the automatic outputs. This pipeline is what makes the dataset's identity labels and annotations trustworthy enough to serve as ground truth.","core_discovery":"On its own terms, the paper claims that the BRIAR dataset, with the BGC3 and BGC4 additions, is the largest and most comprehensive resource for whole-body biometric recognition at altitude and range. Specifically, the dataset comprises 1,173 full subjects plus 587 additional distractors, totaling over 475,000 images and 3,450 hours of video, with field collection at distances up to 720 meters in BGC4 and up to 1,000 meters in the program overall. New elements include group-backpack and pointing activities, a mock-city environment called Hogan's Alley, and a winter-weather collection in a Chicago suburb, all intended to expose recognition models to occlusion, naturalistic behavior, and adverse atmospheric conditions. The paper also details a curation pipeline that generates per-video XML metadata, an automated annotation chain (whole-body detection, pose estimation, re-identification, tracking), and an evaluation protocol with research/test splits, FaceIncluded/FaceRestricted probes, and simple and blended galleries.","pith_inferences":["One could test the annotation-quality assumption by re-verifying every frame of a random sample of group videos; a high identity-switch rate would mean the sparse verification is insufficient.","The mock-city footage likely supports additional computer-vision tasks such as action recognition and multi-person tracking, though the paper only frames it as recognition data.","Because the winter collection damaged sensors and altered viewing conditions, cross-season comparisons using this dataset should first check for systematic differences in image quality between locations.","The reported distance and elevation distributions could be used to benchmark how algorithm performance degrades with range, an analysis the paper does not perform."],"forward_implications":["Researchers gain a common, larger-scale benchmark for evaluating whole-body recognition algorithms at extreme distances, with weather and scenario diversity that prior datasets lacked.","Models developed on this dataset could transfer more readily to operational surveillance systems on rooftops, UAVs, and city streets, where faces are small or absent and body and gait cues matter.","The evaluation protocol, with FaceIncluded and FaceRestricted probe sets and simple versus blended galleries, gives the field a standard way to measure how enrollment quality affects identification performance.","The subject-disjoint research/test split with balanced demographics supports fairness studies and reduces the risk of person-specific overfitting.","The mock-city and group-scenario data create opportunities to study multi-person tracking and occlusion, which are common in real deployments but underrepresented in older face-focused datasets."],"supporting_citations":[{"why":"the original BRIAR dataset that this work extends","marker":"[7]"},{"why":"the evaluation framework that motivates the dataset expansion","marker":"[4]"},{"why":"covariate analysis showing the need for diverse data","marker":"[5]"},{"why":"the whole-body detection model used in the annotation pipeline","marker":"[11]"},{"why":"the multi-object tracker used to generate subject tracks","marker":"[2]"},{"why":"the transformer-based mesh reconstruction model used for pose","marker":"[14]"},{"why":"the distribution-aware coordinate representation for 2D pose estimation","marker":"[21]"},{"why":"the re-identification model used to associate detections with subjects","marker":"[23]"}],"fun_headline_variants":["BRIAR dataset expands to 475K images and 3,450 hours","Winter, mock city, and group scenes join BRIAR biometric dataset","BRIAR set: whole-body biometrics at extreme distances","BRIAR expansion: 475K images, 3,450 hours, 1 km range","Largest whole-body biometric dataset for distance recognition"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that the automated annotation models, together with sparse manual verification of only the first and last frames of each track, produce correct and consistent subject identity labels across hundreds of thousands of clips; if that assumption fails, the dataset's value as a recognition benchmark is undermined.","fun_headline_variants_meta":{"raw":{"variants":["BRIAR dataset expands to 475K images and 3,450 hours","Winter, mock city, and group scenes join BRIAR biometric dataset","BRIAR set: whole-body biometrics at extreme distances","BRIAR expansion: 475K images, 3,450 hours, 1 km range","Largest whole-body biometric dataset for distance recognition"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001214,"raw_usage":{"total_tokens":4945,"prompt_tokens":840,"completion_tokens":4105,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":456,"completion_tokens_details":{"reasoning_tokens":4009}},"tokens_in":456,"tokens_out":4105,"duration_ms":22761,"temperature":1.0,"reasoning_tokens":4009,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T15:23:16.049028+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An independent audit that inspects every frame of a random sample of group and mock-city videos, or that re-runs tracking with a different algorithm and compares label switches, would reveal the rate of identity mislabelling; a nontrivial switch rate in the released annotations would contradict the dataset's claim to be a reliable ground-truth resource.","supporting_citations":[{"cited_title":"Cornett, J","cited_arxiv_id":null,"evidence_quote":"the original BRIAR dataset that this work extends"},{"cited_title":"Aykac, J","cited_arxiv_id":null,"evidence_quote":"the evaluation framework that motivates the dataset expansion"},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"covariate analysis showing the need for diverse data"},{"cited_title":"Jocher, Ayush Chaurasia, A","cited_arxiv_id":null,"evidence_quote":"the whole-body detection model used in the annotation pipeline"},{"cited_title":"Aharon, R","cited_arxiv_id":null,"evidence_quote":"the multi-object tracker used to generate subject tracks"},{"cited_title":"End-to-End Human Pose and Mesh Reconstruction with Transformers","cited_arxiv_id":"2012.09760","evidence_quote":"the transformer-based mesh reconstruction model used for pose"},{"cited_title":"Zhang, X","cited_arxiv_id":null,"evidence_quote":"the distribution-aware coordinate representation for 2D pose estimation"},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"the re-identification model used to associate detections with subjects"}],"review_version":1}