{"id":"86a61e81-6c20-4e6c-8381-43890b80337c","arxiv_id":"2411.19167","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"HOT3D releases 833 minutes of hardware-synchronized, egocentric multi-view video from real headsets with motion-capture ground truth for hands and objects, and shows multi-view baselines outperform single-view baselines on three tracking tasks.","lead":"HOT3D is a public dataset of 833 minutes of egocentric video recorded with Meta's Aria glasses and Quest 3 headset, showing 19 people handling 33 objects with motion-capture ground truth for hands and objects. It gives researchers a benchmark for 3D hand and object tracking from multiple head-mounted cameras, and the paper's baselines show multi-view methods beat single-view ones.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Mocap ground-truth accuracy is never quantified; manual visual inspection alone cannot support the 'high-quality annotations' premise underlying all evaluations.","rationale":"I considered the reader's weakest_assumption and agree that the absence of a quantitative mocap validation is the most load-bearing issue. The central claim has two parts: dataset novelty and demonstrated multi-view advantage. The novelty part is well supported by Table 1: HOT3D is the only dataset with 2-3 egocentric views, hardware sync, and real headsets; ARCTIC uses a helmet mock-up, HO-Cap has only one egocentric view. The multi-view advantage part is supported by experiments. In hand tracking and object pose estimation, the single-/multi-view comparison is controlled (same method, same training), so the relative improvements are credible. The 3D lifting comparison (Table 5) is method-mismatched, but that affects only one of three demonstrations. The weakest link is the annotation quality: the paper states some frames are 'missing or of a lower quality' and filters via visual inspection, yet never quantifies the residual error. Since these poses are the ground truth for both training and evaluation, any uncharacterized error propagates into every reported metric. A concrete independent validation would settle whether this concern lands. If the validation shows errors close to the marker scale (3 mm), the dataset claim is strengthened; if errors are larger, the paper's headline comparison tables may need re-interpretation. I therefore recommend keeping the CONDITIONAL verdict (UNCHANGED), consistent with the reader.","tokens_in":19141,"tokens_out":9240,"duration_ms":79605,"concrete_test":"Independently validate a random sample of, say, 200 frames from the 1.16M manually validated frames. For each frame, manually annotate 2D hand keypoints (or use a high-accuracy multi-view triangulation of hand markers with an independent off-the-shelf tool) and object corners in all synchronized egocentric views; triangulate to 3D and compare with the released OptiTrack-based poses. Report mean/median per-keypoint translation error and object rotation error. If the median hand keypoint error approaches or exceeds the baseline MKPE values (e.g., > 5-10 mm) or the median object rotation error exceeds about 2-3 degrees, the 'high-quality' annotation premise and the numerical comparisons are not supported; if errors are around 1-3 mm, the premise holds.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim that HOT3D provides high-quality ground-truth poses rests on an unquantified OptiTrack pipeline. Sec. 3 states that 1.16M of 1.5M frames passed manual visual inspection, but no numerical accuracy measure is reported. Appendix C describes 3 mm markers, about 19 per hand and 10 per object, semi-automatic registration, and model fitting, but gives no reprojection error, no marker-trajectory residual, and no comparison against an independent ground truth. If marker tracking degrades during fast grasps or when markers are occluded by the hand or object, the remaining 'valid' frames may still contain errors of several millimeters or degrees. Since the same poses are used to train and evaluate the baselines (Tables 2, 3, 5) and to compute in-hand criteria for 2D/3D tasks, any systematic GT error is inherited by every reported number. The novelty claim (first large-scale multi-view egocentric dataset) does not depend on this, but the usefulness claim (high-quality annotations) and the quantitative comparisons do. This is the load-bearing concern because a negative outcome on an independent accuracy check would not invalidate the dataset's existence but would materially weaken its stated purpose.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces HOT3D, a public dataset for egocentric 3D hand and object tracking, containing 833 minutes / 1.5M multi-view frames (3.7M+ images) from Project Aria and Quest 3, with 19 subjects and 33 rigid objects. Ground-truth 6DoF poses of hands (in UmeTrack and MANO formats) and objects are obtained from an optical-marker OptiTrack mocap system; object models come from an in-house scanner with PBR materials. The paper presents three experimental comparisons of multi-view vs single-view baselines: UmeTrack-based 3D hand tracking (Table 2), a multi-view extension of FoundPose for 6DoF object pose estimation (Table 3), and a DINOv2 stereo-matching method for 3D lifting of unknown in-hand objects (Table 5), all reporting that multi-view input substantially improves accuracy over single-view. The dataset also provides validity masks, training/test splits with hidden test annotations, curated clips, and object onboarding sequences.","tokens_in":19360,"tokens_out":9670,"duration_ms":76711,"significance":"If the claims hold, HOT3D fills a clear gap: it is the first large-scale dataset combining multi-view, hardware-time-synchronized egocentric images from real headsets with marker-based mocap ground truth for hands and objects, and it enables benchmarking of multi-view methods on realistic AR/VR hardware. The paper is transparent in several respects: it releases all 1.5M frames, provides validity masks from visual inspection, defines train/test subject splits, uses a public test server for test annotations, and includes object meshes with PBR materials. The proposed baselines, especially the multi-view FoundPose extension and StereoMatch, provide useful starting points. However, the core premise of 'high-quality ground-truth annotations' is not quantitatively validated, and the significance of the multi-view gains is not statistically substantiated; these issues must be addressed before the dataset's value can be fully assessed.","major_comments":[{"comment":"The paper repeatedly characterizes the ground-truth annotations as 'high-quality' (Abstract, Sec. 1, Table 1 caption), but no quantitative accuracy measure of the marker-based mocap pipeline is reported anywhere. The only quality control is manual visual inspection (Sec. 3), which flags gross misalignments but cannot measure sub-centimeter pose errors. Appendix C describes 3 mm optical markers, ~19 markers per hand and ~10 per object, semi-automatic registration, and model fitting, yet gives no reprojection errors, marker-trajectory residuals, or comparison with an independent ground truth. Because the same poses are used to train and evaluate all baselines (Tables 2-5) and to define the in-hand masks and 3D locations (Sec. 4.3-4.4), any systematic GT error is inherited by every reported number. I request a quantitative validation, e.g., per-hand and per-object model-fit residuals, marker reconstruction errors, temporal consistency statistics, and a spot check against manually annotated keypoints or depth-based ICP on a subset of frames, along with the distribution of the visual-inspection rejections.","section":"Sec. 3 and Appendix C"},{"comment":"All experimental results are reported as point estimates without error bars, confidence intervals, or per-sequence/per-subject variation. The abstract and conclusion state that multi-view methods 'significantly outperform' single-view methods; this claim is not statistically supported. The frames in a clip are temporally highly correlated, so the effective sample size is much smaller than the number of frames, and a difference of 8-12 percentage points (Table 3) may not be significant if only a few subjects or sequences drive the effect. Please provide variance estimates (e.g., bootstrap confidence intervals over clips or subjects) for the key comparisons in Tables 2, 3, and 5, and describe how the authors accessed the hidden test annotations (public evaluation server or direct labels), since this affects the reproducibility of the reported numbers.","section":"Tables 2-5, Sec. 4"},{"comment":"The evaluation protocol for the in-hand object tasks relies on threshold definitions stated without justification: an object is 'in-hand' if the minimum distance between object and hand mesh vertices is below 1 cm and the object moves faster than 1 cm/s (Sec. 4.3). The reported mIoU (Table 4) and recall rates (Table 5) are conditional on these thresholds, and the multi-view advantage in Table 5 is computed on this filtered subset. No sensitivity analysis is provided, so it is unclear whether the magnitude of the multi-view gains, or even their sign, would persist under reasonable alternative thresholds (e.g., 0.5 cm or 2 cm, or velocity thresholds of 0.5/2 cm/s). I request a sensitivity analysis for the main results in Table 5, or explicit evidence that the conclusions are robust to these choices.","section":"Sec. 4.3, Tables 4-5"}],"minor_comments":[{"comment":"The title contains 'T racking' with an erroneous space; it should be 'Tracking'.","section":"Title"},{"comment":"The sentence 'manually flagging frames were rendering of hand and object models in the ground-truth poses is not closely aligned with the observed image' is ungrammatical; it should read 'manually flagging frames where the rendering of hand and object models in the ground-truth poses is not closely aligned with the observed image'.","section":"Sec. 3"},{"comment":"The '41% improvement' is a relative reduction in MKPE (13.4/9.5 = 15.4/10.9 = 1.41); please state this explicitly to avoid reader confusion with absolute differences.","section":"Sec. 4.1"},{"comment":"The 'robust mean of the 3D point set' in StereoMatch is undefined; specify the estimator (e.g., trimmed mean, median, or RANSAC-based consensus) and its parameters.","section":"Sec. 4.4"},{"comment":"The code for the proposed multi-view extensions (FoundPose-mv and StereoMatch) is not linked; providing executables or detailed hyperparameters would improve reproducibility of the baseline results.","section":"Sec. 4.2-4.4"},{"comment":"The HandProxy row shows a dash in the 'Views' column; clarify whether this baseline is evaluated in a single-view or multi-view manner.","section":"Table 5"},{"comment":"The phrasing 'over 3.7M+ images' in the abstract and '3.7M+ images' in Sec. 3 is inconsistent; unify the notation.","section":"Abstract and Sec. 3"}],"recommendation":"major_revision","confidential_remarks":"The dataset has clear potential and the multi-view experiments are valuable, but the missing quantitative validation of the mocap ground truth is the most important risk to the dataset's credibility in the community. If the authors can supply a thorough accuracy analysis (e.g., in the supplement) and add error bars to the main comparisons, I would support acceptance. I also note that the authors' affiliation with Meta and their involvement in several of the baseline methods may create an impression of bias, though the use of a held-out test server partially mitigates this; the paper should state precisely how the test labels were accessed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know upfront. The dataset is real and fills a gap no one else has filled: hardware-synchronized multi-view egocentric video from actual shipped/research headsets (Aria and Quest 3), with mocap ground truth for hands and objects, at a scale of 1.5M multi-view frames. The paper is also honest about its own limits: it releases all frames, flags the 1.16M that passed visual inspection, and provides a public test server with hidden annotations. That is a solid, usable resource.\n\nWhat is actually new: the multi-view egocentric capture with real devices, the 33 PBR-scanned objects, the UmeTrack and MANO hand formats, and the curated clips and onboarding sequences. The multi-view baselines are straightforward extensions of existing work (FoundPose with generalized PnP, DINOv2 stereo matching), and that is fine — the paper does not oversell them as novel methods. The experiments show large multi-view gains across all three tasks, and the numbers are reported on held-out subjects with a train/test split.\n\nWhere it gets soft. Most importantly, the mocap ground-truth accuracy is never quantified. Appendix C describes 3 mm markers, semi-automatic registration, and model fitting, but there is no reprojection error, no residual, no comparison against an independent reference. The paper's manual visual inspection filters out badly aligned frames, but that does not tell you the residual error in frames that pass. If marker tracking drifts during fast grasps, the 'valid' frames could still carry several-millimeter errors, and every baseline number inherits that. The stress-test note is right that this is the load-bearing weakness, although I would call it a serious limitation rather than a fatal flaw — the dataset's novelty does not depend on mocap perfection, and most hand-object benchmarks have similar or worse unquantified annotation noise. Still, for a paper claiming 'high-quality' annotations, a few numbers would go a long way.\n\nAlso, Tables 2–5 report point estimates with no error bars or significance tests. The gaps are large enough that the qualitative conclusion is probably robust, but 'significantly outperform' is doing more work than the statistics support. Minor: the in-hand object thresholds (1 cm distance, 1 cm/s velocity) are reasonable but arbitrary; the paper should acknowledge the sensitivity.\n\nBottom line: this is a benchmark paper that will be used. The dataset release, the transparent splits, and the public evaluation server make it a credible resource. A serious referee should engage with it, and the revision should quantify the mocap error and add variance estimates. I'd bring it to reading group and would cite it if I worked in egocentric tracking.","headline":"A genuinely novel, large-scale egocentric multi-view hand-object dataset with mocap ground truth; the experiments are suggestive but the annotation accuracy is under-quantified.","tokens_in":19944,"tokens_out":2003,"would_cite":true,"duration_ms":18178,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"HOT3D introduces a public egocentric multi-view dataset with motion-capture ground truth for hands and objects, and shows that multi-view methods beat single-view baselines on hand tracking, object pose, and 3D lifting.","keywords":["egocentric vision","hand tracking","6DoF object pose","multi-view","dataset","hand-object interaction","3D lifting","motion capture"],"falsifier":"Take a held-out subset of the released training frames, re-run the marker-to-model fitting with an independent procedure, and compare the result with the published ground truth; if the disagreement approaches the 5 cm or 5 degree thresholds used in the recall metrics on a substantial fraction of frames, the supervision quality and the resulting multi-view gains would need to be re-derived. A second check is to train the same single-view baselines on HOT3D training splits and submit to the official test server: if a single-view model matches the reported multi-view recall, the claimed multi-view advantage collapses.","tokens_in":18947,"feed_emoji":"🖐️","tokens_out":7087,"duration_ms":59595,"temperature":0.7,"pith_summary":"HOT3D is a large-scale public dataset for egocentric 3D hand and object tracking, built from hardware-synchronized multi-view video recorded with two real head-mounted devices. The paper's central claim is that this combination of real headsets, multiple simultaneous views, and motion-capture ground truth is new at this scale, and that it matters because multi-view methods for hand tracking, 6DoF object pose estimation, and 3D lifting of unknown in-hand objects clearly outperform single-view baselines. If true, the dataset supplies what AR/VR and contextual-AI systems need: a way to train and benchmark perception that exploits the multiple cameras already on headsets, without relying on power-hungry depth sensors.","feed_headline":"Multi-view beats single-view on egocentric 3D tracking","feed_subtitle":"New public dataset with mocap ground truth from real headsets backs multi-view hand and object tracking","key_machinery":"The load-bearing object is the dataset's capture protocol: two head-mounted devices record several cameras that are triggered by a hardware timecode to produce synchronized views of the same hand-object scene, while a rig of infrared optical-marker cameras tracks small reflective markers glued to the hands and objects. Marker trajectories are fit to scanned 3D models of each hand and object to produce per-frame ground-truth poses. The claim-carrying mechanism is the controlled comparison: the paper deliberately runs single-view and multi-view versions of the same method on the same frames, so the only changed variable is the number of synchronized views. The multi-view baselines are simple, involving two-view training with view masking for the hand tracker, generalized PnP over correspondences from all views for pose estimation, and row-wise nearest-neighbor matching of self-supervised features across a stereo pair for 3D lifting.","core_discovery":"The discovery is the dataset itself and the empirical pattern it exposes. HOT3D offers over 833 minutes of egocentric multi-view video, more than 3.7 million images from 19 subjects interacting with 33 rigid objects, with per-frame ground-truth poses of both hands and objects obtained by an optical-marker motion-capture system, plus 3D object meshes, hand models in two formats, and, for one device, SLAM point clouds and eye gaze. The paper runs three controlled comparisons: a hand tracker trained and tested in single-view versus two-view mode, a pose estimator extended from a single RGB image to multiple synchronized views, and a stereo-match method for finding the 3D location of a handheld object. In all three, the multi-view variant outperforms its single-view counterpart, with lower mean keypoint error, 8 to 12 percentage points higher object-pose recall, and a jump in 5 cm 3D-lifting recall from 14.3 percent to 64.4 percent when ground-truth masks are used.","pith_inferences":["The same multi-view advantage likely transfers to other egocentric tasks the paper does not evaluate, such as hand-object contact estimation or action recognition, because extra views reduce occlusion and provide metric depth cues.","Because the glasses-prototype recordings include eye gaze, one untested extension is gaze-conditioned object lifting, using the gaze ray as an additional correspondence prior to sharpen localization in the coarse 10 cm regime.","The roughly 23 percent of frames released without valid annotations could be exploited by self-supervised multi-view consistency losses; measuring how much of the reported gap closes with unlabeled frames would isolate the value of the labeled ground truth.","A direct comparison of the marker-based ground truth against an independent RGB-D optimization on the same scenes would put a number on the annotation quality, which the paper leaves implicit."],"forward_implications":["HOT3D gives the community a public test bed where multi-view egocentric methods can be trained and compared on real-device data, not just synthetic or exocentric captures.","The two-view consumer-headset configuration being sufficient for large gains means hardware that already ships to consumers can support improved tracking.","The curated clips and the validity mask let researchers use the unannotated frames for self-supervised pretraining while benchmarking on the 1.16 million validated frames.","The multi-view baseline results define a realistic first benchmark for the three tasks, so subsequent methods can be measured against numbers that already show the single-view ceiling."],"supporting_citations":[{"why":"Supplies the glasses-prototype device, its SLAM point clouds, and the eye gaze channel used in the recordings.","marker":"[13]"},{"why":"Supplies the consumer VR headset and its two synchronized monochrome image streams.","marker":"[41]"},{"why":"Supplies the hand tracker used in the single-view versus two-view comparison and the training split combined with HOT3D.","marker":"[26]"},{"why":"Supplies the single-view pose estimator that the paper extends to multiple views.","marker":"[50]"},{"why":"Supplies the visual features used both for pose-template matching and for stereo matching in 3D lifting.","marker":"[49]"},{"why":"Provides the closest prior mocap-annotated hand-object dataset, which HOT3D positions itself against in terms of novelty.","marker":"[15]"}],"fun_headline_variants":["Egocentric dataset shows multi-view 3D tracking edge","HOT3D: multi-view egocentric beat single-view tracking","New egocentric dataset: multi-view tracking wins","3.7M egocentric images boost multi-view tracking"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the optical-marker motion-capture system yields ground-truth poses accurate enough to train and evaluate on; the paper filters out about a quarter of frames as missing or low quality, but reports no quantitative error of the mocap pipeline.","fun_headline_variants_meta":{"raw":{"variants":["Egocentric dataset shows multi-view 3D tracking edge","HOT3D: multi-view egocentric beat single-view tracking","New egocentric dataset: multi-view tracking wins","3.7M egocentric images boost multi-view tracking"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000323,"raw_usage":{"total_tokens":1855,"prompt_tokens":1024,"completion_tokens":831,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":640,"completion_tokens_details":{"reasoning_tokens":764}},"tokens_in":640,"tokens_out":831,"duration_ms":8271,"temperature":1.0,"reasoning_tokens":764,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T10:27:52.799587+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a held-out subset of the released training frames, re-run the marker-to-model fitting with an independent procedure, and compare the result with the published ground truth; if the disagreement approaches the 5 cm or 5 degree thresholds used in the recall metrics on a substantial fraction of frames, the supervision quality and the resulting multi-view gains would need to be re-derived. A second check is to train the same single-view baselines on HOT3D training splits and submit to the official test server: if a single-view model matches the reported multi-view recall, the claimed multi-view advantage collapses.","supporting_citations":[{"cited_title":"Project Aria: A new tool for egocentric multi-modal AI research, 2023","cited_arxiv_id":null,"evidence_quote":"Supplies the glasses-prototype device, its SLAM point clouds, and the eye gaze channel used in the recordings."},{"cited_title":"Quest 3.https://www.meta.com/quest/quest- 3/, 2023","cited_arxiv_id":null,"evidence_quote":"Supplies the consumer VR headset and its two synchronized monochrome image streams."},{"cited_title":"UmeTrack: Unified multi-view end-to-end hand tracking for VR","cited_arxiv_id":null,"evidence_quote":"Supplies the hand tracker used in the single-view versus two-view comparison and the training split combined with HOT3D."},{"cited_title":"FoundPose: Unseen ob- ject pose estimation with foundation features.ECCV, 2024","cited_arxiv_id":null,"evidence_quote":"Supplies the single-view pose estimator that the paper extends to multiple views."},{"cited_title":"DINOv2: Learning robust visual features without supervision.TMLR, 2024","cited_arxiv_id":null,"evidence_quote":"Supplies the visual features used both for pose-template matching and for stereo matching in 3D lifting."},{"cited_title":"ARCTIC: A dataset for dexterous bimanual hand-object manipulation","cited_arxiv_id":null,"evidence_quote":"Provides the closest prior mocap-annotated hand-object dataset, which HOT3D positions itself against in terms of novelty."}],"review_version":1}