{"id":"fab97958-cdb3-4498-880e-6c1e4e800c17","arxiv_id":"2504.17696","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"DARai provides over 200 hours of multi-sensor, three-level annotated daily activity recordings, with benchmarks showing how accuracy varies across sensors, hierarchy levels, and camera or body placement.","lead":"DARai is a new open dataset of over 200 hours of daily-life recordings from 50 people in 10 rooms, using 20 sensor modalities and three levels of activity labels. It lets researchers test how well AI systems recognize, localize, and anticipate everyday actions, and which sensors matter when cameras are limited or absent.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Appendix C.2 states that the full hierarchical labels are only planned for release, which conflicts with the central claim that DARai is publicly available with L1/L2/L3 annotations and reproducible benchmarks.","rationale":"The reader's weakest assumption was host-clock synchronization, citing the 1 ms versus 0.1 s/day drift discrepancy. While that contradiction is real, it is unlikely to be load-bearing: even the larger stated bound of under 0.1 s/day corresponds to under 0.1 s over a 30–60 minute session, which is far smaller than typical L3 procedure durations. The synchronization issue therefore does not threaten the central claim in the way the reader suggested. A more consequential concern is the explicit statement in Appendix C.2 that the full hierarchical labels are only planned for release. This bears directly on the strongest claim that DARai is publicly available with hierarchical annotations and reproducible benchmarks. I am not claiming the dataset is withheld or that the authors are being deceptive; the paper gives a DOI, a project page, and a GitHub repository, and a download check could fully resolve the issue. But as written, the text itself contains a future-tense release statement that contradicts the present-tense availability claims elsewhere. That internal inconsistency is more load-bearing than the synchronization numbers because it affects whether the dataset and its benchmark results can be independently checked at all. The reader's verdict of CONDITIONAL remains appropriate: the paper should be published only after the authors clarify the exact contents of the current release and either provide full hierarchical labels or explicitly state which portions are pending. My recommendation is therefore UNCHANGED rather than a shift to a harsher verdict, because this is a verifiable and fixable issue rather than a demonstrated fatal flaw.","tokens_in":35866,"tokens_out":4959,"duration_ms":54519,"concrete_test":"Download the IEEE Dataport package (DOI 10.21227/ecnr-hy49) and inspect the manifest and README. Verify for a random sample of subjects and sessions (e.g., subject 01, session 03) that L1, L2, and L3 annotation files exist for all claimed modalities, that the provided GitHub loaders can instantiate the cross-subject train/test split, and that the L3 label set contains 98 distinct labels. If any hierarchical level or modality is missing from the release, the central claim is not supported as written.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that DARai is an open, publicly available dataset with 18 L1, 44 L2, and 98 L3 hierarchical labels, released with code, loaders, and cross-subject splits. However, Appendix C.2 ends with: 'We plan to release the losslessly compressed version of the visual data along with the full hierarchical labels.' This sentence indicates that, at the time of writing, the full hierarchical labels were not yet part of the released package. If the downloadable IEEE Dataport artifact lacks complete L3 labels, or if the release is limited to a lossy compressed subset, then the benchmark results in Sections 6 and 7, the counterfactual experiments in Section 7.3, and the claimed 'largest available dataset' status cannot be verified or reproduced. The paper elsewhere asserts that 'the dataset and codes are publicly available,' but Appendix C.2 introduces a direct internal inconsistency about what is actually released. This is a load-bearing concern because the dataset's existence and completeness is the premise for every empirical claim in the paper. Related inconsistencies in sensor counts (12 vs. 16 vs. 20 devices) and shared-procedure percentages (14.2% vs. 20%) compound the uncertainty, but the unresolved release status is the more fundamental issue.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces DARai, a multimodal and hierarchically annotated dataset of daily human activities, and presents a set of machine-learning benchmarks on it. The dataset consists of over 200 hours of continuous, partly unscripted recordings from 50 participants in 10 environments, with 18 L1 activities, 44 L2 actions, and 98 L3 procedures, collected from cameras, depth/radar sensors, wearable IMUs, EMG, insole pressure, biomonitors, and gaze tracking. The authors report unimodal and multimodal recognition results across the three hierarchy levels, cross-view and cross-body robustness experiments, temporal localization, and short- and long-horizon action anticipation, and they release code, loaders, splits, and a dataset DOI via IEEE Dataport. The central claim is that DARai is the largest publicly available dataset in terms of the number of sensors and modalities and that its hierarchical, unscripted design enables studying action-counterfactual variations and fine-grained temporal dependencies. The manuscript is accompanied by extensive appendices on collection hardware, annotation workflow, benchmark details, and failure cases, including an honest account that some multimodal time-series classification attempts were unsuccessful.","tokens_in":36044,"tokens_out":4828,"duration_ms":45556,"significance":"If the release-completeness and synchronization issues are resolved, DARai would be a valuable community resource. Its strengths are concrete: the dataset, code, loaders, and cross-subject splits are public, the hierarchy is large (160 labels across three levels), the recording setup covers an unusually wide set of modalities under one protocol, and the benchmark suite spans recognition, localization, anticipation, and cross-domain robustness. The paper also makes falsifiable empirical observations, such as the sharp visual-model degradation from L1 to L3 and the smaller degradation of wearable modalities, and it compares against external datasets such as Breakfast and 50 Salads. The explicit statement that action counterfactuals are descriptive rather than causal is a useful clarification. However, the paper's central promise of a fully released, hierarchically annotated dataset is currently weakened by internal inconsistencies about what is actually available and about the synchronization accuracy that underpins the cross-modal and temporal benchmarks.","major_comments":[{"comment":"The release-status statements are internally contradictory. Appendix C.2 ends with 'We plan to release the losslessly compressed version of the visual data along with the full hierarchical labels,' which implies that the full L1/L2/L3 annotations are not yet part of the public artifact, whereas the Abstract, Section 3.3, and Appendix A.3 state that the dataset with hierarchical labels and code is publicly available. Because every benchmark in Sections 5–7 and the counterfactual analysis in Section 7.3 presuppose the availability of frame-level L3 labels, the paper must either release those labels in the IEEE Dataport artifact or explicitly restrict all empirical claims to the currently released subset.","section":"Appendix C.2; Appendix A.3; Abstract"},{"comment":"The synchronization figures are inconsistent by two orders of magnitude: Section 1 claims 'a synchronization drift of 1 ms in 24-hour timespan,' while Appendix D states that the NTP-aligned host clocks are 'maintained under 0.1 seconds per day.' Since the cross-modal and temporal benchmarks in Sections 6 and 7 rely on frame-level alignment of all streams with L2/L3 boundaries, the authors should report the actual measured drift, the measurement method, and the worst-case alignment error relative to the shortest L3 procedure duration.","section":"Section 1 (Challenges); Appendix D"},{"comment":"The reported sensor count is inconsistent: the Abstract says 20 sensors, Contribution 1 says 20 data modalities from 12 sensors, Section 3.1 says 12 sensor devices, and Appendix D says 16 sensor devices were used while D.1 says 12 distinct sensor types were selected. The 'largest available dataset in terms of the number of sensors and data modalities' claim depends on these counts, so the authors need to disambiguate device units, sensor types, and derived modalities and use one consistent set of numbers.","section":"Abstract; Section 1; Section 3.1; Appendix D.1"},{"comment":"No inter-annotator agreement or boundary-consistency metric is reported for the L1/L2/L3 frame-level annotations. Given that the dataset's central contribution is a three-level hierarchy with fine-grained L3 boundaries, the authors should report agreement statistics (e.g., kappa or segmental F1) and the number of annotators per clip to support the reliability of the annotation-based benchmarks.","section":"Appendix E (Annotation Workflow)"}],"minor_comments":[{"comment":"The text says '20% of the procedures (L3) are shared between higher-level activities (L1) and actions (L2),' which conflicts with the Abstract's 14.2% of L3 procedures shared between L2 actions; clarify the denominator and definition of sharing.","section":"Section 6.1"},{"comment":"This section states that multimodal time-series classification 'was not successful' on L1 classes, but Section 6.2 and Table 8 report L1 accuracies for EMG + Bio and other sensor groups; the relationship between the failed experiment and the reported results should be reconciled.","section":"Appendix G.1.2"},{"comment":"Table 20 reports results on a 50% subset of L2 and L3 labels, while Table 6 reports all labels; clarify which label subset is used in each benchmark and why.","section":"Table 20 caption"},{"comment":"The sentence 'multi-sensor and and its inclusion' contains a duplicated 'and' and should be corrected.","section":"Abstract"},{"comment":"The example in Figure 2 says the L2 action 'Rest' contains no L3 procedures, but Table 4 lists 98 L3 labels; state how non-procedural activities are represented at L3 (e.g., no label versus an explicit 'other' category).","section":"Section 3.2; Figure 2"}],"recommendation":"major_revision","confidential_remarks":"The core contribution is potentially strong, but the paper currently makes release and completeness claims that Appendix C.2 contradicts. I would ask the editor to verify, before any acceptance, that the IEEE Dataport artifact actually contains the full L1/L2/L3 labels and that the synchronization protocol is described with a single, measured accuracy figure."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The DARai paper is worth a serious look, but the first thing to check is what is actually downloadable. Appendix C.2 says the authors \"plan to release\" the losslessly compressed visual data along with the full hierarchical labels, which contradicts the main text's assertion that the dataset and labels are publicly available on IEEE Dataport. That is a load-bearing ambiguity: if the released artifact is missing L3 labels or only offers lossy compressed video, the hierarchy experiments in Sections 6 and 7 cannot be reproduced, and the \"largest available dataset\" claim is unverifiable.\n\nWhat is genuinely new: a 50-participant, 10-environment, 200-hour multimodal daily activity dataset with 20 modalities, a three-level hierarchy defined before collection (18 L1 activities, 44 L2 actions, 98 L3 procedures), continuous unscripted recordings, and built-in action-counterfactual instances. That combination is new in the surveyed literature. The authors also ship code, loaders, and cross-subject splits, and the benchmark study is honest: it reports visual degradation at fine granularity and under cross-view shift, shows wearables degrade less, and openly admits failures (Appendix G.1.2 says multimodal time-series classification did not work at L1). The anticipation comparison against Breakfast and 50 Salads is a useful reference point.\n\nSoft spots are real but mostly fixable. The synchronization drift is stated as 1 ms/day in the introduction and under 0.1 s/day in Appendix D—a two-order-of-magnitude gap that directly bears on cross-modal fusion and anticipation. The sensor count appears as 12, 16, and 20 in different places; probably 12 sensor types, 16 physical devices, 20 modalities, but it needs to be stated consistently. There is no inter-annotator agreement metric, and some headline results (Figure 7) lack error bars while other tables have them. The shared-procedure percentage also appears as both 14.2% and 20% in different sections; that may be two different definitions, but the paper does not say so.\n\nNone of these issues invalidate the central dataset claim on the face of it, and they are all addressable with a careful revision. This paper deserves peer review, not desk rejection. The authors should be required to reconcile the release status, synchronization numbers, sensor counts, and overlap statistics, and to add annotation reliability metrics. If they do, this becomes a solid benchmark reference for the community. I would bring it to a reading group and would cite it once the release is confirmed.","headline":"A genuinely useful multimodal daily-activity dataset—if the release-status contradiction in the appendix is fixed, it deserves to be a standard benchmark.","tokens_in":36656,"tokens_out":3156,"would_cite":true,"duration_ms":29796,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new public dataset records 200 hours of daily activity from 20 synchronized sensors with three-level hierarchical labels, and its benchmarks show that wearable sensors retain accuracy where vision models collapse.","keywords":["multimodal dataset","hierarchical activity recognition","sensor fusion","action counterfactual","temporal action localization","action anticipation","cross-view robustness","wearable sensors"],"falsifier":"Measure the actual inter-stream timestamp disagreement on a released sample: if the median absolute difference between host clocks and sensor internal clocks at procedure boundaries exceeds the median duration of L3 labels, the hierarchy claims cannot be validated at their stated granularity.","tokens_in":35602,"feed_emoji":"🎯","tokens_out":5549,"duration_ms":46880,"temperature":0.7,"pith_summary":"The paper introduces DARai, a public, open-source dataset of continuous, unscripted daily-activity recordings from 50 people in 10 home environments, captured by 20 sensor streams. The dataset's load-bearing design is its three-level annotation hierarchy—activities, actions, and procedures—together with 'action counterfactuals,' alternative executions of the same activity that let researchers test how sensors and models handle natural variation. The paper argues that if this resource is used, the community gets a testbed where multimodal fusion, hierarchical recognition, temporal localization, and anticipation can be studied on realistic long-range data. Its benchmarks suggest that body-worn sensors such as insole pressure are more robust to task granularity and viewpoint changes than cameras, a finding that matters for privacy-preserving activity understanding.","feed_headline":"Wearables beat video on fine daily tasks in 20-sensor dataset","feed_subtitle":"DARai logs 200 hours of unscripted activity from 50 people; benchmarks show body sensors hold up as cameras drop.","key_machinery":"The central object is the DARai dataset itself: a synchronized, multimodal corpus in which 20 sensor streams (RGB, depth, radar, IMUs, EMG, insole pressure, biomonitors, gaze, microphones, environmental sensors) are time-aligned to frame-level annotation boundaries at three nested levels—L1 activities (goal-level tasks), L2 actions (shared sub-steps), and L3 procedures (exact execution steps). The design includes predefined hierarchical decomposition rules and action-counterfactual instances, where the same activity is performed under different conditions, so that sensor-specific variation can be studied. The benchmark suite built on top—transformer, convolutional, and recurrent models for recognition, localization, and anticipation—is the demonstration mechanism that turns the data into evidence about modality robustness.","core_discovery":"The central claim is that DARai is the largest available daily-activity dataset in number of sensors and modalities, and that its design exposes systematic limitations of vision-centric activity understanding. With 200+ hours from 50 participants, 10 environments, and 20 modalities, annotated at three levels (18 activities, 44 actions, 98 procedures) with shared lower-level labels across higher-level classes, the paper demonstrates that visual models drop sharply from L1 to L2/L3 accuracy while wearable modalities such as insole pressure and EMG degrade less; that cross-view camera testing collapses accuracy (from 89% to 15% in one case) while cross-body wearable transfer degrades less; and that fusion of wearable signals sustains accuracy at fine granularity (over 42% at L3). The paper's benchmarks also show that action anticipation at the procedure level benefits from longer observation, while action-level anticipation plateaus. All of these are claims about what the dataset enables and what it reveals.","pith_inferences":["If the synchronization and annotation boundaries hold at the stated precision, DARai could become a standard evaluation for sensor-fusion and hierarchical-learning research, and the cross-view/cross-body gaps imply that deployment of vision models in homes needs either multi-view setups or modality complementarity.","A testable extension of the paper's findings would be to check whether the L3 robustness of wearables generalizes to activities outside the ten environments, since the environments are all kitchens, living spaces, and offices.","The counterfactual design could support causal-style analyses of what sensor signals actually drive a prediction, moving from accuracy comparisons to attribution comparisons."],"forward_implications":["Non-visual modalities (insole, EMG, IMU, gaze) can sustain fine-grained recognition where RGB and depth models collapse, suggesting privacy-preserving activity understanding without cameras is feasible.","Long-horizon procedural anticipation improves with more observation while action-level anticipation plateaus, meaning models should treat hierarchy levels differently.","The counterfactual design provides paired samples that expose contextual bias, such as models misclassifying counterfactual reading-on-couch actions as phone conversations due to background.","The dataset's fluid, unscripted boundaries make temporal localization harder than in scripted datasets like Breakfast or 50 Salads, providing a stricter test for segmentation methods.","The fixed cross-subject split and released loaders let researchers compare methods on a common benchmark."],"supporting_citations":[{"why":"Supplies the Epic-Kitchen dataset, the benchmark contrast for action anticipation and egocentric daily activities.","marker":"(Damen et al., 2018)"},{"why":"Breakfast Actions, used for anticipation and segmentation comparisons.","marker":"(Kuehne et al., 2014b)"},{"why":"50 Salads, used for anticipation comparisons and hierarchical dataset positioning.","marker":"(Stein and McKenna, 2013)"},{"why":"Ego4D, the large-scale unscripted egocentric baseline that DARai distinguishes itself from in scale and modality count.","marker":"(Grauman et al., 2022)"},{"why":"FUTR, the model used for temporal localization and anticipation benchmarks.","marker":"(Gong et al., 2022)"},{"why":"C2F-TCN, the segmentation model used for L2/L3 localization.","marker":"(Singhania et al., 2021)"},{"why":"Video Swin Transformer, a visual backbone for recognition benchmarks.","marker":"(Liu et al., 2022)"},{"why":"DOI release of DARai itself, the artifact that makes the claims checkable.","marker":"(Kaviani et al., 2024)"}],"fun_headline_variants":["20-sensor DARai exposes camera blind spots on fine actions","Body sensors outlast video in 200-hour daily activity dataset","Wearables win on fine tasks in 20-modality DARai benchmark","DARai: 200 hours, 20 sensors reveal vision's fine-action drop","Multimodal DARai shows wearables beat cameras on detail"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"All conclusions hinge on the host-clock synchronization being accurate enough that every sensor stream's timestamps align with the frame-level L1/L2/L3 boundaries; if alignment error exceeds the duration of short L3 procedures, the hierarchical and temporal benchmarks are built on misaligned inputs.","fun_headline_variants_meta":{"raw":{"variants":["20-sensor DARai exposes camera blind spots on fine actions","Body sensors outlast video in 200-hour daily activity dataset","Wearables win on fine tasks in 20-modality DARai benchmark","DARai: 200 hours, 20 sensors reveal vision's fine-action drop","Multimodal DARai shows wearables beat cameras on detail"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000259,"raw_usage":{"total_tokens":1662,"prompt_tokens":1097,"completion_tokens":565,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":713,"completion_tokens_details":{"reasoning_tokens":469}},"tokens_in":713,"tokens_out":565,"duration_ms":5447,"temperature":1.0,"reasoning_tokens":469,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:33:59.903224+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the actual inter-stream timestamp disagreement on a released sample: if the median absolute difference between host clocks and sensor internal clocks at procedure boundaries exceeds the median duration of L3 labels, the hierarchy claims cannot be validated at their stated granularity.","supporting_citations":[{"cited_title":"Combining embedded accelerometers with com- puter vision for recognizing food preparation activities","cited_arxiv_id":null,"evidence_quote":"50 Salads, used for anticipation comparisons and hierarchical dataset positioning."}],"review_version":1}