{"id":"223363c3-d3f8-4bc4-931d-e9b9e24db231","arxiv_id":"2502.03347","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"DiversityOne is a new multi-country smartphone sensor and self-report dataset, spanning 782 college students in eight countries, intended for cross-cultural behavior modeling.","lead":"This paper introduces and documents DiversityOne, a smartphone sensing dataset collected from 782 college students across eight countries over four weeks. It combines 26 sensor streams with more than 350,000 self-reports and is intended to support cross-country behavior modeling and generalization research.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central 'publicly available' claim is contradicted by the paper's own access-control section: the dataset is gated behind approval, institutional affiliation, and a no-redistribution license.","rationale":"I read the paper as a dataset-description contribution whose central claim is that DiversityOne is a publicly released, reusable corpus. The reader's verdict conditioned acceptance on correcting the 'publicly available' description and reconciling 782 vs 666; my stress-test identifies the same availability contradiction as the single most load-bearing issue, because it undercuts the main contribution and the Table 1 comparison. The paper's own Section 3.7.2 explicitly states the dataset is not publicly available online and access is conditional, and Sections 5.1 through 5.3 describe an approval-based, no-redistribution licensing regime. This is an internal inconsistency, not a disagreement with outside consensus, so it is directly decisive. The self-report comparability concern raised by the reader is real but secondary: the paper acknowledges silver-standard labels and cross-site adaptation, and the dataset can still be used for within-site or sensitivity analyses. By contrast, if the access control is as restrictive as described, researchers cannot use the dataset at all without permission, and the phrase 'publicly available' misrepresents the asset. The 782/666 count is a further internal inconsistency that should be fixed in the same revision. Because the reader already conditioned on these points, my assessment leaves the verdict at CONDITIONAL; I would not reject the paper, since the dataset and documentation appear substantial and the access procedures may be an honest effort to comply with GDPR, but the claims must be reworded and the counts reconciled.","tokens_in":40694,"tokens_out":3060,"duration_ms":30028,"concrete_test":"Independently execute the dataset request process described in Section 5.3 from a non-consortium institution: locate a bundle identifier in the catalog (livepeople.disi.unitn.it), submit the request form with a research proposal, and record whether access is granted without consortium affiliation. Also inspect the Terms and License Agreement for redistribution and public-sharing clauses. If approval is required or redistribution is prohibited, the dataset is not 'publicly available' in the sense used in the abstract and Table 1; the paper should either qualify the claim to 'available on request' or make the data openly downloadable. In parallel, verify whether the catalog links resolve and the request form is active, since an inactive or inaccessible process would make even 'available on request' unverifiable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's Contribution 1 and abstract state DiversityOne is 'publicly released' and 'one of the largest and most diverse publicly available datasets.' However, Section 3.7.2 states verbatim: 'the dataset is not publicly available online and can only be accessed under specific conditions (see Section 5).' Section 5.3 describes a request process requiring research affiliation, a submitted proposal, approval, and a Terms and License Agreement that restricts use to research, forbids redistribution and public sharing, and prohibits re-identification attempts. This is not a minor wording issue: the comparative claim in Table 1 is about 'Public available datasets,' and the central contribution is that the community can reuse the corpus. A gated dataset with individual approval and no redistribution is materially different from publicly released data. Independent researchers cannot verify the data or build on it without the consortium's permission, and derived benchmark datasets are not currently released (Section 5.1). A secondary internal inconsistency compounds this: Table 4 reports 782 'iLog signed' but only 666 with 'iLog data,' while the abstract and Contribution 1 say data from 782 participants. The 'publicly available' claim is the load-bearing assertion; if it fails, the headline contribution narrows to a described but access-restricted collection.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces DiversityOne, a multi-country smartphone sensing dataset collected from college students in eight countries over four weeks, combining passive sensor data from 26 modalities with intensive longitudinal self-reports and questionnaires from over 18,000 respondents. It describes the study design, adaptation procedures, privacy and anonymization pipeline, data catalog, and lessons learned from previous analyses using the dataset. The stated contributions are the public release of this dataset, a detailed description of the cross-country data collection methodology, and a synthesis of findings and recommendations for future smartphone sensing studies.","tokens_in":40942,"tokens_out":2420,"duration_ms":26189,"significance":"If the availability and scale claims hold, DiversityOne would be a valuable community resource for studying cross-country generalization and domain adaptation in mobile sensing, areas that have been limited by the lack of diverse, raw sensor datasets. The paper's strengths include the unusually wide geographic coverage spanning Global North and South, the documentation of local adaptations across sites, the inclusion of raw sensor data rather than only preprocessed features, and the explicit discussion of ethical and privacy procedures. Several prior publications already use subsets of the data, which provides some evidence of usability. However, the central claim that the dataset is publicly available is directly contradicted by the access-control description, and the participant counts are inconsistent between the abstract and the detailed tables; these issues must be resolved before the contribution can be assessed accurately.","major_comments":[{"comment":"The paper's headline claim that DiversityOne is 'publicly released' and 'publicly available' is inconsistent with the access procedure described in Section 5.3 and the explicit statement in Section 3.7.2 that 'the dataset is not publicly available online and can only be accessed under specific conditions.' The request process requires affiliation with a research institution, submission and approval of a research proposal, signing of a Terms and License Agreement that prohibits redistribution and public sharing, and a ban on re-identification attempts. This is controlled access, not public release. The comparison in Table 1 is specifically about 'Public available datasets,' so the characterization materially affects the paper's central novelty claim. Please either revise the abstract and Contribution 1 to describe the dataset as available under controlled conditions, or provide a genuinely open release (including derived benchmark datasets, which Section 5.1 currently defers indefinitely).","section":"Abstract, Contribution 1, Table 1, Section 3.7.2, Section 5.3"},{"comment":"The abstract and Contribution 1 state that the dataset contains 'data from 782 college students' and 'passive smartphone sensor data and self-reports from 782 participants,' but Table 4 reports 782 students who signed into the iLog app and only 666 who actively contributed data. The numbers 666 and 782 are used in different places, and the text should be consistent about which count is being reported. Please correct the abstract and Contribution 1 to '782 signed participants' or '666 participants with iLog data,' and add a sentence in Section 4.1 explaining the difference between signing in and providing at least one data record.","section":"Abstract, Contribution 1, Table 4"},{"comment":"The cross-country comparability of the self-report labels is a load-bearing assumption for the dataset's stated purpose of enabling cross-country behavior modeling and generalization. Section 3.1.1 and Section 3.2.3 describe substantial local adaptations: questionnaire items were modified or removed (e.g., sexuality and religiosity items), response options were changed (e.g., nationality lists, foods, apps), and self-report prompt frequencies differed (IPICYT used half-hour notifications throughout). Section 6.2.5 further acknowledges that self-reported labels are only a 'silver standard.' The paper should provide a per-site summary of which items and response options were adapted, and ideally report basic measurement invariance or per-country response-distribution statistics, so that users can judge whether country differences reflect behavior or instrument differences. Without this, the cross-country generalization analyses that the dataset advertises rest on an unverified comparability assumption.","section":"Section 3.1.1, Section 3.2.3, Section 6.2.5"}],"minor_comments":[{"comment":"There is a typo in the sentence 'Relevant sens r file columns were anonymized'; it should read 'sensor file columns.'","section":"Section 3.7.2"},{"comment":"The phrase 'after transer learning' contains a typo and should read 'after transfer learning.'","section":"Section 6.2.5"},{"comment":"The sentence 'This second will be released soon, hopefully within 2024' is awkward and outdated for a 2025 publication; it should be removed or updated with the actual release status.","section":"Section 6.3"},{"comment":"The text mentions a comparison with data collected 'in 1918' in a study by [40]; this appears to be a typo for 2018 and should be corrected.","section":"Section 6.3"},{"comment":"The incentive table uses dashes for AMRITA and IPICYT; a footnote should clarify whether no incentives were offered or whether the information is unavailable.","section":"Table 3"}],"recommendation":"major_revision","confidential_remarks":"The dataset access model is a substantive issue that the editorial team should weigh carefully. Many sensitive health-related datasets are legitimately released under controlled access, and that can still be a valuable contribution. The problem here is that the paper repeatedly calls the dataset 'publicly available' when the described model is approval-based and no-redistribution. This is fixable by rewording, but the current wording overstates the contribution and misleads readers comparing Table 1. The participant-count inconsistency is also easily fixed but should be caught before publication. I would not recommend rejection, because the underlying dataset and protocol documentation appear real and potentially useful; the issues are in the claims and presentation rather than in the data collection itself."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The DiversityOne paper is worth a serious look. It delivers what the field lacks: raw smartphone sensor data from eight countries, including several in the Global South, with 26 modalities and over 350k self-reports. The protocol documentation is unusually thorough—translation and back-translation, site-specific adaptations, incentive schemes, privacy procedures, and dropout analysis are all described in a level of detail that other dataset papers often skip. The descriptive validation (response rates, sensor coverage, cross-country behavioral patterns) gives a real sense of what the data can support. This is a legitimate contribution to mobile sensing research.\n\nThe soft spots are not fatal, but they are real. The headline claim that the dataset is 'publicly available' is directly contradicted by Section 3.7.2, which states the dataset is 'not publicly available online' and can only be accessed with approval, institutional affiliation, and a no-redistribution license. Section 5.3 confirms this. That turns the main contribution from 'downloadable resource' into 'available to a qualified few by request.' Many gated datasets still get called 'public' in a loose sense, but the paper's own wording is inconsistent, and Table 1 implies the comparison is with openly downloadable datasets. That needs fixing, or the contribution should be re-framed.\n\nThe participant count is also off: the abstract and Contribution 1 say 782 participants, but Table 4 lists 782 iLog sign-ups and only 666 with actual data. If the dataset contains sensor data from 666, then that should be the number advertised. It is not a huge discrepancy, but for a dataset paper, precision matters and reviewers will check.\n\nI am less bothered by the self-report comparability issue. The paper openly acknowledges that labels are 'silver standard' and that instruments were adapted per site. That is honest, and the discussion of label quality (Section 6.2.5) is more candid than most. It is a limitation to note, not a flaw to penalize.\n\nThe paper has no derivation or circularity problems; it is a data release with transparent documentation. For researchers working on cross-country generalization or domain adaptation in mobile sensing, this dataset—even gated—is a potentially important resource. I would send it to peer review, and if it came back to me, I would ask for the access claims and participant counts to be fixed before acceptance.","headline":"A genuinely valuable multi-country sensing corpus, but the 'publicly available' claim is contradicted by the paper's own access-control section, and the participant count needs reconciling.","tokens_in":41554,"tokens_out":1689,"would_cite":true,"duration_ms":19237,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"DiversityOne releases a four-week, eight-country dataset pairing raw smartphone sensor streams from 782 college students with more than 350,000 self-reports, enabling cross-country studies of how behavior-inference models generalize.","keywords":["smartphone sensing","mobile sensing","multi-country dataset","self-reports","behavior modeling","college students","domain generalization","social practices"],"falsifier":"A re-analysis showing that the adapted self-report scales produce systematically different response styles across countries—for example, if controlling for country-specific scale use removes the cross-country differences in model performance—would falsify the claim that the dataset supports fair cross-country behavior comparisons.","tokens_in":40542,"feed_emoji":"📱","tokens_out":4931,"duration_ms":46833,"temperature":0.7,"pith_summary":"This paper presents DiversityOne, a publicly released dataset of everyday life behavior from 782 college students in eight countries over four weeks, combining raw data from 26 smartphone sensor modalities with more than 350,000 in-situ self-reports. The authors aim to overcome a bottleneck in mobile sensing research: existing public datasets are small, sensor-limited, and concentrated in the Global North, so models trained on them may not transfer across countries and cultures. DiversityOne is designed so that researchers can ask cross-country questions about behavior inference, domain adaptation, and model personalization. The paper's own validation shows pronounced cross-site differences in sleep, eating, mood, and app-use patterns, and initial studies on the data indicate that country-specific and individually adapted models outperform generic multi-country ones. If the dataset works as claimed, it gives the field a reusable resource for testing whether mobile behavior models can generalize beyond the populations they were trained on.","feed_headline":"782 students, 26 sensors, 8 countries: new public behavior dataset","feed_subtitle":"Four weeks of raw smartphone sensing plus 350,000+ self-reports let researchers test how behavior models generalize across countries.","key_machinery":"The load-bearing artifact is the dataset itself, produced by a three-part protocol. An invitation questionnaire reached over 18,000 students; a subset of 782 installed a custom logging app (iLog) that passively recorded 26 smartphone sensor modalities at high frequency and prompted in-situ self-reports through morning/evening diaries, half-hourly time diaries, and snack diaries. The protocol operationalizes culture as social practices (material, competence, meaning) and follows HETUS-style time-use diary standards, combining them with experience sampling. To preserve comparability across sites, the central protocol was translated, back-translated, and adapted locally, with an offline mode for regions where cloud notification services or stable connectivity were unavailable. The dataset's modular organization into connectivity, environment, motion, position, app-usage, and device-usage bundles, plus its two GPS anonymization versions (RoundDown and POI), is what lets researchers flexibly assemble the data for behavior inference while respecting privacy constraints.","core_discovery":"The central claim is that DiversityOne is one of the largest and most geographically diverse public datasets that combine questionnaire responses from more than 18,000 students with four weeks of passive smartphone sensing and intensive longitudinal self-reports from 782 participants. Every site followed a shared protocol built on standardized time diaries and experience-sampling questions, with local translations and adaptations; the paper argues that these adaptations preserve enough comparability for meaningful cross-country analysis while capturing genuine behavioral diversity. The dataset ships raw, granular sensor streams (accelerometer, gyroscope, GPS, Bluetooth, WiFi, app usage, notifications, screen and battery events, and more), organized into thematic bundles, alongside structured time diaries. Validation analyses document substantial cross-country variation in daily rhythms and in the distribution of moods and activities, and earlier studies using the data show that country-specific or hybrid personalized models generally beat generic multi-country models. The contribution is therefore not a new algorithm but a new empirical resource with a field-tested collection protocol, intended to make cross-country generalization research in mobile sensing feasible.","pith_inferences":["Editorial inference: the dataset's claim to support fair cross-country comparison depends on the equivalence of the self-report instruments; a promising test would be to compare the psychometric properties of the mood and activity scales across sites before using them as ground truth.","The ranking of apps and daily rhythms already hints at strongly culture-specific behaviors; models that explicitly encode country or cultural context may benefit more than generic transfer methods.","The paper's acknowledgment that self-reported labels are only a silver standard suggests that objective sensor-based labels, such as step count or location patterns, could serve as more stable anchors for validation in future work.","Releasing raw sensor data rather than only pre-processed features opens the door to self-supervised pretraining on mobile sensing, an area the paper notes is largely unexplored."],"forward_implications":["Researchers can benchmark behavior-inference models across eight countries, including Global South sites, testing whether models trained in one country transfer to another.","The dataset enables domain adaptation and domain generalization experiments on raw multimodal time series rather than only pre-computed features.","The rich self-report stream allows study of label shift and the reliability of self-reported ground truth across cultural contexts.","Initial results cited in the paper suggest country-specific or hybrid personalized models will systematically outperform generic multi-country models, guiding deployment choices in real-world applications.","The documented protocol and lessons learned give future multi-country studies a template for ethical, privacy-compliant, adaptive data collection."],"supporting_citations":[{"why":"Defines the continuous vs interaction sensing taxonomy and reviews the landscape that DiversityOne targets.","marker":"[81]"},{"why":"One of the largest earlier public sensing datasets, used as a scale and coverage comparison baseline.","marker":"[71]"},{"why":"Provides the student-population, diary-plus-sensing template that DiversityOne scales up.","marker":"[125]"},{"why":"The closest existing multi-country dataset whose limitations in sample size, country count, and sensors DiversityOne seeks to surpass.","marker":"[133]"},{"why":"Large longitudinal sensing dataset that demonstrates cross-dataset generalization analysis, though geographically concentrated.","marker":"[131]"},{"why":"Initial study using DiversityOne that established country-specific vs multi-country model performance for mood inference.","marker":"[80]"},{"why":"Initial study using DiversityOne for complex daily activity recognition, showing country-model comparisons.","marker":"[3]"},{"why":"Domain adaptation study on DiversityOne data that quantifies cross-country distribution shift and shows gains from adaptation.","marker":"[82]"},{"why":"Supplies the logging app used to collect sensor streams and schedule self-reports.","marker":"[136]"}],"fun_headline_variants":["DiversityOne: public sensor dataset from 8 countries, 782 students","8 nations, 782 students, 26 sensors: everyday behavior data","Multi-country behavior dataset: 782 users, 26 sensors, 350K reports","From 8 countries, 782 participants' smartphone sensors for daily life"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"That self-reported answers about mood, activity, and context remain comparable across countries after translation and local adaptation, even though the questions, incentives, and recruitment differed by site.","fun_headline_variants_meta":{"raw":{"variants":["DiversityOne: public sensor dataset from 8 countries, 782 students","8 nations, 782 students, 26 sensors: everyday behavior data","Multi-country behavior dataset: 782 users, 26 sensors, 350K reports","From 8 countries, 782 participants' smartphone sensors for daily life"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000882,"raw_usage":{"total_tokens":3836,"prompt_tokens":998,"completion_tokens":2838,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":614,"completion_tokens_details":{"reasoning_tokens":2755}},"tokens_in":614,"tokens_out":2838,"duration_ms":20304,"temperature":1.0,"reasoning_tokens":2755,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T05:00:48.254858+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A re-analysis showing that the adapted self-report scales produce systematically different response styles across countries—for example, if controlling for country-specific scale use removes the cross-country differences in model performance—would falsify the claim that the dataset supports fair cross-country behavior comparisons.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the continuous vs interaction sensing taxonomy and reviews the landscape that DiversityOne targets."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"One of the largest earlier public sensing datasets, used as a scale and coverage comparison baseline."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The closest existing multi-country dataset whose limitations in sample size, country count, and sensors DiversityOne seeks to surpass."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Large longitudinal sensing dataset that demonstrates cross-dataset generalization analysis, though geographically concentrated."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Initial study using DiversityOne that established country-specific vs multi-country model performance for mood inference."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the logging app used to collect sensor streams and schedule self-reports."}],"review_version":1}