{"id":"abac208b-d483-4318-8800-eb0887a000de","arxiv_id":"2608.00089","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"DODA is a curated online catalog with metadata and precomputed image statistics for over 120 aesthetics-rated image datasets.","lead":"This paper introduces DODA, a searchable web database that collects more than 120 image datasets used in aesthetics research and adds standardized metadata plus precomputed image statistics. A generalist would use it to find and compare stimulus datasets without downloading every set.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"DODA's 'all important datasets' claim rests on unvalidated manual curation; internal inconsistencies (90 vs 120 datasets, QIP coverage) signal the inventory and metadata need independent audit before the database is relied upon.","rationale":"I read the paper's central claim as conditional on the database being both comprehensive and accurate. The reader's weakest assumption — that the manual curation may be incomplete or inaccurate — is the most load-bearing concern because the entire utility of DODA is as a trusted catalog. The internal inconsistencies (90 vs 120 datasets; QIP coverage statements) strengthen this concern: they indicate that the compilation is not carefully cross-checked. I found no fatal error that would warrant REJECT; the application and dataset list appear real and useful for a first screening. However, without an external audit, ACCEPT would be too strong. The heterogeneity-score binning issue is a separate methodological gap, but it does not affect the core dataset-discovery function as directly as completeness and correctness. A concrete recall and metadata-accuracy audit would settle whether the headline claim is true; until then, CONDITIONAL remains the appropriate verdict. The reader's moderate confidence is consistent with my view.","tokens_in":31853,"tokens_out":6282,"duration_ms":67093,"concrete_test":"Perform an independent completeness and accuracy audit: compile a candidate dataset list by merging the two GitHub overviews cited in the paper (Huckle's art-datasets and dieuroi's Awesome-Image-Aesthetic-Assessment) with datasets extracted from a systematic 2020-2026 literature search for aesthetics and image-quality datasets. Compare this union set to DODA's Table 1 and compute recall. Then randomly sample 20 DODA entries and verify core metadata (number of images, number of raters, scale, source, year) against the original publications. If recall is below about 90% or any sampled entry has an error in a core field, the 'all important datasets' and accurate-metadata claims are falsified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that DODA lets researchers 'browse all important datasets for aesthetics research' with reliable metadata and precomputed image properties. The weakest link is the curation process: entries were compiled from a literature review and Huckle's GitHub overview, not validated against dataset authors or an external registry. No inclusion/exclusion criteria or inter-rater agreement are given. Internal inconsistencies show this risk is real. The manuscript says 'we compiled a list of over 90 datasets' (Section 'Introducing the Database'), whereas the abstract and the introductory section claim 'over 120 datasets'; even adding the acknowledged 25 Huckle sets to 90 gives 115, not 120. Similarly, QIP coverage is described both as 'for many of them' (Abstract), 'whenever images were available and image set size is equal to or below 17,000 images' (Section 'Data Available'), and 'we provide pre-calculated QIPs for all datasets found on DODA' (Section 'Resolution, Image Quality and Image Fidelity'). If the inventory is incomplete or metadata fields (raters, scales, content, sources) are inaccurate, the central promise — fast, reliable dataset discovery — fails, even though the web app itself works. The heterogeneity-score method also omits binning details for continuous QIPs, making Figure 7 non-reproducible, but this is secondary to the curation issue.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces DODA (Database of Datasets for Aesthetics), a web application embedded in the Aesthetics Toolbox that catalogs image datasets annotated for aesthetic variables. It claims to let researchers browse 'all important datasets' in empirical and computational aesthetics, with metadata on dataset size, annotations, raters, image source, resolution, formats, and scales. For many datasets, DODA also provides precomputed quantitative image properties (QIPs) computed via the Aesthetics Toolbox. The paper further proposes a heterogeneity score based on the mean normalized Shannon entropy of selected bounded QIPs, and illustrates it on a subset of datasets. The manuscript includes a supplementary list of datasets and a list of open museum collections, and describes copyright considerations and dataset-reuse arguments.","tokens_in":32219,"tokens_out":2972,"duration_ms":34852,"significance":"If the database is complete and accurate, DODA fills a real gap: there is currently no centralized, filterable registry of aesthetics-annotated image sets spanning both the small controlled stimulus sets of empirical aesthetics and the large machine-learning corpora of computational aesthetics. The paper is pragmatic and useful, and the Shannon-entropy heterogeneity measure is a simple, parameter-light descriptor that could be useful for dataset comparison. The authors openly provide the web application, the dataset list, and precomputed QIPs, which is a concrete open-science contribution. However, the central claim of a reliable, comprehensive catalog rests on manual curation that is not described with enough rigor, and the manuscript contains internal inconsistencies about the number of datasets and the coverage of QIPs. These issues are fixable but must be addressed before the database can be relied upon as the authoritative resource the paper promises.","major_comments":[{"comment":"The paper states 'DODA currently encompasses over 120 datasets' (p. 8) but later says 'Based on a thorough literature review we compiled a list of over 90 datasets' (p. 9). The acknowledgements add that Huckle's GitHub list allowed inclusion of 25 additional image sets; 90 + 25 = 115, not 120. This is not just a wording issue: the central claim is that DODA lets users browse 'all important datasets', so the inventory count and its provenance are load-bearing. Please give the exact number with a date, describe the systematic literature search, inclusion/exclusion criteria, and how completeness is assessed, and specify whether the '90' is an earlier version of the inventory.","section":"Introduction / Introducing the Database (pp. 8–9)"},{"comment":"The manuscript contradicts itself on QIP coverage. The 'Data Available' section says DODA provides QIPs 'whenever images were available and image set size is equal to or below 17,000 images', but the later section says 'we provide pre-calculated QIPs for all datasets found on DODA'. The abstract says 'for many of them'. Since users rely on DODA to know whether QIPs are present, please state unambiguously what fraction of datasets have QIPs, list the cutoff criterion, and make the mapping from datasets to QIP availability explicit in the web interface and in the paper.","section":"Data Available (p. 15) vs. Resolution, Image Quality and Image Fidelity (p. 21)"},{"comment":"The heterogeneity score is not reproducible as written. Equation (2) uses the empirical frequency p_q(v) of 'each observed value v' and defines K_q = |V_q| as the number of possible discrete values, but the selected QIPs (RMS contrast, mean RGB/L, symmetry, balance, homogeneity, DCM distance) are continuous. The paper does not specify the binning scheme, bin width, number of bins, or how K_q is determined for continuous variables. It also does not state how datasets with missing QIPs are treated in the mean over M. Without this information, Figure 7 cannot be reproduced and the stated 'parameter-free' nature of the score is unverifiable. Please add the full discretization protocol and the handling of missing data.","section":"Heterogeneity Score, Eq. (2)"}],"minor_comments":[{"comment":"The text refers to 'Table A in Supplementary Materials' (p. 8), but the supplementary table is labeled 'Table 1'. Please align the cross-reference.","section":"Supplementary Materials, Table 1"},{"comment":"The caption says 'Fifty four exemplary DODA datasets' but does not explain why these 54 were selected or how the others are excluded. State the selection criterion.","section":"Figure 4"},{"comment":"Typo: 'ArtBrench-10' should be 'ArtBench-10', consistent with the main text and the cited reference.","section":"Supplementary Materials, Table 1"},{"comment":"The hierarchical classification rule is described but not formalized. Please specify the exact precedence order and how ties or mixed datasets are resolved, since 'Data Category' is a filterable field.","section":"Data Category, p. 17"},{"comment":"Because DODA is 'continuously growing', please provide a version number, snapshot date, or DOI for the database as described in the paper, so users can cite a stable version and know when the inventory was last updated.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper describes a useful resource, but the dataset-count and QIP-coverage inconsistencies are more than cosmetic: they affect what users can expect from the database. The heterogeneity-score protocol also needs completion. I do not see a circularity problem: DODA's utility does not depend on the heterogeneity formula, and the entropy-based measure uses no fitted constants. The main risk is curation quality, which the authors should address by documenting their workflow and, ideally, by releasing the curatorial notes or an audit trail."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This one is worth a look. DODA is a genuinely useful resource — a filterable catalog of 100+ aesthetics-annotated datasets with standardized metadata and, for many, precomputed image statistics. It fills a real gap, but the manuscript has a few internal inconsistencies that need cleaning up before I'd treat it as the field standard.\n\nWhat's new: the metadata schema (number of images, raters, rating scales, content, source, citation) is more complete than existing GitHub lists, and the inclusion of small empirical-aesthetics datasets is a real improvement. Precomputed QIPs via the Aesthetics Toolbox are a concrete convenience, and the heterogeneity score is a nice add-on for comparing dataset diversity. The paper is clearly written and transparent about its process — they acknowledge Huckle's list and disclose the use of Gemini for code debugging. The web app is openly available, and the supplement lists the datasets with citations.\n\nThe soft spots are the inconsistencies the reader flagged, and they do land on close reading. First, the dataset count: 'over 120' in the introduction vs 'over 90' in Section 'Introducing the Database'; even adding the 25 from Huckle gives 115, not 120. That's a factual error. Second, QIP coverage is described three ways: 'for many', 'whenever images were available and image set size is equal to or below 17,000 images', and 'pre-calculated QIPs for all datasets found on DODA'. The last is false as stated, and users need to know which datasets actually have QIPs. Third, the heterogeneity score (Eq. 2) assumes discrete QIP values, but the chosen QIPs are continuous; the binning is unspecified, so Figure 7 is not reproducible. That's a minor issue, but it should be fixed.\n\nThe bigger, longer-term risk is curation, as the stress-test note says. Entries were compiled from a literature review and Huckle's GitHub list, with no explicit inclusion/exclusion criteria or inter-rater validation. For a first version of a community resource, that's acceptable, but the paper should soften 'all important datasets' to something like 'a curated collection' and be explicit about the process. The authors already invite submissions, which is the right model.\n\nWho this is for: anyone in empirical or computational aesthetics who needs to find or reuse image datasets. It deserves a serious referee — the resource is valuable and the problems are fixable. I'd recommend revise-and-resubmit with the inconsistencies cleaned up and the heterogeneity methodology fully specified.","headline":"Useful catalog for aesthetics dataset discovery, with a few internal inconsistencies that need fixing before it becomes the standard reference.","tokens_in":32670,"tokens_out":3997,"would_cite":true,"duration_ms":60447,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"DODA is a filterable online catalog of over 120 image datasets annotated for aesthetics, with standardized metadata and precomputed image statistics.","keywords":["empirical aesthetics","computational aesthetics","image aesthetics","dataset catalog","image annotation","quantitative image properties","open science","metadata"],"falsifier":"Check DODA's coverage against a systematic literature search of empirical and computational aesthetics papers published in a given window (e.g., 2015–2025): if a clearly relevant dataset is absent, or if a random sample of entries disagrees with the source papers on basic fields like image count or rater number, the catalog's reliability as a find-and-filter tool is called into question.","tokens_in":31738,"feed_emoji":"🖼️","tokens_out":6319,"duration_ms":64332,"temperature":0.7,"pith_summary":"This paper introduces DODA, a web-based catalog that brings together more than 120 image datasets that have been annotated for aesthetics, from small experimental stimulus sets to large computational collections. The aim is to solve a practical problem: researchers currently have to search across papers, code repositories, and file-sharing sites to find datasets that match their study needs, often downloading full collections just to check basic details. DODA standardizes each dataset's metadata—size, resolution, annotation type, number of raters, rating scales, image source, and more—and precomputes quantitative image properties and a heterogeneity score for many entries. The authors argue that this centralized, filterable tool lowers the barrier to dataset reuse, promotes comparison across studies, and supports collaboration between empirical and computational aesthetics. The paper also discusses which dataset properties matter most when selecting stimuli and illustrates the benefits of reusing existing datasets.","feed_headline":"One web app catalogs 120+ aesthetics image datasets","feed_subtitle":"Standardized metadata and precomputed image statistics let researchers compare datasets without downloading them.","key_machinery":"The central object is DODA itself: a structured, filterable table of dataset profiles, each described by a fixed set of metadata columns (e.g., number of images, raters per image, annotation scale, image source, field of common use). The most mathematically distinctive piece is the heterogeneity score: for each dataset, the average Shannon entropy of several bounded quantitative image properties (RMS contrast, color channel means, symmetry, balance, homogeneity, DCM distance) is normalized to [0,1] and averaged, giving a semantics-agnostic measure of how varied the images are in low-level statistics. This score is meant to help users anticipate how dataset diversity could affect statistical","core_discovery":"The central claim is that an openly accessible online database, DODA, can serve as the field's common reference point for finding and selecting image datasets for aesthetics research. For each of the over 120 entries, DODA reports standardized metadata covering year, citation, number of images, number of raters, ratings per image, semantic classes, image style, content, annotation format and scale, source, resolution, and task, plus a link to the original dataset. Where images are available and the set is not prohibitively large, the database also provides precomputed quantitative image properties and a heterogeneity score, computed as the mean normalized Shannon entropy across selected boun","pith_inferences":["If DODA becomes the standard portal for aesthetics datasets, future papers may cite the catalog entry rather than describing dataset characteristics from scratch, making reporting more uniform.","The heterogeneity score could be repurposed as a general-purpose dataset descriptor in other fields that rely on image statistics, such as image quality assessment or material perception.","A natural extension would be to let users upload their own metadata or verify existing entries, turning the catalog into a collaborative registry with version control.","The authors' decision to cover both large computational sets and small empirical stimulus sets positions DODA as a bridge between two research cultures that often do not share data; if adopted, it could uncover regularities that only appear when results are compared across dataset sizes."],"forward_implications":["Researchers can filter across all major aesthetics-annotated image sets at once, cutting the time spent hunting through papers and shared folders for basic dataset facts.","Precomputed image statistics and heterogeneity scores make it possible to assess dataset diversity before committing to download, which is especially useful for large collections.","The standardized metadata schema reveals gaps and overlaps among existing datasets (e.g., which sets share the same source images), supporting conscious reuse.","The invitation for community submissions makes DODA a living catalog that tracks new datasets as they appear.","By lowering the cost of finding suitable datasets, DODA encourages researchers to reuse rather than create one-off stimulus sets, improving comparability across studies."],"fun_headline_variants":["Find the right aesthetics dataset without downloading anything","DODA: one web app to find any aesthetics dataset","Aesthetics dataset search just got a central hub","DODA: 120+ aesthetics datasets, one searchable place","Stop downloading datasets to compare — use DODA"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The entire usefulness of DODA rests on the manual curation being both sufficiently complete and correct; if substantial numbers of aesthetics datasets are missing or their metadata are wrong, the tool cannot reliably guide researchers to the right dataset.","fun_headline_variants_meta":{"raw":{"variants":["Find the right aesthetics dataset without downloading anything","DODA: one web app to find any aesthetics dataset","Aesthetics dataset search just got a central hub","DODA: 120+ aesthetics datasets, one searchable place","Stop downloading datasets to compare — use DODA"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000656,"raw_usage":{"total_tokens":2829,"prompt_tokens":724,"completion_tokens":2105,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":468,"completion_tokens_details":{"reasoning_tokens":2026}},"tokens_in":468,"tokens_out":2105,"duration_ms":17435,"temperature":1.0,"reasoning_tokens":2026,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T01:18:13.089639+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Check DODA's coverage against a systematic literature search of empirical and computational aesthetics papers published in a given window (e.g., 2015–2025): if a clearly relevant dataset is absent, or if a random sample of entries disagrees with the source papers on basic fields like image count or rater number, the catalog's reliability as a find-and-filter tool is called into question.","supporting_citations":[],"review_version":1}