{"id":"038ba7ff-9366-490a-82ea-c7b816d61a74","arxiv_id":"2501.17198","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A publicly released dataset of 6,000 synthetic sound effects in 30 categories, with per-category documentation of the synthesis methods used.","lead":"This paper releases 6,000 synthetic sound-effect samples across 30 categories, each documented with the synthesis methods used to create it. The dataset is a public resource for researchers and sound designers working on procedural audio and audio machine learning.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The release claim rests on label metadata that Table 1 itself contradicts (missing label 36, duplicated label 42) and no manifest exists to check the Zenodo artifact, so the dataset's usability is unverified.","rationale":"The reader's CONDITIONAL verdict and my read align: the strongest_claim (a downloadable, labeled 6,000-sample corpus) is plausible and partially supported, since the arithmetic checks out (200 samples × 30 categories = 6,000; 5 s mono 16-bit at 44.1 kHz totals approximately 2.6 GB, consistent with the stated 2.56 GB), the Zenodo DOI and GitHub repo are concrete artifacts, and the authors disclose the copyright limit on real samples and the subjectivity of Table 2 classifications. My concern is not that the dataset is fabricated; it is that the paper's own metadata apparatus is demonstrably inconsistent at the exact point where the release's usability lives. The label errors in Table 1 are not stylistic: Section 3 embeds labels in filenames, so an erroneous label scheme propagates into the artifact. The absence of a manifest or checksum means no internal check distinguishes a table typo from a mislabeled artifact. The concrete test resolves this by inspecting the Zenodo record directly: if counts, label uniqueness, and the even/odd pairing fail, the dataset annotations are unreliable and the resource claim weakens materially; if they pass, the release survives and only the paper's tables and composition wording need correction. Either way the recommended verdict is unchanged from the reader's CONDITIONAL: the dataset paper should be trusted only after the artifact is checked.","tokens_in":5130,"tokens_out":9962,"duration_ms":81377,"concrete_test":"Download Zenodo record 14517916 and verify programmatically: (1) total file count is 6,000 with 200 per category across 30 categories; (2) every filename parses as category-sample-label and the label set is unique with the even/odd real-synthetic pairing intact, specifically testing whether Whoosh files use 36 or 37 and whether 42 occurs once or twice; (3) per-category method annotations from Table 2 map unambiguously onto the category names appearing in the filenames. If counts or label uniqueness fail, the metadata errors are baked into the artifact and the resource claim weakens; if all checks pass, the release is sound and only Table 1 and the composition wording need correcting.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central deliverable is a publicly released, labeled corpus, but the release's organization depends on metadata that the paper itself undermines. Section 3 fixes filenames as 'Name of the sound category' - 'Sample number' - 'Label', with labels taken from Table 1, and also contradicts itself on counts ('30 samples per category' vs '200 samples per category'). Table 1 contains demonstrable errors: label 36 is absent (Whoosh is listed 37/38 rather than 36/37), and label 42 is assigned twice (Concrete Footsteps S.S. and Wood Footsteps R.S.), breaking both the even/odd real-synthetic pairing and label uniqueness. If the Zenodo files follow Table 1, the released annotations for Whoosh and the two footstep categories are wrong; if the files correct these, then Table 1, the only map to the release, is wrong. No manifest, checksum, or per-category count is supplied; the GitHub link is to preprocessing code. The composition story is also split: the Introduction bullet promises a public dataset containing both synthetic and pre-recorded samples, Section 3 says real samples are not released, and the Conclusion claims the dataset 'integrates' both. Table 2 further shifts category names ('Glass Debris' vs 'Debris Glass', 'Boat' vs 'Boat Engine', 'Fire Embers' vs 'Fire'), so the method annotations cannot be unambiguously joined to released files. Because the paper's stated purpose is to enable evaluation frameworks that depend on correct labels, the headline claim is conditional on the artifact actually matching the description.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces '6KSFx Synth Dataset', a claimed public dataset of 6,000 synthetic audio samples generated with procedural audio techniques, intended to support research on sound-synthesis evaluation and classification. The manuscript describes 30 sound categories, each represented by real and synthetic samples, with a synthesis-method taxonomy provided in Table 2. The dataset is said to be released on Zenodo, with preprocessing code on GitHub. The central claim is that a downloadable, labeled corpus of 6,000 procedural audio samples exists and is publicly available.","tokens_in":5449,"tokens_out":2927,"duration_ms":29786,"significance":"If the dataset is released as described and the labels are correct, this would be a useful community resource: procedural audio research has few public synthetic datasets, and a corpus with per-category synthesis metadata could support reproducible evaluation, machine-learning benchmarks, and comparisons of synthesis methods. The paper also provides explicit Zenodo and GitHub links and a descriptive taxonomy of synthesis methods, which is helpful for orienting users. However, the contribution's value is entirely dependent on the integrity and usability of the released artifact: the paper contains no derivations to verify, and the supporting claims about the dataset's organization contain internal contradictions that must be resolved before the resource can be used reliably.","major_comments":[{"comment":"Table 1 has demonstrable label-numbering errors that break the mapping from file names to categories: the Whoosh category is assigned labels 37 and 38, but the preceding Jet row ends at 35, so label 36 is missing and the even/odd real-synthetic pairing is shifted. More seriously, Concrete Footsteps is labeled 41/42 and Wood Footsteps is labeled 42/43, so label 42 is duplicated and the label-uniqueness assumption underlying the filename scheme ('Name of the sound category' - 'Sample number' - 'Label') fails. The authors must correct the table or, if the released files use different labels, explicitly state that Table 1 is not the key to the release and provide the correct mapping.","section":"Section 3, Table 1"},{"comment":"The sample counts are contradictory. The text first states that the dataset was constructed using 12,000 five-second samples evenly distributed across 30 categories, then says the public release is 6,000 synthetic samples 'divided into 30 samples per category', and then states that 'each category accounts for 1.8% of the dataset (200 samples per category)'. If there are 200 samples per category, 30 categories give 6,000 samples, so the total cannot also be 12,000 with the same per-category balance unless the real subset is also 200 per category. The claim of balanced distribution cannot be interpreted as written; please state the exact number of files per category and per subset.","section":"Section 3"},{"comment":"The paper contradicts itself on what is publicly available. The Introduction bullet says the dataset provides 'a public dataset containing both synthetic and pre-recorded samples', while Section 3 says the real samples are not publicly available due to copyright restrictions and only links to providers are given. The Conclusion then says the dataset 'integrat[es] both real and synthetic sounds'. These statements describe incompatible release scopes; the authors should state unambiguously whether the released Zenodo artifact contains only synthetic samples, and whether the real samples are available at all.","section":"Section 1 bullet list; Section 3; Section 4"},{"comment":"The category names in Table 2 do not match those in Table 1: 'Debris Glass' vs 'Glass Debris', 'Boat Engine' vs 'Boat', 'Fire' vs 'Fire Embers', and 'Bounce Rubber' vs 'Bounce (Rubber)'. Because the synthesis-method metadata in Table 2 is meant to be joined to the released files via the Table 1 labels, this naming mismatch makes the method annotations ambiguous. The authors should align all category names across Table 1, Table 2, Figure 1, and the actual file names on Zenodo.","section":"Tables 1 and 2"},{"comment":"No manifest, file count, checksum, or programmatic verification is provided to confirm that the Zenodo artifact matches the tables. Given that Table 1 already contains label errors and the count text is inconsistent, the absence of a verifiable manifest is a load-bearing omission: a reader cannot tell whether the released files follow the erroneous Table 1 or a corrected version. Please include a manifest or a verification script that checks file counts, labels, durations, and sampling rates against the corrected metadata.","section":"Section 3"}],"minor_comments":[{"comment":"The word 'Footsteps' appears as 'F ootsteps' in both tables; please fix the spacing.","section":"Tables 1 and 2"},{"comment":"'pre recorded' should be 'pre-recorded' throughout, and 'syntheis' in Reference [20] should be 'synthesis'.","section":"Abstract and Section 1"},{"comment":"The text says 'the distribution shown in 1' but should say 'shown in Figure 1'.","section":"Section 3, Figure 1"},{"comment":"Reference [23] cites a paper on microwave resonator filter synthesis, which appears unrelated to sound synthesis; please replace it with a relevant physical-modeling or resonator synthesis reference.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"This is a dataset-release paper whose value rests entirely on the correctness and verifiability of the released artifact, not on mathematical results. The label errors and contradictory counts are fixable, but they currently undermine the central claim that a usable, correctly labeled dataset is being offered. If the authors can supply corrected tables, a clear scope statement, and a manifest, I would view the paper as publishable; without those, the contribution is not verifiable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the paper has a real deliverable in mind—a public, labeled synthetic sound-effects dataset—but the manuscript as written makes it impossible to trust that deliverable. The internal contradictions and labeling errors are concrete and need to be fixed before anyone uses this as a resource.\n\nWhat's genuinely new: 6,000 synthetic samples released on Zenodo, with per-category synthesis-method labels, fills a documented gap—earlier synthetic collections (Moffat and Reiss) were not publicly accessible, and mainstream datasets are real recordings. That's a useful artifact for procedural audio evaluation and for machine learning work on sound effects.\n\nWhat's done well: the pairing of real and synthetic samples across 30 categories is a sensible design for comparative evaluation; the synthesis-method taxonomy is grounded in prior work and applied to each category; and the authors are upfront that classifications are subjective.\n\nNow the soft spots, and they're not minor. The text contradicts itself in three places. The Introduction promises a public dataset with both synthetic and pre-recorded samples; Section 3 says real samples are not released; the Conclusion claims the dataset 'integrates' both. That's a load-bearing confusion about what the artifact actually is. Section 3 also says '30 samples per category' and then says each category accounts for 200 samples—a direct numeric conflict. Table 1, the only map between filenames and categories, skips label 36 (Whoosh is 37/38 instead of 36/37) and assigns label 42 twice. Table 2 uses different category names ('Debris Glass' vs 'Glass Debris', 'Boat Engine' vs 'Boat', 'Fire' vs 'Fire Embers'). No manifest, checksum, or per-category count is supplied to bridge the paper and the Zenodo files.\n\nThese aren't cosmetic issues. The paper's stated purpose is to enable evaluation frameworks that depend on correct labels. If the released files follow Table 1, the annotations for Whoosh and the two footstep categories are wrong; if the files are correct, then Table 1 is wrong and the paper can't be used as a map. Either way, the resource isn't usable as described.\n\nThe core claim—that a dataset exists—is plausible and probably true, but the paper does not give us a way to verify it. That said, the flaws are fixable. The authors need to reconcile the text, fix the label tables, and add a manifest with counts and either checksums or a verification script.\n\nWho this is for: researchers working on procedural audio, sound effect classification/generation, and audio evaluation benchmarks. They'd get value once the dataset is properly documented.\n\nMy recommendation: send it to peer review, but the reviewers should insist on a revision that clears up these contradictions and provides the verification artifacts. A dataset paper's worth is in the artifact; right now the paper undersells it with avoidable errors.","headline":"A genuinely useful dataset idea undermined by avoidable internal contradictions and label errors that make the release impossible to trust as described.","tokens_in":5941,"tokens_out":2543,"would_cite":false,"duration_ms":20947,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper presents a public dataset of 6,000 synthetic audio samples divided evenly across 30 sound-effect categories, with the synthesis methods behind each category documented.","keywords":["procedural audio","synthetic sound effects","digital Foley","sound synthesis","sound-effects dataset","audio dataset","sound classification benchmark","6KSFx"],"falsifier":"Download the public release, count the files in each of the 30 categories, and compare every filename's numeric label with the parity claimed in Table 1; a category with anything other than 200 files, or a file whose label breaks the even/odd pairing, would show that the dataset description is not accurate.","tokens_in":4933,"feed_emoji":"🔊","tokens_out":9555,"duration_ms":84093,"temperature":0.7,"pith_summary":"The paper's central claim is that a large, balanced, and publicly available synthetic sound-effects corpus is now a practical resource: 6,000 computer-generated samples, five seconds each, spread across 30 everyday sound categories such as applause, rain, footsteps, and rocket. Each category contains 200 synthetic samples and is paired with a recorded version under an even/odd label scheme, and the paper documents which synthesis methods were used for each category. The purpose is to remove the biggest barrier to procedural audio research, the absence of a labeled public dataset with known synthesis provenance. If the dataset is as described, researchers and sound designers gain a common benchmark for comparing synthetic effects with recordings and for training and testing audio classifiers. The synthetic samples were processed with reverb and equalization so that comparisons with real recordings are meant to be fair.","feed_headline":"6,000 synthetic sound effects now public with full synthesis metadata","feed_subtitle":"Every one of the 30 sound categories gets 200 labeled samples plus the synthesis methods used to make them.","key_machinery":"The carrying object is the pairing scheme with its label arithmetic and the category-to-method metadata table. Real and synthetic renditions of each sound category receive consecutive labels, even for the recording and odd for the synthesis, so any two adjacent labels are directly comparable. The paper then assigns each odd label one to three methods from its eight-method taxonomy, giving users a route from a sound file back to the code-level recipe that produced it. The preprocessing chain, procedural generation followed by reverb and equalization, is the mechanism intended to put synthetic and recorded samples on equal footing for listening tests or machine-learning evaluation.","core_discovery":"On the paper's own terms, the contribution is the dataset itself rather than a new synthesis technique. The authors generated synthetic audio algorithmically, treated it with reverb and equalization, and organized it into 30 categories of 200 samples at a fixed format of 5 seconds, 44.1 kHz, mono, and 16 bits. A companion table maps each category to one to three synthesis methods chosen from additive, subtractive, granular, physical modeling, physically informed, modal, signal modeling, and frequency modeling. The resulting release is meant to make procedural audio measurable: the same category always appears as a real and a synthetic pair, so quality, realism, and method effectiveness can be compared without confounding on content.","pith_inferences":["Editorial inference: the same even/odd pairing could support a synthetic-audio detection task, in which a model learns to tell generated effects from recorded ones under matched conditions; the paper does not itself train such a detector.","Editorial inference: one could extend the release by asking listeners to rate realism per category and then correlate those ratings with the tabulated synthesis methods, which would turn the dataset into a perceptual benchmark rather than just a collection.","Editorial inference: because the method labels were assigned by reading code, a natural check is to see whether the documented recipes are acoustically distinguishable, for instance by unsupervised clustering of the samples; the paper reports no such analysis.","Editorial inference: the same dataset structure could be scaled to finer-grained subcategories, such as different surfaces or intensities, to support harder evaluation tasks than 30-way classification."],"forward_implications":["A sound-effects classifier can be trained and tested on 30 balanced classes, with the even/odd pairs separating real from synthetic audio within each class.","Synthesis-method comparisons can be made within a category, because the metadata explains which technique generated each label.","The fixed 5-second mono format at 44.1 kHz lets researchers use the corpus as a benchmark without additional normalization.","Procedural audio researchers can cite one public release instead of relying on private or defunct sample libraries.","Sound designers can use the documented recipes to generate new variations of a desired effect rather than recording or licensing it."],"supporting_citations":[{"why":"Representative pre-recorded audio dataset that the paper contrasts with its synthetic release.","marker":"[9]"},{"why":"Large-scale audio-visual dataset used to show that existing corpora contain few or no synthetic samples.","marker":"[3]"},{"why":"Urban sound dataset cited as another pre-recorded resource without procedural samples.","marker":"[18]"},{"why":"Audio captioning dataset that exemplifies the pre-recorded focus of existing public audio corpora.","marker":"[5]"},{"why":"Supplies the procedural-audio categorization framework that the paper uses to assign synthesis methods to categories.","marker":"[13]"},{"why":"Provides the behavioral-abstraction view of procedural audio that motivates the dataset's design.","marker":"[6]"},{"why":"Earlier synthesized-sound-effect study whose samples were no longer publicly accessible, motivating a public release.","marker":"[15]"},{"why":"Earlier perceptual evaluation of sound-effect synthesis that was limited in scope and lacked a public dataset.","marker":"[16]"}],"fun_headline_variants":["6k synthetic sounds for fair audio-quality comparisons","Dataset pairs synthetic and real audio for method testing","6,000 synthetic effects with synthesis method metadata","Procedural audio dataset: 30 categories, 6000 clips"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the files in the public download match the paper's descriptions: 200 samples per category, the stated format, the even/odd real-synthetic labeling, and the synthesis methods listed in Table 2.","fun_headline_variants_meta":{"raw":{"variants":["6k synthetic sounds for fair audio-quality comparisons","Dataset pairs synthetic and real audio for method testing","6,000 synthetic effects with synthesis method metadata","Procedural audio dataset: 30 categories, 6000 clips"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000465,"raw_usage":{"total_tokens":2255,"prompt_tokens":812,"completion_tokens":1443,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":428,"completion_tokens_details":{"reasoning_tokens":1379}},"tokens_in":428,"tokens_out":1443,"duration_ms":10937,"temperature":1.0,"reasoning_tokens":1379,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T12:58:49.695889+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Download the public release, count the files in each of the 30 categories, and compare every filename's numeric label with the parity claimed in Table 1; a category with anything other than 200 files, or a file whose label breaks the even/odd pairing, would show that the dataset description is not accurate.","supporting_citations":[{"cited_title":"Audio set: An ontology and human-labeled dataset for audio events","cited_arxiv_id":null,"evidence_quote":"Representative pre-recorded audio dataset that the paper contrasts with its synthetic release."},{"cited_title":"Vggsound: A large-scale audio-visual dataset","cited_arxiv_id":null,"evidence_quote":"Large-scale audio-visual dataset used to show that existing corpora contain few or no synthetic samples."},{"cited_title":"A dataset and taxonomy for urban sound research","cited_arxiv_id":null,"evidence_quote":"Urban sound dataset cited as another pre-recorded resource without procedural samples."},{"cited_title":"Clotho: an audio captioning dataset","cited_arxiv_id":null,"evidence_quote":"Audio captioning dataset that exemplifies the pre-recorded focus of existing public audio corpora."},{"cited_title":"The state of the art in procedural audio","cited_arxiv_id":null,"evidence_quote":"Supplies the procedural-audio categorization framework that the paper uses to assign synthesis methods to categories."},{"cited_title":"Designing sound","cited_arxiv_id":null,"evidence_quote":"Provides the behavioral-abstraction view of procedural audio that motivates the dataset's design."},{"cited_title":"Perceptual evaluation of synthesized sound effects","cited_arxiv_id":null,"evidence_quote":"Earlier synthesized-sound-effect study whose samples were no longer publicly accessible, motivating a public release."},{"cited_title":"Sound Effect Synthesis, pages 274–299","cited_arxiv_id":null,"evidence_quote":"Earlier perceptual evaluation of sound-effect synthesis that was limited in scope and lacked a public dataset."}],"review_version":1}