{"id":"ed8aa6a3-ab26-414b-ae3d-f849895ccc89","arxiv_id":"2607.20951","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A new public dataset of 231,800 sound-designed vocalizations with seen/unseen preset-style and source-timbre splits, plus a baseline non-human voice-conversion evaluation.","lead":"This paper releases a public dataset of \"designed vocalizations\" — monster growls, robotic voices, and other processed human and animal vocals — with standardized training/test splits and baseline conversion results. It gives voice-conversion researchers a shared benchmark for non-human timbre transfer, an area that previously had no public resource.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Source-seen test clips may overlap training VCTK utterances, undermining the seen/unseen split at the core of the benchmark claim.","rationale":"The reader's verdict was already CONDITIONAL, with two listed concerns: (1) same-content (source, reference) pairs may overstate production generalization, and (2) unstated train/test overlap in source-seen VCTK clips. I agree with CONDITIONAL and with the importance of the overlap issue, but I would elevate it above the same-content concern as the single most load-bearing issue. The same-content pair construction is an explicit, clearly disclosed design choice in Sec. 2.1, and the paper's evaluation is framed around timbre reproduction from the same content, so that concern is more about external validity than about internal correctness. The overlap concern, however, directly threatens the paper's central claim of 'explicit seen/unseen splits.' If the source-seen test set is not disjoint from training at the utterance or speaker level, the reported seen-to-seen and seen-to-unseen numbers cannot support the intended generalization conclusions. The concern is concrete, falsifiable via file-level intersection checks, and not resolved by any statement in the manuscript. Because the reader already conditioned acceptance on addressing this issue, my stress-test does not move the verdict; it sharpens the condition and provides a specific check. The recommendation remains CONDITIONAL until the authors release overlap statistics or a disjoint split.","tokens_in":8222,"tokens_out":2947,"duration_ms":35228,"concrete_test":"Download the released dataset and compute the exact intersection between the source-seen test VCTK sample IDs/utterance indices and the 3,270 training VCTK samples. Also compare the 10 test VCTK speaker IDs against the 109 training speaker IDs. If any exact test utterance or test speaker appears in training, re-run the source-seen conditions in Table 2 with a strictly disjoint split and report the new scores. For completeness, perform the same ID-level disjointness check for the non-linguistic Freesound test sources against the training Freesound sources. If all intersections are empty, the concern is resolved and the seen/unseen claim is clean.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central contribution is a benchmark with 'explicit seen/unseen splits over source timbre groups and preset styles' (Abstract; Sec. 2.4.1). The source-seen condition is therefore load-bearing. Sec. 2.3.1 states the training set contains '30 samples from each of 109 speakers, totaling 3,270 samples' from VCTK. Sec. 2.4.1 states the source-seen linguistic test samples are 'drawn from 10 VCTK speakers.' The paper never states that these 10 speakers are disjoint from the 109 training speakers, nor that the specific test utterances are disjoint from the training utterances. If any test speakers or exact test clips also appear in training, then the seen-source condition is not a held-out condition: the model has already seen the same speaker or even the same utterance, so the seen-to-seen scores in Table 2 (e.g., Cos. Sim. 0.667) are inflated and the seen/unseen comparison is not a valid generalization test. This is not a question of task design—the same-source reference pairs are explicitly defined in Sec. 2.1 and can be defended as a controlled timbre-transfer protocol. The overlap issue, by contrast, is an unstated validity threat to the benchmark itself, and it is directly checkable from the released file lists. The reader flagged this as a secondary concern; I agree it is the most load-bearing because it attacks the headline claim of controlled splits.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces the Designed Vocalizations Dataset, a publicly released resource for non-human voice conversion (H2NH-VC). Raw vocalizations (3,270 VCTK speech clips and 2,384 non-linguistic Freesound clips, plus 120 test sources) are processed with 47 DSP presets—30 built-in Dehumaniser 2 presets and 17 in-house-designed presets—to create 231,800 designed clips. The dataset provides non-parallel training data, a test split of 120 sources × 47 presets with aligned (source, reference) pairs, and explicit seen/unseen splits over source timbre groups and preset styles. The authors report baseline results using their own H2NH-VC model [3] across four conditions (seen/unseen source × seen/unseen preset), with objective metrics (cosine similarity, PCC-E, RMSE-E, CER, WER) and a MOS listening test. The central claim is that this is the first public benchmark enabling controlled evaluation of designed/non-human voice conversion.","tokens_in":8475,"tokens_out":5485,"duration_ms":58014,"significance":"If the dataset is released as described and the splits are clean, this is a valuable contribution. Prior H2NH-VC work relies on internally curated, non-public resources, making fair comparison impossible; a public benchmark with metadata for effect-chain structures and explicit generalization splits directly addresses that gap. The arithmetic is consistent (5,654×40=226,160; 120×47=5,640; total clips including raw sources are 237,574), the construction pipeline is described step-by-step, and the authors appropriately flag the CER/WER interpretation as speculative. The decision to use a self-authored baseline is disclosed and is acceptable for a dataset paper. The significance depends on the integrity of the seen/unseen split and on the ecological validity of the test-pair construction, both of which require scrutiny.","major_comments":[{"comment":"The source-seen test condition is load-bearing for the benchmark claim of ‘explicit seen/unseen splits over source timbre groups.’ The training set uses 30 VCTK samples from each of 109 speakers; the test source-seen linguistic samples are ‘drawn from 10 VCTK speakers.’ The paper does not state whether these 10 speakers are a subset of the 109, nor whether the specific test utterances are disjoint from the 30 training utterances per speaker. If any test speaker or exact test clip also appears in training, the seen-to-seen scores in Table 2 (Cos. Sim. 0.667, MOS 3.81) are inflated and the seen/unseen comparison is not a valid generalization test. Please state the disjointness explicitly, provide a verification script in the released file lists, and if there is any overlap, reselect the test clips and rerun the baseline.","section":"Sec. 2.4.1 and Sec. 2.3.1"},{"comment":"Each test reference is defined as x_r = G_p(x_s), so source and reference always have identical content. This isolates timbre matching, but in typical H2NH-VC production use the reference designed sample comes from a different vocalization with different content. The current test protocol may therefore overstate performance for content-mismatched reference conditions. The paper should explicitly discuss this limitation and, ideally, add a secondary content-mismatched evaluation condition. This does not invalidate the controlled timbre-transfer protocol, but it is an important boundary on what the benchmark claims to measure.","section":"Sec. 2.1"},{"comment":"Objective metrics are reported as point estimates without confidence intervals or significance tests, and the MOS is based on 8 participants / 40 samples. The paper interprets small differences (e.g., 0.15–0.17 MOS) as a monotone trend across conditions. Since benchmark results are intended to support reproducible comparison, the release should include per-sample statistics, bootstrap confidence intervals, or at least error bars so that readers can assess whether the seen/unseen ordering is reliable. This is fixable without changing the dataset.","section":"Table 2 and Sec. 3.3"}],"minor_comments":[{"comment":"The term ‘source timbre group’ is used for both VCTK speaker identities and non-linguistic categories. Please define what constitutes a timbre group and how the 10 non-linguistic types were selected, beyond the tag-based + manual listening description.","section":"Sec. 2.4.1"},{"comment":"The claim that the training set avoids over-representation of any effect module is not quantitatively supported. A table or histogram of effect-module counts would make this checkable.","section":"Sec. 2.4.3"},{"comment":"Please specify the exact training/validation split for the baseline model (number of training clips, validation procedure, hyperparameters) to ensure reproducibility of the baseline numbers in Table 2.","section":"Sec. 3.1"},{"comment":"The dataset is described as publicly available, but the license terms for the released audio—especially regarding Dehumaniser 2 processing and Freesound sources—are not stated. Clarify the licensing conditions in the dataset release.","section":"Sec. 2.2"},{"comment":"The lower CER/WER for unseen-source conditions is interpreted as possibly reflecting weaker conversion leaving the output closer to the source. This speculation is flagged, but it would be useful to report an oracle condition (source only, no conversion) to anchor the ASR metrics.","section":"Sec. 3.3"}],"recommendation":"major_revision","confidential_remarks":"The most important issue is the VCTK overlap question in Sec. 2.4.1. Before acceptance, please verify that the source-seen test clips are disjoint from the training utterances (and decide whether speaker overlap is intended). The same-content reference pairs are a design choice but should be acknowledged. The paper is otherwise a solid dataset contribution within scope. No concerns about the disclosed self-authored baseline."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is the first public dataset and benchmark for human-to-non-human voice conversion, and the authors handle it responsibly. The resource itself is the contribution; the benchmark numbers are useful but preliminary.\n\nWhat's new: they release 231,800 designed vocalizations from 5,774 raw sources, processed through 47 preset DSP chains, with metadata and explicit seen/unseen splits across source timbre types and preset styles. Counts check out. Prior H2NH-VC papers all used internal, unreleased data, so this fills a real gap. The construction pipeline is described in enough detail to reproduce, and they include preset-level effect summaries. They also flag their own speculation about the CER/WER trend, which is good practice.\n\nSoft spots: the evaluation is the weakest part. Table 2 reports a single baseline (their own H2NH-VC model, disclosed) with no variance or significance tests, so you can't tell whether the seen-to-seen vs unseen-to-unseen gaps are real. MOS comes from 8 listeners on 40 samples, which is small. The same-content (source, reference) pairs make the task a timbre-transfer protocol; that's a defensible choice, but real use cases may give a reference with different content, so the scenario scores may overstate generalization. One specific thing the paper should clarify: the source-seen linguistic test samples are drawn from VCTK, and the training set also uses VCTK. It's likely the test speakers are among the 109 training speakers — which is consistent with 'seen' referring to timbre groups — but the paper should explicitly state that the exact test utterances are disjoint from the training utterances. Right now that's an unstated assumption, and it's the only thing that could actually undermine the benchmark claim. I'd ask for that clarification in revision, not treat it as fatal.\n\nBottom line: the dataset is the contribution, and it's a good one. This paper deserves a real peer review. If I were in the voice-conversion area, I'd cite it and bring it to reading group.","headline":"A genuinely useful, honest dataset paper — the first public benchmark for designed-voice conversion — with a thin baseline evaluation and one split-overlap detail that needs clarifying.","tokens_in":9074,"tokens_out":2945,"would_cite":true,"duration_ms":26920,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces a public dataset of 231,800 sound-designed vocalizations for converting human voices into non-human ones, with a standardized benchmark to measure how well systems generalize to unseen styles and sources.","keywords":["designed vocalizations","non-human voice conversion","sound design","timbre transfer","benchmark dataset","seen/unseen generalization","vocal effects processing","audio dataset"],"falsifier":"Take the released test set and compare two conditions: (1) the standard same-source pairs, and (2) pairs where the reference is the designed version of a different source with different content, matched by preset. If the cosine similarity or MOS drops substantially in condition (2), then the benchmark's aligned pairs are not a fair proxy for production use. A second, simpler check is to verify whether the 60 source-seen VCTK test clips are disjoint from the 3,270 VCTK clips used in training; overlapping clips would inflate seen-source scores.","tokens_in":8049,"feed_emoji":"👹","tokens_out":1858,"duration_ms":23347,"temperature":0.7,"pith_summary":"The paper argues that research on converting human voices into non-human, designed vocalizations—monster growls, robotic voices, creature sounds—has been held back by the lack of public data. To close that gap, it introduces a large, publicly available dataset built by taking raw vocal sources (speech, animal sounds, interjections, vocal mimicry) and running them through professional sound-design effect chains. Crucially, the dataset includes a standardized test set with aligned source–reference pairs and explicit seen/unseen splits over both source timbre types and preset styles, so different models can be compared fairly and generalization can be measured. The paper also reports baseline results from a representative voice conversion model, establishing a reference point for future work.","feed_headline":"New dataset turns human voices into monster and robot timbres","feed_subtitle":"231,800 sound-designed vocal clips with standardized test splits open a public benchmark for non-human voice conversion.","key_machinery":"The key mechanism is the preset-specific DSP operator G_p, a composition of effect modules (delay pitch shifting, flanger/chorus, granular, noise generator, pitch shifting, ring modulator, spectral shifting) arranged in serial, parallel, or hybrid chains. These operators define the mapping from raw vocalization to designed vocalization, and they make the benchmark's structure possible: because every reference is G_p applied to the same source, timbre conversion is evaluated independently of content mismatch, while the seen/unseen split over presets and source types probes generalization.","core_discovery":"The central claim is that the Designed Vocalizations Dataset is a viable public resource for training and evaluating human-to-non-human voice conversion. Each designed sample is generated as x^(p) = G_p(x_raw), where G_p is a preset-specific DSP operator built from effect modules, and the test set uses reference samples that are the preset-processed versions of the same source: x_r = G_p(x_s). This construction makes content identical between source and reference, isolating the timbre-transfer task. The test set covers 120 sources and 47 presets—40 seen and 7 unseen—with source timbre groups split into seen and unseen categories, enabling systematic evaluation of how models handle unseen spe","pith_inferences":["A natural extension is to test whether the same-content pairing in the test set masks a key difficulty of real use: production references usually have different content from the input. A variant benchmark where reference and source are different utterances would reveal how much of the reported performance depends on content alignment.","The dataset's structure could support an effect-decoding task: given the raw source and the designed output, can a model predict which preset was used? One could treat the preset labels as a supervision signal for learning interpretable timbre representations.","The benchmark could be extended to a text-conditioned setting where the source is also synthesized or edited, connecting non-human voice conversion to text-to-speech pipelines for character voices.","Because only one baseline model is evaluated, the dataset's difficulty is not yet strongly characterized; running existing zero-shot voice conversion systems on the same splits would show whether the seen/unseen gaps are model-specific or reflect intrinsic dataset difficulty."],"forward_implications":["Research groups can train and compare non-human voice conversion systems on a common public dataset instead of internally curated, unreleased data.","The explicit seen/unseen splits let future work quantify how much performance drops when a model encounters an unseen style or an unseen source type, guiding progress toward production-ready generalization.","The released preset-level effect summaries and full effect-chain parameters could enable research on effect-aware or interpretable voice conversion, where models estimate or control the underlying sound-design chain rather than only matching a target timbre.","The baseline results establish a concrete reference point, so subsequent papers can state whether their method improves on the standard setting.","Because the dataset includes non-linguistic sources like animal sounds and vocal mimicry, it expands voice conversion research beyond speech into broader sound-design use cases for games, film, and interactive media."],"fun_headline_variants":["Monster growls and robot beeps: new dataset for voice conversion","Public vocal-effect dataset to train non-human voice AI","Sound-designed voice clips open benchmark for timbre transfer","Human to monster voice: 231,800 clips for AI conversion","New benchmark dataset for AI to synthesize designed vocals"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The benchmark assumes that a reference designed from the same source as the input is representative of how voice conversion models will actually be used, where the reference usually has different content—if same-content pairs artificially make the task easier, the reported conversion quality may overstate real-world performance.","fun_headline_variants_meta":{"raw":{"variants":["Monster growls and robot beeps: new dataset for voice conversion","Public vocal-effect dataset to train non-human voice AI","Sound-designed voice clips open benchmark for timbre transfer","Human to monster voice: 231,800 clips for AI conversion","New benchmark dataset for AI to synthesize designed vocals"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000275,"raw_usage":{"total_tokens":1452,"prompt_tokens":687,"completion_tokens":765,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":431,"completion_tokens_details":{"reasoning_tokens":683}},"tokens_in":431,"tokens_out":765,"duration_ms":9338,"temperature":1.0,"reasoning_tokens":683,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T08:54:08.701405+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the released test set and compare two conditions: (1) the standard same-source pairs, and (2) pairs where the reference is the designed version of a different source with different content, matched by preset. If the cosine similarity or MOS drops substantially in condition (2), then the benchmark's aligned pairs are not a fair proxy for production use. A second, simpler check is to verify whether the 60 source-seen VCTK test clips are disjoint from the 3,270 VCTK clips used in training; overlapping clips would inflate seen-source scores.","supporting_citations":[],"review_version":1}