{"id":"2462f0e9-6cc5-4650-a513-b92e7d7b64c3","arxiv_id":"2607.16736","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"RealDESED provides a 5,710-clip real-home audio dataset with multi-annotator strong labels and a reviewed test set; a transformer baseline reaches macro PSDS1 0.731.","lead":"RealDESED is a new benchmark of 5,710 short audio recordings made by 652 people in their own homes, each labeled by multiple human annotators for 15 everyday sounds. It matters because most domestic sound-event datasets use synthetic or web-crawled audio, so this gives researchers a more realistic test-bed for hearing systems.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Staged collection and speech ban undermine the 'real-world' representativeness claim; no independent audit verifies authenticity.","rationale":"The reader's weakest assumption is exactly the authenticity of the recordings. This is the most load-bearing premise because the dataset's entire novelty and the transferability of the headline PSDS1 depend on RealDESED representing natural domestic environments. The collection protocol explicitly encourages staged scene production and forbids speech, both of which depart from typical real-home audio. The paper offers no independent verification that the recordings are representative, only self-reported compliance with the protocol. The proposed external transfer test directly checks whether a model trained on RealDESED generalizes to genuinely passive domestic recordings, which would settle the concern. I agree with the reader's conditional verdict because the dataset may still be useful if these limitations are acknowledged and documented, but the central claim needs this caveat.","tokens_in":8678,"tokens_out":8589,"duration_ms":81943,"concrete_test":"Independently collect (or obtain) a set of continuous domestic audio recordings from actual homes without a staged protocol (e.g., from existing smart-home or acoustic-scene corpora). Annotate them with the same 15 classes, run the released baseline checkpoint, and compute macro-PSDS1. If the score is substantially lower than 0.731 (e.g., a drop >0.15), the RealDESED test set is not representative of real homes and the headline number overstates deployment readiness.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.1 describes a collection protocol in which participants were explicitly asked to produce 'eight to ten realistic domestic sound scenes' of 15–35 seconds, to balance target-class coverage, and to exclude speech. This is active staging rather than passive observation: participants decide which actions to perform, when to record, and to avoid speech. The resulting recordings are therefore not representative of natural domestic soundscapes, which typically contain continuous activity, speech, and long periods of silence. The paper provides no independent audit that scenes were not contrived, e.g., isolated actions recorded in quiet rooms. Because the central claim is that RealDESED provides 'realistic variability' and a bridge to 'real-world deployment,' this self-reported authenticity is a load-bearing premise. If the recordings are systematically simpler or cleaner than real homes, a model's PSDS1 of 0.731 may not transfer to actual deployment conditions.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"RealDESED is a new domestic sound event detection benchmark consisting of 5,710 audio recordings collected by 652 participants in their homes, with multi-annotator strong labels for 15 domestic sound classes, a reviewed validation/test split, and rich metadata (device, placement, environment, scene descriptions). The paper describes the data collection, annotation, review, and split procedures, reports dataset statistics including annotation agreement and overlap analysis, and establishes an ATST-F baseline. It further compares annotation aggregation strategies, post-processing methods (median filter vs. cSEBBs), long-form inference settings, and metadata-dependent performance. The strongest reported result is a macro-averaged PSDS1 of 0.731 on the test set using quality-weighted soft labels, cSEBBs post-processing, and a 5-s hop long-form aggregation.","tokens_in":8900,"tokens_out":5149,"duration_ms":52586,"significance":"If the recordings are indeed representative of natural domestic soundscapes, RealDESED would be a valuable community resource: it is large, strongly annotated, multi-annotator, and released with protocols and baselines, thereby addressing a real gap left by synthetic benchmarks such as DESED and by fixed-length web-crawled datasets. The experimental methodology is generally careful: results are averaged over three runs with standard deviations, aggregation and post-processing are compared systematically, and the validation/test labels receive an additional review pass. The main uncertainty is conceptual rather than computational: the 'real-world deployment' contribution rests on the authenticity of the collected scenes, which is self-reported and protocol-driven. The paper should either provide quantitative evidence of naturalness or explicitly narrow its claims; with that clarification, the benchmark has clear value for the SED community.","major_comments":[{"comment":"The central claim that RealDESED comprises 'real-world domestic recordings' and provides a path to 'real-world deployment' is only as strong as the authenticity of the collection protocol. Section 2.1 states that each participant was 'asked to collect eight to ten realistic domestic sound scenes', that the protocol 'encouraged balanced target-class coverage', and that 'recordings containing speech were not permitted'. This is an instructed and explicitly curated setup, not passive observation of everyday home activity. Speech is a ubiquitous component of real domestic soundscapes, and its exclusion, together with the instructions to produce balanced multi-event scenes, likely shifts the event distribution, background conditions, and silence patterns away from natural homes. The manuscript provides no independent audit (e.g., comparisons with unconstrained household recordings, background","section":"§2.1, Abstract, §5"},{"comment":"The reported improvement in annotation agreement after review (Jaccard 0.875 to 0.994; temporal IoU 0.694 to 0.862) is presented as evidence of labeling quality, but the protocol described in §3.3 is not a clean before/after comparison. Because 'approximately 300 low-quality reviews were identified based on remaining disagreement between annotations and reassigned for review', the final numbers may reflect selection or replacement of the hardest annotations rather than the effect of correction alone. The manuscript should specify exactly how the final validation/test annotations are compiled from the reviewed versions and report agreement statistics computed on identical file sets before and after the correction/reassignment step. Without this, the 'reviewed labels' claim is difficult to interpret quantitatively.","section":"§3.3"}],"minor_comments":[{"comment":"The 'Micro' row contains a formatting error: '0.4040.7870.751' should read '0.404 0.787 0.751' (or similar). Please fix the table layout.","section":"Table 2"},{"comment":"The claim 'first large real-world domestic SED benchmark' is unsupported as stated. AudioSet Strong is real-world and includes domestic content; DESED also contains real web-crawled evaluation clips. Please qualify the novelty claim (e.g., 'first large home-collected, multi-annotator domestic SED benchmark') or add a systematic comparison with prior real-world domestic datasets.","section":"§1"},{"comment":"The cSEBBs hyperparameters and the long-form inference settings (triangular floor, hop size, aggregation) are selected on the validation set, but the selected cSEBBs hyperparameter values are not reported. For reproducibility, list them (or point to the released configuration) along with the hyperparameter search ranges.","section":"§4.3–4.4"},{"comment":"The split by collector ID assumes each collector records in a single domestic environment. If collectors could record in multiple rooms or homes, this assumption is not guaranteed. Please state whether the environment metadata was used to verify that all recordings from one collector share the same environment.","section":"§2.4"}],"recommendation":"major_revision","confidential_remarks":"The representativeness issue is the main risk. The dataset and baseline work are otherwise solid and within DCASE scope. I would be comfortable with acceptance after the authors either provide one independent characterization of naturalness (e.g., background/silence statistics or a comparison with a less constrained household corpus) or explicitly revise the 'real-world' claims and discuss the speech ban as a limitation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"RealDESED is a genuinely useful benchmark: 5,710 real domestic recordings, multi-annotator strong labels, reviewed val/test splits, and rich metadata. That's new — DESED is synthetic, MAESTRO Real is weak-label only. The paper reports construction in good detail, including pre/post-review agreement (Jaccard 0.875 to 0.994, temporal IoU 0.694 to 0.862), and runs sensible baseline studies on aggregation, post-processing, long-form inference, and metadata. The ATST-F baseline reaching 0.731 PSDS1-M is a solid starting point, not overclaimed.\n\nThe soft spots are real but not fatal. The 'real-world' claim is based on a collection protocol that asked participants to stage realistic scenes, balance classes, and avoid speech. That is not passive observation of daily home life, and there's no independent audit of how naturally the recordings were made. The speech ban in particular is a departure from real homes. Still, the recordings are from actual homes with varied devices, placements, and backgrounds, so the dataset is far closer to practice than synthetic DESED. The stress-test note overstates the problem; the paper's title and abstract say 'domestic sound event detection,' not full soundscape capture.\n\nA smaller issue: Section 2.3 says about 300 low-quality reviews were reassigned based on remaining disagreement. That's a post hoc selection step; the paper should specify criteria and counts, because it affects the reported agreement numbers.\n\nAlso, the 0.731 number comes after validation-set tuning of alpha, median filter, cSEBBs, and hop size. They report standard deviations over three runs, so it's honest, but it's tuned performance, not out-of-the-box.\n\nOverall, the paper is a solid data contribution. The collection design, annotation pipeline, and benchmark experiments are described well enough for replication. I'd send it to review and expect it to be accepted after minor revisions. The authors should either soften the 'real-world' framing or sample-check a subset of recordings for authenticity, and document the review reassignment.","headline":"RealDESED is a solid new real-home SED benchmark; the 'real-world' claim is slightly overstated but the resource is valuable and honestly reported.","tokens_in":9357,"tokens_out":3501,"would_cite":true,"duration_ms":34428,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"RealDESED is a 5,710-clip benchmark of real home recordings with multi-annotator temporal labels; its transformer baseline reaches a macro-averaged PSDS1 of 0.731, making it a realistic alternative to synthetic domestic SED benchmarks.","keywords":["sound event detection","real-world dataset","domestic environment","multi-annotator labels","temporal annotation","benchmark","transformer baseline","PSDS"],"falsifier":"Audit a random sample of RealDESED test recordings with an independent, opportunistic home-recording campaign matched by environment and time of day, and compare background sound levels, non-target event rates, and event co-occurrence statistics; if the benchmark recordings are systematically quieter or more event-dense than the independent sample, the real-world claim and the 0.731 PSDS1 would not transfer to deployment.","tokens_in":8586,"feed_emoji":"🏠","tokens_out":4871,"duration_ms":46362,"temperature":0.7,"pith_summary":"This paper introduces RealDESED, a benchmark of 5,710 domestic recordings made by 652 people in their own homes, with temporally precise labels for 15 everyday sound classes. Unlike popular SED datasets that rely on synthetic soundscapes or short web-crawled clips, each RealDESED recording is 15–35 seconds long, captured on consumer devices, and annotated by multiple independent annotators, with validation and test labels reviewed for quality. The paper argues that this combination makes RealDESED a more realistic testbed for developing sound event detection systems that work in actual homes. As evidence, it trains a transformer baseline and reports a macro-averaged PSDS1 of 0.731 on the test set, and it shows that annotation aggregation, post-processing, and long-form inference all noticeably affect that number. If the dataset is as natural as claimed, it gives the field a public benchmark that better predicts deployment performance than synthetic or fixed-length alternatives.","feed_headline":"A 5,710-clip real-home benchmark sets sound event detection at 0.731","feed_subtitle":"Grounded in natural homes with multiple annotators, the dataset tests whether detection systems work outside synthetic soundscapes.","key_machinery":"The central object is the dataset itself: 5,710 recordings (about 38 hours) from 652 collectors, each 15–35 seconds long, covering 15 domestic sound classes, with an average of 2.45 independent annotators per file and reviewed labels for the full validation and test splits. Two mechanisms carry the benchmark's value: multi-annotator aggregation, where soft frame-level targets are formed by averaging per-annotator labels with weights derived from pairwise Dice agreement, turning annotator disagreement into a training signal rather than noise; and temporal post-processing with sound event bounding boxes (cSEBBs), which converts raw frame predictions into event-level detections. The dataset als","core_discovery":"The paper's central claim is that a benchmark built from naturally occurring home recordings, with multiple independent annotators and a reviewed evaluation split, is a viable and more realistic alternative to existing domestic SED benchmarks. On RealDESED, a transformer pre-trained on strongly labeled audio and fine-tuned on the new data reaches 0.731 macro-averaged PSDS1 on the test set. The paper further shows that soft labels weighted by annotator quality outperform simple majority or union aggregation, that temporal post-processing with sound event bounding boxes is essential, and that long-form inference with overlapping windows and triangular weighting improves scores. It also documen","pith_inferences":["Going beyond the paper: a natural next test is to use RealDESED's metadata to train a device-robust model, for example via adversarial domain adaptation or device dropout; the paper shows gaps exist but does not attempt to close them.","Going beyond the paper: the weighted-soft-label aggregation recipe is likely to transfer to other audio tasks with noisy or subjective labels, not just sound event detection.","Going beyond the paper: the real-world claim could be stress-tested by comparing ambient noise levels and event co-occurrence rates in RealDESED against independently collected passive home recordings; the paper does not report such a comparison.","Going beyond the paper: because collectors knew they were recording for a benchmark, there may be a self-selection bias toward quiet or staged scenes, so the reported PSDS1 should be treated as an upper bound for fully opportunistic deployment until independent validation."],"forward_implications":["If RealDESED is representative of real homes, SED systems trained on it should transfer to deployment better than those trained on synthetic mixes or fixed 10-second web clips.","The finding that annotator-quality-weighted soft labels match or beat human review suggests future dataset builders can reduce expensive review effort by collecting more cheap annotations and weighting them automatically.","The metadata-driven performance gaps across devices, placements, and environments imply that device-robust and domain-adaptive SED should be evaluated on RealDESED to see whether proposed gains actually generalize.","The strong effect of long-form inference with overlapping windows indicates that scores on short clips may understate performance on longer, real-world recordings.","Reviewed validation and test labels, combined with multi-annotator coverage, make RealDESED a reliable common ground for comparing SED systems in the domestic domain."],"fun_headline_variants":["Real-home audio benchmark: 5,710 clips, 0.731 PSDS1","New benchmark for SED in real homes: 5,710 clips, 0.731 PSDS1","5,710 real domestic recordings set SED baseline at 0.731","Multi-annotator home recordings benchmark hits 0.731 PSDS1","Real-world SED benchmark from 5,710 home clips achieves 0.731"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The core premise is that the course-collected recordings are natural domestic scenes and not noticeably staged or simplified; no independent audit verifies that the homes, background conditions, and event sequences are representative of real everyday life.","fun_headline_variants_meta":{"raw":{"variants":["Real-home audio benchmark: 5,710 clips, 0.731 PSDS1","New benchmark for SED in real homes: 5,710 clips, 0.731 PSDS1","5,710 real domestic recordings set SED baseline at 0.731","Multi-annotator home recordings benchmark hits 0.731 PSDS1","Real-world SED benchmark from 5,710 home clips achieves 0.731"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000378,"raw_usage":{"total_tokens":1860,"prompt_tokens":770,"completion_tokens":1090,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":514,"completion_tokens_details":{"reasoning_tokens":977}},"tokens_in":514,"tokens_out":1090,"duration_ms":9392,"temperature":1.0,"reasoning_tokens":977,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T20:04:30.827761+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Audit a random sample of RealDESED test recordings with an independent, opportunistic home-recording campaign matched by environment and time of day, and compare background sound levels, non-target event rates, and event co-occurrence statistics; if the benchmark recordings are systematically quieter or more event-dense than the independent sample, the real-world claim and the 0.731 PSDS1 would not transfer to deployment.","supporting_citations":[],"review_version":1}