{"id":"9c8cc1b1-687c-4a91-b530-ce8f40c99f33","arxiv_id":"2511.20544","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"New York Smells is an in-the-wild dataset of 7,000 co-captured image–e-nose smell pairs covering 3,500 objects, and contrastive vision-smell training on it yields olfactory representations that outperform hand-crafted features.","lead":"Researchers collected 7,000 picture-and-smell pairs from 3,500 everyday objects across New York City using an electronic nose mounted with cameras, creating a natural-world dataset for machine olfaction. The paper shows that pairing smell with images lets AI learn smell features that beat hand-crafted olfaction features on retrieval, scene, and grass-species tasks.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Baseline ambient block may be doing the work: raw-signal gains over smellprint could come from scene-level background odor, not target-object smell; an ablation isolating sample-minus-baseline is needed.","rationale":"The reader's weakest assumption and my concern coincide: the dataset's collection protocol does not separate the target object's odor from the ambient environment, and the raw signal fed to the model explicitly includes a baseline period. The paper's central empirical claim—that visual supervision enables olfactory representations that outperform smellprints—does not require that claim to be false to be vulnerable; it only requires that the observed gains could be explained by scene-level or session-level confounds. The near-ceiling scene-recognition numbers (99.5% for raw scratch vs 42.2% for smellprint) strongly suggest the baseline/scene cue is present and easy to exploit. A baseline-only ablation would directly test whether the raw-signal advantage survives when the ambient block is removed or when the sample stage is normalized against baseline. This is an addressable experimental omission rather than a logical inconsistency, so the appropriate verdict remains CONDITIONAL rather than ACCEPT or REJECT. I agree with the reader that the paper's artifacts and labels also warrant attention, but the baseline confound is the most load-bearing because it threatens the central mechanism of the representation-learning claim.","tokens_in":14368,"tokens_out":4991,"duration_ms":54588,"concrete_test":"Train the same COIP encoders on (a) baseline-only (first 14 timesteps) and (b) per-sensor baseline-subtracted sample stage (sample block minus mean baseline), then evaluate Table 1 retrieval and Table 2 scene/object/material probes. If baseline-only reproduces a large share of the full-signal results—or if (b) loses the reported margin over smellprint—the central object-level claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is that the 28×32 raw signal isolates the targeted object's odor. The Sec. 3.1 protocol concatenates a 10-s ambient baseline (purge inlet) with two 10-s snout samples. Because the snout is open to the same environment, the sample stage is a mixture of object-emitted VOCs and ambient air; for low-odor objects (metal, stone, plastic) it may be indistinguishable from baseline. The smellprint baseline (Eq. 3–5) is explicitly designed to remove this ambient component via relative response, which is why the paper frames raw > smellprint as richer information. But the raw encoder can instead exploit the baseline block to encode the scene/session, and the contrastive loss (Eq. 1) will then align images with scene-level background odor. This is strongly consistent with Table 2's near-ceiling scene accuracy (99.5% raw scratch vs 42.2% smellprint), and it could inflate retrieval and object/material numbers without any object-level smell learning. The paper does not ablate the baseline, report within-scene vs cross-scene retrieval, or control for co-occurring background in the contrastive pairs.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces New York Smells, a claimed in-the-wild paired vision-olfaction dataset of 7,000 smell-image pairs from 3,500 objects, with 70x more objects than prior olfactory datasets. The authors mount a Cyranose 320 e-nose with a camera, record a 10-second ambient baseline followed by two 10-second snout samples per object, and concatenate them into a 28x32 raw signal. Object and material labels are generated by GPT-4o from the accompanying images. Using a contrastive learning objective (COIP, Eq. 1-2), they train smell and image encoders and evaluate on three tasks: smell-to-image retrieval, scene/object/material classification from smell, and grass-species discrimination. They report that raw-signal encoders outperform hand-crafted smellprint features across all tasks.","tokens_in":14708,"tokens_out":2332,"duration_ms":23437,"significance":"If the dataset and results hold, this is a substantial community resource: it is much larger and more naturalistic than existing e-nose datasets, it is the first to pair in-the-wild olfaction with images, and it demonstrates a plausible route to self-supervised olfactory representations. The authors commit to releasing data and code, and the dataset collection protocol is clearly described. However, the central scientific claim—that the learned representations capture object-level odor rather than scene/background odor—is not yet established because the raw signal includes the ambient baseline block and no ablation isolates the target-object contribution. The evaluation also lacks error bars, and the labels are machine-generated without verification. These issues are load-bearing for the claimed advances.","major_comments":[{"comment":"The raw olfactory signal is defined as the concatenation of a 10-second ambient baseline (purge inlet) and two 10-second snout samples, yielding a 28x32 matrix. The contrastive encoder thus has direct access to the ambient background block. The near-ceiling scene accuracy of the raw CNN (99.5% scratch) versus the smellprint (42.2%) is consistent with the hypothesis that the model aligns images with scene-level background odor rather than with object-emitted VOCs. This is a load-bearing confound for the retrieval and object/material results. Please add ablations: (a) train raw encoders on sample-minus-baseline, on baseline-only, and on sample-only inputs; (b) report retrieval and classification metrics split by whether the distractor/query share the same scene; and (c) quantify the signal-to-noise ratio between the sample and baseline windows for low-odor materials (e.g., metal, stone, pl","section":"Sec. 3.1, Fig. 5 and Sec. 5.2, Table 2"},{"comment":"Object and material labels are generated exclusively by GPT-4o from the visual stream, with no reported human verification or agreement measure. These labels define the held-out evaluation sets for Tables 2 and 3. If the VLM labels are noisy or biased (e.g., guessing 'plant shrub' from visual context rather than the probed object), the classification accuracies are not a valid measure of olfactory discriminability. Please report a human-verified subset (even a few hundred samples), per-category label reliability, or an inter-annotator agreement between GPT-4o and human raters. Also report how many samples were labeled 'unlabeled' and how these are handled in evaluation.","section":"Sec. 3.1 'Labeling the dataset' and Sec. 5.2"},{"comment":"No error bars, confidence intervals, or significance tests are reported anywhere. The retrieval test set is N=933; the fine-grained grass task uses only 42 held-out samples. Differences between architectures (e.g., CNN vs Transformer retrieval recall @20: 32.6 vs 43.1) and between raw and smellprint may be within noise, especially given the small fine-grained test set. Please run multiple training seeds or bootstrap over test samples and report mean +/- std (or CIs). This is required to support the quantitative claims of superiority over hand-crafted features.","section":"Tables 1-3 and Sec. 5.1"},{"comment":"The distractor sampling procedure is underspecified. The text says 'we sample a distractor set of images' but does not state whether distractors are drawn from the whole test set, whether same-scene images are excluded, or how N=933 is derived. If distractors include images from the same scene, retrieval can succeed by matching scene-level background, which would inflate recall. Specify the distractor distribution and, ideally, report retrieval after removing same-scene distractors and after ablating the baseline block.","section":"Sec. 5.1 retrieval protocol"}],"minor_comments":[{"comment":"Typographical errors in labels: 'Planets Shrub' should be 'Plants Shrub', 'Treet Parts' should be 'Tree Parts'. The color-coding description in Sec. 7.2 says blue for objects and green for materials, but Fig. 8 caption may be inconsistent with the actual rendering; please check.","section":"Fig. 8 and Sec. 7.2"},{"comment":"The smellprint definition is clear, but the Savitzky-Golay filter parameters (window length w, polynomial order p) are never specified. Since the baseline comparison depends on these parameters, please provide the exact values used in all experiments.","section":"Sec. 4.2 and Eq. 5"},{"comment":"The dataset split description says 'uniformly split' but also requires both samples of an object to be in the same split. Clarify whether the split is by object or by scene/session, and report the number of distinct scenes and objects in train vs test.","section":"Sec. 3.1"},{"comment":"Reference [15] is cited as 'concurrent, unpublished work' and appears twice in the related work and once in the introduction. If it has been published or updated, please use the final version and avoid repeating the same citation in adjacent sentences.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The dataset itself is potentially valuable, but the baseline confound is central. I would like the editor to ensure the revised version includes the baseline ablation and label verification; without those, the paper's headline claims are not yet credible. Also note that the paper's self-identified 'unpublished' concurrent work [15] deserves a careful check for overlap, though this is not a basis for rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. This is the first large in-the-wild paired olfaction-vision dataset I know of, and it's a real contribution on scale and diversity alone: 7,000 smell-image pairs from 3,500 objects, roughly 70× the concurrent SmellNet set. The paper also builds three sensible benchmark tasks on top of it, and the qualitative retrievals (book→book, foliage→foliage) suggest the learned representations have some cross-modal structure. So the dataset itself deserves attention.\n\nThe problem is the paper's headline result. The authors claim that contrastive learning on raw e-nose signals beats the standard hand-crafted smellprint. But the raw signal is a concatenation of a 10-second ambient baseline and two 10-second snout samples, and the smellprint is explicitly normalized relative to that same baseline. That means the raw encoder can encode the scene/session directly from the baseline block, while the smellprint cannot. Given that the raw model hits 99.5% scene accuracy while the smellprint gets 42.2%, it looks like the raw model is using room-level background smell. That confound is load-bearing: without an ablation that trains on only the sample block (or sample-minus-baseline), the claim that raw signals are richer for object smell is not established. The paper doesn't report such an ablation, nor within-scene vs cross-scene retrieval.\n\nThe other soft spots are minor. Labels are generated by GPT-4o with no reported human verification; for a dataset paper that's acceptable if the labels are released and the noise is characterized, but I'd want a small human-evaluation subset. No error bars in any table, which makes it hard to judge how robust the margins are. And the data and code aren't out yet—the paper promises them, but the value of this work is almost entirely in the artifact.\n\nI think the right call is to send this to peer review. The dataset is novel and useful, and the benchmark tasks are a good starting point for machine olfaction. But I'd make acceptance conditional on a baseline ablation and label-quality check. If the raw-signal advantage evaporates once you remove the ambient baseline, then the paper becomes a solid dataset paper, not a big representation-learning claim. Either way, it's worth a serious referee.","headline":"Genuinely novel in-the-wild smell-vision dataset, but the raw-vs-smellprint claim needs a baseline ablation before I'd trust it.","tokens_in":15173,"tokens_out":4288,"would_cite":true,"duration_ms":38168,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new 7,000-pair dataset links smell to sight and shows that vision can teach machines to recognize odors in the wild.","keywords":["olfaction","electronic nose","multimodal learning","contrastive learning","dataset","smell-to-image retrieval","representation learning","in-the-wild sensing"],"falsifier":"Query the trained model with olfactory recordings taken with the snout sealed or pointed at empty air in the same scenes; if retrieval accuracy remains comparable to the results with real object samples, the embedding is exploiting scene background rather than object odor. Alternatively, swapping the baseline segment for one from a different scene while keeping the sample segment should shift predictions if the object odor is the true signal.","tokens_in":14299,"feed_emoji":"👃","tokens_out":3721,"duration_ms":36649,"temperature":0.7,"pith_summary":"This paper attempts to establish that co-located vision can serve as a teacher for olfaction. It introduces a dataset of 7,000 smell-image pairs from 3,500 objects recorded in natural indoor and outdoor settings with a handheld 32-sensor electronic nose and a camera. Using contrastive learning on the raw sensor signals, the authors show that the learned olfactory representations outperform the widely used hand-crafted smellprint feature across smell-to-image retrieval, scene/object/material recognition from smell alone, and fine-grained grass species discrimination. If the claim holds, machine olfaction can move beyond controlled laboratories into everyday environments, and training can rely on synchronized sight instead of costly chemical analyses.","feed_headline":"Vision teaches machines to recognize smells in the wild","feed_subtitle":"A new 7,000-pair dataset of sight and scent beats hand-crafted smell features on retrieval and classification.","key_machinery":"The central object is the raw olfactory signal matrix from the Cyranose 320 electronic nose: 10 seconds of ambient baseline followed by two 10-second samples of the target object, concatenated over 32 sensors into a 28×32 time-series. The argument is carried by a contrastive learning objective (termed COIP) that aligns this signal with synchronized images, learning a shared smell-sight embedding. The baseline comparator is the smellprint, a hand-crafted 32-dimensional feature computed as relative sensor response (sample peak minus baseline, divided by baseline) after Savitzky–Golay filtering.","core_discovery":"The paper's central claim is that paired visual and olfactory signals, captured together in the wild, enable cross-modal olfactory representation learning that beats hand-crafted features. Concretely, training a contrastive joint embedding between raw e-nose time series and synchronized images yields a smell encoder that substantially outperforms the standard smellprint descriptor on three benchmark tasks: retrieving the matching image from a smell query, recognizing scenes, objects, and materials from smell alone, and discriminating between two co-located grass species. The authors attribute this to the richer information present in the raw 28×32 sensor matrix compared to the 32-dimensional","pith_inferences":["If ambient scene odor dominates the 28×32 signal, the contrastive loss may align images with background smell rather than the probed object; a direct test would be to query with samples of empty air or to ablate the baseline stage.","The high scene-recognition accuracy could partly reflect this scene-level confound, so downstream users should report object classification with environmental variation held out.","The grass discrimination result is the strongest evidence for genuine object-level olfactory signal, since the two species are co-located; extending the benchmark to multiple co-located objects per scene would further validate the object-level claim.","A transfer test to a different e-nose or sensor array would clarify whether the learned embedding captures generic chemical properties or device-specific artifacts."],"forward_implications":["Smell-to-image retrieval becomes feasible in the wild: raw-signal encoders reach roughly 43% recall@20 compared to about 6% for the smellprint baseline.","Scene recognition from smell alone reaches around 99.5% accuracy with a CNN, showing that ambient olfactory context is strongly encoded in the raw signal.","Learned representations beat hand-crafted smellprints across all three benchmark tasks, including fine-grained discrimination between two grass species coexisting on the same lawn.","The dataset's scale—about 70 times more distinct objects than existing lab-collected olfaction datasets—opens the door to data-driven olfaction research outside controlled settings.","Visual supervision supplies a label-free training signal for olfaction, avoiding the need for costly perceptual descriptors or molecular analyses."],"fun_headline_variants":["AI learns smells from sights: 7,000 pairs beat hand-crafted","Sight-scent pairing beats traditional smell features","New smell dataset leverages vision to teach machines","Vision boosts machine olfaction beyond smellprint descriptors","Cross-modal smell learning: images boost scent recognition"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The 10-second sample stage, recorded after the ambient baseline, is assumed to reflect the target object's odor rather than the surrounding scene, and the co-located images are assumed to correspond to that same object; if ambient background dominates, the contrastive learning mostly aligns images with scene-level smell.","fun_headline_variants_meta":{"raw":{"variants":["AI learns smells from sights: 7,000 pairs beat hand-crafted","Sight-scent pairing beats traditional smell features","New smell dataset leverages vision to teach machines","Vision boosts machine olfaction beyond smellprint descriptors","Cross-modal smell learning: images boost scent recognition"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000176,"raw_usage":{"total_tokens":1078,"prompt_tokens":648,"completion_tokens":430,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":392,"completion_tokens_details":{"reasoning_tokens":353}},"tokens_in":392,"tokens_out":430,"duration_ms":4764,"temperature":1.0,"reasoning_tokens":353,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T06:43:14.357400+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Query the trained model with olfactory recordings taken with the snout sealed or pointed at empty air in the same scenes; if retrieval accuracy remains comparable to the results with real object samples, the embedding is exploiting scene background rather than object odor. Alternatively, swapping the baseline segment for one from a different scene while keeping the sample segment should shift predictions if the object odor is the true signal.","supporting_citations":[],"review_version":1}