{"id":"b8e8a4b7-9d93-4129-8013-310019c3c24c","arxiv_id":"2506.00398","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A call to add olfaction, with standardized data and benchmarks, to the list of core modalities that embodied AI systems should sense and reason about.","lead":"This position paper argues that artificial intelligence research neglects the sense of smell, and that standardizing olfactory data and benchmarks is essential for building embodied AI systems. It identifies five systemic gaps, from unresolved smell theories to missing datasets, and proposes a research agenda including new benchmark tasks for olfaction in robotics and multimodal AI.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The §2.1 parallel-progress claim is load-bearing: if the right olfactory data abstraction depends on the unresolved STO/VTO coding question, standardization now could encode a wrong representation.","rationale":"The reader identified this same assumption, and I agree it is the weakest point. The paper's empirical evidence for neglect (Fig. 1, title-only queries, no linked code) is thinner, but the central claim does not hinge on exact counts. The parallel-progress premise is load-bearing because the paper's first call to action is to define a data standard now; if the standard must encode a theory of olfaction, the call is premature. The analogy to JPEG and aircraft standards is not sufficient: both analogies concern standards that are either downstream of settled physics or do not attempt to represent the sensory signal in a canonical way. The paper deserves credit for a concrete benchmark taxonomy, explicit treatment of sensor heterogeneity and sub-perceptual signals, and candid discussion of dataset limitations. The concern is not that the position is wrong, but that it has an unproven keystone; a conditional acceptance requiring the authors to either demonstrate theory-neutrality or propose a staged standardization strategy would be appropriate. This does not change the reader's conditional verdict.","tokens_in":21886,"tokens_out":5476,"duration_ms":57852,"concrete_test":"Run a transfer experiment: train an odor classifier on the Vergara MOx gas-sensor dataset (§2.4) for the 10-gas identification task, then test on data from a sensor with a different transduction principle (e.g., optical or vibrational gas sensor) for the same gases, applying only the calibration/normalization specified in the proposed standard. If cross-principle accuracy drops to chance while within-MOx transfer remains high, the standardized representation is sensor-theory-bound and the §2.1 parallel-progress assumption fails. If no suitable optical dataset exists, simulate VTO-style sensor responses from molecular vibrational spectra and repeat the test.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.1 asks how to standardize a modality without scientific consensus and answers that \"scientific progress and standardization process can move forward in parallel,\" citing JPEG/PNG and aircraft standards. This analogy is weaker than the paper needs. JPEG/PNG did not require understanding visual cortex, but they presupposed settled front-end physics: trichromatic sampling and RGB/YUV color spaces. Aircraft airworthiness standards are functional safety rules, not sensory-data representations, so they do not address the wrong-abstraction risk. In olfaction the analog of a pixel is contested: MOx arrays, optical/vibrational sensors (VTO), and receptor-activation models (STO) imply different feature spaces. Section 2.4 recommends collecting raw digitized sensor data, but raw data from one transducer is device-specific, not a modality-level standard. If the correct olfactory representation depends on the unresolved receptor coding question, benchmarks and datasets built now may encode a wrong abstraction, and the title claim that standardization is essential is not established. The paper does not show how a standard can be theory-neutral; this is the central premise that would need defense.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This position paper argues that olfaction has been systematically neglected in AI research, not because it is irrelevant to embodied intelligence but because of five structural gaps: unresolved competing theories of olfactory coding (STO vs. VTO), heterogeneous sensor technologies and the absence of a data standard, the subjectivity of semantic labels, a lack of large-scale peer-reviewed datasets, and the absence of AI-oriented benchmarks. After reviewing bandwidth estimates and sensor properties, the paper proposes a parallel-progress model in which standardization and scientific understanding of olfaction develop together, and it sketches three classes of benchmark tasks (foundational perception, static-scene olfaction, dynamic-scene olfaction) plus ethical considerations. The central recommendation is a community-wide investment in olfactory data standards, datasets, and benchmarks comparable to those available for vision and language.","tokens_in":22023,"tokens_out":12501,"duration_ms":118506,"significance":"This is a timely and clearly written position piece that identifies a real gap in embodied AI research and proposes a concrete set of actions. If the main argument holds, it could help refocus community resources toward olfactory datasets, benchmarks, and standards, in the same way that earlier position pieces catalyzed investment in other modalities. The paper's strengths include: a literature-grounded enumeration of five structural gaps; a transparent first-order bandwidth model with explicit anatomical assumptions; reproducible publication-count queries with code and a stated limitation (Supplementary A); and a detailed taxonomy of benchmark tasks that goes beyond slogans. The central claim is not circular: the argument does not depend on the authors' own prior results, and self-citations appear only as supporting examples. The main weakness is the unexamined feasibility of standardizing olfaction before the receptor-coding question is resolved, which is the focus of the major comments below.","major_comments":[{"comment":"The paper's load-bearing premise — 'scientific progress and standardization process can move forward in parallel' — is not established by the analogies given. JPEG/PNG standardization did not require understanding the visual cortex, but it did presuppose a settled front-end physical representation: images are sampled spatially and trichromatically (RGB/YUV), a convention that was not in dispute. In olfaction, the analogous 'pixel' is exactly what is contested: Section 2.2 lists MOx, electrochemical, optical, acoustic, and carbon-nanotube sensors, and Sections 2.1 and 2.2 describe the unresolved STO/VTO debate. Aircraft airworthiness standards are functional safety criteria (e.g., margins, inspection intervals), not sensory-data representations; they do not address the risk that a standard encodes a wrong abstraction. The paper's own remedy in Section 2.4 (collecting 'raw, digitized sensor data') yields device-specific time series, not a modality-level standard, unless an additional theory-neutral layer (e.g., calibration metadata, transducer-agnostic event descriptors, and task-specific evaluation protocols) is specified. Without such a layer, the paper has not shown that a standard can be both built now and robust to the eventual resolution of the receptor-coding question. If the right representation depends on the STO/VTO outcome, datasets and benchmarks built on the wrong representation would partly waste the investment the paper calls for. The authors should either specify a theory-neutral standard layer or soften the claim to 'standardize multiple candidate representations in parallel,' which would change the title's assertion.","section":"Section 2.1 (parallel-progress claim); Section 2.4 (raw sensor data)"},{"comment":"The paper never defines what 'standard' it is calling for. Gap 2 in Section 1 lists both 'standardized data representation' and 'hardware specification,' while Section 2.4 recommends 'standardized protocols for sensor calibration, data acquisition, and comprehensive annotation' and Section 2.5 calls for benchmark tasks. These are different objects: a data format, a hardware interface, an annotation protocol, and an evaluation metric are not interchangeable, and a single 'olfactory data standard' cannot serve all of them simultaneously. In particular, the recommendation to collect 'raw, digitized sensor data' is a per-instrument convention, not a modality-level standard: raw MOx conductance values and raw optical spectra share no common vector space. The authors should specify which layer they propose to standardize (transducer output encoding, calibration metadata, task-level evaluation, or all three) and how that layer interacts with the unresolved representation question raised in Section 2.1. Without this specification, the central call to action is ambiguous and difficult to evaluate.","section":"Section 2.4 (object of standardization)"}],"minor_comments":[{"comment":"The bandwidth calculation does not match the stated assumptions. With 400 ORN types, 2 glomeruli each, 25 mitral cells per glomerulus, 4 bits per cell, and 1 Hz sniffing, the rate is 400 × 2 × 25 × 4 = 80,000 bit/s = 10 kB/s, not '>5 kB/s' as printed. Use the correct value or revise the assumptions; the qualitative claim that olfaction is a high-bandwidth channel is unaffected.","section":"Section 2.2 (bandwidth)"},{"comment":"The publication-count comparison is not apples-to-apples: olfaction is counted by title keywords, while the comparator fields are counted by arXiv category. This likely undercounts olfaction (e.g., papers using 'electronic nose' or 'gas sensor array' without the listed keywords). State in the caption and text that the reported ratios are lower bounds, not precise measurements.","section":"Figure 1 and Section A (publication counts)"},{"comment":"The claim that 'Emerging olfaction-vision-language models (OVLMs)' are exemplified by references [61,133] relies on a commercial mobile application rather than a peer-reviewed model; replace these citations with a scholarly reference or rephrase as 'early commercial demonstrations.'","section":"Section 3 (OVLM citations)"},{"comment":"The text says 'Luca Turin re-popularized the idea in 2001' but the cited reference [145] is a 2015 PNAS correspondence; add the original 1996 Nature paper or a 2001 source to support the date.","section":"Section 2.1 (Turin citation)"},{"comment":"The paragraph on encoding claims olfaction 'forms a discrete space' and later emphasizes 'episodic, rather than continuous, sampling'; these two notions (discrete molecular-space encoding vs. temporal sparsity) should be explicitly separated, because a standard must handle both dimensions.","section":"Section 2.2 (encoding paragraph)"},{"comment":"Fix inconsistent spacing in 'W A V', 'A VI', 'V on Neumann', and 'ArXiV' (Sections 1 and 2.2, Figure 1).","section":"Throughout (typos)"}],"recommendation":"major_revision","confidential_remarks":"The paper is better suited to a position/opinion track than a technical-results journal. The major obstacle is the parallel-progress premise; if the authors can specify a theory-neutral standard layer (e.g., raw transducer outputs plus calibration metadata plus task-specific evaluation wrappers) and show it is robust to STO/VTO outcomes, I would support acceptance after revision. The publication-count figure should be presented as a lower bound. Also note the self-citation to a commercial app in Section 3; it would be safer to cite a peer-reviewed source."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: this is a serious agenda-setting position paper for a genuinely neglected modality. It argues that olfaction is excluded from embodied AI not because it doesn't matter but because of five structural gaps, and it proposes a concrete benchmark taxonomy. If you work on multimodal or embodied systems, this is worth a read even if you disagree with the conclusion.\n\nCredit where earned. The five-gap framing is clean and well supported by citations. The proposed taxonomy—foundational perception, static scenes, dynamic scenes—is a sensible way to organize future work, and the ethics section is not an afterthought. The bandwidth calculation is a nice back-of-envelope: 400 ORN types × 2 glomeruli × 25 mitral cells × 4 bits ≈ 10 kB/s, and the paper's \">5 kB/s\" is actually conservative given its own assumptions. The publication gap figure is striking, though the analysis is title-only and the query code isn't linked in the text—easy to fix, and worth doing.\n\nThe soft spot is Section 2.1. The claim that standardization can proceed in parallel with unresolved science is load-bearing, and the JPEG/aircraft analogies don't fully carry it. JPEG presupposed settled color physics; aircraft standards are about functional safety, not data representation. In olfaction, the 'pixel' is contested: MOx arrays, optical sensors, and vibrational/receptor models imply different feature spaces. The paper partly answers this by recommending raw digitized sensor data plus metadata, but it doesn't argue that such a standard can be theory-neutral. If the receptor coding question resolves in a way that makes today's benchmark representation wrong, early investment could be partly wasted. This doesn't sink the paper—it's a position paper, and the position is still defensible—but the authors should acknowledge the risk explicitly and say more about what a theory-neutral standard would look like.\n\nNet: a well-argued, well-cited agenda that deserves serious peer review. I'd send it to referees with a request to push on Section 2.1 and the data-query reproducibility. For reading group, yes—it will generate good debate.","headline":"A well-cited, coherent agenda paper that deserves peer review; the main soft spot is the Section 2.1 claim that standardization can outpace scientific consensus, which the authors should defend more explicitly.","tokens_in":22594,"tokens_out":2554,"would_cite":true,"duration_ms":24594,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This position paper argues that olfaction is missing from AI because of five structural, fixable gaps—no settled smell science, no data standard, no objective labels, scarce datasets, and no benchmarks—and that closing them should make…","keywords":["artificial olfaction","embodied AI","olfactory data standard","machine olfaction benchmarks","olfactory datasets","multimodal perception","neuromorphic olfaction","AI ethics"],"falsifier":"A concrete test: build the proposed standardized olfactory dataset from raw sensor recordings with calibrated molecular ground truth, train equivalent models on it and on today's ad hoc small datasets, and compare both on the same real-world scent-source localization task; if the standardized data yields no measurable advantage in accuracy or generalization, the claim that missing standards are the key bottleneck would be weakened.","tokens_in":21644,"feed_emoji":"👃","tokens_out":10067,"duration_ms":98698,"temperature":0.7,"pith_summary":"This paper argues that AI's neglect of the sense of smell is not a sign that smell is unimportant, but the result of five structural gaps that can be fixed. Those gaps are the absence of settled science on how odor perception works, the lack of a standard format for olfactory data, the difficulty of labeling signals that humans may not consciously detect, the scarcity of large datasets, and the absence of AI benchmarks for olfaction. The authors' position is that standardization can proceed in parallel with open scientific debates, just as image and audio formats were standardized without a complete model of perception. If they are right, closing these gaps would let machine olfaction grow at the pace vision and language did, and would make smell a core sensory input for embodied and multimodal AI. The paper's proposed first step is the creation of an olfactory data standard built from raw, digitized sensor recordings of molecular signatures, collected under controlled conditions and paired with consensus-based human annotations.","feed_headline":"AI needs a smell standard to become truly embodied","feed_subtitle":"Five fixable gaps—unsettled science, no data format, few datasets, no benchmarks—keep machines from smelling.","key_machinery":"The load-bearing mechanism is the proposed olfactory data standard, modeled on the way image and audio standards turned raw physical measurements into shared digital formats. Receptor arrays are treated as the analogue of pixels, signal intensity as the analogue of dynamic range, and the standard itself as the substrate that makes datasets from different labs compatible. Around that substrate the paper organizes its case into five named gaps—scientific understanding, data standard, objective annotation, datasets, and benchmarks—with the standard and benchmark suite doing the causal work of catalyzing progress. A supporting mechanism is the paper's bandwidth calculation, which ranks olfaction as a high-throughput sense and motivates event-based, neuromorphic processing for the sparse, intermittent plumes that carry odor stimuli in natural environments.","core_discovery":"The central claim is that the exclusion of olfaction from AI architectures is a correctable infrastructure failure rather than a sign of irrelevance. The paper asserts that objective progress can be made without first resolving the debate between shape-based and vibration-based theories of odor detection: raw, digitized sensor data capturing molecular signatures can serve as a common substrate, with human semantic labels added as layered, consensus-based annotations. It estimates human olfactory bandwidth above five kilobytes per second, placing smell third behind vision and hearing, and notes that canines likely exceed that by about twenty times, which makes superhuman machine olfaction a plausible engineering goal. The conclusion the paper draws is that a coordinated investment in standards, datasets, and benchmarks would let olfactory embeddings join visual and language embeddings in multimodal models, enabling embodied systems to locate odor sources, navigate by scent, and reason about scenes with chemical information.","pith_inferences":["An editorial inference: the paper's own bandwidth arithmetic sets a concrete engineering target it does not name—a canine-level machine nose needs roughly twenty times human olfactory throughput, so sensor arrays and processors should be designed against that budget from the start.","An editorial inference: because odor signals arrive as short bursts separated by clean air, an event-based representation may fit olfaction better than fixed-frame recordings, and the standard should arguably be built around plume structure rather than continuous sampling.","An editorial inference: the paper identifies breath and body odor as future personal data but stops short of designing for it; a natural extension is that olfactory benchmarks should include consent, privacy, and audit protocols as first-class components from the outset."],"forward_implications":["A shared olfactory data format would let research groups pool raw sensor recordings, ending the current fragmentation in which datasets from different labs cannot be compared or combined.","Benchmark suites modeled on large-scale vision and language evaluation templates would make machine smell measurable and attract research effort comparable to that given to other modalities.","Embodied systems could take on scent-based tasks such as gas-leak localization, plume tracking, food and crop quality assessment, breath-based diagnostics, and olfaction-visual reasoning in static and moving scenes.","Treating a person's odor profile as sensitive personal data would expand AI ethics beyond vision, language, and audio, requiring consent rules and auditing procedures for olfactory models."],"supporting_citations":[{"why":"Supplies the large-scale benchmark-and-dataset template the paper argues olfaction lacks an equivalent of.","marker":"[45]"},{"why":"Grounds the bandwidth argument with measured human sensory rates, establishing smell as the third-highest-throughput sense.","marker":"[165]"},{"why":"Shows an existing attempt to digitize odor space and documents the non-unanimous human descriptors that complicate labeling.","marker":"[91]"},{"why":"Provides one of the few machine-learning-ready gas sensor datasets, recorded under turbulent flow in open-air conditions.","marker":"[146]"},{"why":"Links 480 molecules to human perceptual ratings, serving as an anchor dataset for modeling structure-percept correlations.","marker":"[85]"},{"why":"Documents drift-related reliability limitations in a widely used metal-oxide sensor dataset, supporting the case for new standards.","marker":"[48]"},{"why":"Represents the unresolved scientific debate on odor coding that the paper says standardization should not wait for.","marker":"[17]"},{"why":"Shows a general AI benchmark with no olfactory component, illustrating the absence of smell from mainstream AI evaluation.","marker":"[32]"}],"fun_headline_variants":["AI's missing sense is smell, and it's fixable","Give AI a nose: standardize smell data","Smell gap in AI is fixable, not fundamental","To smell, AI needs standardized datasets","Olfaction is AI's fixable infrastructure gap"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument rests on the premise that data standards, datasets, and benchmarks can be built and remain useful before the science of smell is settled; if the right way to represent odor turns out to depend on that unsettled science, early standards could encode a wrong abstraction and much of the proposed investment would be wasted.","fun_headline_variants_meta":{"raw":{"variants":["AI's missing sense is smell, and it's fixable","Give AI a nose: standardize smell data","Smell gap in AI is fixable, not fundamental","To smell, AI needs standardized datasets","Olfaction is AI's fixable infrastructure gap"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001068,"raw_usage":{"total_tokens":4483,"prompt_tokens":963,"completion_tokens":3520,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":579,"completion_tokens_details":{"reasoning_tokens":3445}},"tokens_in":579,"tokens_out":3520,"duration_ms":22330,"temperature":1.0,"reasoning_tokens":3445,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:05:44.855572+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete test: build the proposed standardized olfactory dataset from raw sensor recordings with calibrated molecular ground truth, train equivalent models on it and on today's ad hoc small datasets, and compare both on the same real-world scent-source localization task; if the standardized data yields no measurable advantage in accuracy or generalization, the claim that missing standards are the key bottleneck would be weakened.","supporting_citations":[{"cited_title":"The unbearable slowness of being: Why do we live at 10 bits/s?Neuron, 113(2):192–204, Jan 2025","cited_arxiv_id":null,"evidence_quote":"Grounds the bandwidth argument with measured human sensory rates, establishing smell as the third-highest-throughput sense."},{"cited_title":"Gas sensor arrays in open sampling settings","cited_arxiv_id":null,"evidence_quote":"Provides one of the few machine-learning-ready gas sensor datasets, recorded under turbulent flow in open-air conditions."}],"review_version":1}