{"id":"9ce8593d-9d70-42af-9de1-9691da32bb7d","arxiv_id":"2606.27672","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Empirical tests show TSFMs like Chronos-2 and MOMENT need fine-tuning for E-Nose data, with fusion to specialized models further improving gas identification and concentration prediction.","lead":"This paper empirically tests time-series foundation models on electronic nose sensor data for gas identification and concentration tasks. A smart generalist might read it to learn whether general-purpose AI models are ready for specialized sensing applications or still require domain-specific adaptation.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Generalization from two TSFMs (Chronos-2, MOMENT) and specific E-Nose datasets to the full class rests on unverified representativeness","rationale":"The reader's weakest_assumption directly identifies the load-bearing generalization step for an empirical study whose headline claim is about an entire model class. No internal inconsistency, missing control, or formal gap is visible from the provided abstract; the concern is external validity rather than execution error. Full-text details on dataset statistics or additional ablations could mitigate but are not supplied here.","tokens_in":1670,"tokens_out":340,"duration_ms":17030,"concrete_test":"Re-run the zero-shot vs. fine-tuned vs. fusion experiments on the same E-Nose tasks using two additional TSFMs (e.g., TimesFM and Lag-Llama); if any model achieves satisfactory performance without fine-tuning, the necessity claim does not hold for the class.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim asserts that fine-tuning is necessary for satisfactory E-Nose performance and that fusion with specialized models yields further gains, implying limited readiness of current TSFMs. This rests on the assumption that Chronos-2 and MOMENT (plus the chosen datasets) are representative of the broader TSFM family and gas-sensing domain. If other TSFMs exhibit stronger zero-shot transfer or if the E-Nose collections share unstated biases (sensor drift, limited gas variety, small sample regimes), the necessity of fine-tuning and the value of fusion would not generalize. The abstract provides no evidence of broader model coverage or dataset characterization that would secure this step.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript presents an empirical assessment of time-series foundation models (TSFMs) on electronic nose (E-Nose) data. Using embeddings from Chronos-2 and MOMENT, it evaluates their utility for gas identification and concentration prediction tasks. The central claims are that fine-tuning these embeddings is necessary to reach satisfactory performance and that fusing them with representations from specialized predictive models yields further gains, indicating both potential and limitations of current TSFMs for gas-sensing applications.","tokens_in":1799,"tokens_out":414,"duration_ms":26736,"significance":"If the reported empirical patterns hold under broader testing, the work would document a concrete domain where zero-shot TSFM transfer underperforms and would illustrate a practical hybrid strategy that improves results. This could usefully inform subsequent TSFM development for sensor-derived time series.","major_comments":[{"comment":"Abstract: The evaluation is restricted to Chronos-2 and MOMENT with the assertion that they are 'representative TSFMs,' yet no justification, selection criteria, or comparison to other models (e.g., additional TSFMs) is supplied. This assumption is load-bearing for the general claim that fine-tuning is necessary across the TSFM class for E-Nose data.","section":"Abstract"},{"comment":"Abstract: The E-Nose datasets are not characterized with respect to sample size, sensor types, gas variety, or potential biases such as drift or small-sample regimes. Without this information it is impossible to determine whether the necessity of fine-tuning and the value of fusion are robust findings or specific to the chosen collections.","section":"Abstract"}],"minor_comments":[{"comment":"The abstract states conclusions ('we show that...') without any quantitative metrics, baseline comparisons, or dataset identifiers; even a high-level summary of key numbers would strengthen readability.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback on our manuscript. We address each major comment below and outline the revisions we will make to improve clarity and transparency.","responses":[{"response":"We agree that the abstract would be strengthened by explicit justification for selecting Chronos-2 and MOMENT. These were chosen as representative due to their recency, public availability of embeddings, and reported strong performance on general time-series benchmarks. In the revised manuscript we will add selection criteria to the abstract and a short paragraph in the introduction, while qualifying that our findings apply to these models and do not claim universality across all TSFMs without further evaluation.","revision_made":"yes","referee_comment":"[Abstract] Abstract: The evaluation is restricted to Chronos-2 and MOMENT with the assertion that they are 'representative TSFMs,' yet no justification, selection criteria, or comparison to other models (e.g., additional TSFMs) is supplied. This assumption is load-bearing for the general claim that fine-tuning is necessary across the TSFM class for E-Nose data."},{"response":"We will add a concise characterization of the E-Nose datasets (sample sizes, sensor types, gas varieties, and notes on drift mitigation) directly into the abstract. The full experimental section already contains these details; elevating a summary to the abstract will make the scope of our empirical findings on fine-tuning and fusion more transparent to readers.","revision_made":"yes","referee_comment":"[Abstract] Abstract: The E-Nose datasets are not characterized with respect to sample size, sensor types, gas variety, or potential biases such as drift or small-sample regimes. Without this information it is impossible to determine whether the necessity of fine-tuning and the value of fusion are robust findings or specific to the chosen collections."}],"tokens_in":1302,"tokens_out":400,"duration_ms":28909,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The one or two things to know are that this work applies Chronos-2 and MOMENT to electronic nose data for the first time and reports that their embeddings require fine-tuning to work well for gas identification and concentration prediction, with additional gains from fusing them with specialized models.\n\nThe new part is the domain extension to gas sensing, which fills an explicit gap noted in the abstract. The paper does a solid job of running the empirical assessment and pointing out both where the TSFMs fall short in zero-shot settings and where they can be adapted. This kind of targeted check is useful for seeing how well the generalization claims hold up in practice.\n\nWhere it gets soft is the reliance on just those two models to speak for time-series foundation models as a whole. If other TSFMs handle the data better out of the box, the necessity of fine-tuning wouldn't apply broadly. The datasets chosen might also share traits like sensor issues or limited scope that aren't representative of all E-Nose setups. The abstract is light on specifics like exact dataset descriptions, metrics used, or comparison baselines, which makes it harder to evaluate how convincing the results are. The stress-test concern about unverified representativeness holds up here.\n\nNo formal math or proofs are involved, so that's not an issue. The citation pattern looks standard and relevant to TSFM work.\n\nThis paper would interest people in sensor applications or foundation model adaptation for time series. A practitioner in gas sensing or someone exploring new domains for TSFMs could pick up practical insights from the findings.\n\nIt deserves a serious referee because it tackles an open question with experiments, even if revisions for more breadth would help.\n\nRecommendation: Send it out for peer review to get feedback on strengthening the experimental design and scope.","headline":"The paper tests two TSFMs on E-Nose data and finds zero-shot embeddings fall short while fine-tuning plus fusion improves results, but the narrow model coverage and missing experimental details limit how far the conclusions reach.","tokens_in":2424,"tokens_out":449,"would_cite":false,"duration_ms":53076,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Time-series foundation model embeddings need fine-tuning to work on electronic nose gas data.","keywords":["time-series foundation models","electronic nose","embeddings","gas identification","fine-tuning","sensor data","Chronos","MOMENT"],"falsifier":"Demonstration of a new time-series foundation model that reaches high accuracy on several E-Nose gas-identification and concentration tasks without any fine-tuning or fusion steps.","tokens_in":2561,"feed_emoji":"","tokens_out":623,"duration_ms":19729,"temperature":0.7,"pith_summary":"The paper tests whether embeddings from broad time-series foundation models can serve as effective features for identifying gases and estimating their concentrations in electronic nose sensor readings. It evaluates models including Chronos-2 and MOMENT on standard E-Nose datasets and finds that the raw embeddings yield weak results on these tasks. Fine-tuning the foundation models on the gas-sensing data produces usable performance, while further gains occur when the tuned embeddings are combined with features from models built specifically for the prediction tasks. The findings point to both the adaptability of current foundation models to new sensor domains and their need for domain-specific adjustment before reliable use.","feed_headline":"Time-series foundation models need tuning for E-nose data","feed_subtitle":"Off-the-shelf embeddings underperform on gas sensor tasks, but fine-tuning plus fusion with specialized models improves results.","key_machinery":"Embeddings extracted from pre-trained time-series foundation models, used as feature inputs for downstream classification and regression on multivariate E-Nose sensor time series.","core_discovery":"Embeddings produced by representative time-series foundation models do not directly deliver satisfactory performance for gas identification and concentration prediction on E-Nose data. Fine-tuning the models on the target data is required to reach acceptable accuracy, and fusing the resulting embeddings with representations learned by specialized predictive models yields additional improvement.","pith_inferences":["The same pattern of needing adaptation may appear when applying current time-series foundation models to other chemical or environmental sensor streams.","Pre-training future foundation models on larger collections of sensor-array data could reduce reliance on per-domain fine-tuning.","The fusion approach could be examined in other time-series domains where off-the-shelf foundation embeddings underperform."],"forward_implications":["Fine-tuning is required before time-series foundation model embeddings support reliable gas identification or concentration estimates from E-Nose readings.","Fusion of fine-tuned foundation model embeddings with outputs from task-specific models produces higher performance than fine-tuning alone.","Current time-series foundation models show limitations for direct application to gas-sensing data but retain value once adapted.","Hybrid use of foundation-model and specialized representations offers a practical route for deploying these models in sensor domains."],"fun_headline_variants":["TSFMs require fine-tuning to handle E-nose data","Off-the-shelf TSFM embeddings fall short on gas sensors","E-nose gas tasks demand tuned time-series foundation models","Fusing models enhances TSFM performance for E-nose"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The specific foundation models and E-Nose datasets examined stand in for the full range of time-series foundation models and gas-sensing applications.","fun_headline_variants_meta":{"raw":{"variants":["TSFMs require fine-tuning to handle E-nose data","Off-the-shelf TSFM embeddings fall short on gas sensors","E-nose gas tasks demand tuned time-series foundation models","Fusing models enhances TSFM performance for E-nose"]},"model":"grok-4.3","cost_usd":0.003745,"raw_usage":{"total_tokens":1823,"prompt_tokens":596,"num_sources_used":0,"completion_tokens":68,"cost_in_usd_ticks":37453000,"prompt_tokens_details":{"text_tokens":596,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1159,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":596,"tokens_out":68,"duration_ms":7703,"temperature":1.0,"reasoning_tokens":1159,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T04:51:58.025953+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Demonstration of a new time-series foundation model that reaches high accuracy on several E-Nose gas-identification and concentration tasks without any fine-tuning or fusion steps.","supporting_citations":[],"review_version":1}