{"id":"13659996-cd6e-4eb5-91df-8de1b19132a8","arxiv_id":"2412.06472","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Curated exogenous regressors, especially climate and geopolitical series, modestly improve Canadian food CPI forecasts, while LLM-based data curation shows mixed results.","lead":"This technical report from Canada's Food Price Report team compares machine learning and statistical models for forecasting Canadian food prices, and tests whether adding carefully chosen climate, geopolitical, and economic data improves accuracy. It finds modest gains from curated data, especially climate and geopolitical indicators, while noting that simple statistical models remain competitive.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The curation advantage in Table 2 may be an artifact of selecting ensembles on the same 2018–2023/2024 windows used to report MAPE; a strict temporal holdout is needed.","rationale":"I read the paper as an applied forecasting report whose main empirical contribution is the comparison of data-curation strategies for food-price CPI forecasting. The authors are transparent about many implementation details and about the volatility of the evaluation period, and the result that 'all regressors' tends to hurt while some curated subsets help is plausible and consistent with related work. However, the central claim in Section 5 depends on the Table 2 comparison, and that comparison is made on the same windows used to select the final ensembles. Section 3.4 explicitly says the top-performing ensemble was selected 'during the same evaluation period.' With only six annual windows, this creates a real risk that the reported curation advantage is partly an artifact of picking the best condition on the test set. The reader's weakest assumption identifies the same issue: the evaluation windows are small, internally inconsistent between Sections 3.1 and 3.4, and not followed by a genuine holdout after model selection. I therefore agree with the reader's assessment. I do not see a reason to move away from CONDITIONAL: the paper should be publishable with the requested holdout evaluation, code/data release, and consistency fixes, but the headline numbers should not yet be treated as unbiased estimates of forecast error. My concrete test is the minimal check that would settle whether the curation advantage survives out-of-sample selection.","tokens_in":14687,"tokens_out":4742,"duration_ms":47230,"concrete_test":"Re-run the full pipeline with a strictly nested temporal split: train on 1986–2017, use 2018–2021 as the validation set for selecting curated subsets and ensemble members, and reserve 2022–2024 as the final untouched test set. Recompute Table 2 and the final-ensemble MAPE only on that holdout, after first resolving whether the evaluation period is 2018–2024 or 2018–2023 as stated in Sections 3.1 and 3.4. If curated groups do not beat both the all-regressor and no-regressor baselines for a majority of the eight food categories on the untouched holdout, the Section 5 claim that curation improves performance is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in Section 5 is that simple curation techniques, such as thematic groupings of exogenous regressors, can improve model performance. The supporting evidence is Table 2, where curated subsets (climate, geopolitical, LLM-suggested) appear to beat both all-regressor and no-regressor baselines. The load-bearing condition is that these gains reflect predictive skill rather than selection on the evaluation period. That condition is not secured by the paper's protocol. Section 3.4 states that final forecasts were generated by exploring all combinations of the top 10 performing models for each category and selecting the top-performing ensemble during the same evaluation period. Those same windows are then used to report the Table 2 MAPE averages and the Section 4/5 conclusions. With only six annual evaluation windows, ten candidate models, and multiple curation conditions, choosing the best combination on these windows can create an optimistic bias. The reported advantages are also small relative to the reported variability: many curated-vs-none differences in Table 2 are on the order of 0.005–0.01 MAPE, comparable to the stated standard deviations, and no significance test or holdout is provided. The protocol inconsistency between Section 3.1 (evaluation 2018–2024) and Section 3.4 (evaluation 2018–2023) further obscures which years were used for selection and which, if any, were left untouched. Because the paper's central empirical conclusion rests on this comparison, the lack of a selection-free holdout is the single most load-bearing concern. This is not a claim of misconduct; it is a standard selection-on-evaluation risk that the paper does not address.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports an empirical study supporting the University of Guelph/Vector Institute contribution to the 2025 Canada's Food Price Report. It evaluates statistical, deep-learning, transformer, foundation, and LLM forecasting models on nine Canadian food CPI categories, using monthly data from 1986 onward. The central methodological contribution is data curation: exogenous regressors are grouped by human-suggested themes (economic, climate, geopolitical, manufacturing) or by LLM persona-based selection, and these curated sets are compared to using all regressors or none. Performance is measured by MAPE over 18-month forecasts for recent evaluation windows. The paper's main claim, stated in Section 5, is that simple curation techniques can improve model performance. Secondary claims concern which model families win for which categories and the relation to intrinsic time-series complexity metrics.","tokens_in":14906,"tokens_out":3096,"duration_ms":31279,"significance":"If the central claim is secured, the paper would be a useful practical contribution to food-price forecasting and to data-centric AI: it introduces a new multi-source dataset, a replicable LLM-based curation protocol, and a comparison of many model families in a real forecasting setting. The work also connects to a concrete deployment (the 2025 CFPR), which gives the evaluation ecological validity. The LLM-persona curation method and the public-data aggregation are the most valuable parts. However, the evidence for the central claim is currently weakened by the evaluation protocol, which selects final ensembles on the same windows used to report performance, and by the absence of significance testing or a strict holdout. These issues are fixable within the scope of the paper.","major_comments":[{"comment":"The final ensembles are selected by 'exploring all combinations of the top 10 performing models for each category and selecting the top-performing ensemble during the same evaluation period' (Section 3.4), and the performance of those ensembles is then reported in Table 2 and used to support the Section 5 conclusion that curation improves performance. This is a selection-on-the-evaluation-window protocol: with only six annual evaluation windows, ten candidate models, and multiple curation conditions, choosing the best combination on the same windows used to compute the reported MAPE values creates an optimistic bias. The authors should either hold out a strict temporal test period (e.g., select on 2018-2021 and test on 2022-2023/2024) or clearly report selection-aware, nested estimates. Without this, the reported curation gains are not a reliable estimate of predictive skill.","section":"Section 3.4 and Section 4.2"},{"comment":"The paper gives two inconsistent descriptions of the data split. Section 3.1 states that training uses 1986-2017 and evaluation covers 18-month forecasts for 'the six most recent years (2018 to 2024)', while Section 3.4 states training uses 1986-2016 and evaluation covers 2018 to 2023. These differences change every reported number and also affect the interpretation of the context length (36 months versus 75 months for LLMs). The authors must state exactly which split and which evaluation years were used, and reconcile the 'six most recent years (2018 to 2024)' phrasing, which appears to list seven calendar years. This is load-bearing because the entire empirical comparison depends on the evaluation windows.","section":"Section 3.1 vs Section 3.4"},{"comment":"The differences between curated subsets and the 'None' baseline are often small relative to the reported variability: for example, Bakery climate 0.039±0.00 vs None 0.044±0.01, Meat geopolitical 0.029±0.01 vs None 0.031±0.01, and Vegetables climate 0.050±0.01 vs None 0.060±0.02. No significance tests, paired comparisons, or confidence intervals are provided, and with only a handful of evaluation windows it is unclear whether these differences reflect systematic gains or noise. The paper should report per-window paired errors and at least a paired test or a bootstrap confidence interval for the curated-versus-baseline comparison, otherwise the conclusion that curation 'can improve model performance' is not statistically supported.","section":"Table 2 and Section 4"},{"comment":"The complexity metrics in Table 3 are computed using '3-year overlapping windows from 1986-2024', which includes the evaluation period (2018-2023/2024). Using a period that overlaps with the evaluation to explain which model family wins introduces leakage into the explanatory analysis: the 'intrinsic complexity' ranking is not purely intrinsic if it is partly determined by the same years used to evaluate forecasting performance. The authors should recompute these metrics using only data up to the start of the evaluation period (e.g., 1986-2017 or 1986-2016) to support the claim that intrinsic properties predict model-family success.","section":"Section 4.1 and Table 3"}],"minor_comments":[{"comment":"The report uses both 'CPFR' and 'CFPR' inconsistently; the abbreviation for Canada's Food Price Report should be fixed (the title and abstract use CFPR, while the introduction uses CPFR).","section":"Throughout"},{"comment":"There is a typo in 'wich depend heavily on signal stationarity' (should be 'which').","section":"Section 4.1"},{"comment":"The model name 'TemporoSpatialTransformer' appears to be a misnomer; the text refers to the Temporal Fusion Transformer, and the reference [Lim et al., 2021] is correct, but the model name should match the reference.","section":"Section 3.3.2"},{"comment":"In 'LLMs are a form of foundation model that leverage extensive corpa of text-based data', 'corpa' should be 'corpora'.","section":"Section 3.3"},{"comment":"The sentence beginning 'Consistent with findings from similar approaches, such as those reported by Kristina L. Kupferschmidt, Cody Kupferschmidt, Joshua A. Skorburg, ...' lists authors but gives no year or reference entry, so it is impossible for readers to locate the cited work.","section":"Section 4.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a technical report accompanying a real forecasting exercise, which gives it practical relevance but also means the evaluation protocol is not as rigorous as a standalone methods paper. The main issue is the selection-on-the-evaluation-window protocol; this is fixable with a strict holdout or at least a clearly reported nested selection procedure, so I would not reject. The inconsistency between Sections 3.1 and 3.4 must be resolved before the numbers can be trusted. The LLM-curation and dataset contributions are worth preserving. I would not require a fully independent replication, but the authors should either provide code and data or clearly state availability, since the paper's value depends on the reproducibility of the empirical comparisons."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What's new here is the first direct comparison I've seen between hand-curated thematic regressor sets and LLM-persona-curated sets for Canadian food CPI, using 165 scraped series and several model families, tied to a real annual forecast. That makes it a useful applied benchmark, and the authors are transparent that this is a technical report. The appendix material (prompts, variable lists, category groupings) is unusually complete for this genre. They also deserve credit for stating early on that, averaged over categories, curation gave only marginal improvements; the later conclusion that curation \"can improve\" performance is in tension with that, but the tension is visible, not hidden.\n\nThe core evidence in Table 2 is suggestive rather than conclusive. Curated conditions (climate, geopolitical, CPI) do beat the no-regressor baseline in most categories, and the all-regressor baseline is worst nearly everywhere. But the differences are small, standard deviations overlap, there are multiple conditions with no multiple-testing correction, and no significance test is reported. With six evaluation windows, a consistent but small edge could easily be noise. Paired tests across windows, or an additional out-of-sample year now that 2024 CPI is known, would have settled this.\n\nThe stress-test note worries that Table 2 reflects ensemble selection on the evaluation period. I don't think that's right: Table 2 is an average over model types and conditions, not over the selected ensembles. The selection-on-evaluation issue is real, but it affects the final 2025 forecast choices described in Section 3.4, not the per-condition MAPE averages in Table 2. Still, the broader concern stands: no holdout period is left untouched for the overall pipeline, so the reported differences should be treated as exploratory.\n\nTwo concrete inconsistencies need fixing: training data is 1986–2017 in Section 3.1 but 1986–2016 in Section 3.4, and the evaluation period is 2018–2024 in one place and 2018–2023 in another. No code or data are shipped; the public sources and prompts are listed, but exact series and preprocessing are not. For a numerical benchmark, that limits reproducibility.\n\nThis paper is for people doing applied food-price or CPI forecasting, especially those using LLM-curated covariates. It is not a new method and it does not settle a long-open question, but it is an honest, readable empirical report. I would send it to peer review at an applied ML or forecasting venue, but with major revision: release data/code or at least a precise data appendix, add significance tests, fix the train/test inconsistencies, and ideally validate the curation advantage on a period not used for any model or ensemble selection. If the curation advantage survives that, it becomes a genuinely useful published benchmark. As is, I'd cite it only as a technical report with promising but unverified results.","headline":"A solid applied benchmark of food-price forecasting with a plausible but unproven curation advantage; needs significance tests and a clean holdout before the headline numbers are trusted.","tokens_in":15559,"tokens_out":5706,"would_cite":false,"duration_ms":59953,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Curated data, not bigger models, drives better Canadian food price forecasts","keywords":["food price forecasting","Consumer Price Index","data curation","exogenous regressors","time series forecasting","large language models","ensemble forecasting","Canada's Food Price Report"],"falsifier":"Refit the same curation conditions using data through 2022 and evaluate only on untouched 2023, 2024, and 2025 realized CPI data; if curated climate and geopolitical sets no longer beat the all-regressor and no-regressor baselines on that holdout, the reported curation advantage is an artifact of the chosen evaluation windows.","tokens_in":14397,"feed_emoji":"📈","tokens_out":5920,"duration_ms":56541,"temperature":0.7,"pith_summary":"The paper asks whether machine learning can forecast Canadian food inflation more accurately when the data fed to the models is curated rather than used wholesale. It tests 18-month forecasts of nine food Consumer Price Index categories under three settings: all 165 external regressors, none, and curated subsets grouped by expert themes (climate, geopolitical, economic, manufacturing) or selected by LLM personas. It reports that curated subsets, especially climate and geopolitical, often beat both baselines, that including all regressors consistently hurt, and that no single model family dominated. The paper also finds that simple complexity metrics of each CPI series track which model class performs best. The result matters because it suggests that cheap, interpretable data curation, rather than larger models, is the main lever for improving annual food price predictions.","feed_headline":"Curated regressors beat all-in models for food CPI forecasts","feed_subtitle":"Themed climate and geopolitical sets outperform both all-regressor and no-regressor baselines across nine food categories.","key_machinery":"The carrying objects are the curated regressor sets: 165 monthly time series collected from public sources and grouped into four expert themes (economic, climate, geopolitical, manufacturing) plus LLM-generated selections made by persona-prompted GPT-4o. The evaluation mechanism is mean absolute percentage error (MAPE) over 18-month forecast horizons across annual windows, comparing each curation condition to an all-regressor baseline and a target-only baseline. A second mechanism, the category complexity ranking, uses five metrics computed in three-year windows to rank the nine CPI categories, showing that high-complexity categories such as Vegetables, Meat, and Fruit favor transformer and foundation models while low-complexity categories favor Exponential Smoothing or simple feed-forward networks.","core_discovery":"The paper's central claim is that simple curation techniques, such as thematic groupings of exogenous regressors, can improve forecasting performance relative to both using every available regressor and using none. Averaged over categories and models the gains are modest, but in every food category the best-performing configuration used a curated set, and climate and geopolitical regressors were the most frequent winners. It also claims that intrinsic properties of a time series, including trend and seasonality strength, residual variance, residual MAD, and Shannon entropy, predict whether a high-capacity transformer or foundation model will outperform a simple statistical baseline. In the resulting 2025 Canada's Food Price Report forecasts, most categories were served by ensembles dominated by transformer, foundation, or LLM models, with only Fish and Other using statistical models.","pith_inferences":["The curation effect is plausibly portable: any country's food CPI forecasting exercise could test whether climate and geopolitical subsets beat full-regressor sets before investing in larger models, which is a cheap and direct extension.","The complexity metrics suggest a Mixture-of-Experts style routing rule for time series, where residual variance and entropy pick the model class at inference time; this is testable on other forecasting benchmarks.","A stricter validation protocol, with model selection on one period and evaluation on a later untouched period, would clarify whether the curation gains persist outside the volatile 2018-2024 windows; the paper's numbers do not yet establish that.","LLM persona curation could be made more robust by treating personas as an ensemble rather than taking consensus at a fixed threshold of 7, a variant the paper does not run."],"forward_implications":["Curated regressor sets, particularly climate and geopolitical proxies, consistently match or beat models that see all 165 regressors or none; no food category improved by including everything.","The best model family shifts with measured time-series complexity: low-complexity categories do well with Exponential Smoothing or simple networks, while high-complexity categories do best with transformers and the Chronos foundation model.","LLM-suggested regressors can perform as well as human-expert themes, but feeding LLMs additional future forecasts tends to degrade their output, while including the previous CFPR helped Bakery and Meat but not Vegetables.","For the 2025 report, the chosen ensembles were mostly ML-based, using transformer, foundation, and LLM models rather than statistical baselines, marking a shift from earlier editions.","The curation advantage appears in per-category results even where the averaged gains are small, suggesting that selecting the right subset of context can be more valuable than the choice of model family."],"supporting_citations":[{"why":"Establishes that food price drivers are not independent and interactions among multiple regressors must be considered, which motivates the whole curated-regressor approach.","marker":"Kalkuhl et al., 2016"},{"why":"Identifies climate events, supply chain disruptions, and policy changes as key factors in Canadian food pricing, underpinning the expert-identified themes.","marker":"Charlebois et al., 2024a,b"},{"why":"Documents the Russia-Ukraine war's effect on grain and food prices, supporting the geopolitical regressor theme.","marker":"Carter and Steinbach, 2024"},{"why":"Supplies the human-centric DelphAI framework that the paper adapts for consulting experts and guiding forecast development.","marker":"Kupferschmidt et al., 2022"},{"why":"Provides the Direct Prompt framework used to make LLMs output forecasts directly with contextual inputs.","marker":"Williams et al., 2024"},{"why":"Provides the LLM Processes approach that motivates using repeatedly probed LLMs as time-series forecasters.","marker":"Requeima et al., 2024"},{"why":"Supplies the persona-generation technique used to prompt GPT-4o to rate the value of external regressors.","marker":"Schuller et al., 2024"},{"why":"Describes Chronos, the pre-trained time-series foundation model used as a zero-shot baseline and in final ensembles.","marker":"Ansari et al., 2024"},{"why":"Defines the Temporal Fusion Transformer, the transformer model used for multivariate forecasting with exogenous regressors.","marker":"Lim et al., 2021"}],"fun_headline_variants":["Curated regressors beat all-in models for food CPI","Themed regressor sets win across all food categories","Climate and geopolitical data boost food CPI forecasts","Simple data curation improves food price predictions","Time series traits predict best model for food prices"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the evaluation windows, described as 2018 to 2024 in Section 3.1 and 2018 to 2023 in Section 3.4 and including the volatile COVID and supply-shock years, represent the conditions the 2025 forecast must handle, even though the models and ensembles were selected using those same windows.","fun_headline_variants_meta":{"raw":{"variants":["Curated regressors beat all-in models for food CPI","Themed regressor sets win across all food categories","Climate and geopolitical data boost food CPI forecasts","Simple data curation improves food price predictions","Time series traits predict best model for food prices"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000676,"raw_usage":{"total_tokens":3036,"prompt_tokens":868,"completion_tokens":2168,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":484,"completion_tokens_details":{"reasoning_tokens":2097}},"tokens_in":484,"tokens_out":2168,"duration_ms":17972,"temperature":1.0,"reasoning_tokens":2097,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T19:36:58.880758+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Refit the same curation conditions using data through 2022 and evaluate only on untouched 2023, 2024, and 2025 realized CPI data; if curated climate and geopolitical sets no longer beat the all-regressor and no-regressor baselines on that holdout, the reported curation advantage is an artifact of the chosen evaluation windows.","supporting_citations":[{"cited_title":"Food Price Volatility and Its Implications for Food Security and Policy","cited_arxiv_id":null,"evidence_quote":"Establishes that food price drivers are not independent and interactions among multiple regressors must be considered, which motivates the whole curated-regressor approach."},{"cited_title":"Did grain futures prices overreact to the russia--ukraine war due to herding? Journal of Commodity Markets, 35: 0 100422, 2024","cited_arxiv_id":null,"evidence_quote":"Documents the Russia-Ukraine war's effect on grain and food prices, supporting the geopolitical regressor theme."},{"cited_title":"Generating personas using llms and assessing their viability","cited_arxiv_id":null,"evidence_quote":"Supplies the persona-generation technique used to prompt GPT-4o to rate the value of external regressors."}],"review_version":1}