{"id":"0e12efb4-b42e-4969-83be-6ff0ca6437e2","arxiv_id":"2502.02028","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Fine-tuning small language models for recipe generation produces mixed results: Phi-2 degrades on the authors' custom quality scores while SmolLM-360M and 1.7B perform similarly.","lead":"A team of researchers fine-tuned several small language models to generate cooking recipes and built custom scoring rules to judge them. They found that fine-tuning did not always help, and that model size did not consistently predict recipe quality under their metrics.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claims rest on the four custom recipe metrics (§4.2, Appendix E), which are specified only as bullet-point heuristics with no formulas, code, or human calibration; if those metrics do not track actual recipe quality, the Phi-2 degradation and SmolLM comparability findings lose support.","rationale":"The reader's weakest_assumption identifies the same load-bearing concern: the custom domain-specific metrics are under-specified and unvalidated, and the paper's main conclusions depend on them. The manuscript's own Appendix E confirms the metrics are defined as high-level heuristics with no formulas, thresholds, or released code, and nowhere in the paper are they calibrated against human or expert judgment. The implausibly low coherence scores across all systems are a symptom, not proof, of invalidity; they motivate the proposed human-calibration test. The concern is not about the authors' integrity or the paper's internal consistency, but about whether the evaluation instrument measures what it claims. Even if significance testing and error bars were added, invalid metrics would still undermine the central claims. The paper is honest about several limitations, but it omits the absence of metric validation, which is the most serious gap. The recommended verdict remains CONDITIONAL as the reader originally stated, so no adjustment is needed.","tokens_in":14800,"tokens_out":3759,"duration_ms":37230,"concrete_test":"Release the evaluation code and run a human calibration study on a stratified sample of 300 recipes (50 per condition × 6 model/evaluation conditions: baseline and fine-tuned Phi-2, SmolLM-360M, SmolLM-1.7B). Have at least two independent annotators with cooking expertise score each recipe on ingredient coverage, step complexity, coherence, and temperature/time specification on a 0–1 scale. Compute per-metric Spearman rank correlations between the custom metric and human scores, and compare the metric's ranking of the six conditions against the human ranking. If the correlation is below ~0.5 or the Phi-2 baseline-vs-fine-tuned order reverses on ingredient coverage or temp/time, then Table 3's central claims are unsupported and should be reframed as metric-development results.","verdict_should_be":"UNCHANGED","load_bearing_attack":"To support the central claim that fine-tuning degrades Phi-2's recipe quality and that the custom framework reveals what traditional metrics miss, the four domain-specific metrics must be valid. Appendix E defines each metric by a short list of operations ('count distinct operations,' 'build step dependency graph,' 'validate ranges per method') with no scoring formulas, thresholds, or released implementation, and no calibration against human or expert judgment. The near-universal coherence scores of 0.02–0.12 across all systems, including fluent baselines, suggest the coherence metric may be scoring formatting artifacts (numbered steps, line breaks, specific connectives) rather than culinary logic. This is load-bearing because the paper's most striking result—Phi-2 baseline ingredient coverage 0.59 and temp/time 0.329 dropping to 0.30 and 0.24 after fine-tuning, while step complexity rises from 0.79 to 0.99—can be explained by a metric that rewards longer, more structured but less correct output. The paper's own sample outputs in Appendices D and G show fine-tuned Phi-2 generations that are visibly less readable and less relevant; if the coherence metric still gives them scores (0.07–0.12) comparable to baselines (0.08–0.09), the metric is not tracking the qualitative difference humans see. The Limitations section acknowledges the 500-sample evaluation and stochastic LLM judge, but does not acknowledge the absence of validation for the custom auto-metrics that carry the central claims.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper fine-tunes several small language models—T5-small, GPT-2 (small/medium), SmolLM-135M/360M/1.7B, and Phi-2—on the Food.com recipe dataset for the task of generating cooking instructions from recipe names and ingredient lists. It evaluates the resulting generations with traditional metrics (BLEU, ROUGE, perplexity), four newly proposed domain-specific auto-metrics (ingredient coverage, step complexity, recipe coherence, temperature/time specification), and a Qwen2.5-7B LLM-as-a-judge. It also develops prompt-based and RAG-assisted allergen substitution systems. The main conclusions are that fine-tuning Phi-2 degrades its domain-specific scores (ingredient coverage dropping from 0.59 to 0.30 and temperature/time from 0.329 to 0.24 in Table 3), that SmolLM-360M and SmolLM-1.7B perform comparably despite the size difference, and that the multi-dimensional evaluation framework reveals limitations of traditional overlap-based metrics for creative generation.","tokens_in":15150,"tokens_out":6149,"duration_ms":51715,"significance":"The paper addresses a relevant problem—domain-specific evaluation for creative NLG—and its broad comparison across model scales and architectures is useful empirical groundwork. The appendices are reasonably detailed, including hyperparameters (J–M), sample generations (B, D, G), and an allergen substitution database (F). However, the central findings rest entirely on four custom metrics whose operational definitions are not provided (Section 4.2 and Appendix E), and which are not validated against human judgment. Some reported scores are inconsistent with the paper's own sample outputs (e.g., fine-tuned SmolLM-360M in Table 10 vs. Table 3). Because of this, the significance is currently conditional: if the metrics were made precise, released, and calibrated against human ratings, the results could be an interesting contribution; as written, they are not reproducible and the main claims are not supported.","major_comments":[{"comment":"The four recipe-specific metrics are described only as lists of operations (e.g., \"build step dependency graph\", \"validate ranges per method\") with no scoring formulas, thresholds, or implementation. The scores in Tables 3, 4, and 6 are therefore not reproducible, and the central claims about Phi-2 degradation and SmolLM comparability are untestable. The authors should provide the exact algorithms, release the code, and specify how each sub-score is aggregated.","section":"§4.2 and Appendix E"},{"comment":"Recipe coherence scores fall in a narrow band of 0.02–0.12 for every model and condition, including fluent baseline outputs and degenerate fine-tuned outputs. A metric with such a compressed range cannot support claims of \"marginal improvements\" in coherence or of fine-tuning \"degradation\" in coherence. The authors should report the distribution of coherence scores and demonstrate that the metric tracks human ratings of logical flow.","section":"Table 3 and Tables 4/6"},{"comment":"All results are point estimates over a single 500-sample evaluation, with no standard errors, confidence intervals, or significance tests. Since generation is stochastic (temperature 0.75, top-p 0.95/0.8; Appendices L and M), the observed differences—e.g., SmolLM-360M vs. 1.7B ingredient coverage 0.21 vs. 0.29 in Table 3—could be noise. The claim that the two SmolLM models are \"comparable\" requires repeated sampling or an appropriate statistical test.","section":"Tables 2–7"},{"comment":"The fine-tuned SmolLM-360M output shown in Table 10 is largely random characters, yet Table 3 reports step complexity of 0.98 and ingredient coverage of 0.16 for this model's fine-tuned version. This internal inconsistency suggests the step complexity and coverage metrics are capturing surface formatting (e.g., numbered lines, length) rather than the intended content quality. The authors must reconcile the quantitative scores with their own qualitative examples.","section":"Appendix G, Table 10 vs. Table 3"},{"comment":"The LLM-as-a-judge scores are presented as evidence about allergen safety and recipe quality, but no evidence is given that Qwen2.5-7B's judgments correlate with human or expert assessments. The Limitations section acknowledges stochasticity but does not address validity; a judge that is not calibrated cannot support the allergen substitution conclusions.","section":"§4.3 and Tables 5/7"}],"minor_comments":[{"comment":"The T5-small fine-tuned BLEU-1 and BLEU-2 scores of 0.00 are suspiciously low; please verify the decoding and scoring setup for this model.","section":"Table 1"},{"comment":"The description of QLoRA fine-tuning is incomplete (only rank 8 is given); include quantization bit-width and other LoRA hyperparameters to improve reproducibility.","section":"§3.4 and Appendix K"},{"comment":"Several references are missing URLs or venue details, e.g., \"Microsoft Research. 2023. Phi-2\" and \"Qwen Team. 2024. Qwen2.5\"; please complete them.","section":"References"},{"comment":"The abstract and title refer to a \"Benchmark Study,\" but no benchmark dataset or evaluation code is publicly released; consider adding a link to the code and data.","section":"Abstract and title"},{"comment":"\"Mixed Precision* fp16 or fp32\" is ambiguous; state which precision was used for each small model.","section":"Appendix J"}],"recommendation":"major_revision","confidential_remarks":"The paper is an empirical study with a potentially interesting finding about fine-tuning large LMs for a specialized domain, but the evaluation methodology is too underspecified for the claims as written. If the authors can make the metrics precise, release code, and provide a calibration study with human judges, the contribution could be valuable. For a journal venue, the current form is not sufficient; the paper would be more appropriate for a workshop or demo track in its present state."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis paper does a respectable sweep: it fine-tunes several small LMs (T5, SmolLM, Phi-2) on Food.com and evaluates with standard plus recipe-specific metrics, plus prompt-based and RAG allergen substitution. That combination is the real contribution, and the paper is transparent about its compute constraints and sample size.\n\nThe central finding—Phi-2's ingredient coverage and temp/time scores drop after fine-tuning while step complexity rises—is interesting if the metrics are trustworthy. But they aren't, not yet. The four custom metrics in §4.2/Appendix E are described as bullet-point heuristics: count distinct operations, build dependency graph, validate ranges. No formulas, no thresholds, no released implementation, no calibration against human or expert judgment. Without that, the low coherence scores (0.02–0.12 for everything) look like the metric is capturing formatting, not culinary sense. The appendix samples show fine-tuned Phi-2 output that is visibly worse, yet gets coherence scores comparable to baselines. If the metric can't distinguish what a human sees, the fine-tuning degradation claim is on shaky ground.\n\nAlso, all comparisons are point estimates with no error bars or significance tests. The limitations section mentions the 500-sample evaluation and that LLM-as-a-judge is stochastic, but doesn't address the missing validation of the auto-metrics, which is the load-bearing piece.\n\nNone of this is fatal to the paper as an exploratory study. The allergen substitution systems and the idea of a multi-dimensional recipe evaluation are worth thinking about. The authors are honest about the exploratory nature. But the abstract and discussion frame the results as if the framework is already reliable, and that's an overreach.\n\nFor a serious venue, I'd send it to review only with the expectation of major revisions: validate the metrics against human judgments, release code and data, add error bars or significance tests, and temper the conclusions. As is, it's more of a workshop paper or a useful starting point for the authors' future work. I wouldn't cite it yet, but I'd point students to it as an example of domain-specific evaluation design.","headline":"A well-intentioned recipe-generation benchmark whose load-bearing custom metrics are unvalidated; results are suggestive, not conclusive.","tokens_in":15651,"tokens_out":1922,"would_cite":false,"duration_ms":19485,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This study reports that fine-tuning the larger model Phi-2 degraded ingredient coverage and temperature/time accuracy in generated recipes, while smaller SmolLM models held their own.","keywords":["recipe generation","fine-tuning","language models","domain-specific evaluation metrics","ingredient coverage","allergen substitution","retrieval-augmented generation","LLM-as-a-judge"],"falsifier":"Have a panel of human cooks rate a blinded sample of the baseline and fine-tuned recipes and compare their ratings with the four domain-specific scores; if human rankings do not reproduce the reported Phi-2 degradation and SmolLM comparability, the central claim is not supported.","tokens_in":14619,"feed_emoji":"🍳","tokens_out":9907,"duration_ms":82330,"temperature":0.7,"pith_summary":"The paper tries to establish that standard fine-tuning is not a dependable route to better domain-specific recipe generation: for Phi-2, the largest model tested, fine-tuning cut ingredient coverage from 0.59 to 0.30 and temperature/time specification from 0.329 to 0.24, while SmolLM-360M and SmolLM-1.7B performed comparably despite their size difference. The paper argues that traditional metrics such as BLEU and ROUGE cannot reveal this because they reward overlap with a single ground-truth recipe and punish creative divergence. It therefore builds a multi-dimensional evaluation with recipe-specific scores for ingredient coverage, step complexity, coherence, and temperature/time checks, plus an LLM-based judge, and applies it to baseline and fine-tuned models with and without allergen substitution. If the framework holds, evaluation of creative generation tasks should include domain-specific quality signals rather than relying on overlap metrics alone.","feed_headline":"Fine-tuning a larger recipe model hurt recipe quality","feed_subtitle":"Domain-specific scores fell for Phi-2 even as step structure rose; overlap metrics miss it.","key_machinery":"The load-bearing machinery is the paper's four recipe-specific auto-evaluation metrics: ingredient coverage (does the generated text use the listed ingredients?), step complexity (how detailed and parameterized are the instructions?), recipe coherence (does the step dependency graph make logical and temporal sense?), and temperature/time specification (are cooking parameters present and in plausible ranges?). These scores, not the overlap metrics, are what generate the paper's central contrast between Phi-2's degradation and SmolLM's comparability. A supplementary LLM-as-a-judge rubric covering clarity, completeness, consistency, practicality, relevance, and allergen safety provides a second lens on the same generated recipes.","core_discovery":"On the paper's own terms, the central discovery is that fine-tuning changes what a recipe model does rather than simply improving it. Phi-2's fine-tuned version scored higher on step complexity (from 0.79 to 0.99) but lower on ingredient coverage (from 0.59 to 0.30), recipe coherence (from 0.08 to 0.07), and temperature/time specification (from 0.329 to 0.24), which the authors interpret as a trade-off between producing complete step-by-step instructions and preserving the semantic relations among ingredients. The two SmolLM sizes behaved similarly to each other both before and after fine-tuning, suggesting that parameter count is not the main driver of recipe quality. The paper also reports that prompt-based and retrieval-based allergen substitution both lower some quality scores, and that the multi-dimensional evaluation exposes discrepancies that BLEU and ROUGE scores hide.","pith_inferences":["A direct human-rating study on the same 500 test recipes would settle whether the domain-specific scores track actual culinary quality; if human rankings do not reproduce the reported Phi-2 degradation and SmolLM comparability, the benchmark would need revision before the central finding could be trusted.","The same metric structure, coverage of input items, step dependency, and parameter specification, could transfer to other structured instruction-generation tasks such as workout plans, medication instructions, or DIY repair guides, where missing a detail is costly.","The Phi-2 pattern suggests that standard language-model fine-tuning may teach a model to emit recipe-like scaffolding, numbered steps, temperatures, and times, while weakening its link to the specific ingredient list; a testable extension is whether instruction-tuning or reinforcement-learning objectives recover ingredient coverage without sacrificing step detail.","The fact that retrieval-based substitution lowered ingredient coverage while improving step complexity implies that post-hoc substitution fixes allergens but not the model's underlying planning; an editing pass with faithfulness constraints on the final ingredient list would be a natural next system to test."],"forward_implications":["Fine-tuning on domain text can improve surface structure, such as step-by-step formatting, while eroding fidelity to the input ingredients and cooking parameters; recipe systems should track both dimensions.","Model scale alone does not determine post-fine-tuning recipe quality: the SmolLM-1.7B and SmolLM-360M models landed close together, so smaller, cheaper models can be a sensible choice for this task.","Overlap-based metrics like BLEU and ROUGE should not be the primary yardstick for creative generation; the paper's domain-specific scores and LLM judge give a different, more practical picture.","Allergen substitution, whether prompt-driven or retrieval-driven, changes the quality profile of generated recipes; substitution is not a free add-on and needs its own evaluation.","The step-complexity versus coherence trade-off suggests that conventional fine-tuning objectives may need to be rethought for specialized domains where semantic correctness matters as much as fluency."],"supporting_citations":[{"why":"Supplies the large recipe dataset used for training, validation, and test splits, as well as the encoder-decoder recipe-generation approach the study treats as a baseline.","marker":"(Majumder et al., 2019)"},{"why":"Provides the SmolLM-135M, SmolLM-360M, and SmolLM-1.7B models whose fine-tuning behavior is a main comparison in the paper.","marker":"(Allal et al., 2024)"},{"why":"Provides Phi-2, the model whose post-fine-tuning degradation is the paper's central finding.","marker":"(Research, 2023)"},{"why":"Supplies the GPT-2 small and medium models used in the early small-scale model comparisons that motivate scaling up.","marker":"(Radford et al., 2019)"},{"why":"Defines BLEU, one of the traditional overlap metrics the paper argues misjudges creative generation.","marker":"(Papineni et al., 2002)"},{"why":"Defines ROUGE, the other traditional overlap metric the paper finds insufficient for evaluating recipe quality.","marker":"(Lin, 2004)"},{"why":"Introduced ingredient coverage in recipe evaluation, which the paper adapts into one of its four recipe-specific metrics.","marker":"(Salvador et al., 2019)"},{"why":"Provides the retrieval-augmented generation approach used in the paper's allergen-substitution system.","marker":"(Lewis et al., 2021)"},{"why":"Provides the Qwen2.5-7B model used as the LLM judge for the six-category recipe-quality assessment.","marker":"(Team, 2024)"}],"fun_headline_variants":["Fine-tuning steps up, ingredient coverage down in recipes","Phi-2 fine-tuning trades steps for recipe coherence","Small recipe models match large ones after tuning","Domain metrics expose recipe model trade-offs","Allergen substitution cuts recipe quality, not just allergens"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claims stand on the assumption that the paper's hand-built recipe-quality scores, especially step complexity and recipe coherence, really measure culinary quality, but those scores are never calibrated against human judgment.","fun_headline_variants_meta":{"raw":{"variants":["Fine-tuning steps up, ingredient coverage down in recipes","Phi-2 fine-tuning trades steps for recipe coherence","Small recipe models match large ones after tuning","Domain metrics expose recipe model trade-offs","Allergen substitution cuts recipe quality, not just allergens"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000356,"raw_usage":{"total_tokens":1926,"prompt_tokens":931,"completion_tokens":995,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":547,"completion_tokens_details":{"reasoning_tokens":923}},"tokens_in":547,"tokens_out":995,"duration_ms":10824,"temperature":1.0,"reasoning_tokens":923,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T13:37:13.792318+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have a panel of human cooks rate a blinded sample of the baseline and fine-tuned recipes and compare their ratings with the four domain-specific scores; if human rankings do not reproduce the reported Phi-2 degradation and SmolLM comparability, the central claim is not supported.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the SmolLM-135M, SmolLM-360M, and SmolLM-1.7B models whose fine-tuning behavior is a main comparison in the paper."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides Phi-2, the model whose post-fine-tuning degradation is the paper's central finding."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines BLEU, one of the traditional overlap metrics the paper argues misjudges creative generation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines ROUGE, the other traditional overlap metric the paper finds insufficient for evaluating recipe quality."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduced ingredient coverage in recipe evaluation, which the paper adapts into one of its four recipe-specific metrics."}],"review_version":1}