Pith. sign in

REVIEW 5 major objections 7 minor 52 references

Evaluating Large Language Models for Real-World Engineering Tasks

T0 review · 5 major / 7 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read On a new benchmark of more than 100 real production-engineering questions, four LLMs show reliable performance on basic temporal and structural reasoning but fail on abstract reasoning, formal modeling, and context-sensitive engineering…

desk verdict A genuinely useful engineering benchmark whose headline claims are currently under-supported by the scoring procedure; worth reviewing, but tables need a methodological overhaul. read the letter →

arxiv 2505.13484 v1 pith:23GR6VAQ submitted 2025-05-12 cs.AI cs.CL

classification cs.AIcs.CL
keywords largelanguagemodelsengineeringbenchmarkworldcausalreasoningtemporaltimeseriesforecastingdesignintentioninferenceproductionsystems
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether large language models can handle real engineering work, not just exam-style questions. It builds a curated set of over 100 questions drawn from an actual glass-foaming plant and a simulated production line, together with time-series forecasting tasks, and uses them to test four current LLMs. The central result is that the models handle basic temporal and structural reasoning well but fail on abstract reasoning, formal modeling, and context-sensitive engineering logic. The authors conclude that LLMs are useful as supervised engineering assistants but are not ready to act as autonomous agents on complex engineering problems.

What carries the argument

The evaluative machinery is a purpose-built benchmark: a set of more than 100 domain-specific questions tied to two production-system models, a physical glass-foaming plant and a simulated seven-stage manufacturing line, plus 200 forecasting samples from two time-series datasets. Each question is tagged to one of five research questions and to measurable capability categories such as iterative consistency, non-locality, type-level abstraction, causal inference, and design-intention inference. The scoring protocol is the second component: answers are graded on a 0, 0.5, 1 scale anchored to junior- and senior-engineer competence, and each question is posed multiple times with only the least accurate response counted. This design is meant to reflect engineering reliability standards by measuring worst-case rather than typical performance.

What would settle it

A reader could settle the central claim by taking a random sample of the recorded responses, having several engineers re-grade them with a written rubric, and checking inter-rater agreement; if agreement is low, the category scores are not stable. Alternatively, if a future evaluation shows a current LLM scoring near 1 on strong-fault model generation and multi-variable dependency resolution under the same protocol, the paper's main weakness claim would be contradicted.

Watch

Extended reading notes

Core claim

The paper's central claim is that current LLMs show partial competence in localized, sequential, and pattern-based engineering tasks, but their performance deteriorates sharply when tasks require non-local dependencies, type-level abstraction, formal causal modeling, or reasoning about nonlinear dynamic systems. The evidence comes from category scores across five research questions: high scores appear for state-transition comprehension, sequential understanding, and closed-world consistency, while low scores appear for multi-variable dependency resolution, optimization trade-offs, recognizing design absurdities, and constructing strong-fault logical or physical models. In forecasting, specialized baselines with far fewer parameters beat all LLMs on both tested datasets, and adding detailed system descriptions did not reliably improve the LLMs' predictions. The authors interpret the pattern as a fundamental mismatch between the statistical foundations of LLMs and the structured, constraint-driven reasoning that engineering requires.

Load-bearing premise

The load-bearing premise is that the three-point scoring system (0, 0.5, 1) anchored to junior- and senior-engineer competence measures actual engineering capability, even though no detailed rubric or inter-rater reliability check is reported; every category average and model ranking depends on that scoring.

Editorial extensions

If this is right

  • Under the paper's evidence, an LLM can be trusted as a drafting or hypothesis-generation aid for engineers, but not as an unsupervised agent for tasks requiring formal models or cross-component reasoning.
  • Tasks that stay local, concrete, and sequential are within reach; tasks that require abstraction, optimization trade-offs, or far-reaching causal chains are not.
  • For time-series forecasting, the paper's numbers imply that purpose-built specialist models with far fewer parameters are preferable to zero-shot LLM prompting.
  • The gap between larger cloud and smaller local models is non-linear, so local deployment may be adequate for easy categories but more likely to fail on hard ones.
  • Practical guidance for practitioners: avoid failure-mode reasoning and backward diagnosis with current LLMs; forward reasoning about normal operations is the more reliable direction.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The benchmark could serve as a reusable stress test for future models, because its questions are tied to concrete production scenarios rather than textbook exercises.
  • If the pattern survives further scaling, model size alone will not fix the weakest categories, since recognizing design absurdities and weighting trade-offs requires rejecting plausible-but-wrong answers rather than generating more text.
  • A direct extension would test whether the worst-response protocol materially changes the conclusions: rescoring with average or best responses would reveal how much of the reported weakness is due to reliability rather than competence.
  • A testable next step is to fine-tune models on strong-fault logical model examples and see whether the formal-modeling deficit shrinks, which would separate missing knowledge from architectural limits.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 7 minor

Summary. The paper introduces a curated dataset of over 100 engineering questions drawn from a real glass foaming plant and a simulated manufacturing environment, plus time-series forecasting tasks, and uses it to evaluate four LLMs (Qwen2.5-72B, DeepSeek-R1-Distill-Llama-70B, LLaMA 3.3-70B, and GPT-4o) under five research questions spanning world-model consistency, transitional and non-local relations, implicit design intention, temporal/causal reasoning, and forecasting. Responses to RQ1–RQ4 are graded on a 0/0.5/1 scale anchored to junior- and senior-engineer competence, with only the least accurate response retained from repeated prompts per question; category means are reported in Tables 1–4. RQ5 compares zero-shot LLM forecasting with LSTM and DLinear baselines using MSE. The paper concludes that LLMs show strengths in basic temporal and structural reasoning but struggle with abstract reasoning, formal modeling, and context-sensitive engineering logic, and therefore are best used as supervised assistants rather than autonomous agents.

Significance. If the central claim is established, the paper would provide a useful, real-world-grounded benchmark for engineering-oriented LLM evaluation and support the practical recommendation that general-purpose LLMs currently serve as assistant tools rather than autonomous engineers. The strengths include the publicly available repository with rated answers, the use of a real industrial plant and a simulation model rather than examination-style items, the systematic organization around five research questions, and the inclusion of specialized baselines in the forecasting study. However, the measurement pipeline behind Tables 1–4 is load-bearing: rubric-free subjective scoring, no inter-rater reliability check, small per-cell item counts, and an unreported minimum-of-repetitions aggregation rule. Because the Section 4 qualitative summary is read directly off these averages, the paper's central claim is not yet established at the level of precision the tables imply. The stress-test concern about the measurement pipeline is therefore valid and lands on the core of the paper.

major comments (5)
  1. [Section 3 (scoring paragraph), Tables 1–4] The entire quantitative basis for RQ1–RQ4 is a 0/0.5/1 scale anchored to 'junior engineer' and 'senior engineer' competence, but no rubric is provided, the raters are not described as blinded, and no inter-rater reliability statistic is reported. Because every table entry and the Section 4 qualitative summary depend on these ratings, the measurement pipeline is load-bearing and currently unsupported.
  2. [Section 3 (aggregation rule)] The sentence 'only the least accurate response for each question was considered for evaluation' introduces a minimum-of-repetitions statistic, but the manuscript does not report the number of repetitions, the sampling temperature, or the response variance for RQ1–RQ4. For a random per-question score with mean μ and variance σ², the expectation of the minimum over n draws is below μ and the shortfall grows with n and σ²; if n or σ² differs across models or categories, the table entries are not comparable. This is a concrete statistical threat to the cross-model and cross-category comparisons.
  3. [Tables 2 and 3] With 20 questions split across five categories, each cell in Tables 2 and 3 is the average of only four questions. A one-question change moves a cell by 0.25 on the 0–1 scale, so differences such as 0.38 vs. 0.63 in Multi-Variable Dependency Resolution are within the resolution of the measurement, and no confidence intervals or significance tests are provided. The same limitation affects Table 1 (45 questions across four categories, roughly 11 per cell) and Table 4, whose per-category question counts are not stated.
  4. [Section 3.4, Table 4] The number of questions for RQ4 is never reported, so the denominators underlying Logical Diagnosis, Detecting Causalities, Logical Modeling, and Physical Modeling are unknown. This is especially problematic because the table contains extreme values (e.g., LLaMA 3.3 at 0.00 and GPT-4o at 1.00 for Logical Modeling) that could each arise from a single item or a very small item set.
  5. [Section 3, footnote 4] The model referred to as 'DeepSeek-R1' is actually DeepSeek-R1-Distill-Llama-70B, a distilled 70B model, not the full DeepSeek-R1. The paper's model comparisons and any statement about 'DeepSeek-R1' performance should be relabeled and qualified; as written, the naming overstates the model evaluated and can mislead readers about the capabilities of the full R1 system.
minor comments (7)
  1. [Figure 1 caption] The caption 'The productions sells Primer Application, Polyurethane Foaming, and Trimming & Final Inspection are highlighted' is ungrammatical and should be rewritten.
  2. [Section 3.1] The phrase 'reasoning extends local scope seems to remain challenging' should be revised to something like 'reasoning beyond local scope seems to remain challenging.'
  3. [Section 3.1] The sentence 'their current capacity to construct reliable, generalizable world models do not meet the requirements' has a subject-verb agreement error and should read 'does not meet.'
  4. [Section 4] 'the fith research question' should be corrected to 'the fifth research question.'
  5. [Section 5] 'publically available' should be 'publicly available,' and 'overgeneralized or speculative completions' should be 'overgeneralize or speculate.'
  6. [Section 4] The claim that GPT-4o has 'about three times more parameters' than the locally hosted models is unsupported, since parameter counts of proprietary models are not publicly disclosed; this sentence should be removed or replaced with a comparison based on observable behavior.
  7. [Section 3.4] The characterization of high Logical Diagnosis performance as 'a trivial result given the large amount of literature' is an unsupported gloss; the existence of literature does not trivially imply high LLM performance, and the sentence should be revised or removed.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper is an empirical benchmark whose conclusions are read directly from scored LLM responses, with no fitted parameter, imported uniqueness theorem, or definitional reduction.

full rationale

The paper's inference chain is an empirical measurement pipeline: curated engineering questions are posed to four LLMs, responses are scored on a 0/0.5/1 scale, category averages are computed, and the qualitative conclusions in the abstract and Section 4 summarize those averages. There is no derivation in which an output quantity is defined in terms of the claimed result, nor any fitted parameter that is later renamed as a prediction. The only self-references are to prior work supplying the Three-Tank testbed [34], the repeated-question protocol [40], and general design-principles methodology [20]; none of these citations asserts the paper's central conclusion, and none forbids alternative interpretations. The RQ5 forecasting comparison is anchored to external baselines (LSTM, DLinear) and the public ETTh1 dataset, so that result is not guaranteed by construction. Concerns about the unrubricated scoring scale, missing inter-rater reliability, the worst-response aggregation rule, and the small number of questions per category are threats to measurement validity and should be treated as correctness risk, but they are not circular reasoning: the scores are not defined in terms of the conclusions, and no specific reduction of a claimed result to an input could be exhibited from the paper's text.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper's claims depend on domain assumptions about how to measure engineering competence, not on fitted parameters or new entities. The main assumptions are: (1) the three-point scoring scale anchored to junior/senior engineer competence is a valid measure; (2) taking the minimum score across repeated queries is an appropriate reliability metric; (3) the curated questions are representative of real-world engineering tasks. No free parameters or invented entities were introduced.

assumptions (3)
  • domain assumption The three-point scoring scale anchored to junior/senior engineer competence is a valid measure of engineering capability.
    Section 3 defines the scale but provides no rubric, grader identity, or inter-rater reliability, so the validity of all score tables rests on this assumption.
  • domain assumption Taking the minimum score across repeated queries is an appropriate metric for reliability.
    Section 3 states that only the least accurate response is considered, but the number of repetitions and score distributions are not reported, so the metric cannot be audited.
  • domain assumption The curated questions are representative of real-world engineering tasks.
    The questions are derived from two specific systems (a glass foaming plant and a simulated factory) plus 20 product-design questions; no external validation shows they capture the breadth of engineering practice.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Evaluating Large Language Models for Real-World Engineering Tasks." pith.science (2026). https://pith.science/paper/23GR6VAQ

@misc{pith2026250513484,
  author       = {Pith},
  title        = {Pith review of: Evaluating Large Language Models for Real-World Engineering Tasks},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/23GR6VAQ}},
  note         = {Machine review of arXiv:2505.13484}
}
read the original abstract

Large Language Models (LLMs) are transformative not only for daily activities but also for engineering tasks. However, current evaluations of LLMs in engineering exhibit two critical shortcomings: (i) the reliance on simplified use cases, often adapted from examination materials where correctness is easily verifiable, and (ii) the use of ad hoc scenarios that insufficiently capture critical engineering competencies. Consequently, the assessment of LLMs on complex, real-world engineering problems remains largely unexplored. This paper addresses this gap by introducing a curated database comprising over 100 questions derived from authentic, production-oriented engineering scenarios, systematically designed to cover core competencies such as product design, prognosis, and diagnosis. Using this dataset, we evaluate four state-of-the-art LLMs, including both cloud-based and locally hosted instances, to systematically investigate their performance on complex engineering tasks. Our results show that LLMs demonstrate strengths in basic temporal and structural reasoning but struggle significantly with abstract reasoning, formal modeling, and context-sensitive engineering logic.

Figures

Figures reproduced from arXiv: 2505.13484 by the authors.

Figure 1
Figure 1. Modular glass foaming plant that served as first use case. The [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

52 extracted references · 43 canonical work pages

  1. [1]

    L. S. Balhorn, M. Caballero, and A. M. Schweidtmann. Toward auto- correction of chemical process flowsheets using large language models. In Computer Aided Chemical Engineering, volume 53, pages 3109–3114. Elsevier, 2024

  2. [2]

    T. Ban, L. Chen, D. Lyu, X. Wang, Q. Zhu, and H. Chen. Llm-driven causal discovery via harmonized prior. IEEE Transactions on Knowl- edge and Data Engineering , 2025

  3. [3]

    A. Bochman. A logical theory of causality . MIT Press, 2021

  4. [4]

    A. A. Chuganskaya, A. K. Kovalev, and A. Panov. The Problem of Concept Learning and Goals of Reasoning in Large Language Mod- els. In P. Garc´ ıa Bringas, H. P´ erez Garc´ ıa, F. J. Mart´ ınez de Pis´ on, F. Mart´ ınez ´Alvarez, A. Troncoso, ´A. Herrero, J. L. Calvo Rolle, H. Quinti´ an, and E. Corchado, editors, Hybrid Artificial Intelligent Systems, Lec...

  5. [5]

    N. Crilly. The roles that artefacts play: technical, social and aesthetic functions. Design Studies, 31(4):311–344, 2010

  6. [6]

    Guo, and e

    DeepSeek-AI, D. Guo, and e. a. Yang. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning, Jan. 2025

  7. [7]

    Efatmaneshnik, J

    M. Efatmaneshnik, J. Bradley, and M. J. Ryan. Complexity and fragility in system of systems. International Journal of System of Sys- tems Engineering, 7(4):294–312, 2016

  8. [8]

    ElMaraghy, H

    W. ElMaraghy, H. ElMaraghy, T. Tomiyama, and L. Monostori. Complexity in engineering design and manufacturing. CIRP annals , 61(2):793–814, 2012. 19

Show all 52 references
  1. [9]

    Feldman, G

    A. Feldman, G. Provan, and A. Van Gemund. Approximate model- based diagnosis using greedy stochastic search. Journal of Artificial Intelligence Research, 38:371–413, 2010

  2. [10]

    J. Feng, S. Russell, and J. Steinhardt. Monitoring latent world states in language models with propositional probes, 2024

  3. [11]

    C. Fields. Object permanence. In Encyclopedia of evolutionary psycho- logical science, pages 5505–5510. Springer, 2021

  4. [12]

    Frisk, M

    E. Frisk, M. Krysander, and D. Jung. A toolbox for analysis and de- sign of model based diagnosis systems for large scale models. IFAC- PapersOnLine, 50(1):3287–3293, 2017

  5. [13]

    Grattafiori and e

    A. Grattafiori and e. a. Dubey. The Llama 3 Herd of Models, Nov. 2024

  6. [14]

    Hirtreiter, L

    E. Hirtreiter, L. Schulze Balhorn, and A. M. Schweidtmann. Toward automatic generation of control structures for process flow diagrams with large language models. AIChE Journal, 70(1):e18259, 2024

  7. [15]

    Hochreiter and J

    S. Hochreiter and J. Schmidhuber. Long Short-Term Memory | Neu- ral Computation. https://dl.acm.org/doi/10.1162/neco.1997.9.8.1735, 1997

  8. [16]

    Huang and K

    J. Huang and K. C.-C. Chang. Towards reasoning in large language models: A survey, 2023

  9. [17]

    Hurst, A

    A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radford, et al. Gpt-4o system card. arXiv preprint arXiv:2410.21276 , 2024

  10. [18]

    S. Kang, G. An, and S. Yoo. A preliminary evaluation of llm-based fault localization. arXiv preprint arXiv:2308.05487 , 2023

  11. [19]

    S. Kato, C. Zhang, and M. Kano. Simple algorithm for judging equiv- alence of differential-algebraic equation systems. Scientific reports , 13(1):11534, 2023

  12. [20]

    K¨ ocher, R

    A. K¨ ocher, R. Heesch, N. Widulle, A. Nordhausen, J. Putzke, A. Wind- mann, and O. Niggemann. A research agenda for ai planning in the field of flexible production systems. In 2022 IEEE 5th International Confer- ence on Industrial Cyber-Physical Systems (ICPS) , pages 1–8. IEEE, 2022

  13. [21]

    K. Li, A. K. Hopkins, D. Bau, F. Vi´ egas, H. Pfister, and M. Watten- berg. Emergent world representations: Exploring a sequence model trained on a synthetic task. 2023 International Conference on Learning Representations, 2024. 20

  14. [22]

    Meier, F

    S. Meier, F. T¨ oper, J. Gebele, B. Rachinger, S. Klarmann, and J. Franke. Knowledge mining using generative ai for causal discovery in electronics production. In 2024 47th International Spring Seminar on Electronics Technology (ISSE), volume 2024, pages 1–10, 2024

  15. [23]

    Merkelbach, A

    S. Merkelbach, A. Diedrich, A. Sztyber-Betley, L. Trav´ e-Massuy` es, E. Chanthery, O. Niggemann, and R. Dumitrescu. Using multi-modal llms to create models for fault diagnosis (short paper). In 35th Inter- national Conference on Principles of Diagnosis and Resilient Systems (...

  16. [24]

    Ogundare, G

    O. Ogundare, G. Q. Araya, I. Akrotirianakis, and A. Shukla. Resiliency analysis of llm generated models for industrial automation. In 2023 In- ternational Conference on Modeling & E-Information Research, Artifi- cial Learning and Digital Applications (ICMERALDA), pages 113–116...

  17. [25]

    G. Pahl, W. Beitz, L. Blessing, J. Feldhusen, K.-H. Grote, and K. Wal- lace. Engineering Design: A Systematic Approach . Springer-Verlag London Limited, London, third edition edition, 2007

  18. [26]

    Patel and E

    R. Patel and E. Pavlick. Mapping language models to grounded concep- tual spaces. In International Conference on Learning Representations , 2022

  19. [27]

    J. Pearl. Causality. Cambridge university press, 2009

  20. [28]

    Peifeng, L

    L. Peifeng, L. Qian, X. Zhao, and B. Tao. Joint knowledge graph and large language model for fault diagnosis and its application in aviation assembly. IEEE Transactions on Industrial Informatics , 2024

  21. [29]

    Pill and J

    I. Pill and J. De Kleer. Challenges for model-based diagnosis. In 35th International Conference on Principles of Diagnosis and Resilient Sys- tems (DX 2024) , pages 6–1. Schloss Dagstuhl–Leibniz-Zentrum f¨ ur In- formatik, 2024

  22. [30]

    Plaat, A

    A. Plaat, A. Wong, S. Verberne, J. Broekens, N. van Stein, and T. Back. Reasoning with large language models, a survey, 2024

  23. [31]

    L. M. Reinpold, M. Schieseck, L. P. Wagner, F. Gehlhoff, and A. Fay. Exploring llms for verifying technical system specifications against re- quirements. In 2024 IEEE 3rd Industrial Electronics Society Annual On-Line Conference (ONCON), page 1–6. IEEE, Dec. 2024

  24. [32]

    A. Saba, R. Hantach, and M. Benslimane. Text detection and recog- nition from piping and instrumentation diagrams. In 2023 8th Inter- 21 national Conference on Image, Vision and Computing (ICIVC) , pages 711–716. IEEE, 2023

  25. [33]

    Sowa and A

    K. Sowa and A. Przegalinska. From expert systems to generative arti- ficial experts: A new concept for human-ai collaboration in knowledge work. Journal of Artificial Intelligence Research , 82:2101–2124, 2025

  26. [34]

    Steude, A

    H. Steude, A. Windmann, and O. Niggemann. Learning Physical Concepts in CPS: A Case Study with a Three-Tank System. IFAC- PapersOnLine, 55(6):15–22, Jan. 2022

  27. [35]

    R. B. Stone and K. L. Wood. Development of a Functional Basis for Design. In K. Otto, editor, 11th International Conference on Design Theory and Methodology , Proceedings of the 1999 ASME Design En- gineering Technical Conferences, pages 261–275, New York, NY, 1999. ASME

  28. [36]

    Struss and O

    P. Struss and O. Dressler. ” physical negation” integrating fault models into the general diagnostic engine. In IJCAI, volume 89, pages 1318– 1323, 1989

  29. [37]

    Suresh and S

    V. Suresh and S. Rawat. Gpt takes the sat: Tracing changes in test difficulty and students’ math performance. Technical report, working paper, Social Sciences Research Network, 2024

  30. [38]

    H. Tang, C. Zhang, M. Jin, Q. Yu, Z. Wang, X. Jin, Y. Zhang, and M. Du. Time Series Forecasting with LLMs: Understanding and En- hancing Model Capabilities. SIGKDD Explor. Newsl. , 26(2):109–118, Jan. 2025

  31. [39]

    Vicente, J

    M. Vicente, J. Guarda, F. Batista, et al. Gutenbrain: An architecture for equipment technical attributes extraction from piping & instrumen- tation diagrams. In KDIR, pages 204–211, 2022

  32. [40]

    Vranjeˇ s, J

    D. Vranjeˇ s, J. Ehrhardt, R. Heesch, L. Moddemann, H. S. Steude, and O. Niggemann. Design principles for falsifiable, replicable and repro- ducible empirical machine learning research. In 35th International Con- ference on Principles of Diagnosis and Resilient Systems (DX 202...

  33. [41]

    Vukovi´ c and S

    M. Vukovi´ c and S. Thalmann. Causal discovery in manufacturing: A structured literature review. Journal of Manufacturing and Materials Processing, 6(1):10, 2022

  34. [42]

    L. Wang, C. Ma, X. Feng, Z. Zhang, H. Yang, J. Zhang, Z. Chen, J. Tang, X. Chen, Y. Lin, W. X. Zhao, Z. Wei, and J. Wen. A survey on large language model based autonomous agents. Frontiers of Computer Science, 18(6):186345, Mar. 2024. 22

  35. [43]

    Windmann, P

    A. Windmann, P. Wittenberg, M. Schieseck, and O. Niggemann. Ar- tificial intelligence in industry 4.0: A review of integration challenges for industrial systems. In 2024 IEEE 22nd International Conference on Industrial Informatics (INDIN) , pages 1–8, 2024

  36. [44]

    Y. Wu, Z. Li, J. M. Zhang, M. Papadakis, M. Harman, and Y. Liu. Large language models in fault localisation. arXiv preprint arXiv:2308.15276, 2023

  37. [45]

    Y. Xia, N. Jazdi, J. Zhang, M. Weyrich, et al. Control indus- trial automation system with large language models. arXiv preprint arXiv:2409.18009, Sept. 2024

  38. [46]

    F. Xu, Q. Lin, J. Han, T. Zhao, J. Liu, and E. Cambria. Are large lan- guage models really good logical reasoners? a comprehensive evaluation and beyond. IEEE Transactions on Knowledge and Data Engineering , 2025

  39. [47]

    Yamada, Y

    Y. Yamada, Y. Bao, A. K. Lampinen, J. Kasai, and I. Yildirim. Eval- uating spatial understanding of large language models. arXiv preprint arXiv:2310.14540, 2023

  40. [48]

    A. Yang, B. Yang, and et. al. Qwen2 Technical Report, Sept. 2024

  41. [49]

    A. Zeng, M. Chen, L. Zhang, and Q. Xu. Are Transformers Effective for Time Series Forecasting? Proceedings of the AAAI Conference on Artificial Intelligence, 37(9):11121–11128, June 2023

  42. [50]

    Zhang, X

    Z. Zhang, X. Wang, Z. Zhang, H. Li, Y. Qin, and W. Zhu. Llm4dyg: can large language models solve spatial-temporal problems on dynamic graphs? In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining , pages 4350–4361, 2024

  43. [51]

    H. Zhou, S. Zhang, J. Peng, S. Zhang, J. Li, H. Xiong, and W. Zhang. Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting. Proceedings of the AAAI Conference on Artificial Intelli- gence, 35(12):11106–11115, May 2021. 23

  44. [2023]

    Springer Nature Switzerland and Imprint Springer

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.