REVIEW 3 major objections 3 minor 48 references
This paper argues that structural-engineering LLM agents should be judged by the consistency of an executable evidence chain, not by the plausibility of final text or artifacts; it reports that artifact presence alone overstates reliability
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-04 04:14 UTC pith:EQRT7ENW
load-bearing objection A genuinely useful, carefully scoped benchmark for engineering LLM agents, whose headline gap is real under its own fixture rules but must be read as convention-conformity until independently reviewed alternatives are shown to be rare. the 3 major comments →
StructureClaw: Traceable LLM Agents and an Executable Benchmark for Structural Engineering Workflows
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On the paper's own terms, the central discovery is a measurement: when success requires an unbroken evidence chain—strict one-to-one structural-model matching to a reference fixture and numerical-response agreement with frozen responses from the selected analysis engine—generic LLM execution collapses from 87.0% artifact-presence to 22.0% end-to-end success, while the artifact-centered automatic workflow retains 82.9%. This gap, the paper argues, shows that plausible models do not imply reliable engineering workflows. The same measurements localize the remaining bottlenecks: semantic-state consistency in interactive settings and executable model reconstruction from images or DXF drawings in
What carries the argument
The carrying mechanism is the evidence chain, defined as the linked set of artifacts, tool executions, validation records, and terminal decisions that support an engineering result. StructureClaw realizes it with persistent artifact state, governed skills, typed tools, a shared structural-model protocol, and local analysis backends; StructureClaw-Bench operationalizes it with fixture-defined typed assertions. The central evaluative machinery is strict one-to-one matching of nodes, elements, restraints, properties, and loads against a reference fixture, followed by numerical matching of mapped solver responses to frozen same-engine references, with a trial succeeding only when every required
Load-bearing premise
The benchmark's own reference fixtures and tolerance policy are treated as ground truth, so the end-to-end numbers measure agreement with StructureClaw's conventions; if those fixtures do not cover the space of valid engineering idealizations, the reliability claims shift accordingly.
What would settle it
An independent structural engineer could blind-review randomized trials from the generic-only condition, comparing those that produce a plausible model but fail strict matching against the reference fixture. If a large share of those failing trials is judged professionally sound, the benchmark's binary would be measuring convention-following rather than engineering reliability; if most are judged non-equivalent or unsafe, the evidence-chain requirement is necessary.
If this is right
- Benchmarks that score final text or artifact presence overstate an LLM agent's reliability for structural engineering; executable, evidence-chain assertions are needed to expose missing or inconsistent intermediates.
- The automatic workflow raised end-to-end success for every one of the nine evaluated text-agent configurations, with gains ranging from roughly 41 to 85 percentage points, indicating the benefit is not specific to one model.
- High local performance on clarification, safe abstention, or artifact presence does not guarantee end-to-end success, so workflow-level measures should be reported alongside stage-level diagnostics.
- Multimodal structural reconstruction is bottlenecked less by recognizing the structure type than by converting perception into a complete executable model with correct topology, restraints, properties, and loads.
- Reporting single-attempt reliability and three-trial stability is more deployment-relevant than retry-oriented metrics such as pass@3.
Where Pith is reading between the lines
- Editorial extension: the evidence-chain evaluation design could transfer to other engineering domains, such as mechanical or geotechnical workflows, where a final report is only trustworthy if model, solver trace, checks, and report remain linked.
- Editorial extension: the paper explicitly scopes its strict fixture matching in its reference-fixture appendix and validity-boundary table; an independent human-engineer review of rejected cases would test whether the benchmark measures engineering reliability or agreement with its own conventions.
- Editorial extension: because the automatic-versus-generic comparison changes routing, guidance, artifact expectations, and validation together, the 60.9-point lift is a system-level result; component-level ablations would be a natural follow-up to attribute the gains.
- Editorial extension: feeding the benchmark's positive-and-negative evidence conjunction back into agent training—rewarding safe non-execution only when paired with a clarification or recovery trace—might reduce the stable-failure cases, which remain common in multimodal reconstruction.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents StructureClaw, an LLM-agent workbench for structural engineering that persists typed engineering artifacts (interpreted requirements, structural model, validation records, solver outputs, checks, report) under governed skills and typed tools, and StructureClaw-Bench, a 150-scenario executable benchmark spanning standard workflows, interactive robustness, and multimodal model reconstruction. Success requires a conjunctive evidence chain: strict one-to-one structural-model matching against author-generated reference fixtures plus numerical agreement with frozen same-engine responses for analyzable cases, and positive clarification/recovery evidence with safe non-execution for interactive cases. Across nine text-agent configurations, automatic StructureClaw reaches 82.9% pooled E2E Success versus 22.0% for generic-only execution; the generic-only model-artifact rate is 87.0% but strict structural fidelity is only 25.0%. Interactive robustness reaches 82.6% and multimodal reconstruction 53.2%. The central claim is that artifact presence alone overstates reliability and that failure bottlenecks differ by task family.
Significance. The contribution is substantial if the results hold: it is one of the few agentic-engineering benchmarks that executes and checks the full artifact chain rather than scoring final text, and it is methodologically careful — three retained trials with stability partitions, scenario-clustered bootstrap intervals with sign-flip tests and Holm correction, leakage guards, frozen evaluation snapshots, and cross-validation of 53 canonical fixtures against an independent direct-stiffness implementation plus six pairwise engine records. The paper is also unusually candid about its own boundaries (App. A.6, B.2, B.3, Table 8). The evidence ladder (87.0% artifact presence vs. 22.0% E2E in generic mode) is internally compelling. The principal risks are external validity: the strict matcher encodes the authors' own idealizations, the fixtures are author-generated, and the semantic judge is one of the evaluated models, so the absolute rates and the Auto–Generic gap currently measure agreement with the benchmark's conventions, pending professional-validation evidence.
major comments (3)
- [§5.1, §B.4, Table 10, App. B.2] The headline 60.9-point Auto–Generic gap assumes the strict one-to-one matcher is a fair proxy for engineering validity. §B.4 requires node/element/load precision and recall of exactly 1 with 0.05 m and 5% tolerances, so valid alternative discretizations, section groupings, or load idealizations are scored as failures. App. B.2 concedes this; Table 10 shows Generic collapsing from 87.0% artifact presence to 25.0% strict fidelity, with 838/1,053 Generic failures first failing strict-model fidelity (App. E.4). Since the Auto skills encode the same design-basis conventions as the author-generated fixtures, part of the gap may measure convention-conformity rather than engineering reliability. The internal evidence-ladder claim is supported; the externally readable reliability comparison needs support. Request: an engineer-reviewed sample of rejected Generic models, a matching-sensitivity ana
- [§B.3, §5.2, Table 2] The semantic judge is GLM-5.2 — itself one of the nine evaluated agents — run uncalibrated against professional raters (§B.3). Semantic evaluation is the most common first failed stage in interactive tasks (96/235; App. E.4), so interactive E2E rates and the 'semantic state consistency as dominant bottleneck' finding partly measure one evaluated model's interpretation of the rubric. The disclosure that judge outcomes 'measure consistency with the fixed benchmark rubric rather than professional correctness' mitigates but does not remove the same-model circularity: the top scorer (GLM-5.2, 90.0%) is also the arbiter of semantic success. Request: calibrate the judge on a sample against professional raters with agreement reported, and re-score a subset of interactive scenarios with a non-evaluated judge to show robustness.
- [§C.1, Eq. (5)] The primary metric pools only 'retained' trials, and §C.1 records 88 infrastructure reruns (78/3/7 across batches) that add no scoring rows. However, the rule by which a run is classified as an infrastructure failure — versus a scored model outcome — is not specified, nor are per-condition rerun counts and the rerun decision procedure. Without a pre-specified, outcome-independent exclusion rule, the clean 4,950-outcome matrix is not fully auditable and the stability partitions (Table 9) could be sensitive to exclusion choices. Request: state the rerun criterion, its pre-registration status, and per-condition rerun counts; confirm exclusions are outcome-independent.
minor comments (3)
- [§B.2 vs §B.5] The paper states there are 59 model fixtures (App. B.2) yet reports that all 50 standard and 47 multimodal scenarios evaluate numerical fidelity (App. B.5), implying 97 fixture-requiring scenarios. Clarify the fixture–scenario mapping: which scenarios share fixtures, which lack a reference fixture, and why the counts differ.
- [§5.1] The claim that 'every paired 95% scenario-bootstrap interval is strictly positive' is reported without interval bounds. Since Figure 4 shows pooled rates rather than paired differences, please include the interval limits (e.g., in a supplementary table) so readers can see the magnitude of uncertainty.
- [§C.1, Table 3] Kimi-K2.6's vision model runs at temperature 1 while all other vision configurations use temperature 0 (§C.1). This is a confound for its unusually low multimodal E2E (13.3%) and should be noted in the Table 3 caption or discussed as a limitation.
Circularity Check
No significant circularity: the benchmark's relative comparisons are symmetric under identical fixtures, no fitted parameter is renamed as a prediction, and the acknowledged fixture/judge limitations are validity scoping, not circular reductions.
full rationale
I walked the paper's derivation chain from the evidence-chain formulation, through Benchmark assertions (Eq. 1), to the Auto-vs-Generic comparison and the conclusion that 'artifact presence alone overstates reliability' (§5.1). No load-bearing step reduces to its inputs by construction. The main quantitative result (82.9% vs 22.0% E2E) is an empirical measurement under a fixed evaluation protocol, not a quantity fitted to the same data. The structural-fidelity and numerical-response assertions are deterministic and defined independently of agent outputs; the fixtures are generated with cross-process and cross-engine validation (§B.2). The Auto-vs-Generic and OpenClaw comparisons apply identical fixtures, tools, and assertions to both conditions, so the relative claims are symmetric and not convention-conforming by construction. The semantic judge (GLM-5.2) is used for a subset of natural-language assertions and is disclosed as uncalibrated (§B.3); this is a measurement limitation, not a circular derivation. The paper explicitly scopes its claims to agreement with its declared fixtures and operational rules rather than professional design validity (§B.2, Table 8), and lists external review and externally authored challenge cases as extensions. Self-citations appear only as contextual references in Related Work and are not load-bearing. Thus no specific circular step can be exhibited under the required evidence standard.
Axiom & Free-Parameter Ledger
free parameters (3)
- coordinate matching tolerance =
0.05 m
- property/load relative tolerance =
5%
- numerical response tolerance =
1% relative, 1e-8 absolute
axioms (6)
- standard math Euler-Bernoulli closed-form and an independent direct-stiffness implementation are valid cross-checks for canonical fixtures.
- domain assumption One-to-one strict model matching is the correctness criterion for structural-model fidelity.
- domain assumption Frozen solver responses from the requested engine define numerical ground truth.
- domain assumption The 150 author-designed scenarios represent the intended structural-engineering agent workflows.
- ad hoc to paper A fixed, uncalibrated LLM (GLM-5.2) can judge natural-language semantic assertions.
- domain assumption Safe non-execution is operationalized as no analysis artifact plus positive clarification or semantic evidence.
read the original abstract
Addressing a structural-engineering request requires more than a single answer; it requires a chain of interdependent artifacts: interpreted requirements, a computable model, validation records, solver outputs, applicable engineering checks, and a final report. Evaluations centered on question answering or script generation may therefore reward fluent outputs even when the underlying workflow is incomplete, inconsistent, or non-executable. We present StructureClaw, an artifact-centered workbench in which LLM agents operate through governed engineering skills, typed tools, shared artifact state, and local analysis backends, together with StructureClaw-Bench, an executable benchmark of 150 controlled scenarios spanning standard workflows, interactive robustness, and multimodal structural-model reconstruction. Its analyzable standard and multimodal cases require both strict one-to-one structural-model matching and numerical-response agreement with frozen reference responses from the selected analysis engine; interactive cases instead require positive clarification or recovery evidence together with safe non-execution when appropriate. A trial succeeds only when every fixture-required assertion passes. Across nine text-agent configurations, generic-only execution passed the model-artifact check in 87.0% of retained outcomes but achieved only 22.0% E2E Success, whereas automatic StructureClaw reached 82.9%. Interactive and multimodal evaluations further identify semantic state consistency and executable model reconstruction as the dominant remaining bottlenecks. The code and benchmark are available at https://github.com/structureclaw/structureclaw.
Figures
Reference graph
Works this paper leans on
-
[1]
Multi-agent large language model framework for code-compliant automated design of rein- forced concrete structures.Automation in Construction, 177: 106331, 2025
Jinxin Chen and Yi Bao. Multi-agent large language model framework for code-compliant automated design of rein- forced concrete structures.Automation in Construction, 177: 106331, 2025. 2
2025
-
[2]
RAMASC: A retrieval-augmented multi-agent framework for automated structural calculation.Advanced Engineering Informatics, 74:104698, 2026
Kichang Choi, Minwoo Jeong, Taegeon Kim, Seokhwan Kim, Seungwon Baek, and Hongjo Kim. RAMASC: A retrieval-augmented multi-agent framework for automated structural calculation.Advanced Engineering Informatics, 74:104698, 2026. 2
2026
-
[3]
BIMgent: Towards autonomous building modeling via computer-use agents
Zihan Deng, Changyu Du, Stavros Nousias, and Andr ´e Bor- rmann. BIMgent: Towards autonomous building modeling via computer-use agents. InICML 2025 Workshop on Com- puter Use Agents, 2025. 2
2025
-
[4]
Text2BIM: Generating building models using a large language model-based multiagent framework.Journal of Computing in Civil Engineering, 40(2):04025142, 2026
Changyu Du, Sebastian Esser, Stavros Nousias, and Andr ´e Borrmann. Text2BIM: Generating building models using a large language model-based multiagent framework.Journal of Computing in Civil Engineering, 40(2):04025142, 2026. 2
2026
-
[5]
PAL: Program-aided language models
Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. PAL: Program-aided language models. InProceedings of the 40th International Conference on Machine Learning, pages 10764–10799. PMLR, 2023. 2
2023
-
[6]
Ziheng Geng, Jiachen Liu, Ran Cao, Lu Cheng, Haifeng Wang, and Minghui Cheng. A lightweight large language model-based multi-agent system for 2D frame structural analysis.arXiv preprint arXiv:2510.05414, 2025. 2
arXiv 2025
-
[7]
Ziheng Geng, Ian Franklin, Santiago Martinez, Jiachen Liu, Yunhe Zhao, and Minghui Cheng. Agentic large language models for automated structural analysis of 3D frame sys- tems.arXiv preprint arXiv:2606.06525, 2026
Pith/arXiv arXiv 2026
-
[8]
Ziheng Geng, Jiachen Liu, Ian Franklin, Ran Cao, Dan M. Frangopol, and Minghui Cheng. Automating structural anal- ysis across multiple software platforms using large language models.arXiv preprint arXiv:2604.09866, 2026. 2
Pith/arXiv arXiv 2026
-
[9]
Intel- ligent design of shear wall layout based on diffusion mod- els.Computer-Aided Civil and Infrastructure Engineering, 39(23):3610–3625, 2024
Yi Gu, Yuli Huang, Wenjie Liao, and Xinzheng Lu. Intel- ligent design of shear wall layout based on diffusion mod- els.Computer-Aided Civil and Infrastructure Engineering, 39(23):3610–3625, 2024. 2
2024
-
[10]
Toward engineering AGI: Benchmarking the engineering design capabilities of LLMs
Xingang Guo, Yaxin Li, Xiangyi Kong, Yilan Jiang, Xiayu Zhao, Zhihua Gong, Yufan Zhang, Daixuan Li, Tianle Sang, Beixiao Zhu, et al. Toward engineering AGI: Benchmarking the engineering design capabilities of LLMs. InAdvances in Neural Information Processing Systems, 2025. Datasets and Benchmarks Track. 2, 3
2025
-
[11]
Enhancing reliability and automation of LLM- based structural analysis using a hybrid multi-agent pipeline
Seokjae Heo. Enhancing reliability and automation of LLM- based structural analysis using a hybrid multi-agent pipeline. Scientific Reports, 16:19458, 2026. 2
2026
-
[12]
Leveraging large language models for BIM-based automated compliance checking.Au- tomation in Construction, 182:106707, 2026
Odin Iversen and Lizhen Huang. Leveraging large language models for BIM-based automated compliance checking.Au- tomation in Construction, 182:106707, 2026. 2
2026
-
[13]
SoK: Agentic skills – beyond tool use in LLM agents.arXiv preprint arXiv:2602.20867, 2026
Yanna Jiang, Delong Li, Haiyu Deng, Baihe Ma, Xu Wang, Qin Wang, and Guangsheng Yu. SoK: Agentic skills – beyond tool use in LLM agents.arXiv preprint arXiv:2602.20867, 2026. 2
Pith/arXiv arXiv 2026
-
[14]
Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan
Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. SWE- bench: Can language models resolve real-world GitHub is- sues? InInternational Conference on Learning Representa- tions, 2024. 2
2024
-
[15]
The framework and implementation of using large language mod- els to answer questions about building codes and standards
Isaac Joffe, George Felobes, Youssef Elgouhari, Moham- mad Talebi Kalaleh, Qipei Mei, and Ying Hei Chui. The framework and implementation of using large language mod- els to answer questions about building codes and standards. Journal of Computing in Civil Engineering, 39(4):05025004,
-
[16]
A review of LLMs and their applications in the architecture, engineering and construction industry.Artificial Intelligence Review, 58:250,
Dimitrios Kampelopoulos, Athina Tsanousa, Stefanos Vrochidis, and Ioannis Kompatsiaris. A review of LLMs and their applications in the architecture, engineering and construction industry.Artificial Intelligence Review, 58:250,
-
[17]
Aleksei Kondratenko, Mussie Birhane, Houssame E. Hsain, and Guido Maciocci. AECV-Bench: Benchmarking multi- modal models on architectural and engineering drawings un- derstanding.arXiv preprint arXiv:2601.04819, 2026. 3
arXiv 2026
-
[18]
Yinsheng Li, Zhen Dong, and Yi Shao. DrafterBench: Benchmarking large language models for tasks automation in civil engineering.arXiv preprint arXiv:2507.11527, 2025. 3
Pith/arXiv arXiv 2025
-
[19]
AECBench: A hierarchical benchmark for knowledge evaluation of large language models in the AEC field.Advanced Engineering Informatics, 71:104314, 2026
Chen Liang, Zhaoqi Huang, Haofen Wang, Fu Chai, Chuny- ing Yu, Huanhuan Wei, Zhengjie Liu, Yanpeng Li, Hongjun Wang, Ruifeng Luo, and Xianzhong Zhao. AECBench: A hierarchical benchmark for knowledge evaluation of large language models in the AEC field.Advanced Engineering Informatics, 71:104314, 2026. 3
2026
-
[20]
Haoran Liang, Mohammad Talebi Kalaleh, and Qipei Mei. Integrating large language models for automated structural analysis.arXiv preprint arXiv:2504.09754, 2025. 2
Pith/arXiv arXiv 2025
-
[21]
Haoran Liang, Yufa Zhou, Mohammad Talebi Kalaleh, and Qipei Mei. Automating structural engineering work- flows with large language model agents.arXiv preprint arXiv:2510.11004, 2025. 2
arXiv 2025
-
[22]
Automated structural design of shear wall residential buildings using generative adversarial networks
Wenjie Liao, Xinzheng Lu, Yuli Huang, Zhe Zheng, and Yuanqing Lin. Automated structural design of shear wall residential buildings using generative adversarial networks. Automation in Construction, 132:103931, 2021. 2
2021
-
[23]
Generative AI design for building structures.Au- tomation in Construction, 157:105187, 2024
Wenjie Liao, Xinzheng Lu, Yifan Fei, Yi Gu, and Yuli Huang. Generative AI design for building structures.Au- tomation in Construction, 157:105187, 2024. 2
2024
-
[24]
SkillAct: Using skill abstractions im- proves LLM agents
Anthony Zhe Liu, Jongwook Choi, Sungryull Sohn, Yao Fu, Jaekyeom Kim, Dong-Ki Kim, Xinhe Wang, Jaewon Yoo, and Honglak Lee. SkillAct: Using skill abstractions im- proves LLM agents. InICML 2024 Workshop on LLMs and Cognition, 2024. 2
2024
-
[25]
A large language model- empowered agent for reliable and robust structural analysis
Jiachen Liu, Ziheng Geng, Ran Cao, Lu Cheng, Paolo Bocchini, and Minghui Cheng. A large language model- empowered agent for reliable and robust structural analysis. Structure and Infrastructure Engineering, pages 1–16, 2026. 2
2026
-
[26]
AgentBench: Evaluating LLMs as agents
Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhengx- iao Du, Chenhui Zhang, Sheng Shen, Tianjun Zhang, Yu Su, Huan Sun, Minlie Huang, Yuxiao Dong, and Jie Tang. AgentBench: Evaluating LLMs as agents. InInternational 9 Conference on Learning Represent...
2024
-
[27]
Xinzheng Lu, Wenjie Liao, Yu Zhang, and Yuli Huang. Intel- ligent structural design of shear wall residence using physics- enhanced generative adversarial networks.Earthquake En- gineering & Structural Dynamics, 51(7):1657–1676, 2022. 2
2022
-
[28]
Large lan- guage model-driven code compliance checking in building information modeling.Electronics, 14(11):2146, 2025
Soumya Madireddy, Lu Gao, Zia Ud Din, Kinam Kim, Ahmed Senouci, Zhe Han, and Yunpeng Zhang. Large lan- guage model-driven code compliance checking in building information modeling.Electronics, 14(11):2146, 2025. 2
2025
-
[29]
Harsh Mankodiya, Chase Gallik, Theodoros Galanos, and Andriy Mulyar. AEC-Bench: A multimodal benchmark for agentic systems in architecture, engineering, and construc- tion.arXiv preprint arXiv:2603.29199, 2026. 2, 3
arXiv 2026
-
[30]
OpenSees: A framework for earthquake en- gineering simulation.Computing in Science & Engineering, 13(4):58–66, 2011
Frank McKenna. OpenSees: A framework for earthquake en- gineering simulation.Computing in Science & Engineering, 13(4):58–66, 2011. 4
2011
-
[31]
Bharathi Kannan Nithyanantham, Tobias Sesterhenn, Ash- win Nedungadi, Sergio Peral Garijo, Janis Zenkner, Chris- tian Bartelt, and Stefan L¨udtke. MCP4IFC: IFC-based build- ing design using large language models.arXiv preprint arXiv:2511.05533, 2025. 2
arXiv 2025
-
[32]
Patil, Tianjun Zhang, Xin Wang, and Joseph E
Shishir G. Patil, Tianjun Zhang, Xin Wang, and Joseph E. Gonzalez. Gorilla: Large language model connected with massive APIs. InAdvances in Neural Information Process- ing Systems, pages 126544–126565, 2024. 2
2024
-
[33]
Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng- Jie Ji, Vishnu Suresh, Ion Stoica, and Joseph E
Shishir G. Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng- Jie Ji, Vishnu Suresh, Ion Stoica, and Joseph E. Gonzalez. The Berkeley function calling leaderboard (BFCL): From tool use to agentic evaluation of large language models. In Proceedings of the 42nd International Conference on Ma- chine Learning, pages 48371–48392. PMLR, 2025. 2
2025
-
[34]
Sizhong Qin, Hong Guan, Wenjie Liao, Yi Gu, Zhe Zheng, Hongjing Xue, and Xinzheng Lu. Intelligent design and opti- mization system for shear wall structures based on large lan- guage models and generative artificial intelligence.Journal of Building Engineering, 95:109996, 2024. 2
2024
-
[35]
Leveraging data-driven artificial intelligence in optimization design for building structures: A review.Engineering Struc- tures, 341:120810, 2025
Sizhong Qin, Yifan Fei, Wenjie Liao, and Xinzheng Lu. Leveraging data-driven artificial intelligence in optimization design for building structures: A review.Engineering Struc- tures, 341:120810, 2025. 2
2025
-
[36]
ToolLLM: Facilitating large language models to mas- ter 16000+ real-world APIs
Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, Sihan Zhao, Lauren Hong, Runchu Tian, Ruobing Xie, Jie Zhou, Mark Gerstein, Dahai Li, Zhiyuan Liu, and Maosong Sun. ToolLLM: Facilitating large language models to mas- ter 16000+ real-world APIs. InInternational Conference on Learning Represent...
2024
-
[37]
GPT models in construction industry: Oppor- tunities, limitations, and a use case validation.Developments in the Built Environment, 17:100300, 2024
Abdullahi Saka, Ridwan Taiwo, Nurudeen Saka, Ba- batunde Abiodun Salami, Saheed Ajayi, Kabiru Akande, and Hadi Kazemi. GPT models in construction industry: Oppor- tunities, limitations, and a use case validation.Developments in the Built Environment, 17:100300, 2024. 1
2024
-
[38]
Toolformer: Lan- guage models can teach themselves to use tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dess `ı, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. Toolformer: Lan- guage models can teach themselves to use tools. InAdvances in Neural Information Processing Systems, pages 68539– 68551, 2023. 2
2023
-
[39]
A survey on large language model based autonomous agents
Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, Wayne Xin Zhao, Zhewei Wei, and Jirong Wen. A survey on large language model based autonomous agents. Frontiers of Computer Science, 18:186345, 2024. 2
2024
-
[40]
OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments
Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, and Tao Yu. OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments. InAd- vances in Neural Information...
2024
-
[41]
LLM-based agents for tool learning: A survey.Data Science and Engineering, 10:533–563, 2025
Weikai Xu, Chengrui Huang, Shen Gao, and Shuo Shang. LLM-based agents for tool learning: A survey.Data Science and Engineering, 10:533–563, 2025. 2
2025
-
[42]
Narasimhan, and Yuan Cao
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R. Narasimhan, and Yuan Cao. ReAct: Synergizing reasoning and acting in language models. InInternational Conference on Learning Representations, 2023. 2, 4
2023
-
[43]
Narasimhan.τ-bench: A benchmark for tool-agent-user in- teraction in real-world domains
Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik R. Narasimhan.τ-bench: A benchmark for tool-agent-user in- teraction in real-world domains. InInternational Conference on Learning Representations, 2025. 2, 3
2025
-
[44]
Sompote Youwai, David Phim, Vianne Gayl Murcia, and Ri- anne Clair Onas. Large language model-based multi-agent systems for automated foundation design: Router-driven task classification and expert selection framework.AI in Civil En- gineering, 5:5, 2026. 2
2026
-
[45]
Intelligent design of shear wall layout based on graph neural networks.Advanced Engineering Informatics, 55:101886,
Pengju Zhao, Wenjie Liao, Yuli Huang, and Xinzheng Lu. Intelligent design of shear wall layout based on graph neural networks.Advanced Engineering Informatics, 55:101886,
-
[46]
Dynamic prompt-based virtual assistant framework for BIM information search.Au- tomation in Construction, 155:105067, 2023
Junwen Zheng and Martin Fischer. Dynamic prompt-based virtual assistant framework for BIM information search.Au- tomation in Construction, 155:105067, 2023. 2
2023
-
[47]
StructDiffusion: End-to-end intelligent shear wall structure layout generation and analysis using diffusion model.Engineering Structures, 309:118068, 2024
Ying Zhou, Hao Leng, Shiqiao Meng, Hao Wu, and Zheng Zhang. StructDiffusion: End-to-end intelligent shear wall structure layout generation and analysis using diffusion model.Engineering Structures, 309:118068, 2024. 2
2024
-
[48]
Minjie Zhu, Frank McKenna, and Michael H. Scott. OpenSeesPy: Python library for the OpenSees finite element framework.SoftwareX, 7:6–11, 2018. 4 10 A. Implementation and Trace Semantics StructureClaw is implemented around an explicit evidence chain rather than a single conversational output. The chain records what the agent was asked to do, which governed...
2018
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.