Pith. sign in

REVIEW 2 major objections 1 cited by

Learning the energy structure of distributed systems from data lets boundary controllers keep trajectories bounded even when the model is wrong.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-13 10:20 UTC

load-bearing objection Wrong full text was supplied for 2604.04266; only the abstract is usable, so the probabilistic dPHS claim cannot be audited. the 2 major comments →

arxiv 2604.04266 v1 submitted 2026-04-05 eess.SY cs.SYmath-phmath.MPmath.OC

Data-Driven Boundary Control of Distributed Port-Hamiltonian Systems

classification eess.SY cs.SYmath-phmath.MPmath.OC
keywords distributed port-Hamiltonian systemsGaussian process learningboundary control by interconnectionenergy-based robustnessprobabilistic boundednessmodel mismatchshallow water system
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Distributed port-Hamiltonian models give a clean energy-based way to control systems governed by partial differential equations through their boundaries, but they normally need an accurate Hamiltonian. This paper shows how to replace that unknown structure with a Gaussian-process model learned from data, then folds the model’s own posterior uncertainty into the energy balance of the closed loop. The result is a set of probabilistic conditions that guarantee the controlled trajectories stay bounded despite mismatch between the learned model and the true plant. The approach is demonstrated on a simulated shallow-water system, indicating that data-driven energy models can still support rigorous boundary control.

Core claim

When a Gaussian-process distributed port-Hamiltonian model is learned from data and its posterior uncertainty is inserted into an energy-based interconnection analysis, one obtains explicit probabilistic conditions under which the closed-loop trajectories of the true infinite-dimensional system remain bounded even though the Hamiltonian is only approximately known.

What carries the argument

GP-dPHS posterior uncertainty embedded in energy-based robustness analysis: the learned Hamiltonian and its covariance are treated as part of the storage function so that passivity-like inequalities hold with high probability, yielding boundedness certificates for the interconnected closed loop.

Load-bearing premise

The Gaussian-process model of the Hamiltonian, trained on available data, must be faithful enough that its stated posterior uncertainty correctly covers the true infinite-dimensional plant; if the uncertainty is miscalibrated, the probabilistic boundedness guarantees do not transfer.

What would settle it

On a shallow-water or similar distributed plant, train the GP-dPHS model, close the loop with the proposed boundary interconnection, and check whether trajectories that the theory predicts remain bounded with high probability actually leave any prescribed energy ball when the true Hamiltonian differs from the learned mean by an amount inside the claimed posterior.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Boundary controllers for PDE systems can be designed without a first-principles Hamiltonian, relying instead on data-driven energy models.
  • Uncertainty quantification becomes an explicit design parameter: larger posterior variance tightens or relaxes the probabilistic boundedness region.
  • The same energy-interconnection architecture can be reused across different physical domains once a GP-dPHS surrogate is available.
  • Simulation evidence on shallow water suggests the method is immediately testable on laboratory fluid or flexible-structure testbeds.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the probabilistic certificates remain valid under modest sensor noise, the approach could reduce the modeling burden for industrial distributed-parameter control.
  • The same uncertainty-aware energy analysis might extend to collocation or finite-element discretizations, giving a bridge between infinite-dimensional theory and practical finite-dimensional implementation.
  • A natural next measurement is how sample complexity of the GP scales with spatial dimension before the boundedness probability collapses.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 0 minor

Summary. The submission under review (arXiv:2604.04266) claims, per its abstract, that Gaussian Process distributed Port-Hamiltonian system (GP-dPHS) learning can be combined with boundary control by interconnection: the GP-dPHS model infers unknown Hamiltonian structure from data, and its posterior uncertainty is folded into an energy-based robustness analysis to obtain probabilistic conditions under which closed-loop trajectories remain bounded despite model mismatch, with illustration on a simulated shallow-water system. The material supplied as the full manuscript, however, is an unrelated position paper on agentic information retrieval (arXiv:2604.04269, “Beyond Fluency…”). No GP-dPHS construction, energy-balance inequalities, uncertainty-embedding lemmas, sampling assumptions, or shallow-water experiments appear. Consequently only the abstract of the claimed contribution is available for assessment.

Significance. If the abstract’s claims hold for the true infinite-dimensional plant, the work would be a meaningful bridge between data-driven Hamiltonian learning and classical boundary control by interconnection, giving probabilistic robustness certificates that pure model-based dPHS methods lack when dynamics are nonlinear or partially unknown. That significance cannot be confirmed or quantified from the supplied text: there are no machine-checked proofs, reproducible code, parameter-free derivations, or falsifiable simulation metrics to credit.

major comments (2)
  1. The full-text block provided for review is a different manuscript (Agentic IR / arXiv:2604.04269). No section, equation, table, or figure of the claimed GP-dPHS + interconnection paper is present. The central claim—that GP posterior uncertainty can be rigorously embedded into an energy-based interconnection analysis so that probabilistic closed-loop boundedness holds for the true plant, not merely the learned surrogate—cannot be audited. A technical referee report on soundness is therefore impossible until the correct manuscript is supplied.
  2. Even restricted to the abstract, the load-bearing premise remains unchecked: that a GP-dPHS posterior trained on the unknown Hamiltonian structure is sufficiently faithful, and that its uncertainty can be converted into probabilistic trajectory-boundedness conditions for the infinite-dimensional plant under model mismatch. Without the derivation, kernel/prior assumptions, sampling hypotheses on the PDE state, or the energy-balance inequalities, this premise cannot be verified or refuted.

Circularity Check

0 steps flagged

No circularity found; supplied full text is an unrelated position paper, so the dPHS derivation chain cannot be inspected for self-definitional or fitted-input reductions.

full rationale

The target abstract claims that a GP-dPHS posterior is learned from data and its uncertainty is then plugged into an energy-based interconnection analysis to obtain probabilistic closed-loop boundedness conditions. Nothing in that abstract equates the boundedness statement to a fitted quantity by construction, nor does it invoke a self-citation uniqueness theorem that forces the result. The CACHEABLE full-manuscript block, however, is the completely different position paper “Beyond Fluency: Toward Reliable Trajectories in Agentic IR” (arXiv 2604.04269). That text contains only a taxonomy of agentic failure modes, proposals for verification gates, and qualitative metrics; it has no GP-dPHS model, no energy-balance inequalities, no uncertainty embedding lemmas, and no shallow-water simulation. Consequently no equation-level reduction (self-definitional, fitted-input-called-prediction, or load-bearing self-citation) can be exhibited. Per the hard rules, absence of quotable circular steps yields score 0 and an empty steps list. Residual model-class risk (whether the GP posterior is rich enough for the true infinite-dimensional plant) is ordinary scientific uncertainty, not circularity.

Axiom & Free-Parameter Ledger

2 free parameters · 3 axioms · 1 invented entities

Abstract-only review. Load-bearing background includes standard dPHS modeling of PDE systems, boundary control by interconnection / energy shaping, and Gaussian Process regression for structured system ID. Free parameters (kernel hyperparameters, noise models, data density, interconnection gains) and any invented technical objects (specific GP-dPHS prior, exact probabilistic certificate) cannot be enumerated from the abstract. No independent formal verification or shipped artifacts are mentioned.

free parameters (2)
  • GP kernel / prior hyperparameters for Hamiltonian structure
    Any GP-dPHS fit depends on kernel choice and hyperparameters; these are free relative to the robustness claim and are not fixed by the abstract.
  • Interconnection / boundary feedback gains
    Control-by-interconnection designs typically leave damping and interconnection parameters as design choices that affect the energy balance and thus the probabilistic bounds.
axioms (3)
  • domain assumption Distributed Port-Hamiltonian structure is an appropriate model class for the target PDE plant and admits boundary control by interconnection.
    Stated as the starting framework in the abstract; effectiveness of classical methods is assumed when the model is accurate.
  • domain assumption A Gaussian Process can represent the unknown Hamiltonian structure well enough that posterior uncertainty is meaningful for energy-based robustness.
    Core methodological premise of GP-dPHS learning as used for control guarantees.
  • ad hoc to paper Energy-based robustness analysis can convert GP posterior uncertainty into probabilistic closed-loop trajectory boundedness conditions under model mismatch.
    This is the paper's claimed technical step; it is not a standard theorem cited in the abstract.
invented entities (1)
  • GP-dPHS (Gaussian Process distributed Port-Hamiltonian system) model used for control no independent evidence
    purpose: Infer unknown Hamiltonian structure from data and supply posterior uncertainty for robustness analysis of interconnection boundary control.
    Named as the learning object combined with boundary control; may build on prior GP-dPHS work, but the control-oriented uncertainty use is the paper's vehicle. No independent evidence beyond the claimed simulation is given in the abstract.

pith-pipeline@v1.1.0-grok45 · 14519 in / 2673 out tokens · 28425 ms · 2026-07-13T10:20:59.864717+00:00 · methodology

0 comments
read the original abstract

Distributed Port-Hamiltonian (dPHS) theory provides a powerful framework for modeling physical systems governed by partial differential equations and has enabled a broad class of boundary control methodologies. Their effectiveness, however, relies heavily on the availability of accurate system models, which may be difficult to obtain in the presence of nonlinear and partially unknown dynamics. To address this challenge, we combine Gaussian Process distributed Port-Hamiltonian system (GP-dPHS) learning with boundary control by interconnection. The GP-dPHS model is used to infer the unknown Hamiltonian structure from data, while its posterior uncertainty is incorporated into an energy-based robustness analysis. This yields probabilistic conditions under which the closed-loop trajectories remain bounded despite model mismatch. The method is illustrated on a simulated shallow water system.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Commercial Persuasion in AI-Mediated Conversations

    cs.CY 2026-04 conditional novelty 7.0

    LLM agents nearly triple sponsored-product selection versus search placement (61.2% vs 22.4%), with detection near chance and labels failing to protect users.

Reference graph

Works this paper leans on

56 extracted references · 2 canonical work pages · cited by 1 Pith paper

  1. [1]

    Marcia J. Bates. 1989. The Design of Browsing and Berrypicking Techniques for the Online Search Interface.Online Review13, 5 (1989), 407–424. https: //eric.ed.gov/?id=EJ404172

  2. [2]

    Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big?. InProceedings of the 2021 ACM conference on fairness, accountability, and transparency. 610–623

  3. [3]

    Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer. 2015. Sched- uled Sampling for Sequence Prediction with Recurrent Neural Networks. arXiv:1506.03099 [cs.NE] https://arxiv.org/abs/1506.03099

  4. [4]

    Brown et al

    T. Brown et al. 2020. Language Models are Few-Shot Learners. InNeurIPS 2020

  5. [5]

    Qingpeng Cai, Xiangyu Zhao, Ling Pan, Xin Xin, Jin Huang, Weinan Zhang, Li Zhao, Dawei Yin, and Grace Hui Yang. 2024. Agentir: 1st workshop on agent- based information retrieval. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. 3025–3028

  6. [6]

    Vikrant Chaugule, D Abhishek, Aadheeshwar Vijayakumar, Pravin Bhaskar Ramteke, and Shashidhar G Koolagudi. 2016. Product review based on optimized facial expression detection. In2016 Ninth International Conference on Contempo- rary Computing (IC3). IEEE, 1–6

  7. [7]

    X Chen, A Zeng, et al. 2024. A survey on large language model based autonomous agents. InCCL 2024–23rd Chinese Natl Conf Comput Linguist, Vol. 2. 141–150

  8. [8]

    Jeffrey Dalton, Chenyan Xiong, and Jamie Callan. 2020. TREC CAsT 2019: The Conversational Assistance Track Overview. arXiv:2003.13624 [cs.IR] https: //arxiv.org/abs/2003.13624

  9. [9]

    Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Samuel Stevens, Boshi Wang, Huan Sun, and Yu Su. 2023. Mind2Web: Towards a Generalist Agent for the Web. arXiv:2306.06070 [cs.CL] https://arxiv.org/abs/2306.06070

  10. [10]

    Abhishek Dharmaratnakar, Srivaths Ranganathan, Debanshu Das, and Anushree Sinha. 2026. Generative AI for Video Trailer Synthesis: From Extractive Heuristics to Autoregressive Creativity.Authorea Preprints(2026)

  11. [11]

    Jiaxuan Gao, Wei Fu, Minyang Xie, Shusheng Xu, Chuyi He, Zhiyu Mei, Banghua Zhu, and Yi Wu. 2025. Beyond ten turns: Unlocking long-horizon agentic search with large-scale asynchronous rl.arXiv preprint arXiv:2508.07976(2025)

  12. [12]

    Yonatan Geifman and Ran El-Yaniv. 2017. Selective Classification for Deep Neural Networks. InAdvances in Neural Information Processing Systems. arXiv:1705.08500 https://arxiv.org/abs/1705.08500

  13. [13]

    Weinberger

    Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017. On Calibration of Modern Neural Networks. arXiv:1706.04599 [cs.LG] https://arxiv.org/abs/ 1706.04599

  14. [14]

    Xu Huang, Weiwen Liu, Xiaolong Chen, Xingmei Wang, Hao Wang, Defu Lian, Yasheng Wang, Ruiming Tang, and Enhong Chen. 2024. Understanding the planning of llm agents: A survey.arXiv preprint arXiv:2402.02716(2024)

  15. [15]

    Zhidian Huang, Zijun Yao, Ji Qi, Shangqing Tu, Junxian Ma, Jinxin Liu, Weichuan Liu, Xiaoyin Che, Lei Hou, and Juanzi Li. 2026. MM-THEBench: Do Reasoning MLLMs Think Reasonably?arXiv preprint arXiv:2601.22735(2026)

  16. [16]

    Price, Lois M

    Kalervo J"arvelin, Susan L. Price, Lois M. L. Delcambre, and Marianne Lykke Nielsen. 2008. Discounted Cumulated Gain Based Evaluation of Multiple-Query IR Sessions. InProceedings of the 30th European Conference on Information Retrieval (ECIR). 4–15

  17. [17]

    Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation.ACM computing surveys55, 12 (2023), 1–38

  18. [18]

    Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan

    Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. 2023. SWE-bench: Can Language Models Resolve Real- World GitHub Issues? arXiv:2310.06770 [cs.SE] https://arxiv.org/abs/2310.06770

  19. [19]

    Subbarao Kambhampati, Karthik Valmeekam, Lin Guan, Mudit Verma, Kaya Stechly, Siddhant Bhambri, Lucas Paul Saldyt, and Anil B Murthy. 2024. Position: LLMs can’t plan, but can help planning in LLM-modulo frameworks. InForty-first International Conference on Machine Learning

  20. [20]

    Diane Kelly. 2009. Methods for Evaluating Interactive Information Retrieval Systems with Users. Foundations and Trends in Information Retrieval. https: //dl.acm.org/doi/10.1561/1500000012

  21. [21]

    Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. 2017. Sim- ple and Scalable Predictive Uncertainty Estimation using Deep Ensembles. arXiv:1612.01474 [stat.ML] https://arxiv.org/abs/1612.01474

  22. [22]

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing systems33 (2020), 9459–9474

  23. [23]

    Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. TruthfulQA: Measuring How Models Mimic Human Falsehoods. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL). https://aclanthology.org/ 2022.acl-long.229/

  24. [24]

    Xuannan Liu, Xiao Yang, Zekun Li, Peipei Li, and Ran He. 2026. AgentHallu: Benchmarking Automated Hallucination Attribution of LLM-based Agents.arXiv preprint arXiv:2601.06818(2026)

  25. [25]

    Potsawee Manakul, Adian Liusie, and Mark J. F. Gales. 2023. SelfCheckGPT: Zero- Resource Black-Box Hallucination Detection for Generative Large Language Models. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP). https://aclanthology.org/2023.emnlp-main.557/

  26. [26]

    N. Nanda. 2023. A Comprehensive Mechanistic Interpretability Explainer. Neel Nanda’s Blog

  27. [27]

    Ouyang et al

    L. Ouyang et al. 2022. Training language models to follow instructions. InNeurIPS 2022

  28. [28]

    Shishir G Patil, Tianjun Zhang, Xin Wang, and Joseph E Gonzalez. 2024. Go- rilla: Large language model connected with massive apis.Advances in Neural Information Processing Systems37 (2024), 126544–126565

  29. [29]

    Yujia Qin, Yongliang Liang, Yifan Wang, et al. 2024. ToolLLM: Facilitating Large Language Models to Master 16000+ APIs. InProceedings of the International Conference on Learning Representations (ICLR)

  30. [30]

    Srivaths Ranganathan, Abhishek Dharmaratnakar, Anushree Sinha, and Deban- shu Das. 2026. Multi-Agent Video Recommenders: Evolution, Patterns, and Open Challenges.arXiv preprint arXiv:2604.02211(2026)

  31. [31]

    Srivaths Ranganathan, Chieh Lo, Bernardo Cunha, Nikhil Khani, Li Wei, Aniruddh Nath, Shawn Andrews, Gergo Varady, Yanwei Song, Jochen Klingenhoefer, et al

  32. [32]

    InProceedings of the Nineteenth ACM Conference on Recommender Systems

    Zero-shot Cross-domain Knowledge Distillation: A Case study on YouTube Music. InProceedings of the Nineteenth ACM Conference on Recommender Systems. 1122–1125

  33. [33]

    Shaina Raza, Ranjan Sapkota, Manoj Karkee, and Christos Emmanouilidis. 2025. Trism for agentic ai: A review of trust, risk, and security management in llm-based agentic multi-agent systems.arXiv preprint arXiv:2506.04133(2025)

  34. [34]

    Gordon, and J

    St’ephane Ross, Geoffrey J. Gordon, and J. Andrew Bagnell. 2011. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning. InProceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics (AISTATS). https://proceedings.mlr.press/v15/ross11a.html

  35. [35]

    Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language models can teach themselves to use tools.Advances in neural information processing systems36 (2023), 68539–68551

  36. [36]

    Noah Shinn, Jeffrey Labash, Ashwin Gopinath, Manya Wadhwa, Pratyusha Kumar, Yiming Yang, Ameya Joshi, Shunyu Yao, et al. 2023. Reflexion: Language Agents with Verbal Reinforcement Learning. InAdvances in Neural Information Processing Systems

  37. [37]

    Smucker and Charles L

    Mark D. Smucker and Charles L. A. Clarke. 2012. Time-based Calibration of Effectiveness Measures. InProceedings of the 35th International ACM SI- GIR Conference on Research and Development in Information Retrieval (SIGIR). https://dl.acm.org/doi/10.1145/2348283.2348300

  38. [38]

    Touvron et al

    H. Touvron et al . 2023. Llama: Open and Efficient Foundation Models. arXiv:2302.13971(2023)

  39. [39]

    Aadheeshwar Vijayakumar, D Abhishek, and K Chandrasekaran. 2016. DSL approach for development of gaming applications. InInformation Systems Design and Intelligent Applications: Proceedings of Third International Conference INDIA 2016, Volume 1. Springer India New Delhi, 199–211

  40. [40]

    Tu Vu, Mohit Iyyer, Xuezhi Wang, Noah Constant, Jerry Wei, Jason Wei, Chris Tar, Yun-Hsuan Sung, Denny Zhou, Quoc Le, et al. 2024. Freshllms: Refreshing large language models with search engine augmentation. InFindings of the Association for Computational Linguistics: ACL 2024. 13697–13720

  41. [41]

    Hanlin Wang, Chak Tou Leong, Jiashuo Wang, Jian Wang, and Wenjie Li. 2025. Spa-rl: Reinforcing llm agents via stepwise progress attribution.arXiv preprint SIGIR 2026, July 2026, Gold Coast, Australia Anushree Sinha, Srivaths Ranganathan, Debanshu Das, and Abhishek Dharmaratnakar arXiv:2505.20732(2025)

  42. [42]

    Wang et al

    J. Wang et al. 2025. UltraHorizon: Evaluating Agents in Long-Horizon Tasks. arXiv:2509.21766(2025)

  43. [43]

    Jerry Wei, Chengrun Yang, Xinying Song, Yifeng Lu, Nathan Hu, Jie Huang, Dustin Tran, Daiyi Peng, Ruibo Liu, Da Huang, et al. 2024. Long-form factuality in large language models.Advances in Neural Information Processing Systems37 (2024), 80756–80827

  44. [44]

    Fangzhi Xu, Hang Yan, Qiushi Sun, Jinyang Wu, Zixian Huang, Muye Huang, Jingyang Gong, Zichen Ding, Kanzhi Cheng, Yian Wang, et al . 2026. Odysse- yArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions.arXiv preprint arXiv:2602.05843(2026)

  45. [45]

    Ziwei Xu, Sanjay Jain, and Mohan Kankanhalli. 2024. Hallucination is inevitable: An innate limitation of large language models.arXiv preprint arXiv:2401.11817 (2024)

  46. [46]

    Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. 2022. WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents. InAdvances in Neural Information Processing Systems. arXiv:2207.01206 https: //arxiv.org/abs/2207.01206

  47. [47]

    Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. 2022. React: Synergizing reasoning and acting in language models. InThe eleventh international conference on learning representations

  48. [48]

    Chenlong Yin, Zeyang Sha, Shiwen Cui, and Changhua Meng. 2025. The reasoning trap: How enhancing LLM reasoning amplifies tool hallucination.arXiv preprint arXiv:2510.22977(2025)

  49. [49]

    Chenglin Yu, Yuchen Wang, Songmiao Wang, Hongxia Yang, and Ming Li. 2026. InfiAgent: An Infinite-Horizon Framework for General-Purpose Autonomous Agents.arXiv preprint arXiv:2601.03204(2026)

  50. [50]

    Weinan Zhang, Junwei Liao, Ning Li, Kounianhua Du, and Jianghao Lin. 2024. Agentic information retrieval.arXiv preprint arXiv:2410.09713(2024)

  51. [51]

    Yuxiang Zhang, Jing Chen, Junjie Wang, Yaxin Liu, Cheng Yang, Chufan Shi, Xinyu Zhu, Zihao Lin, Hanwen Wan, Yujiu Yang, et al. 2024. Toolbehonest: A multi-level hallucination diagnostic benchmark for tool-augmented large lan- guage models. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 11388–11422

  52. [52]

    Ying Zhang, Maarten de Rijke, and Evangelos Kanoulas. 2017. A General Formal Framework for IR Evaluation. InProceedings of the 8th International Conference on the Theory of Information Retrieval (ICTIR). https://liususan091219.github.io/ pdf/ictir17_zhang.pdf

  53. [53]

    Yuxiang Zhang, Jiangming Shu, Ye Ma, Xueyuan Lin, Shangxi Wu, and Jitao Sang

  54. [54]

    Memory as action: Autonomous context curation for long-horizon agentic tasks.arXiv preprint arXiv:2510.12635(2025)

  55. [55]

    Jialong Zhou, Lichao Wang, and Xiao Yang. 2025. Guardian: Safeguarding llm multi-agent collaborations with temporal graph modeling.arXiv preprint arXiv:2505.19234(2025)

  56. [56]

    Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig

    Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig. 2023. WebArena: A Realistic Web Environment for Building Autonomous Agents. arXiv:2307.13854 [cs.CL] https://arxiv.org/abs/2307.13854