Pith. sign in

REVIEW 3 major objections 5 minor 68 references

Towards Adaptive Mechanism Activation in Language Agent

T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Fine-tuning on self-explored trajectories lets an 8B language agent choose the right problem-solving mechanism for each task, beating fixed-mechanism agents.

desk verdict A useful data-efficient agent fine-tuning recipe, but the 'adaptive mechanism activation' claim needs behavior-level evidence before it can be taken at face value. read the letter →

arxiv 2412.00722 v1 pith:D7USVUNQ submitted 2024-12-01 cs.CL cs.AI

classification cs.CLcs.AI
keywords languageagentsadaptivemechanismactivationself-explorationUniActKTOLLMfine-tuningmathematicalreasoningknowledge-intensive
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that language agents should not be locked into one problem-solving mechanism or a fixed sequence of them; instead, an agent can learn to activate the mechanism that fits each task. To show this, it introduces ALAMA, which first turns five known mechanisms (Reason, Plan, Memory, Reflection, External-Augmentation) into explicit actions within a shared framework called UniAct, then trains an 8B model on self-explored trajectories using supervised fine-tuning followed by binary-reward preference optimization. The reported results are that the trained agent beats every fixed-mechanism baseline and the average of all mechanisms on held-in tasks, transfers to held-out math and knowledge tasks, and approaches a much larger fine-tuned baseline while training on far less data. A sympathetic reader would care because the method points toward a cheap, self-contained route to adaptive agent behavior: no expert demonstrations, no hand-crafted routers, and no pairwise preference data.

What carries the argument

The load-bearing object is UniAct, a harmonized agent framework that re-expresses each mechanism as explicit action tokens in one shared space: Make_plan and Carry_out_plan stand for Plan, Retrieve_memory stands for Memory, Reflect stands for Reflection, Calculate/Search/Lookup stand for External-Augmentation, and Finish ends the trajectory, with Thoughts and Observations surrounding these actions. The argument is carried by two training stages over trajectories collected by self-exploration with manual mechanism activation: IMAO applies standard supervised fine-tuning on positive UniAct trajectories (with observation tokens masked from the loss), and MAAO applies KTO, a preference-learning objective that requires only a binary desirable/undesirable label per trajectory, using final-answer correctness as that label. The action space makes mechanism choice observable and learnable as ordinary next-token prediction, and the binary-reward optimization biases the model toward mechanisms that succeed on a given task and away from those that fail.

What would settle it

Record the action choices (Make_plan, Carry_out_plan, Retrieve_memory, Reflect, Calculate) made by the trained ALAMA on GSM8K and compare them to an oracle mechanism selector; if the trained agent's choices match the oracle no better than the IMAO-only model's choices, or no better than chance, the central adaptive-mechanism-activation claim is refuted.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that adaptive mechanism activation can be learned rather than engineered. ALAMA's pipeline collects diverse solution trajectories by manually activating each mechanism during self-exploration, converts them into the UniAct action format, and optimizes the agent in two stages: IMAO supervised fine-tuning on successful trajectories teaches the format and implicit preferences, and MAAO, built on KTO, uses only binary success/failure signals to push the agent toward mechanisms that work for a task and away from those that do not. The authors report that this yields accuracy gains over all single-mechanism baselines and over the average of mechanisms on GSM8K and HotpotQA, plus held-out gains on NumGLUE and SVAMP. They also report an oracle analysis showing that a perfect mechanism selector would solve 96.89% of GSM8K tasks while only 42.61% are solvable by all mechanisms, which they read as evidence that the task space contains real mechanism specificity and that adaptive activation has a high ceiling. Section 4.2 states the mechanism-learning effect directly: behavior contrastive learning enables the model to preferentially activate certain mechanisms while refusing to activate the remaining ones.

Load-bearing premise

The paper assumes that a single binary reward saying whether the final answer is correct, applied to self-explored trajectories, is enough to teach the agent which mechanism to activate for each task; if the accuracy gains come only from imitating successful trajectory formats, the adaptive mechanism activation claim is unsupported.

Editorial extensions

If this is right

  • An 8B open-weight agent trained only on GSM8K self-exploration data reaches 82.18% on GSM8K, beating the average of its five fixed-mechanism baselines by about 6 points and approaching a much larger fine-tuned agent trained on far more data.
  • The learned mechanism preference transfers zero-shot: on NumGLUE and SVAMP, ALAMA beats the best fixed single mechanism by 3.95 and 2.3 accuracy points respectively, and Self-Adapt Consistency extends those margins.
  • Adding self-consistency sampling on top of ALAMA yields further gains, so adaptive activation and voting over diverse generated trajectories are complementary rather than redundant.
  • The oracle analysis implies that mechanism sensitivity is common, with more than half of GSM8K tasks not solvable by all mechanisms, so any method that improves mechanism selection has real headroom, with 96.89% as an upper bound on GSM8K.
  • Because training needs only self-explored trajectories and binary correctness labels, the same recipe should apply to any set of mechanisms and any domain where the environment can score answers.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the accuracy gains survive a direct audit of the emitted actions, the method constitutes an implicit, fully learned router; if they do not, the gains could reflect format-following from supervised fine-tuning on successful trajectories, which is my inference to test.
  • The paper restricts itself to activating a single mechanism per trajectory; a natural extension is simultaneous or sequenced activation of multiple mechanisms, which the authors themselves flag as future work and which might close part of the gap to the oracle ceiling.
  • A testable extension would be to run the same ALAMA pipeline with an explicit oracle action label in MAAO; if that label does not improve over the current reward-only version, the binary-reward signal is already capturing mechanism preference, and if it does, the current objective leaves selection information on the table.
  • The approach could connect to a broader design space where mechanisms are not fixed prompts but composable skills; the UniAct action space already gives a generic interface for adding new mechanisms without changing the training objective.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes ALAMA, a method for training a language agent to activate different problem-solving mechanisms (Reason, Plan, Memory, Reflection, External-Augmentation) adaptively rather than using a fixed mechanism or a predefined sequence. The authors introduce UniAct, a unified action-based format for these mechanisms, collect trajectories by self-exploration with manually activated mechanisms, and train with two stages: IMAO (SFT on positive trajectories) and MAAO (KTO on binary success/failure labels). Experiments on GSM8K and HotpotQA as held-in tasks and NumGLUE, SVAMP, TriviaQA, and Bamboogle as held-out tasks report accuracy/EM scores comparing against fixed-mechanism, majority-voting, and fine-tuning baselines. The paper also reports an oracle mechanism-selection analysis on GSM8K showing a high ceiling (96.89%) for adaptive activation.

Significance. If the central claim were fully supported, the work would be a useful step toward data-efficient, self-explored training of language agents that select solution strategies per task. The UniAct unification of several mechanisms into a shared action space is a sensible and potentially reusable design. The use of KTO with binary rewards rather than DPO-style pairwise preferences is a reasonable efficiency choice, and the comparison against strong fine-tuning baselines (Husky, MAmmoTH2) with only GSM8K self-exploration data is informative. However, the paper's central claim of adaptive mechanism activation currently rests on aggregate accuracy gains; no behavioral evidence shows that the trained agent actually emits different mechanisms across tasks. The held-out results also do not uniformly support the generalization claim. For these reasons the contribution is plausible but not yet established.

major comments (3)
  1. The central claim that ALAMA learns adaptive mechanism activation is not directly supported by any experiment. The paper never reports the distribution of actions (Make_plan, Carry_out_plan, Retrieve_memory, Reflect, Calculate/Search, Finish) emitted by ALAMA at inference, nor does it compare that distribution with the oracle mechanism selector of §5.1. Equations (7)–(9) train on a binary correctness reward only; KTO can improve accuracy by making the model imitate successful mixed-format trajectories and avoid failed rollouts without learning any task-conditioned mechanism-selection policy. In fact, §5.1 shows an oracle selector reaches 96.89% on GSM8K while ALAMA reaches 82.18%, so the adaptive-selection component is exactly the part left unvalidated. Please add an action-level analysis: per-dataset action frequencies, agreement with the oracle mechanism choice on mechanism-sensitive tasks, or a controlled experiment where the trained model is forced to use a single mechanism. Without such evidence, the observed gains are equally consistent with a data-efficient fine-tuning recipe on self-explored trajectories.
  2. The held-out generalization claim is overstated. On TriviaQA, ALAMA (IMAO+MAAO) achieves 43.60 EM, which is below the Average of 44.28 and far below the best fixed mechanism Reflection at 55.80. On Bamboogle, ALAMA achieves 32.80, below Memory at 44.80. The sentence in §4.2 that ALAMA "also outperforms most baselines, including Average, on TriviaQA and Bamboogle" is not accurate for TriviaQA. The paper should report these results with error bars or significance tests and revise the generalization claim to reflect the mixed held-out performance.
  3. The comparison with fine-tuning baselines in Table 2 is presented as evidence of data efficiency, but the comparison is not apples-to-oranges in an important aspect: ALAMA is trained only on GSM8K self-exploration data, whereas several baselines use additional datasets or larger supervision sources. This is acknowledged in the text, but the conclusion "fully demonstrating the data efficiency of ALAMA" is stronger than what the evidence supports, because ALAMA also benefits from the base model Llama-3-8B-Instruct and from the in-context mechanism demonstrations used during self-exploration. Please temper the claim or provide an ablation that isolates the contribution of the self-exploration data itself.
minor comments (5)
  1. There are typographical and grammatical issues: "proposesAdaptive" is missing a space, "Language Agent could be endowed" should be "Language agents can be endowed," and "superior performance" is misspelled as "suprior performance" in §4.2.
  2. The notation "(-1)1(u∈...)" is unclear; it should be written as (-1)^{1(u∈...)} or with an explicit indicator function. Also, the role of λ_pos/λ_neg in the equation should be defined more carefully.
  3. The row for λ_pos/λ_neg is garbled: "λDnD λU nU 4/3" does not clearly state the values. Please write λ_pos and λ_neg explicitly.
  4. The title "The Effects of Mixing Different Mechanism Data" promises a comparison of mixed-mechanism subsets, but the section only compares full mixed data with single-mechanism data. The Limitation paragraph acknowledges that mixing effects are omitted, which is fine, but the section title should be aligned with the actual content.
  5. The term "Mammoth2-Plus" is written inconsistently as "MAmmoTH2-Plus" in the same subsection; please standardize.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: ALAMA's gains are measured on held-out benchmarks, and the adaptive-activation claim is under-validated rather than circular.

full rationale

ALAMA's training pipeline is not circular in the derivation-chain sense. The mechanism labels (Make_plan, Carry_out_plan, Retrieve_memory, Reflect, Calculate) are defined by the UniAct prompt templates in Appendix F and by the manually activated in-context demonstrations, not by the training outcome. The training signal is a binary correctness reward on self-explored trajectories (Equations 7-9), and the reported improvements are evaluated on held-out datasets (NumGLUE, SVAMP, TriviaQA, Bamboogle) that are not used for training. No equation equates the adaptive-mechanism-activation conclusion with an input of the method, and no fitted parameter is renamed as a prediction. The closest issue is evidentiary, not circular: Section 4.2 asserts that 'Behavior contrastive learning enables the model to preferentially activate certain specific mechanisms while refusing to activate the remaining ones,' but Section 5 does not report the distribution of actions emitted by the trained agent or compare it with the oracle selector described in Section 5.1, so the adaptive-activation claim is under-validated. That is a measurement/interpretation gap rather than a circularity. Self-citations such as ExpNote (Sun et al., 2023, with overlapping authors) are used only as one implementation of the Memory mechanism and are not load-bearing for the central claim. Accordingly, this paper receives a circularity score of 0.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The central claim rests on the correctness-only reward and the UniAct transform being faithful, plus the untested assumption that KTO induces genuine mechanism selection. These are domain assumptions rather than fitted constants or invented physical entities.

free parameters (4)
  • KTO beta (β)
    KTO loss in Eq. 8 uses β as the temperature; the paper never reports its value, making the exact update un-reproducible.
  • lambda_pos/lambda_neg = 4/3
    Table 5 lists λpos/λneg = 4/3; this manual ratio controls the balance between positive and negative KTO updates.
  • Self-exploration sampling temperature
    The number of samples per task and the decoding temperature used to generate si,j (Eq. 1) are not reported; these determine which trajectories are deemed mechanism-sensitive.
  • Number of self-exploration trajectories per mechanism per task = 1 (implied)
    Algorithm 1 generates one trajectory per mechanism per task; with only one sample, a task may be incorrectly labeled sensitive due to sampling noise.
assumptions (4)
  • domain assumption The base model's final-answer correctness is a sufficient reward signal for mechanism activation quality in KTO optimization.
    MAAO assigns positive/negative labels solely from the gold answer match (Section 3, Eq. 7-9); no per-step or mechanism-level reward is used.
  • ad hoc to paper The UniActTransform deterministically maps each ICL trajectory into an equivalent UniAct trajectory whose inserted actions faithfully represent the mechanism.
    Appendix E defines the transformation; if actions like Make_plan or Reflect are inserted into trajectories that did not actually contain those steps, the mechanism supervision is an artifact.
  • domain assumption Each hand-written in-context demonstration activates a distinct, identifiable mechanism in the base model.
    Section 2 and Appendix D assume the five prompts (CoT, Plan-and-Solve, ExpNote, Reflexion, ReAct) elicit distinct solution structures; no behavioral check is provided.
  • domain assumption KTO preference optimization remains valid when desirable/undesirable labels are automatic binary rewards rather than human judgments.
    The paper applies KTO from Ethayarajh et al. (2024) to trajectory rewards (correct/incorrect) without discussing violations of KTO's assumptions about the reward distribution.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Towards Adaptive Mechanism Activation in Language Agent." pith.science (2026). https://pith.science/paper/D7USVUNQ

@misc{pith2026241200722,
  author       = {Pith},
  title        = {Pith review of: Towards Adaptive Mechanism Activation in Language Agent},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/D7USVUNQ}},
  note         = {Machine review of arXiv:2412.00722}
}
read the original abstract

Language Agent could be endowed with different mechanisms for autonomous task accomplishment. Current agents typically rely on fixed mechanisms or a set of mechanisms activated in a predefined order, limiting their adaptation to varied potential task solution structures. To this end, this paper proposes \textbf{A}daptive \textbf{L}anguage \textbf{A}gent \textbf{M}echanism \textbf{A}ctivation Learning with Self-Exploration (\textbf{ALAMA}), which focuses on optimizing mechanism activation adaptability without reliance on expert models. Initially, it builds a harmonized agent framework (\textbf{UniAct}) to \textbf{Uni}fy different mechanisms via \textbf{Act}ions. Then it leverages a training-efficient optimization method based on self-exploration to enable the UniAct to adaptively activate the appropriate mechanisms according to the potential characteristics of the task. Experimental results demonstrate significant improvements in downstream agent tasks, affirming the effectiveness of our approach in facilitating more dynamic and context-sensitive mechanism activation.

Figures

Figures reproduced from arXiv: 2412.00722 by the authors.

Figure 1
Figure 1. Illustration of Language Agent with different [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. The UniAct trajectory examples for five mechanisms. The underlined contents are generated by the vanilla [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. The illustration of ALAMA process. The UniAct trajectories are collected by Self-Exploration with [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Mechanism specificity analysis results on [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

68 extracted references · 8 canonical work pages

  1. [1]

    AI@Meta. 2024. https://github.com/meta-llama/llama3/blob/main/MODEL_CARD.md Llama 3 model card

  2. [2]

    Anonymous. 2024. Samoyed: Towards generalized llm agents via fine-tuning on 50000+ interaction trajectories

  3. [3]

    Anthropic. 2023. https://www.anthropic.com/news/introducing-claude Introducing claude

  4. [4]

    Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, Ilya Sutskever, and Jeff Wu. 2023. https://arxiv.org/abs/2312.09390 Weak-to-strong generalization: Eliciting strong capabilities with weak supervision . Preprint, arXiv:2312.09390

  5. [5]

    Boxi Cao, Keming Lu, Xinyu Lu, Jiawei Chen, Mengjie Ren, Hao Xiang, Peilin Liu, Yaojie Lu, Ben He, Xianpei Han, Le Sun, Hongyu Lin, and Bowen Yu. 2024. https://arxiv.org/abs/2406.01252 Towards scalable automated alignment of llms: A survey . Preprint, arXiv:2406.01252

  6. [6]

    Baian Chen, Chang Shu, Ehsan Shareghi, Nigel Collier, Karthik Narasimhan, and Shunyu Yao. 2023. https://arxiv.org/abs/2310.05915 Fireact: Toward language agent fine-tuning . Preprint, arXiv:2310.05915

  7. [7]

    Zehui Chen, Kuikun Liu, Qiuchen Wang, Wenwei Zhang, Jiangning Liu, Dahua Lin, Kai Chen, and Feng Zhao. 2024 a . https://arxiv.org/abs/2403.12881 Agent-flan: Designing data and methods of effective agent tuning for large language models . Preprint, arXiv:2403.12881

  8. [8]

    Zixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji, and Quanquan Gu. 2024 b . https://arxiv.org/abs/2401.01335 Self-play fine-tuning converts weak language models to strong language models . Preprint, arXiv:2401.01335

Show all 68 references
  1. [9]

    Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vino...

  2. [10]

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. https://arxiv.org/abs/2110.14168 Training verifiers to solve math word problems . ...

  3. [11]

    Zhang, Hanwei Xu, Hao Yang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J

    DeepSeek-AI, Aixin Liu, Bei Feng, Bin Wang, Bingxuan Wang, Bo Liu, Chenggang Zhao, Chengqi Dengr, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei Li, H. Zhang, Hanwei Xu, Hao Yang, Haowei...

  4. [12]

    Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, Lei Li, and Zhifang Sui. 2023. https://arxiv.org/abs/2301.00234 A survey on in-context learning . Preprint, arXiv:2301.00234

  5. [13]

    Tenenbaum, and Igor Mordatch

    Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mordatch. 2023. https://arxiv.org/abs/2305.14325 Improving factuality and reasoning in language models through multiagent debate . Preprint, arXiv:2305.14325

  6. [14]

    Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky, and Douwe Kiela. 2024. https://arxiv.org/abs/2402.01306 Kto: Model alignment as prospect theoretic optimization . Preprint, arXiv:2402.01306

  7. [15]

    Chen Gao, Xiaochong Lan, Nian Li, Yuan Yuan, Jingtao Ding, Zhilun Zhou, Fengli Xu, and Yong Li. 2023. https://arxiv.org/abs/2312.11970 Large language models empowered agent-based modeling and simulation: A survey and perspectives . Preprint, arXiv:2312.11970

  8. [16]

    Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Meng Wang, and Haofen Wang. 2024. https://arxiv.org/abs/2312.10997 Retrieval-augmented generation for large language models: A survey . Preprint, arXiv:2312.10997

  9. [17]

    Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang, Minlie Huang, Nan Duan, and Weizhu Chen. 2024. https://arxiv.org/abs/2309.17452 Tora: A tool-integrated reasoning agent for mathematical problem solving . Preprint, arXiv:2309.17452

  10. [18]

    Caglar Gulcehre, Tom Le Paine, Srivatsan Srinivasan, Ksenia Konyushkova, Lotte Weerts, Abhishek Sharma, Aditya Siddhant, Alex Ahern, Miaosen Wang, Chenjie Gu, Wolfgang Macherey, Arnaud Doucet, Orhan Firat, and Nando de Freitas. 2023. https://arxiv.org/abs/2308.08998 Reinforced...

  11. [19]

    Namgyu Ho, Laura Schmid, and Se-Young Yun. 2023. https://doi.org/10.18653/v1/2023.acl-long.830 Large language models are reasoning teachers . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 14852--14882,...

  12. [20]

    Jiaxin Huang, Shixiang Shane Gu, Le Hou, Yuexin Wu, Xuezhi Wang, Hongkun Yu, and Jiawei Han. 2023. https://openreview.net/forum?id=uuUQraD4XX Large language models can self-improve . In The 2023 Conference on Empirical Methods in Natural Language Processing

  13. [21]

    Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou. 2024. https://openreview.net/forum?id=IkmD3fKBPQ Large language models cannot self-correct reasoning yet . In The Twelfth International Conference on Learning Representations

  14. [22]

    Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. https://doi.org/10.18653/v1/P17-1147 T rivia QA : A large scale distantly supervised challenge dataset for reading comprehension . In Proceedings of the 55th Annual Meeting of the Association for Computational...

  15. [23]

    Joongwon Kim, Bhargavi Paranjape, Tushar Khot, and Hannaneh Hajishirzi. 2024. https://arxiv.org/abs/2406.06469 Husky: A unified, open-source language agent for multi-step reasoning . Preprint, arXiv:2406.06469

  16. [24]

    Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023. https://doi.org/10.1145/3600006.3613165 Efficient memory management for large language model serving with pagedattention . In Proceedings of the 29...

  17. [25]

    Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Zhaopeng Tu, and Shuming Shi. 2023. Encouraging divergent thinking in large language models through multi-agent debate. arXiv preprint arXiv:2305.19118

  18. [26]

    Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2021. https://arxiv.org/abs/2107.13586 Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing . Preprint, arXiv:2107.13586

  19. [27]

    Tengxiao Liu, Qipeng Guo, Yuqing Yang, Xiangkun Hu, Yue Zhang, Xipeng Qiu, and Zheng Zhang. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.169 Plan, verify and switch: Integrated reasoning with diverse X -of-thoughts . In Proceedings of the 2023 Conference on Empirical Meth...

  20. [28]

    Jianqiao Lu, Wanjun Zhong, Wenyong Huang, Yufei Wang, Qi Zhu, Fei Mi, Baojun Wang, Weichao Wang, Xingshan Zeng, Lifeng Shang, Xin Jiang, and Qun Liu. 2024. https://arxiv.org/abs/2310.00533 Self: Self-evolution with language feedback . Preprint, arXiv:2310.00533

  21. [29]

    Haipeng Luo, Qingfeng Sun, Can Xu, Pu Zhao, Jianguang Lou, Chongyang Tao, Xiubo Geng, Qingwei Lin, Shifeng Chen, and Dongmei Zhang. 2023. https://arxiv.org/abs/2308.09583 Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct . Pre...

  22. [30]

    Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark. 2023. https://openrev...

  23. [31]

    Swaroop Mishra, Arindam Mitra, Neeraj Varshney, Bhavdeep Sachdeva, Peter Clark, Chitta Baral, and Ashwin Kalyan. 2022. Numglue: A suite of fundamental yet challenging mathematical reasoning tasks. ACL

  24. [32]

    OpenAI. 2022. Introducing chatgpt

  25. [33]

    OpenAI. 2024. https://openai.com/index/hello-gpt-4o/ Hello gpt-4o

  26. [34]

    Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021. https://doi.org/10.18653/v1/2021.naacl-main.168 Are NLP models really able to solve simple math word problems? In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Ling...

  27. [35]

    Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah Smith, and Mike Lewis. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.378 Measuring and narrowing the compositionality gap in language models . In Findings of the Association for Computational Linguistics: EMNLP 20...

  28. [36]

    Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. 2018. Improving language understanding by generative pre-training

  29. [37]

    Manning, and Chelsea Finn

    Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn. 2023. https://arxiv.org/abs/2305.18290 Direct preference optimization: Your language model is secretly a reward model . Preprint, arXiv:2305.18290

  30. [38]

    Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020. https://doi.org/10.1145/3394486.3406703 Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters . In Proceedings of the 26th ACM SIGKDD International Conferenc...

  31. [39]

    Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, and Yuxiong He. 2021. https://www.usenix.org/conference/atc21/presentation/ren-jie ZeRO-Offload : Democratizing Billion-Scale model training . In 2021 USENIX Annual Tec...

  32. [40]

    Timo Schick, Jane Dwivedi-Yu, Roberto Dessi, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. https://openreview.net/forum?id=Yacmpz84TH Toolformer: Language models can teach themselves to use tools . In Thirty-seventh C...

  33. [41]

    Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik R Narasimhan, and Shunyu Yao. 2023. https://openreview.net/forum?id=vAElhFcKW6 Reflexion: language agents with verbal reinforcement learning . In Thirty-seventh Conference on Neural Information Processing Systems

  34. [42]

    Yifan Song, Da Yin, Xiang Yue, Jie Huang, Sujian Li, and Bill Yuchen Lin. 2024. https://arxiv.org/abs/2403.02502 Trial and error: Exploration-based trajectory optimization for llm agents

  35. [43]

    Theodore Sumers, Shunyu Yao, Karthik Narasimhan, and Thomas Griffiths. 2024. https://openreview.net/forum?id=1i6ZCvflQJ Cognitive architectures for language agents . Transactions on Machine Learning Research. Survey Certification

  36. [44]

    Wangtao Sun, Xuanqing Yu, Shizhu He, Jun Zhao, and Kang Liu. 2023. https://openreview.net/forum?id=1Xht3SKAoY Expnote: Black-box large language models are better task solvers with experience notebook . In The 2023 Conference on Empirical Methods in Natural Language Processing

  37. [45]

    Zhengwei Tao, Ting-En Lin, Xiancai Chen, Hangyu Li, Yuchuan Wu, Yongbin Li, Zhi Jin, Fei Huang, Dacheng Tao, and Jingren Zhou. 2024. https://arxiv.org/abs/2404.14387 A survey on self-evolution of large language models . Preprint, arXiv:2404.14387

  38. [46]

    Leandro von Werra, Younes Belkada, Lewis Tunstall, Edward Beeching, Tristan Thrush, Nathan Lambert, and Shengyi Huang. 2020. Trl: Transformer reinforcement learning. https://github.com/huggingface/trl

  39. [47]

    Lei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu, Yunshi Lan, Roy Ka-Wei Lee, and Ee-Peng Lim. 2023 a . https://doi.org/10.18653/v1/2023.acl-long.147 Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models . In Proceedings of the 61st Annua...

  40. [48]

    Renxi Wang, Haonan Li, Xudong Han, Yixuan Zhang, and Timothy Baldwin. 2024 a . https://arxiv.org/abs/2402.11651 Learning from failure: Integrating negative examples when fine-tuning large language models as agents . Preprint, arXiv:2402.11651

  41. [49]

    Ruiyi Wang, Haofei Yu, Wenxin Zhang, Zhengyang Qi, Maarten Sap, Graham Neubig, Yonatan Bisk, and Hao Zhu. 2024 b . https://arxiv.org/abs/2403.08715 Sotopia- : Interactive learning of socially intelligent language agents . Preprint, arXiv:2403.08715

  42. [50]

    Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou

    Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023 b . https://openreview.net/forum?id=1PL1NIMMrw Self-consistency improves chain of thought reasoning in language models . In The Eleventh International Confer...

  43. [51]

    Chi, Quoc V Le, and Denny Zhou

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou. 2022. https://openreview.net/forum?id=_VjQlMeSB_J Chain of thought prompting elicits reasoning in large language models . In Advances in Neural Information Proc...

  44. [52]

    Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, Rui Zheng, Xiaoran Fan, Xiao Wang, Limao Xiong, Yuhao Zhou, Weiran Wang, Changhao Jiang, Yicheng Zou, Xiangyang Liu, Zhangyue Yin, Shihan Dou, Rongxiang Weng, W...

  45. [53]

    Lillicrap, Kenji Kawaguchi, and Michael Shieh

    Yuxi Xie, Anirudh Goyal, Wenyue Zheng, Min-Yen Kan, Timothy P. Lillicrap, Kenji Kawaguchi, and Michael Shieh. 2024. https://arxiv.org/abs/2405.00451 Monte carlo tree search boosts reasoning via iterative preference learning . Preprint, arXiv:2405.00451

  46. [54]

    Yiheng Xu, Hongjin SU, Chen Xing, Boyu Mi, Qian Liu, Weijia Shi, Binyuan Hui, Fan Zhou, Yitao Liu, Tianbao Xie, Zhoujun Cheng, Siheng Zhao, Lingpeng Kong, Bailin Wang, Caiming Xiong, and Tao Yu. 2024. https://openreview.net/forum?id=hNhwSmtXRh Lemur: Harmonizing natural langua...

  47. [55]

    Aiyuan Yang, Bin Xiao, Bingning Wang, Borong Zhang, Ce Bian, Chao Yin, Chenxu Lv, Da Pan, Dian Wang, Dong Yan, Fan Yang, Fei Deng, Feng Wang, Feng Liu, Guangwei Ai, Guosheng Dong, Haizhou Zhao, Hang Xu, Haoze Sun, Hongda Zhang, Hui Liu, Jiaming Ji, Jian Xie, JunTao Dai, Kun Fa...

  48. [56]

    Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. https://doi.org/10.18653/v1/D18-1259 H otpot QA : A dataset for diverse, explainable multi-hop question answering . In Proceedings of the 2018 Conference...

  49. [57]

    Zonghan Yang, Peng Li, Ming Yan, Ji Zhang, Fei Huang, and Yang Liu. 2024. https://arxiv.org/abs/2403.14589 React meets actre: When language agents enjoy training data autonomy . Preprint, arXiv:2403.14589

  50. [58]

    Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. 2023. https://openreview.net/forum?id=WE_vluYUL-X React: Synergizing reasoning and acting in language models . In The Eleventh International Conference on Learning Representations

  51. [59]

    Da Yin, Faeze Brahman, Abhilasha Ravichander, Khyathi Chandu, Kai-Wei Chang, Yejin Choi, and Bill Yuchen Lin. 2024 a . https://arxiv.org/abs/2311.05657 Agent lumos: Unified and modular training for open-source language agents . Preprint, arXiv:2311.05657

  52. [60]

    Da Yin, Faeze Brahman, Abhilasha Ravichander, Khyathi Chandu, Kai-Wei Chang, Yejin Choi, and Bill Yuchen Lin. 2024 b . https://openreview.net/forum?id=VmnWoLbzCS LUMOS : Towards language agents that are unified, modular, and open source

  53. [61]

    Lifan Yuan, Ganqu Cui, Hanbin Wang, Ning Ding, Xingyao Wang, Jia Deng, Boji Shan, Huimin Chen, Ruobing Xie, Yankai Lin, Zhenghao Liu, Bowen Zhou, Hao Peng, Zhiyuan Liu, and Maosong Sun. 2024 a . https://arxiv.org/abs/2404.02078 Advancing llm reasoning generalists with preferen...

  54. [62]

    Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li, Sainbayar Sukhbaatar, Jing Xu, and Jason Weston. 2024 b . https://arxiv.org/abs/2401.10020 Self-rewarding language models . Preprint, arXiv:2401.10020

  55. [63]

    Xiang Yue, Tuney Zheng, Ge Zhang, and Wenhu Chen. 2024. https://arxiv.org/abs/2405.03548 Mammoth2: Scaling instructions from the web . Preprint, arXiv:2405.03548

  56. [64]

    Aohan Zeng, Mingdao Liu, Rui Lu, Bowen Wang, Xiao Liu, Yuxiao Dong, and Jie Tang. 2023. https://arxiv.org/abs/2310.12823 Agenttuning: Enabling generalized agent abilities for llms . Preprint, arXiv:2310.12823

  57. [65]

    Denny Zhou, Nathanael Sch \"a rli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc V Le, and Ed H. Chi. 2023. https://openreview.net/forum?id=WZH7099tgfM Least-to-most prompting enables complex reasoning in large language mode...

  58. [66]

    Le, Ed H

    Pei Zhou, Jay Pujara, Xiang Ren, Xinyun Chen, Heng-Tze Cheng, Quoc V. Le, Ed H. Chi, Denny Zhou, Swaroop Mishra, and Huaixiu Steven Zheng. 2024. https://arxiv.org/abs/2402.03620 Self-discover: Large language models self-compose reasoning structures . Preprint, arXiv:2402.03620

  59. [67]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...

  60. [68]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.