REVIEW 3 major objections 5 minor 2 cited by
Multi-agent system prompts can be optimized within 50 evaluations by a bandit search whose surrogate reads the workflow's graph, and this beats existing single-agent and multi-agent prompt optimizers across six LLM benchmarks.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 19:17 UTC pith:TUDUROAJ
load-bearing objection Useful new combination of GNN surrogate and bandit search for prompt optimization in frozen-topology MAS, but the headline overclaims and the comparison against MIPRO is confounded by candidate-prompt domain. the 3 major comments →
MASPOB: Bandit-Based Prompt Optimization for Multi-Agent Systems with Graph Neural Networks
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central discovery claim is that the bottleneck in optimizing fixed-topology multi-agent systems is not the prompt search per se, but the way the search handles inter-agent coupling—and that this coupling can be modeled with a graph. MASPOB represents each agent as a node in the workflow DAG, feeds the candidate prompts' embeddings through a graph-neural-network surrogate to predict system performance, and adds a linear upper-confidence-bound bonus that grows with how unexplored a prompt combination is in embedding space. A coordinate-ascent loop then updates one agent's prompt at a time against the UCB score, cutting the per-round search from exponential to linear in the number o
What carries the argument
Three components carry the argument. (1) A graph neural network (GNN) surrogate: the workflow is a DAG; each agent is a node whose feature is the embedding of its current prompt, and attention-based message passing lets the surrogate predict how a prompt change ripples downstream. (2) A linear upper-confidence-bound (UCB) rule: an information matrix accumulates the embeddings of evaluated combinations; the term sqrt(φ(c)ᵀ M⁻¹ φ(c)) estimates uncertainty, and the acquisition score adds this to the GNN's predicted score, balancing exploitation and exploration. (3) Coordinate ascent: starting from the incumbent best combination, each agent's prompt is greedily replaced with the one maximizing U
Load-bearing premise
MASPOB can only pick from fixed, pre-generated candidate prompts per agent, so if the best prompt for an agent was never drafted, the search cannot find it and all reported gains are bounded by the quality and diversity of that initial candidate pool (a limitation the paper states explicitly in Section 4 and Appendix A.3).
What would settle it
Run MASPOB on the same budgets but with a deliberately degraded candidate pool (e.g., 20 near-duplicate paraphrases of one prompt) and compare to random selection from that pool: if random search matches or beats MASPOB, the bandit/GNN machinery is not the source of the reported gains. The complementary test is to give the same 50-evaluation budget to a simple evolutionary search over an identical candidate pool: if the simple baseline matches MASPOB's scores, the topology-aware surrogate and UCB exploration contribute nothing beyond ordinary search.
If this is right
- With a budget of 50 end-to-end evaluations, MASPOB improves average test accuracy by about 12 percentage points over plain prompting and about 2 points over the strongest multi-agent prompt baseline.
- The improvements appear across all six benchmarks—QA, code generation, and math—suggesting the benefit comes from coordination rather than task-specific prompt content.
- Removing the graph surrogate costs about 2.3 average points, indicating that topology-aware modeling is a measurable source of the gain, not a cosmetic addition.
- Coordinate ascent matches exhaustive search within roughly 0.3–0.5 points but runs 98–99.8% faster, so the search for good prompt combinations is computationally practical.
- The benefit transfers to a different backbone LLM and to an alternative, independently generated candidate pool, which the paper interprets as evidence that the gains come from the optimization procedure rather than from a single model or a single prompt-domain recipe.
Where Pith is reading between the lines
- Because MASPOB only selects among pre-written candidates, its ceiling is set by the prompt generator; a natural next step is to let the bandit's uncertainty signal trigger the drafting of new variants in unexplored regions, converting selection into closed-loop generation.
- The reported scaling suggests the GNN surrogate's advantage should grow with workflow size and coupling: a testable prediction is that on workflows with more agents or with feedback edges (currently excluded by the DAG assumption), a topology-aware surrogate will separate from structure-blind baselines by a larger margin than the roughly 2.3 points seen here.
- The near-tie between coordinate ascent and global search suggests the UCB bonus itself may be supplying the global exploration that makes coordinate-wise greedy updates safe; ablating the exploration coefficient or removing the bonus on a fixed budget would test whether the safety comes from the bonus or from the smoothness of the performance landscape.
- For practitioners in regulated settings, the practical implication is that prompt tuning can substitute for workflow restructuring up to a point: the paper's protocol preserves the audited topology and still delivers gains, which aligns with deployment constraints where re-validation of the workflow is expensive or forbidden.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes MASPOB, a prompt optimizer for multi-agent systems (MAS) with a fixed workflow topology. The method combines a GAT-based surrogate, a LinUCB-style exploration bonus, and coordinate ascent to search a discrete Cartesian product of per-agent prompt variants under a tight evaluation budget (50 end-to-end executions). The authors evaluate on six benchmarks spanning QA, code generation, and mathematical reasoning, comparing against single-agent baselines (IO, CoT, ReAct, PromptBreeder, Instinct) and multi-agent methods (AFlow, MIPRO). They report an average score of 80.58% and claim consistent state-of-the-art performance. Ablations investigate the GNN, uncertainty estimation, warm-up, embedding model, coordinate ascent, and exploration coefficient.
Significance. If the results hold, MASPOB addresses a practically important and under-studied problem: sample-efficient prompt optimization for already-validated, frozen MAS workflows. The combination of a topology-aware GNN surrogate with a linear UCB bonus is well motivated, and the paper provides a fairly extensive experimental suite with a fixed 50-evaluation budget, ablations on several design components, robustness checks across LLMs and prompt domains, and a complexity comparison against exhaustive search. The public code link and detailed hyperparameters (Appendix A.3) support reproducibility. The main contribution is empirical rather than theoretical: no regret or convergence guarantees are given, but the algorithmic structure is sensible and the experimental setup is largely careful. The principal weakness is that the headline comparison against the closest prompt-optimization baseline, MIPRO, is not fully controlled with respect to how the candidate prompt pool is constructed, and a stated claim of 'best on every benchmark' is contradicted by a tie in Table 1.
major comments (3)
- [§4.2, Table 1, Abstract] The sentence 'MASPOB achieves the best result on every benchmark' is not supported by Table 1. On MBPP, MASPOB and MIPRO both report 80.65, i.e., a tie, and the table's bolding gives MASPOB sole credit. The abstract's 'consistently outperforming' is likewise too strong. Please revise to 'matches or outperforms' or otherwise qualify the claim, and correct the bolding.
- [§4.2, Table 5, Appendix A.3] The comparison between MASPOB and MIPRO is confounded with prompt-domain construction. MASPOB builds its candidate pool via 20 GPT-4o-mini style-controlled paraphrases per agent (Appendix A.3), while MIPRO uses its own data/program/fewshot/tip-aware strategies. Table 5 only swaps MIPRO's domain into MASPOB, showing similar scores; it never runs MIPRO on MASPOB's domain. Section 4 also concedes that 'the quality and diversity of candidate prompts can still affect absolute performance.' Therefore the reported 1.71-point average advantage over MIPRO could be partly or wholly due to the candidate pool rather than the GNN/UCB/coordinate-ascent selection. Please run MIPRO on the MASPOB candidate pool (or otherwise hold the candidate domain fixed across optimizers) and report both directions before claiming that the gains 'mainly come from topology-aware contextual-bandit optimization.'
- [§6 Related Work, Tables 1-4] The paper identifies MAPRO (Zhang et al., 2025c) as 'the closest prior work to ours' and as a principled multi-agent prompt optimizer, yet MAPRO is never included as a baseline. For a state-of-the-art claim, omitting the closest competitor is a significant gap. Please add MAPRO to the main comparisons if its code/API allows, or provide a concrete reason (e.g., incompatibility with the fixed-budget protocol) why it cannot be included.
minor comments (5)
- [Table 5 caption] The caption has a typo: 'We report mean accuracy (standard deviation over three runs' is missing the '±' and an opening bracket. Also, Table 5 reports only DROP and MATH, while the text in §4.2 says 'as shown in Table 5' without noting the limited coverage; please state the scope explicitly.
- [§3.2, Algorithm 1, Table 8] The hyperparameter table lists a 'Fisher matrix update coefficient' of 10, but Algorithm 1 and Eq. (10) update the information matrix as M ← M + Φ(c)Φ(c)⊤ with no coefficient. Either the algorithm description is missing a scaling factor or the table entry is unused. Please clarify.
- [§4.1, Metrics] MATH is referred to as 'MATHlv5*' in the metrics paragraph, which is inconsistent with the dataset name 'MATH' elsewhere. Please unify terminology.
- [§4.2, Figure 3] The convergence figure reports test accuracy at checkpoints every 5 rounds and validation as a binned average, but the caption does not define how the test checkpoints are averaged (the text says 'evaluated and averaged over three runs'). Please state whether the three runs are the three final test repetitions or separate optimization runs.
- [§4.2, Table 1] The 12.02% average improvement over IO is emphasized as a headline result, but IO is a single-LLM-call baseline while MASPOB uses a multi-agent workflow with multiple LLM calls. The evaluation budget is matched in number of full-workflow executions, not in inference cost. Please add a sentence clarifying this distinction so readers do not interpret the gain as being achieved at equal total LLM inference cost.
Circularity Check
No significant circularity: MASPOB's GNN/UCB surrogate is trained on observed validation scores and evaluated on held-out test benchmarks; self-citations to neural-bandit prior work are not load-bearing.
full rationale
MASPOB's central claim is an empirical comparison on six external benchmarks. The GNN surrogate is trained on validation-set scores obtained by end-to-end MAS executions (Eq. 1, Algorithm 1 lines 5-9 and 22-25), and the final prompt combination is selected by validation performance and then evaluated on a held-out test split (Section 2 and Appendix A.1); the test labels are never used to fit the surrogate or the UCB information matrix. The LinUCB uncertainty term (Eqs. 11-12) is a standard exploration bonus over the same fitted representation, not a renamed version of the objective being predicted. Coordinate ascent (Eq. 13) only reduces acquisition-search cost, and Table 6 shows it tracks global search on the same acquisition function. The paper cites prior work by its own authors (e.g., Lin et al. 2023; Wu et al. 2024b; Kong et al. 2025) for the bandit formulation and for the Instinct baseline, but no load-bearing step invokes those papers as an external fact that forces the result; the 'NeuralLinear-style' decomposition in Appendix B is presented as a design choice and is validated against a neural-uncertainty ablation (Table 9). The prompt-domain robustness experiment (Table 5) only swaps MIPRO's candidate domain into MASPOB, so the MIPRO comparison may be confounded by candidate-prompt quality, but that is an experimental-design/validity concern, not a definitional or self-citational reduction; Section 4 explicitly concedes that 'the quality and diversity of candidate prompts can still affect absolute performance, since the optimizer can only select from the provided prompt domain.' Overall, no equation or fitted parameter is constructed so that the reported benchmark improvement is true by definition.
Axiom & Free-Parameter Ledger
free parameters (5)
- Exploration coefficient α =
0.2
- Regularization coefficient λ =
1.0
- Warm-up rounds T0 =
5
- Fisher matrix update coefficient =
10 (unexplained)
- Prompt variants per agent =
20
axioms (6)
- domain assumption MAS inter-agent information flow is a static DAG
- domain assumption Small validation splits are reliable proxies for test performance
- ad hoc to paper GAT surrogate trained on ≤50 observations provides useful performance predictions
- ad hoc to paper LinUCB uncertainty in combined prompt-embedding space is a valid exploration signal
- ad hoc to paper Coordinate ascent on a non-concave UCB acquisition function finds near-optimal combinations
- domain assumption Pretrained prompt embeddings preserve task-relevant semantic differences
read the original abstract
Large Language Models (LLMs) have achieved great success in many real-world applications, especially the one serving as the cognitive backbone of Multi-Agent Systems (MAS) to orchestrate complex workflows in practice. Since many deployment scenarios preclude MAS workflow modifications and its performance is highly sensitive to the input prompts, prompt optimization emerges as a more natural approach to improve its performance. However, real-world prompt optimization for MAS is impeded by three key challenges: (1) the need of sample efficiency due to prohibitive evaluation costs, (2) topology-induced coupling among prompts, and (3) the combinatorial explosion of the search space. To address these challenges, we introduce MASPOB (Multi-Agent System Prompt Optimization via Bandits), a novel sample-efficient framework based on bandits. By leveraging Upper Confidence Bound (UCB) to quantify uncertainty, the bandit framework balances exploration and exploitation, maximizing gains within a strictly limited budget. To handle topology-induced coupling, MASPOB integrates Graph Neural Networks (GNNs) to capture structural priors, learning topology-aware representations of prompt semantics. Furthermore, it employs coordinate ascent to decompose the optimization into univariate sub-problems, reducing search complexity from exponential to linear. Extensive experiments across diverse benchmarks demonstrate that MASPOB achieves state-of-the-art performance, consistently outperforming existing baselines.
Figures
Forward citations
Cited by 2 Pith papers
-
MAS-PromptBench: When Does Prompt Optimization Improve Multi-Agent LLM Systems?
A new benchmark study finds that prompt optimization can deliver significant gains in multi-agent LLM systems but its effectiveness varies strongly with task, workflow, communication protocol, and team size.
-
PYTHALAB-MERA: Validation-Grounded Memory, Retrieval, and Acceptance Control for Frozen-LLM Coding Agents
An external controller for frozen LLMs raises strict validation success on three RL coding tasks from 0/9 to 8/9 by selecting memory records and skills, running fail-fast checks, and propagating credit via eligibility traces.
Reference graph
Works this paper leans on
-
[1]
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y., P \'a l, D., and Szepesv \'a ri, C. Improved algorithms for linear stochastic bandits. Advances in neural information processing systems, 24, 2011
2011
-
[2]
Program synthesis with large language models
Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., et al. Program synthesis with large language models. arXiv preprint arXiv:2108.07732, 2021
Pith/arXiv arXiv 2021
-
[3]
and Travis, J
Bodnari, A. and Travis, J. Scaling enterprise ai in healthcare: the role of governance in risk mitigation frameworks. npj Digital Medicine, 8 0 (1): 0 272, 2025
2025
-
[4]
D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems, 33: 0 1877--1901, 2020
1901
-
[5]
Instructzero: Efficient instruction optimization for black-box large language models
Chen, L., Chen, J., Goldstein, T., Huang, H., and Zhou, T. Instructzero: Efficient instruction optimization for black-box large language models. arXiv preprint arXiv:2306.03082, 2023 a
Pith/arXiv arXiv 2023
-
[6]
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. D. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021
Pith/arXiv arXiv 2021
-
[7]
Agentverse: Facilitating multi-agent collaboration and exploring emergent behaviors
Chen, W., Su, Y., Zuo, J., Yang, C., Yuan, C., Chan, C.-M., Yu, H., Lu, Y., Hung, Y.-H., Qian, C., et al. Agentverse: Facilitating multi-agent collaboration and exploring emergent behaviors. In The Twelfth International Conference on Learning Representations, 2023 b
2023
-
[8]
Contextual bandits with linear payoff functions
Chu, W., Li, L., Reyzin, L., and Schapire, R. Contextual bandits with linear payoff functions. In Proceedings of the fourteenth international conference on artificial intelligence and statistics, pp.\ 208--214. JMLR Workshop and Conference Proceedings, 2011
2011
-
[9]
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021
Pith/arXiv arXiv 2021
-
[10]
Rlprompt: Optimizing discrete text prompts with reinforcement learning
Deng, M., Wang, J., Hsieh, C.-P., Wang, Y., Guo, H., Shu, T., Song, M., Xing, E., and Hu, Z. Rlprompt: Optimizing discrete text prompts with reinforcement learning. In Proceedings of the 2022 conference on empirical methods in natural language processing, pp.\ 3369--3391, 2022
2022
-
[11]
Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dua, D., Wang, Y., Dasigi, P., Stanovsky, G., Singh, S., and Gardner, M. Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs. arXiv preprint arXiv:1903.00161, 2019
Pith/arXiv arXiv 1903
-
[12]
Promptbreeder: Self-referential self-improvement via prompt evolution
Fernando, C., Banarse, D., Michalewski, H., Osindero, S., and Rockt \"a schel, T. Promptbreeder: Self-referential self-improvement via prompt evolution. arXiv preprint arXiv:2309.16797, 2023
Pith/arXiv arXiv 2023
-
[13]
Connecting large language models with evolutionary algorithms yields powerful prompt optimizers
Guo, Q., Wang, R., Guo, J., Li, B., Song, K., Tan, X., Liu, G., Bian, J., and Yang, Y. Connecting large language models with evolutionary algorithms yields powerful prompt optimizers. arXiv preprint arXiv:2309.08532, 2023
Pith/arXiv arXiv 2023
-
[14]
Measuring mathematical problem solving with the math dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J. Measuring mathematical problem solving with the math dataset. arXiv preprint arXiv:2103.03874, 2021
Pith/arXiv arXiv 2021
-
[15]
Hong, S., Zhuge, M., Chen, J., Zheng, X., Cheng, Y., Wang, J., Zhang, C., Wang, Z., Yau, S. K. S., Lin, Z., et al. Metagpt: Meta programming for a multi-agent collaborative framework. In The twelfth international conference on learning representations, 2023
2023
-
[16]
Automated design of agentic systems
Hu, S., Lu, C., and Clune, J. Automated design of agentic systems. arXiv preprint arXiv:2408.08435, 2024
Pith/arXiv arXiv 2024
-
[17]
P., Perelman, A., Ramesh, A., Clark, A., Ostrow, A., Welihinda, A., Hayes, A., Radford, A., et al
Hurst, A., Lerer, A., Goucher, A. P., Perelman, A., Ramesh, A., Clark, A., Ostrow, A., Welihinda, A., Hayes, A., Radford, A., et al. Gpt-4o system card. arXiv preprint arXiv:2410.21276, 2024
Pith/arXiv arXiv 2024
-
[18]
Khattab, O., Singhvi, A., Maheshwari, P., Zhang, Z., Santhanam, K., Vardhamanan, S., Haq, S., Sharma, A., Joshi, T. T., Moazam, H., et al. Dspy: Compiling declarative language model calls into self-improving pipelines. arXiv preprint arXiv:2310.03714, 2023
Pith/arXiv arXiv 2023
-
[19]
S., Reid, M., Matsuo, Y., and Iwasawa, Y
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y. Large language models are zero-shot reasoners. Advances in neural information processing systems, 35: 0 22199--22213, 2022
2022
-
[20]
Meta-prompt optimization for llm-based sequential decision making
Kong, M., Wang, Z., Shu, Y., and Dai, Z. Meta-prompt optimization for llm-based sequential decision making. arXiv preprint arXiv:2502.00728, 2025
Pith/arXiv arXiv 2025
-
[21]
Camel: Communicative agents for" mind" exploration of large language model society
Li, G., Hammoud, H., Itani, H., Khizbullin, D., and Ghanem, B. Camel: Communicative agents for" mind" exploration of large language model society. Advances in Neural Information Processing Systems, 36: 0 51991--52008, 2023
2023
-
[22]
Li, L., Chu, W., Langford, J., and Schapire, R. E. A contextual-bandit approach to personalized news article recommendation. In Proceedings of the 19th international conference on World wide web, pp.\ 661--670, 2010
2010
-
[23]
Combinatorial optimization with graph convolutional networks and guided tree search
Li, Z., Chen, Q., and Koltun, V. Combinatorial optimization with graph convolutional networks and guided tree search. Advances in neural information processing systems, 31, 2018
2018
-
[24]
Lin, X., Wu, Z., Dai, Z., Hu, W., Shu, Y., Ng, S.-K., Jaillet, P., and Low, B. K. H. Use your instinct: Instruction optimization using neural bandits coupled with transformers. arXiv preprint arXiv:2310.02905, 2023
Pith/arXiv arXiv 2023
-
[25]
Lin, X., Dai, Z., Verma, A., Ng, S.-K., Jaillet, P., and Low, B. K. H. Prompt optimization with human feedback. arXiv preprint arXiv:2405.17346, 2024
Pith/arXiv arXiv 2024
-
[26]
Dynamic llm-agent network: An llm-agent collaboration framework with agent team optimization
Liu, Z., Zhang, Y., Li, P., Liu, Y., and Yang, D. Dynamic llm-agent network: An llm-agent collaboration framework with agent team optimization. arXiv preprint arXiv:2310.02170, 2023
Pith/arXiv arXiv 2023
-
[27]
Sop-bench: Complex industrial sops for evaluating llm agents
Nandi, S., Datta, A., Vichare, N., Bhattacharya, I., Raja, H., Xu, J., Ray, S., Carenini, G., Srivastava, A., Chan, A., et al. Sop-bench: Complex industrial sops for evaluating llm agents. arXiv preprint arXiv:2506.08119, 2025
arXiv 2025
-
[29]
J., Purtell, J., Broman, D., Potts, C., Zaharia, M., and Khattab, O
Opsahl-Ong, K., Ryan, M. J., Purtell, J., Broman, D., Potts, C., Zaharia, M., and Khattab, O. Optimizing instructions and demonstrations for multi-stage language model programs. arXiv preprint arXiv:2406.11695, 2024 b
Pith/arXiv arXiv 2024
-
[30]
Flow-of-action: Sop enhanced llm-based multi-agent system for root cause analysis
Pei, C., Wang, Z., Liu, F., Li, Z., Liu, Y., He, X., Kang, R., Zhang, T., Chen, J., Li, J., et al. Flow-of-action: Sop enhanced llm-based multi-agent system for root cause analysis. In Companion Proceedings of the ACM on Web Conference 2025, pp.\ 422--431, 2025
2025
-
[31]
Grips: Gradient-free, edit-based instruction search for prompting large language models
Prasad, A., Hase, P., Zhou, X., and Bansal, M. Grips: Gradient-free, edit-based instruction search for prompting large language models. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pp.\ 3845--3864, 2023
2023
-
[32]
Pryzant, R., Iter, D., Li, J., Lee, Y. T., Zhu, C., and Zeng, M. Automatic prompt optimization with" gradient descent" and beam search. arXiv preprint arXiv:2305.03495, 2023
Pith/arXiv arXiv 2023
-
[33]
Chatdev: Communicative agents for software development
Qian, C., Liu, W., Liu, H., Chen, N., Dang, Y., Li, J., Yang, C., Chen, W., Su, Y., Cong, X., et al. Chatdev: Communicative agents for software development. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp.\ 15174--15186, 2024
2024
-
[34]
Verifai: verified generative ai
Tang, N., Yang, C., Fan, J., Cao, L., Luo, Y., and Halevy, A. Verifai: verified generative ai. arXiv preprint arXiv:2307.02796, 2023
Pith/arXiv arXiv 2023
-
[35]
Veli c kovi \'c , P., Cucurull, G., Casanova, A., Romero, A., Lio, P., and Bengio, Y. Graph attention networks. arXiv preprint arXiv:1710.10903, 2017
Pith/arXiv arXiv 2017
-
[36]
Voyager: An open-ended embodied agent with large language models
Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., and Anandkumar, A. Voyager: An open-ended embodied agent with large language models. arXiv preprint arXiv:2305.16291, 2023 a
Pith/arXiv arXiv 2023
-
[37]
Self-consistency improves chain of thought reasoning in language models, 2023 b
Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., and Zhou, D. Self-consistency improves chain of thought reasoning in language models, 2023 b . URL https://arxiv.org/abs/2203.11171
Pith/arXiv arXiv 2023
-
[38]
V., Zhou, D., et al
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35: 0 24824--24837, 2022
2022
-
[39]
J., Nguyen, H
Wells, B. J., Nguyen, H. M., McWilliams, A., Pallini, M., Bovi, A., Kuzma, A., Kramer, J., Chou, S.-H., Hetherington, T., Corn, P., et al. A practical framework for appropriate implementation and review of artificial intelligence (fair-ai) in healthcare. NPJ digital medicine, 8 0 (1): 0 514, 2025
2025
-
[40]
Medreason: Eliciting factual medical reasoning steps in llms via knowledge graphs
Wu, J., Deng, W., Li, X., Liu, S., Mi, T., Peng, Y., Xu, Z., Liu, Y., Cho, H., Choi, C.-I., et al. Medreason: Eliciting factual medical reasoning steps in llms via knowledge graphs. arXiv preprint arXiv:2504.00993, 2025
Pith/arXiv arXiv 2025
-
[41]
Autogen: Enabling next-gen llm applications via multi-agent conversations
Wu, Q., Bansal, G., Zhang, J., Wu, Y., Li, B., Zhu, E., Jiang, L., Zhang, X., Zhang, S., Liu, J., et al. Autogen: Enabling next-gen llm applications via multi-agent conversations. In First Conference on Language Modeling, 2024 a
2024
-
[42]
Wu, Z., Lin, X., Dai, Z., Hu, W., Shu, Y., Ng, S.-K., Jaillet, P., and Low, B. K. H. Prompt optimization with ease? efficient ordering-aware automated selection of exemplars. Advances in Neural Information Processing Systems, 37: 0 122706--122740, 2024 b
2024
-
[43]
Yang, A., Li, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Gao, C., Huang, C., Lv, C., et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025
Pith/arXiv arXiv 2025
-
[44]
V., Zhou, D., and Chen, X
Yang, C., Wang, X., Lu, Y., Liu, H., Le, Q. V., Zhou, D., and Chen, X. Large language models as optimizers. In The Twelfth International Conference on Learning Representations, 2023
2023
-
[45]
Finrobot: An open-source ai agent platform for financial applications using large language models
Yang, H., Zhang, B., Wang, N., Guo, C., Zhang, X., Lin, L., Wang, J., Zhou, T., Guan, M., Zhang, R., et al. Finrobot: An open-source ai agent platform for financial applications using large language models. arXiv preprint arXiv:2405.14767, 2024
Pith/arXiv arXiv 2024
-
[46]
Yang, Z., Qi, P., Zhang, S., Bengio, Y., Cohen, W., Salakhutdinov, R., and Manning, C. D. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. In Proceedings of the 2018 conference on empirical methods in natural language processing, pp.\ 2369--2380, 2018
2018
-
[47]
R., and Cao, Y
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K. R., and Cao, Y. React: Synergizing reasoning and acting in language models. In The eleventh international conference on learning representations, 2022
2022
-
[48]
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T., Cao, Y., and Narasimhan, K. Tree of thoughts: Deliberate problem solving with large language models. Advances in neural information processing systems, 36: 0 11809--11822, 2023 a
2023
-
[49]
React: Synergizing reasoning and acting in language models, 2023 b
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y. React: Synergizing reasoning and acting in language models, 2023 b . URL https://arxiv.org/abs/2210.03629
Pith/arXiv arXiv 2023
-
[50]
Sop-agent: Empower general purpose ai agent with domain-specific sops
Ye, A., Ma, Q., Chen, J., Li, M., Li, T., Liu, F., Mai, S., Lu, M., Bao, H., and You, Y. Sop-agent: Empower general purpose ai agent with domain-specific sops. arXiv preprint arXiv:2501.09316, 2025
Pith/arXiv arXiv 2025
-
[51]
Thought propagation: An analogical approach to complex reasoning with large language models
Yu, J., He, R., and Ying, R. Thought propagation: An analogical approach to complex reasoning with large language models. arXiv preprint arXiv:2310.03965, 2023
Pith/arXiv arXiv 2023
-
[52]
Yuksekgonul, M., Bianchi, F., Boen, J., Liu, S., Huang, Z., Guestrin, C., and Zou, J. Textgrad: Automatic" differentiation" via text. arXiv preprint arXiv:2406.07496, 2024
Pith/arXiv arXiv 2024
-
[53]
Aflow: Automating agentic workflow generation
Zhang, J., Xiang, J., Yu, Z., Teng, F., Chen, X., Chen, J., Zhuge, M., Cheng, X., Hong, S., Wang, J., et al. Aflow: Automating agentic workflow generation. arXiv preprint arXiv:2410.10762, 2024
Pith/arXiv arXiv 2024
-
[54]
Gnns as predictors of agentic workflow performances
Zhang, Y., Hou, Y., Tang, B., Chen, S., Zhang, M., Dong, X., and Chen, S. Gnns as predictors of agentic workflow performances. arXiv preprint arXiv:2503.11301, 2025 a
Pith/arXiv arXiv 2025
-
[55]
Qwen3 embedding: Advancing text embedding and reranking through foundation models
Zhang, Y., Li, M., Long, D., Zhang, X., Lin, H., Yang, B., Xie, P., Yang, A., Liu, D., Lin, J., et al. Qwen3 embedding: Advancing text embedding and reranking through foundation models. arXiv preprint arXiv:2506.05176, 2025 b
Pith/arXiv arXiv 2025
-
[56]
Mapro: Recasting multi-agent prompt optimization as maximum a posteriori inference, 2025 c
Zhang, Z., Ge, L., Li, H., Zhu, W., Zhang, C., and Ye, Y. Mapro: Recasting multi-agent prompt optimization as maximum a posteriori inference, 2025 c . URL https://arxiv.org/abs/2510.07475
arXiv 2025
-
[57]
Neural contextual bandits with ucb-based exploration
Zhou, D., Li, L., and Gu, Q. Neural contextual bandits with ucb-based exploration. In International conference on machine learning, pp.\ 11492--11502. PMLR, 2020
2020
-
[58]
I., Han, Z., Paster, K., Pitis, S., Chan, H., and Ba, J
Zhou, Y., Muresanu, A. I., Han, Z., Paster, K., Pitis, S., Chan, H., and Ba, J. Large language models are human-level prompt engineers. In The eleventh international conference on learning representations, 2022
2022
-
[59]
Gptswarm: Language agents as optimizable graphs
Zhuge, M., Wang, W., Kirsch, L., Faccio, F., Khizbullin, D., and Schmidhuber, J. Gptswarm: Language agents as optimizable graphs. In Forty-first International Conference on Machine Learning, 2024
2024
-
[60]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.