REVIEW 4 major objections 4 minor 16 references
Creativity in LLM-based Multi-Agent Systems: A Survey
T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This survey argues that creativity in LLM-based multi-agent systems can be organized by three workflow phases, a proactivity spectrum, persona granularity, and three generation techniques, and it offers the field's first unified framework…
desk verdict A useful first map of creative MAS, but the 'first survey' claim needs a tighter scope and explicit inclusion criteria before it holds up. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is a three-part descriptive framework: a workflow decomposed into Planning, Process, and Decision Making; a proactivity spectrum spanning reactive to fully autonomous agents, built from two facets of initiative and control; and a taxonomy of generation techniques and persona granularity that ranges from coarse role labels to fine-grained data-derived profiles. This framework carries the argument by letting the authors place each surveyed system on a map, read off why it is or is not creative, and locate missing benchmarks and evaluation standards. The persona granularity axis also serves as a design lever, with coarse profiles favoring breadth and spontaneity while fine profiles yield predictable, controllable, but potentially biased behavior.
What would settle it
A systematic literature search with explicit inclusion criteria across major HCI and AI venues would falsify the coverage claim if it surfaced a sizable body of creativity-oriented MAS work that does not fit the three-technique taxonomy, or if it found a dedicated creativity-in-MAS survey published before this one; concretely, counting omitted published systems and checking whether they introduce additional generation mechanisms or evaluation dimensions would test whether the framework maps the field or only a sample.
Extended reading notes
Core claim
The paper's central claim is that creative output in LLM-based multi-agent systems is not an accidental emergent effect but a consequence of three controllable design axes: where agents sit on a proactivity spectrum defined by initiative and control, how their personas are specified along a granularity spectrum and by profiling method, and which generation technique is used (divergent exploration, iterative refinement, or collaborative synthesis). It further claims that evaluation should combine objective artifact metrics, subjective creativity assessments, and interaction-level user studies. The paper asserts this is the first survey devoted to creativity in MAS and grounds the taxonomy in an examination of representative text- and image-generation systems, using the framework to expose open problems such as inconsistent evaluation standards, persona bias, coordination conflicts, and the lack of unified benchmarks.
Load-bearing premise
The taxonomy and 'first survey' status rest on the assumption that the systems the authors chose to review fairly represent the whole field of creative LLM multi-agent research, even though the survey states no systematic search or inclusion protocol and limits itself to English-language, text- and image-based, largely Western-centric work.
Editorial extensions
If this is right
- Researchers can position new systems along the workflow, proactivity, and persona axes, making it possible to compare how autonomous and how detailed agents are across otherwise unrelated papers.
- Evaluation practice can move toward a standard bundle: objective diversity metrics such as Distinct-n, Self-BLEU, FID, and semantic similarity, combined with subjective creativity tests such as TTCT and Boden's criteria and user-study instruments such as the Creativity Support Index.
- The documented trade-off between agent proactivity and user trust implies that future creative MAS should adaptively calibrate agent initiative to the task and the user rather than defaulting to maximum autonomy.
- Persona granularity becomes a deliberate design choice with predictable consequences: coarse personas promote divergent exploration, while fine-grained, data-derived personas produce stable and realistic but potentially biased behavior.
- The identified gaps, including no unified benchmark, inconsistent evaluation, amplified bias, coordination conflicts, and ambiguous authorship, define a concrete roadmap for the next wave of creative MAS research.
Reading between the lines
- The paper leaves implicit that the proactivity spectrum could be re-read as a control dial for mixed-initiative creativity, so future systems might modulate agent initiative in real time based on user engagement signals, extending the CoQuest finding that processing delays shape the co-creative process.
- The persona granularity spectrum suggests a testable hypothesis the survey does not run: task type moderates the optimal granularity, with coarse personas helping open-ended divergent tasks and fine-grained personas helping constrained convergent tasks; a controlled comparison could settle this.
- Because the surveyed literature is English-centric and Western-centric, the taxonomy may misdescribe creativity in multilingual or non-Western settings, where criteria such as originality and usefulness themselves differ; extending the survey to those settings is a direct test of its generality.
- The 'first survey' status depends on the authors' selection of representative works; a more systematic review with explicit inclusion criteria could reorganize the taxonomy without contradicting its internal logic.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This survey claims to be the first dedicated to creativity in LLM-based multi-agent systems (MAS). It proposes a three-phase workflow (Planning, Process, Decision Making), a proactivity spectrum, a persona-granularity taxonomy, and three generation techniques (Divergent Exploration, Iterative Refinement, Collaborative Synthesis). It also reviews datasets, evaluation metrics (objective and subjective), interaction/user-study methods, and challenges such as bias, conflict, authorship, and resource efficiency. The paper focuses on text and image generation and includes a Limitations section acknowledging the exclusion of audio/video/embodied modalities and the English-centric, Western-centric nature of the surveyed work.
Significance. If the surveyed corpus is representative, this would be a useful organizing framework for a rapidly growing area that existing MAS surveys mostly treat as infrastructure. The paper's strengths include a broad coverage of HCI and AI work, a clear three-way technique taxonomy, detailed tables mapping frameworks to personae and evaluation methods, and a candid Limitations section. The authors also provide a GitHub repository and name concrete example systems throughout, which helps orient new researchers. However, the survey's central promise of a 'unified framework' and its 'first survey' status are weakened by the absence of a reproducible search/inclusion methodology and by the inclusion of systems whose objectives are correctness- or utility-driven rather than creativity-driven under the paper's own definition.
major comments (4)
- [Abstract and Section 1] The paper claims to 'systematically map' the landscape and to be 'the first survey dedicated to creativity in MAS,' but it never states its search strategy, inclusion/exclusion criteria, time frame, databases, or screening process. A reader cannot determine whether the corpus is representative or complete, and the novelty claim is not verifiable against prior surveys (e.g., computational creativity surveys or LLM-based MAS surveys). This is a load-bearing issue for a survey whose value is its roadmap. Please add a methodology section or, at minimum, an explicit description of how the corpus was assembled and how it relates to existing surveys.
- [Section 5.5 and Table 1/Table 4] Section 5.5 states that the evaluation summary 'focuses exclusively on studies where the primary objective is to assess the creativity of the generated content, deliberately excluding those centered on accuracy, precision, or similar metrics.' Yet Table 1 classifies ProofNet (formal theorem proving), TheAgentCompany (consequential real-world tasks), Agent4Rec (recommendation), LawLuo (legal consultation), and Baby-AIGS-MLer (ML research) under creative technique categories, and Table 4 includes LawLuo and Azerbayev et al. with no creativity-oriented evaluation. The paper does not explain why these systems satisfy its own definition of creativity ('novel and valuable'). This inconsistency makes the proposed taxonomy look like a generic MAS taxonomy rather than a creativity-specific one. Either justify these inclusions with explicit creativity-relevant reasoning or exclude them and revise the claims accordingly.
- [Section 3 (core techniques)] The three techniques—Divergent Exploration, Iterative Refinement, and Collaborative Synthesis—are described in terms of brainstorming, feedback loops, and role-based cooperation, which are general MAS collaboration patterns that apply to any problem-solving system. The paper does not establish what is creativity-specific about these techniques, e.g., how they operationalize novelty and value, or how they differ from generic MAS coordination. To support the claim of a 'unified framework' for creativity, the authors should explicitly connect each technique to the creativity definition and to the evaluation criteria discussed in Section 5, or acknowledge that the taxonomy captures general collaboration mechanisms and position the creativity contribution in the persona and evaluation dimensions.
- [Appendix A and Figure 2] The proactivity classification that underlies Figure 2 is described qualitatively ('red' fully agent-driven, 'purple' low proactivity), but the paper provides no coding rubric, decision rules, or inter-rater reliability measure. For example, MaCTG is classified as fully agent-driven while Co-Scientist and CollabStory are placed at lower levels, but the criteria for these placements are not operationalized. Since the proactivity spectrum is a central contribution, the classification should be reproducible. Please provide explicit indicators for each level and, ideally, a table showing how each included system was coded.
minor comments (4)
- [Section 3.2] The text refers to the 'CIAR benchmark (He et al., 2020)'; the benchmark introduced in He et al. (2020) is the CIA (Commonsense Inference in machine trAnslation) benchmark, so the abbreviation should be corrected.
- [Section 5.1] The sentence 'creativity are inherently subjective' should read 'creativity is inherently subjective' or 'creative outputs are inherently subjective.'
- [Table 4] The entry 'Insightfullness' for Shaer et al. (2024) should be 'Insightfulness'.
- [Table 3] The table lists 'CoAGent (2023b)' but the reference list entry for Zheng et al. (2023b) has the title 'Synergizing Human-AI Agency: A Guide of 23 Heuristics for Service Co-Creation with LLM-Based Agents,' which does not obviously match the 'CoAGent' name; please reconcile the citation or add the correct reference.
Circularity Check
No circularity: the paper is a survey organizing external works; self-citations are illustrative examples, not load-bearing evidence.
full rationale
This paper is a survey, not a derivation. Its central outputs are a taxonomy of proactivity (Sec. 2), three technique families (Sec. 3), persona granularity (Sec. 4), evaluation methods (Sec. 5), and dataset categories (Sec. 6). These are organizing frameworks imposed on a corpus of external published systems, not quantities derived from the paper's own assumptions. There are no fitted parameters, no equations whose outputs equal their inputs, and no 'prediction' that is statistically forced by a prior fit. The self-citations that appear (e.g., Lu et al. 2024 in Table 4; Lin et al. 2024a in the Limitations section) function only as surveyed examples or as a scope caveat about multilingual settings; they are not invoked to justify the taxonomy, to forbid alternatives, or to provide the survey's structure. The paper's claim to be 'the first survey dedicated to creativity in MAS' is a scope assertion, not a derivation, and the absence of a stated systematic search protocol is a representativeness or correctness risk, not circularity. The skeptical observation that some tabulated systems (ProofNet, TheAgentCompany, Agent4Rec, LawLuo) are correctness- or utility-driven does not make the framework circular; it bears on whether the survey's inclusion matches its own creativity definition, which is an external-validity concern. No load-bearing step reduces to its own inputs, so the appropriate circularity score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption Creativity is defined as production of artifacts that are both novel and valuable.
- domain assumption The creative workflow can be decomposed into Planning, Process, and Decision Making phases.
- domain assumption Multi-agent collaboration enhances creativity beyond single-LLM generation.
- domain assumption The reviewed systems are representative of the field of creative MAS.
Cite this review
Pith. "Pith review of Creativity in LLM-based Multi-Agent Systems: A Survey." pith.science (2026). https://pith.science/paper/YATWBUEH
@misc{pith2026250521116,
author = {Pith},
title = {Pith review of: Creativity in LLM-based Multi-Agent Systems: A Survey},
year = {2026},
howpublished = {\url{https://pith.science/paper/YATWBUEH}},
note = {Machine review of arXiv:2505.21116}
}
read the original abstract
Large language model (LLM)-driven multi-agent systems (MAS) are transforming how humans and AIs collaboratively generate ideas and artifacts. While existing surveys provide comprehensive overviews of MAS infrastructures, they largely overlook the dimension of \emph{creativity}, including how novel outputs are generated and evaluated, how creativity informs agent personas, and how creative workflows are coordinated. This is the first survey dedicated to creativity in MAS. We focus on text and image generation tasks, and present: (1) a taxonomy of agent proactivity and persona design; (2) an overview of generation techniques, including divergent exploration, iterative refinement, and collaborative synthesis, as well as relevant datasets and evaluation metrics; and (3) a discussion of key challenges, such as inconsistent evaluation standards, insufficient bias mitigation, coordination conflicts, and the lack of unified benchmarks. This survey offers a structured framework and roadmap for advancing the development, evaluation, and standardization of creative MAS.
Figures
Reference graph
Works this paper leans on
-
[5]
A survey on LLM-based multi-agent sys- tems: workflow, infrastructure, and challenges. Vici- nagearth, 1(1):9. Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Shuming Shi, and Zhaopeng Tu. 2024. Encouraging divergent thinking in large language models through multi-agent debate. In Proceedings of the 2024 Conference on Empi...
arXiv 2024
-
[6]
AI Idea Bench 2025: AI Research Idea Gener- ation Benchmark. Preprint, arXiv:2504.14191. Marissa Radensky. 2024. Mixed-Initiative Methods for Co-Creation in Scientific Research. In Proceedings of the 16th Conference on Creativity & Cognition , C&C ’24, page 1–7. Association for Computing Ma- chinery. Nils Reimers and Iryna Gurevych. 2019. Sentence-bert: S...
arXiv 2025
-
[7]
Personality Traits in Large Language Mod- els. Preprint, arXiv:2307.00184. Orit Shaer, Angelora Cooper, Osnat Mokryn, Andrew L Kun, and Hagit Ben Shoshan. 2024. Ai-augmented brainwriting: Investigating the use of llms in group ideation. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, CHI ’24, New York, NY , USA. Associatio...
arXiv 2024
-
[8]
arXiv preprint arXiv:2412.04629
Argumentative Experience: Reducing Con- firmation Bias on Controversial Issues through LLM- Generated Multi-Persona Debates. arXiv preprint arXiv:2412.04629. Aakriti Singh, Shipra Saraswat, and Neetu Faujdar
-
[10]
Lean copilot: Large language models as copilots for theorem proving in lean. Preprint, arXiv:2404.12534. Aarohi Srivastava et al. 2023. Beyond the imitation game: Quantifying and extrapolating the capabili- ties of language models. Transactions on Machine Learning Research. Haoyang Su, Renqi Chen, Shixiang Tang, Zhenfei Yin, Xinzhe Zheng, Jinzhe Li, Biqin...
arXiv 2023
-
[11]
It Felt Like Hav- ing a Second Mind
CREA: A Collaborative Multi-Agent Frame- work for Creative Content Generation with Diffusion Models. Preprint, arXiv:2504.05306. Saranya Venkatraman, Nafis Irtiza Tripto, and Dongwon Lee. 2024. Collabstory: Multi-llm collaborative story generation and authorship analysis. In Proc. NAACL ’25. M. A. Wallach and N. Kogan. 1965. Modes of Think- ing in Young C...
arXiv 2024
-
[12]
Bingyu Yan, Xiaoming Zhang, Litian Zhang, Lian Zhang, Ziyi Zhou, Dezhuang Miao, and Chaozhuo Li
TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks.Preprint, arXiv:2412.14161. Bingyu Yan, Xiaoming Zhang, Litian Zhang, Lian Zhang, Ziyi Zhou, Dezhuang Miao, and Chaozhuo Li
-
[13]
Beyond self-talk: A communication-centric survey of llm-based multi-agent systems. Preprint, arXiv:2502.14321. Kaixun Yang, Yixin Cheng, Linxuan Zhao, Mladen Rakovi´c, Zachari Swiecki, Dragan Gaševi ´c, and Guanliang Chen. 2024. Ink and algorithm: Exploring temporal dynamics in human-ai collaborative writing. Preprint, arXiv:2406.14885. 18 Lyumanshan Ye, ...
arXiv 2024
Show all 16 references
-
[14]
arXiv preprint arXiv:2504.12735
The Athenian Academy: A Seven-Layer Ar- chitecture Model for Multi-Agent Systems. arXiv preprint arXiv:2504.12735. An Zhang, Yuxin Chen, Leheng Sheng, Xiang Wang, and Tat-Seng Chua. 2024a. On generative agents in recommendation. In Proceedings of the 47th Inter- national ACM S...
2024 arXiv
-
[15]
arXiv preprint arXiv:2410.19245
Mactg: Multi-agent collaborative thought graph for automatic programming. arXiv preprint arXiv:2410.19245. Lianmin Zheng et al. 2023a. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. In Advances in Neural Information Processing Systems. Qingxiao Zheng, Zhongwei Xu, Abh...
-
[16]
Colin (Ye et al., 2024) exhibits the agent proactivity through a different way
not only receives iterative user requests dur- ing refinement but also incorporates environmental data collected from its sensors, such as weather con- ditions, camera input, and audio input. Colin (Ye et al., 2024) exhibits the agent proactivity through a different way. The s...
2025
-
[2010]
Ma- chine Bias
Conflict resolution in multi-agent systems based on negotiation and arbitrage. In 2010 2nd IEEE International Conference on Information Man- agement and Engineering, pages 304–307. Junseok Kim, Nakyeong Yang, and Kyomin Jung. 2024. Persona is a double-edged sword: Mitigating t...
2010 arXiv
-
[2017]
In 2017 International Confer- ence on Computing, Communication and Automation (ICCCA)
Analyzing Titanic disaster using machine learning algorithms. In 2017 International Confer- ence on Computing, Communication and Automation (ICCCA). Peiyang Song, Kaiyu Yang, and Anima Anandkumar
2017
-
[2021]
Preprint, arXiv:2102.08317
Resource allocation in dynamic multiagent systems. Preprint, arXiv:2102.08317. Mike D’Arcy, Tom Hope, Larry Birnbaum, and Doug Downey. 2024. Marg: Multi-agent review generation for scientific papers. Preprint, arXiv:2401.04259. Allegra De Filippo, Michela Milano, et al. 2024. ...
2024 arXiv
-
[2024]
Naomi Imasato, Kazuki Miyazawa, Takayuki Nagai, and Takato Horii
A Collaborative, Interactive and Context- Aware Drawing Agent for Co-Creative Design.IEEE Transactions on Visualization and Computer Graph- ics. Naomi Imasato, Kazuki Miyazawa, Takayuki Nagai, and Takato Horii. 2024. Creative agents: Simulat- ing the systems model of creativit...
2024 arXiv
-
[2025]
IEEE Access, 13:10499–10512
An Innovative Solution to Design Problems: Applying the Chain-of-Thought Technique to Inte- grate LLM-Based Agents With Concept Generation Methods. IEEE Access, 13:10499–10512. Juraj Gottweis et al. 2025. Towards an AI co-scientist. Preprint, arXiv:2502.18864. Aaron Grattafior...
2025 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.