Pith. sign in

REVIEW 4 major objections 4 minor 16 references

Creativity in LLM-based Multi-Agent Systems: A Survey

T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This survey argues that creativity in LLM-based multi-agent systems can be organized by three workflow phases, a proactivity spectrum, persona granularity, and three generation techniques, and it offers the field's first unified framework…

desk verdict A useful first map of creative MAS, but the 'first survey' claim needs a tighter scope and explicit inclusion criteria before it holds up. read the letter →

arxiv 2505.21116 v1 pith:YATWBUEH submitted 2025-05-27 cs.HC cs.AIcs.CL

classification cs.HCcs.AIcs.CL
keywords multi-agentsystemsLLMagentscomputationalcreativitydivergentthinkingagentproactivitypersonadesigncreativeevaluationhuman-AIco-creativity
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Existing surveys of LLM-based multi-agent systems catalog infrastructure, communication, and collaboration mechanisms but largely leave creativity out. This paper supplies what it positions as the first survey dedicated to creativity in such systems, decomposing creative workflows into Planning, Process, and Decision Making and placing agents on a proactivity spectrum from reactive to fully autonomous. It also identifies three generation techniques that reliably shape creative output, organizes persona design by granularity and profiling method, and maps the datasets and evaluation metrics used to measure creativity. For a reader, the payoff is a shared vocabulary and roadmap that turns creativity into a designable dimension of multi-agent systems rather than an incidental byproduct. If the framework holds, future systems can be compared, evaluated, and standardized along concrete axes.

What carries the argument

The central machinery is a three-part descriptive framework: a workflow decomposed into Planning, Process, and Decision Making; a proactivity spectrum spanning reactive to fully autonomous agents, built from two facets of initiative and control; and a taxonomy of generation techniques and persona granularity that ranges from coarse role labels to fine-grained data-derived profiles. This framework carries the argument by letting the authors place each surveyed system on a map, read off why it is or is not creative, and locate missing benchmarks and evaluation standards. The persona granularity axis also serves as a design lever, with coarse profiles favoring breadth and spontaneity while fine profiles yield predictable, controllable, but potentially biased behavior.

What would settle it

A systematic literature search with explicit inclusion criteria across major HCI and AI venues would falsify the coverage claim if it surfaced a sizable body of creativity-oriented MAS work that does not fit the three-technique taxonomy, or if it found a dedicated creativity-in-MAS survey published before this one; concretely, counting omitted published systems and checking whether they introduce additional generation mechanisms or evaluation dimensions would test whether the framework maps the field or only a sample.

Watch

Extended reading notes

Core claim

The paper's central claim is that creative output in LLM-based multi-agent systems is not an accidental emergent effect but a consequence of three controllable design axes: where agents sit on a proactivity spectrum defined by initiative and control, how their personas are specified along a granularity spectrum and by profiling method, and which generation technique is used (divergent exploration, iterative refinement, or collaborative synthesis). It further claims that evaluation should combine objective artifact metrics, subjective creativity assessments, and interaction-level user studies. The paper asserts this is the first survey devoted to creativity in MAS and grounds the taxonomy in an examination of representative text- and image-generation systems, using the framework to expose open problems such as inconsistent evaluation standards, persona bias, coordination conflicts, and the lack of unified benchmarks.

Load-bearing premise

The taxonomy and 'first survey' status rest on the assumption that the systems the authors chose to review fairly represent the whole field of creative LLM multi-agent research, even though the survey states no systematic search or inclusion protocol and limits itself to English-language, text- and image-based, largely Western-centric work.

Editorial extensions

If this is right

  • Researchers can position new systems along the workflow, proactivity, and persona axes, making it possible to compare how autonomous and how detailed agents are across otherwise unrelated papers.
  • Evaluation practice can move toward a standard bundle: objective diversity metrics such as Distinct-n, Self-BLEU, FID, and semantic similarity, combined with subjective creativity tests such as TTCT and Boden's criteria and user-study instruments such as the Creativity Support Index.
  • The documented trade-off between agent proactivity and user trust implies that future creative MAS should adaptively calibrate agent initiative to the task and the user rather than defaulting to maximum autonomy.
  • Persona granularity becomes a deliberate design choice with predictable consequences: coarse personas promote divergent exploration, while fine-grained, data-derived personas produce stable and realistic but potentially biased behavior.
  • The identified gaps, including no unified benchmark, inconsistent evaluation, amplified bias, coordination conflicts, and ambiguous authorship, define a concrete roadmap for the next wave of creative MAS research.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper leaves implicit that the proactivity spectrum could be re-read as a control dial for mixed-initiative creativity, so future systems might modulate agent initiative in real time based on user engagement signals, extending the CoQuest finding that processing delays shape the co-creative process.
  • The persona granularity spectrum suggests a testable hypothesis the survey does not run: task type moderates the optimal granularity, with coarse personas helping open-ended divergent tasks and fine-grained personas helping constrained convergent tasks; a controlled comparison could settle this.
  • Because the surveyed literature is English-centric and Western-centric, the taxonomy may misdescribe creativity in multilingual or non-Western settings, where criteria such as originality and usefulness themselves differ; extending the survey to those settings is a direct test of its generality.
  • The 'first survey' status depends on the authors' selection of representative works; a more systematic review with explicit inclusion criteria could reorganize the taxonomy without contradicting its internal logic.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. This survey claims to be the first dedicated to creativity in LLM-based multi-agent systems (MAS). It proposes a three-phase workflow (Planning, Process, Decision Making), a proactivity spectrum, a persona-granularity taxonomy, and three generation techniques (Divergent Exploration, Iterative Refinement, Collaborative Synthesis). It also reviews datasets, evaluation metrics (objective and subjective), interaction/user-study methods, and challenges such as bias, conflict, authorship, and resource efficiency. The paper focuses on text and image generation and includes a Limitations section acknowledging the exclusion of audio/video/embodied modalities and the English-centric, Western-centric nature of the surveyed work.

Significance. If the surveyed corpus is representative, this would be a useful organizing framework for a rapidly growing area that existing MAS surveys mostly treat as infrastructure. The paper's strengths include a broad coverage of HCI and AI work, a clear three-way technique taxonomy, detailed tables mapping frameworks to personae and evaluation methods, and a candid Limitations section. The authors also provide a GitHub repository and name concrete example systems throughout, which helps orient new researchers. However, the survey's central promise of a 'unified framework' and its 'first survey' status are weakened by the absence of a reproducible search/inclusion methodology and by the inclusion of systems whose objectives are correctness- or utility-driven rather than creativity-driven under the paper's own definition.

major comments (4)
  1. [Abstract and Section 1] The paper claims to 'systematically map' the landscape and to be 'the first survey dedicated to creativity in MAS,' but it never states its search strategy, inclusion/exclusion criteria, time frame, databases, or screening process. A reader cannot determine whether the corpus is representative or complete, and the novelty claim is not verifiable against prior surveys (e.g., computational creativity surveys or LLM-based MAS surveys). This is a load-bearing issue for a survey whose value is its roadmap. Please add a methodology section or, at minimum, an explicit description of how the corpus was assembled and how it relates to existing surveys.
  2. [Section 5.5 and Table 1/Table 4] Section 5.5 states that the evaluation summary 'focuses exclusively on studies where the primary objective is to assess the creativity of the generated content, deliberately excluding those centered on accuracy, precision, or similar metrics.' Yet Table 1 classifies ProofNet (formal theorem proving), TheAgentCompany (consequential real-world tasks), Agent4Rec (recommendation), LawLuo (legal consultation), and Baby-AIGS-MLer (ML research) under creative technique categories, and Table 4 includes LawLuo and Azerbayev et al. with no creativity-oriented evaluation. The paper does not explain why these systems satisfy its own definition of creativity ('novel and valuable'). This inconsistency makes the proposed taxonomy look like a generic MAS taxonomy rather than a creativity-specific one. Either justify these inclusions with explicit creativity-relevant reasoning or exclude them and revise the claims accordingly.
  3. [Section 3 (core techniques)] The three techniques—Divergent Exploration, Iterative Refinement, and Collaborative Synthesis—are described in terms of brainstorming, feedback loops, and role-based cooperation, which are general MAS collaboration patterns that apply to any problem-solving system. The paper does not establish what is creativity-specific about these techniques, e.g., how they operationalize novelty and value, or how they differ from generic MAS coordination. To support the claim of a 'unified framework' for creativity, the authors should explicitly connect each technique to the creativity definition and to the evaluation criteria discussed in Section 5, or acknowledge that the taxonomy captures general collaboration mechanisms and position the creativity contribution in the persona and evaluation dimensions.
  4. [Appendix A and Figure 2] The proactivity classification that underlies Figure 2 is described qualitatively ('red' fully agent-driven, 'purple' low proactivity), but the paper provides no coding rubric, decision rules, or inter-rater reliability measure. For example, MaCTG is classified as fully agent-driven while Co-Scientist and CollabStory are placed at lower levels, but the criteria for these placements are not operationalized. Since the proactivity spectrum is a central contribution, the classification should be reproducible. Please provide explicit indicators for each level and, ideally, a table showing how each included system was coded.
minor comments (4)
  1. [Section 3.2] The text refers to the 'CIAR benchmark (He et al., 2020)'; the benchmark introduced in He et al. (2020) is the CIA (Commonsense Inference in machine trAnslation) benchmark, so the abbreviation should be corrected.
  2. [Section 5.1] The sentence 'creativity are inherently subjective' should read 'creativity is inherently subjective' or 'creative outputs are inherently subjective.'
  3. [Table 4] The entry 'Insightfullness' for Shaer et al. (2024) should be 'Insightfulness'.
  4. [Table 3] The table lists 'CoAGent (2023b)' but the reference list entry for Zheng et al. (2023b) has the title 'Synergizing Human-AI Agency: A Guide of 23 Heuristics for Service Co-Creation with LLM-Based Agents,' which does not obviously match the 'CoAGent' name; please reconcile the citation or add the correct reference.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper is a survey organizing external works; self-citations are illustrative examples, not load-bearing evidence.

full rationale

This paper is a survey, not a derivation. Its central outputs are a taxonomy of proactivity (Sec. 2), three technique families (Sec. 3), persona granularity (Sec. 4), evaluation methods (Sec. 5), and dataset categories (Sec. 6). These are organizing frameworks imposed on a corpus of external published systems, not quantities derived from the paper's own assumptions. There are no fitted parameters, no equations whose outputs equal their inputs, and no 'prediction' that is statistically forced by a prior fit. The self-citations that appear (e.g., Lu et al. 2024 in Table 4; Lin et al. 2024a in the Limitations section) function only as surveyed examples or as a scope caveat about multilingual settings; they are not invoked to justify the taxonomy, to forbid alternatives, or to provide the survey's structure. The paper's claim to be 'the first survey dedicated to creativity in MAS' is a scope assertion, not a derivation, and the absence of a stated systematic search protocol is a representativeness or correctness risk, not circularity. The skeptical observation that some tabulated systems (ProofNet, TheAgentCompany, Agent4Rec, LawLuo) are correctness- or utility-driven does not make the framework circular; it bears on whether the survey's inclusion matches its own creativity definition, which is an external-validity concern. No load-bearing step reduces to its own inputs, so the appropriate circularity score is 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The survey introduces no numeric parameters or new physical or conceptual entities. Its central framework rests on domain assumptions about how creativity can be defined and how creative workflows decompose; these are borrowed from cited prior work or asserted as organizing principles.

assumptions (4)
  • domain assumption Creativity is defined as production of artifacts that are both novel and valuable.
    Invoked in Section 1 via Wiggins (2006) and Veale and Cardoso (2019); the survey uses this definition to decide what counts as creative MAS.
  • domain assumption The creative workflow can be decomposed into Planning, Process, and Decision Making phases.
    Section 2.1 adopts this three-phase structure from Xie and Zou (2024) and Mukobi et al. (2023); the survey's taxonomy depends on this decomposition.
  • domain assumption Multi-agent collaboration enhances creativity beyond single-LLM generation.
    Stated in the Introduction and Section 3 without formal proof; it is the premise that motivates the entire survey.
  • domain assumption The reviewed systems are representative of the field of creative MAS.
    The paper acknowledges in Limitations that it excludes audio, video, robotics, and multilingual and Western-centric contexts; representativeness is assumed but not demonstrated via a systematic search.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Creativity in LLM-based Multi-Agent Systems: A Survey." pith.science (2026). https://pith.science/paper/YATWBUEH

@misc{pith2026250521116,
  author       = {Pith},
  title        = {Pith review of: Creativity in LLM-based Multi-Agent Systems: A Survey},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YATWBUEH}},
  note         = {Machine review of arXiv:2505.21116}
}
read the original abstract

Large language model (LLM)-driven multi-agent systems (MAS) are transforming how humans and AIs collaboratively generate ideas and artifacts. While existing surveys provide comprehensive overviews of MAS infrastructures, they largely overlook the dimension of \emph{creativity}, including how novel outputs are generated and evaluated, how creativity informs agent personas, and how creative workflows are coordinated. This is the first survey dedicated to creativity in MAS. We focus on text and image generation tasks, and present: (1) a taxonomy of agent proactivity and persona design; (2) an overview of generation techniques, including divergent exploration, iterative refinement, and collaborative synthesis, as well as relevant datasets and evaluation metrics; and (3) a discussion of key challenges, such as inconsistent evaluation standards, insufficient bias mitigation, coordination conflicts, and the lack of unified benchmarks. This survey offers a structured framework and roadmap for advancing the development, evaluation, and standardization of creative MAS.

Figures

Figures reproduced from arXiv: 2505.21116 by the authors.

Figure 1
Figure 1. Overview of multi-agent creativity systems. Given user inputs in text or image form, agents engage in a [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. MAS frameworks positioned along a two-dimensional spectrum reflecting levels of proactivity in [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Categories of Persona Granularity: A conceptual framework illustrated with selected attributes, accompanied by a concise example representing each defined persona. Coarse-Grained Persona Agents carry only high-level identity or expertise labels (e.g. “mar￾keting strategist,” “data analyst”). This minimal specification tolerates ambiguity, fostering diverse idea generation across fewer constraints. For exam￾ple, Solo… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

16 extracted references · 3 canonical work pages

  1. [5]

    Vici- nagearth, 1(1):9

    A survey on LLM-based multi-agent sys- tems: workflow, infrastructure, and challenges. Vici- nagearth, 1(1):9. Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Shuming Shi, and Zhaopeng Tu. 2024. Encouraging divergent thinking in large language models through multi-agent debate. In Proceedings of the 2024 Conference on Empi...

  2. [6]

    Preprint, arXiv:2504.14191

    AI Idea Bench 2025: AI Research Idea Gener- ation Benchmark. Preprint, arXiv:2504.14191. Marissa Radensky. 2024. Mixed-Initiative Methods for Co-Creation in Scientific Research. In Proceedings of the 16th Conference on Creativity & Cognition , C&C ’24, page 1–7. Association for Computing Ma- chinery. Nils Reimers and Iryna Gurevych. 2019. Sentence-bert: S...

  3. [7]

    Preprint, arXiv:2307.00184

    Personality Traits in Large Language Mod- els. Preprint, arXiv:2307.00184. Orit Shaer, Angelora Cooper, Osnat Mokryn, Andrew L Kun, and Hagit Ben Shoshan. 2024. Ai-augmented brainwriting: Investigating the use of llms in group ideation. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, CHI ’24, New York, NY , USA. Associatio...

  4. [8]

    arXiv preprint arXiv:2412.04629

    Argumentative Experience: Reducing Con- firmation Bias on Controversial Issues through LLM- Generated Multi-Persona Debates. arXiv preprint arXiv:2412.04629. Aakriti Singh, Shipra Saraswat, and Neetu Faujdar

  5. [10]

    Preprint, arXiv:2404.12534

    Lean copilot: Large language models as copilots for theorem proving in lean. Preprint, arXiv:2404.12534. Aarohi Srivastava et al. 2023. Beyond the imitation game: Quantifying and extrapolating the capabili- ties of language models. Transactions on Machine Learning Research. Haoyang Su, Renqi Chen, Shixiang Tang, Zhenfei Yin, Xinzhe Zheng, Jinzhe Li, Biqin...

  6. [11]

    It Felt Like Hav- ing a Second Mind

    CREA: A Collaborative Multi-Agent Frame- work for Creative Content Generation with Diffusion Models. Preprint, arXiv:2504.05306. Saranya Venkatraman, Nafis Irtiza Tripto, and Dongwon Lee. 2024. Collabstory: Multi-llm collaborative story generation and authorship analysis. In Proc. NAACL ’25. M. A. Wallach and N. Kogan. 1965. Modes of Think- ing in Young C...

  7. [12]

    Bingyu Yan, Xiaoming Zhang, Litian Zhang, Lian Zhang, Ziyi Zhou, Dezhuang Miao, and Chaozhuo Li

    TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks.Preprint, arXiv:2412.14161. Bingyu Yan, Xiaoming Zhang, Litian Zhang, Lian Zhang, Ziyi Zhou, Dezhuang Miao, and Chaozhuo Li

  8. [13]

    Preprint, arXiv:2502.14321

    Beyond self-talk: A communication-centric survey of llm-based multi-agent systems. Preprint, arXiv:2502.14321. Kaixun Yang, Yixin Cheng, Linxuan Zhao, Mladen Rakovi´c, Zachari Swiecki, Dragan Gaševi ´c, and Guanliang Chen. 2024. Ink and algorithm: Exploring temporal dynamics in human-ai collaborative writing. Preprint, arXiv:2406.14885. 18 Lyumanshan Ye, ...

Show all 16 references
  1. [14]

    arXiv preprint arXiv:2504.12735

    The Athenian Academy: A Seven-Layer Ar- chitecture Model for Multi-Agent Systems. arXiv preprint arXiv:2504.12735. An Zhang, Yuxin Chen, Leheng Sheng, Xiang Wang, and Tat-Seng Chua. 2024a. On generative agents in recommendation. In Proceedings of the 47th Inter- national ACM S...

  2. [15]

    arXiv preprint arXiv:2410.19245

    Mactg: Multi-agent collaborative thought graph for automatic programming. arXiv preprint arXiv:2410.19245. Lianmin Zheng et al. 2023a. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. In Advances in Neural Information Processing Systems. Qingxiao Zheng, Zhongwei Xu, Abh...

  3. [16]

    Colin (Ye et al., 2024) exhibits the agent proactivity through a different way

    not only receives iterative user requests dur- ing refinement but also incorporates environmental data collected from its sensors, such as weather con- ditions, camera input, and audio input. Colin (Ye et al., 2024) exhibits the agent proactivity through a different way. The s...

  4. [2010]

    Ma- chine Bias

    Conflict resolution in multi-agent systems based on negotiation and arbitrage. In 2010 2nd IEEE International Conference on Information Man- agement and Engineering, pages 304–307. Junseok Kim, Nakyeong Yang, and Kyomin Jung. 2024. Persona is a double-edged sword: Mitigating t...

  5. [2017]

    In 2017 International Confer- ence on Computing, Communication and Automation (ICCCA)

    Analyzing Titanic disaster using machine learning algorithms. In 2017 International Confer- ence on Computing, Communication and Automation (ICCCA). Peiyang Song, Kaiyu Yang, and Anima Anandkumar

  6. [2021]

    Preprint, arXiv:2102.08317

    Resource allocation in dynamic multiagent systems. Preprint, arXiv:2102.08317. Mike D’Arcy, Tom Hope, Larry Birnbaum, and Doug Downey. 2024. Marg: Multi-agent review generation for scientific papers. Preprint, arXiv:2401.04259. Allegra De Filippo, Michela Milano, et al. 2024. ...

  7. [2024]

    Naomi Imasato, Kazuki Miyazawa, Takayuki Nagai, and Takato Horii

    A Collaborative, Interactive and Context- Aware Drawing Agent for Co-Creative Design.IEEE Transactions on Visualization and Computer Graph- ics. Naomi Imasato, Kazuki Miyazawa, Takayuki Nagai, and Takato Horii. 2024. Creative agents: Simulat- ing the systems model of creativit...

  8. [2025]

    IEEE Access, 13:10499–10512

    An Innovative Solution to Design Problems: Applying the Chain-of-Thought Technique to Inte- grate LLM-Based Agents With Concept Generation Methods. IEEE Access, 13:10499–10512. Juraj Gottweis et al. 2025. Towards an AI co-scientist. Preprint, arXiv:2502.18864. Aaron Grattafior...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.