REVIEW 2 major objections 5 minor 1 cited by
TeamCraft: A Benchmark for Multi-Modal Multi-Agent Systems in Minecraft
T0 review · 2 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read TeamCraft introduces a 55,000-task Minecraft benchmark and shows that current vision-language-action models struggle on unseen goals, scenes, and larger teams.
desk verdict A genuinely useful benchmark for multi-modal multi-agent Minecraft tasks, but the 'models struggle' conclusion should be read with a grain of salt until the planner-generated demonstrations are independently validated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the multi-modal prompt: a language instruction interleaved with three orthographic view images (top, left, front) that specify goal or initial states. Demonstrations are generated by a planner that uses privileged environment information to minimize a cost function $C = w_1T + w_2\sum_i E_i + w_3D + w_4\sum_i\sum_{j\in A_i} c_{ij} + w_5U$, where weights are hand-tuned per task family. The TeamCraft-VLA model combines a vision encoder, a projector, and a language model to map these prompts, per-agent first-person views, and inventories to high-level skills. Generalization splits (novel shape/material, novel crop, novel goal, novel scene, four agents) provide the evaluation machinery that makes the paper's difficulty claims measurable.
What would settle it
A concrete falsifier would be to retrain TeamCraft-VLA on demonstrations collected from human players or from a planner using different cost weights, and evaluate on the same generalization splits; if novel-goal or four-agent success rates rise substantially, the reported difficulty is an artifact of the expert-demonstration policy rather than an inherent property of multi-modal multi-agent tasks. A second falsifier: give the model the ground-truth goal text instead of the orthographic views on the novel-crop farming split; if success jumps from 0.00 to high values, the bottleneck is visual grounding of the goal rather than multi-agent coordination.
Extended reading notes
Core claim
The core claim is that multi-modal multi-agent generalization can be benchmarked in a visually rich 3D world, and that current models do not generalize. The benchmark covers four task families—building, clearing, farming, and smelting—with more than 55,000 procedurally generated task variants. Agents receive first-person RGB images and inventory information, and emit high-level actions; training uses planner-generated expert demonstrations that minimize a weighted cost over completion time, idle actions, dependencies, action costs, and redundancy. Centralized 7B and 13B vision-language-action models reach task success rates as high as 0.64 on in-distribution clearing tasks, but fall to near zero on novel goals and four-agent settings; the novel-crop farming split is 0.00 for every model tested. The proprietary multimodal model scores near zero everywhere, with failures concentrated in 3D spatial reasoning and object-state recognition.
Load-bearing premise
The load-bearing premise is that the planner-generated demonstrations, produced under a hand-tuned cost function with privileged environment information, are a good enough proxy for expert collaborative behavior that imitation learning on them measures multi-modal multi-agent generalization rather than reproducing one particular planning policy.
Editorial extensions
If this is right
- If TeamCraft's results are correct, current multimodal vision-language-action models are not ready for zero-shot or few-shot multi-agent planning in 3D worlds; they need explicit mechanisms for unseen goals and for coordinating teams of unseen size.
- The uniform 0.00 on novel farming crops indicates that object-state recognition, not just spatial reasoning, is a bottleneck for embodied multi-agent generalization.
- Centralized control outperforms decentralized control and produces far fewer redundant actions, suggesting that models without inter-agent communication or joint inference degrade as team size grows.
- The grid-world ablation, where the same tasks are described in text instead of images, achieves substantially higher success rates, implying that a large part of the failure is visual grounding rather than high-level task planning alone.
- Scaling training data helps in-distribution and scene splits but does not repair novel-goal or novel-agent generalization, and larger models do not automatically close that gap.
Reading between the lines
- A testable extension would be to add explicit communication channels to decentralized TeamCraft agents; comparing redundancy rates and task success with the current implicit-communication setting would show how much of the gap is due to missing coordination.
- Because demonstrations come from a planner with hand-set cost weights, human-collected or differently weighted demonstrations on the same tasks could reveal whether the observed generalization failures are inherent to multimodal multi-agent learning or an artifact of one particular expert policy.
- The scene splits change textures, lighting, and biome layouts together, so part of the reported scene generalization may be appearance robustness rather than task-level generalization; disentangling these factors would sharpen the benchmark's diagnostic value.
- The 0.00 novel-crop result offers a sharp probe: a model that can read the crop name from the orthographic blueprint and map it to its inventory should at least sow the correct seed, so failure on that single step is a strong test of grounded multimodal understanding.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript presents TeamCraft, a benchmark for multi-modal multi-agent collaboration in Minecraft. It contains 55,257 procedurally generated demonstrations across four task types (Building, Clearing, Farming, Smelting), specified by multi-modal prompts that combine language instructions with orthographic-view images. The paper introduces a planner that uses privileged information to generate 'expert' demonstrations, then evaluates several baselines: TeamCraft-VLA (7B and 13B, centralized and decentralized), GPT-4o (one-shot), and a text-based Grid-World ablation. The evaluation protocol defines task success rate, subgoal success rate, and redundancy rate. Results show that all models achieve low success rates on the test set and on generalization splits (novel goals, novel scenes, and four-agent teams), with Farming Crop generalization at 0.00 for every model. The authors conclude that current multi-modal multi-agent models face significant challenges in generalization.
Significance. TeamCraft addresses a real gap in the literature: few benchmarks combine multi-modal task specification, vision-based multi-agent control, and diverse generalization splits in a visually rich 3D environment. The released dataset, code, and checkpoints are concrete contributions, and the dataset statistics in Table 2 sum consistently (55,257 demonstrations), with a datasheet reporting 57,207 total instances including 1,000 validation and 950 test cases. The evaluation metrics are clearly defined. If the demonstrations are validated, the benchmark could become a valuable testbed for studying multi-modal multi-agent generalization. However, the central conclusion that 'existing models face significant challenges' depends on the quality and representativeness of the planner-generated demonstrations, and the current evidence for that quality is thin.
major comments (2)
- [Section 3.6, Eq. (1), Appendix D] The demonstration generation algorithm is underspecified. The cost function C = w1T + w2*sum(E_i) + w3D + w4*sum(c_ij) + w5U is given, but the paper never explains how the planner optimizes this cost (e.g., exact search, greedy assignment, or heuristic scheduling) or how ties are broken. This is load-bearing for two reasons. First, without this algorithmic detail the benchmark cannot be reproduced, even with released code, because the design decisions are not documented. Second, all VLA baselines are trained solely on these trajectories, which are labeled 'expert demonstrations,' yet the only quality evidence is the datasheet statement that the data were 'verified by the team via manual inspection.' There is no quantitative validation that the trajectories are near-optimal, human-comparable, or even consistent with the partial observations available to agents at test time. The observed poor generalization, especially the 0.00 Farming Crop result for all models, could therefore reflect a mismatch between the planner's privileged-information policy and the first-person observations, rather than intrinsic task difficulty. Please provide the optimization algorithm, and add at least one quantitative validation, such as comparing demonstration length against a theoretical lower bound, training a privileged-information oracle on the demonstrations, or reporting a human-collected demonstration baseline.
- [Section 4.3, Tables 6-8] The test and generalization splits contain only 50 cases per condition, which makes the point estimates very imprecise for binary success metrics. For example, a reported success rate of 0.00 has a 95% confidence interval upper bound of about 0.07, and differences such as 0.02 vs 0.04 (Building Agents, 7B vs 13B) are within sampling noise. The paper draws conclusions like 'scaling up model sizes blindly does not guarantee success' from such comparisons without reporting confidence intervals or significance tests. To make the benchmark results more trustworthy, please report bootstrap confidence intervals or a significance test for the headline results in Figure 5 and Table 8.
minor comments (5)
- [Section 3.5] The phrase 'More than 30 target object or resource are used' should be 'More than 30 target objects or resources are used.'
- [Section 3.7, Appendix K] The validation set (1,000 instances) appears only in the datasheet; state its split in the main text for completeness, alongside the training (55,257) and test (950) sizes.
- [Appendix K] The datasheet says 'The dataset contain all possible instances'; the verb should be 'contains.'
- [Figure 2] The caption contains a stray token 'Object 232233' that should be removed.
- [Figure 5] The figure is dense with multiple panels and lines; consider adding a grayscale-readable legend and ensuring y-axis labels are visible in every panel.
Circularity Check
No significant circularity: the benchmark's claims are empirical measurements from held-out evaluation, not derivations from the construction.
full rationale
TeamCraft is an empirical benchmark paper. The central claims — that multimodal multi-agent models have low success rates and generalize poorly to novel goals, scenes, and agent counts — are experimental results obtained by training TeamCraft-VLA on the released demonstrations and evaluating on deliberately held-out splits, and by prompting GPT-4o in a one-shot setting. No equation or protocol in the paper defines the reported success rates in terms of the demonstration-generation cost function or the benchmark's construction; the planner in Section 3.6 (Eq. 1) generates training labels, not test predictions, and its hand-set weights are stated rather than fitted to the evaluation outcome. The farming 'Crop' 0.00 result is an empirical generalization-failure measurement on a hold-out crop, not a tautology: the models were trained on other crops and the benchmark measures whether they transfer. The paper's self-citations (e.g., [13]-[17], [31]-[33]) appear only as related-work context and architectural inspiration, not as load-bearing justification for the difficulty conclusions. The concern that planner-generated 'expert' demonstrations may be biased or not human-comparable is a validity/fairness limitation, correctly confined to Appendix D and the datasheet's manual-inspection note; it does not make the evaluation circular. Overall, the derivation chain is self-contained and empirically grounded.
Assumptions & free parameters
free parameters (1)
- Planner cost weights w1-w5 =
Building: w2=1.4, others=0.9; Clearing: w4=1.8, others=0.8; Farming: all 1.0; Smelting: w3=1.8, others=0.8
assumptions (3)
- domain assumption Planner-generated demonstrations are expert behavior that teach robust collaboration.
- domain assumption Rejection sampling guarantees all tasks are solvable under Minecraft world rules.
- domain assumption Minecraft is a valid testbed for real-world multi-agent embodied collaboration.
Cite this review
Pith. "Pith review of TeamCraft: A Benchmark for Multi-Modal Multi-Agent Systems in Minecraft." pith.science (2026). https://pith.science/paper/MUONEPAB
@misc{pith2026241205255,
author = {Pith},
title = {Pith review of: TeamCraft: A Benchmark for Multi-Modal Multi-Agent Systems in Minecraft},
year = {2026},
howpublished = {\url{https://pith.science/paper/MUONEPAB}},
note = {Machine review of arXiv:2412.05255}
}
read the original abstract
Collaboration is a cornerstone of society. In the real world, human teammates make use of multi-sensory data to tackle challenging tasks in ever-changing environments. It is essential for embodied agents collaborating in visually-rich environments replete with dynamic interactions to understand multi-modal observations and task specifications. To evaluate the performance of generalizable multi-modal collaborative agents, we present TeamCraft, a multi-modal multi-agent benchmark built on top of the open-world video game Minecraft. The benchmark features 55,000 task variants specified by multi-modal prompts, procedurally-generated expert demonstrations for imitation learning, and carefully designed protocols to evaluate model generalization capabilities. We also perform extensive analyses to better understand the limitations and strengths of existing approaches. Our results indicate that existing models continue to face significant challenges in generalizing to novel goals, scenes, and unseen numbers of agents. These findings underscore the need for further research in this area. The TeamCraft platform and dataset are publicly available at https://github.com/teamcraft-bench/teamcraft.
Figures
Figures from the paper (39 more)
Forward citations
Cited by 1 Pith paper
-
Large Language Models for Planning: A Comprehensive and Systematic Survey
A structured survey of LLM planning methods, benchmarks, and interpretability work, organized around a three-way taxonomy.
Reference graph
Works this paper leans on
-
[1]
A Minecraft based sim- ulated task environment for human AI teaming,
A. Amresh, N. Cooke, and A. Fouse, “A Minecraft based sim- ulated task environment for human AI teaming,” in Proceed- ings of the 23rd ACM International Conference on Intelligent Virtual Agents, 2023, pp. 1–3. 9
work page 2023
-
[2]
P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. Sün- derhauf, I. Reid, S. Gould, and A. Van Den Hengel, “Vision- and-language navigation: Interpreting visually-grounded nav- igation instructions in real environments,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recog- nition, 2018, pp. 3674–3683. 2
work page 2018
-
[3]
Video Pre- Training (VPT): Learning to act by watching unlabeled online videos,
B. Baker, I. Akkaya, P. Zhokov, J. Huizinga, J. Tang, A. Ecof- fet, B. Houghton, R. Sampedro, and J. Clune, “Video Pre- Training (VPT): Learning to act by watching unlabeled online videos,” Advances in Neural Information Processing Systems, vol. 35, pp. 24 639–24 654, 2022. 9
work page 2022
-
[4]
ROCKET-1: Master open-world interaction with visual-temporal context prompting,
S. Cai, Z. Wang, K. Lian, Z. Mu, X. Ma, A. Liu, and Y . Liang, “ROCKET-1: Master open-world interaction with visual-temporal context prompting,” arXiv preprint arXiv:2410.17856, 2024. 1
arXiv 2024
-
[5]
On the utility of learning about humans for human-AI coordination,
M. Carroll, R. Shah, M. K. Ho, T. L. Griffiths, S. A. Seshia, P. Abbeel, and A. Dragan, “On the utility of learning about humans for human-AI coordination,” 2020. 2, 9
work page 2020
-
[6]
PARTNR: A benchmark for planning and reasoning in embodied multi-agent tasks,
M. Chang, G. Chhablani, A. Clegg, M. D. Cote, R. De- sai, M. Hlavac, V . Karashchuk, J. Krantz, R. Mottaghi, P. Parashar et al., “PARTNR: A benchmark for planning and reasoning in embodied multi-agent tasks,” arXiv preprint arXiv:2411.00081, 2024. 2
arXiv 2024
-
[7]
B. Chen, S. Song, H. Lipson, and C. V ondrick, “Visual hide and seek,” inArtificial Life Conference Proceedings 32. MIT Press One Rogers Street, Cambridge, MA 02142-1209, USA journals-info . . . , 2020, pp. 645–655. 1
work page 2020
-
[8]
A. Das, S. Datta, G. Gkioxari, S. Lee, D. Parikh, and D. Batra, “Embodied question answering,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 1–10. 2
work page 2018
Show all 67 references
-
[9]
TarMAC: Targeted multi-agent communi- cation,
A. Das, T. Gervet, J. Romoff, D. Batra, D. Parikh, M. Rabbat, and J. Pineau, “TarMAC: Targeted multi-agent communi- cation,” in Proceedings of the International Conference on Machine Learning. PMLR, 2019, pp. 1538–1546. 1
2019
-
[10]
VillagerAgent: A Graph-Based Multi-Agent Framework for Coordinating Complex Task Dependencies in Minecraft,
Y . Dong, X. Zhu, Z. Pan, L. Zhu, and Y . Yang, “VillagerAgent: A Graph-Based Multi-Agent Framework for Coordinating Complex Task Dependencies in Minecraft,” in Findings of the Association for Computational Linguistics: ACL 2024 , L.-W. Ku, A. Martins, and V . Srikumar, Eds. B...
2024
-
[11]
Mine- Dojo: Building open-ended embodied agents with internet- scale knowledge,
L. Fan, G. Wang, Y . Jiang, A. Mandlekar, Y . Yang, H. Zhu, A. Tang, D.-A. Huang, Y . Zhu, and A. Anandkumar, “Mine- Dojo: Building open-ended embodied agents with internet- scale knowledge,” 2022. 2, 9
2022
-
[12]
Alexa arena: A user-centric interactive platform for embodied AI,
Q. Gao, G. Thattai, X. Gao, S. Shakiah, S. Pansare, V . Sharma, G. Sukhatme, H. Shi, B. Yang, D. Zheng et al., “Alexa arena: A user-centric interactive platform for embodied AI,”arXiv preprint arXiv:2303.01586, 2023. 2
2023 arXiv
-
[13]
DialFRED: Dialogue-enabled agents for embod- ied instruction following,
X. Gao, Q. Gao, R. Gong, K. Lin, G. Thattai, and G. S. Sukhatme, “DialFRED: Dialogue-enabled agents for embod- ied instruction following,” IEEE Robotics and Automation Letters, vol. 7, no. 4, pp. 10 049–10 056, 2022. 2
2022
-
[14]
Joint mind modeling for explanation generation in complex human-robot collaborative tasks,
X. Gao, R. Gong, Y . Zhao, S. Wang, T. Shu, and S.-C. Zhu, “Joint mind modeling for explanation generation in complex human-robot collaborative tasks,” in Proceedings of the 29th IEEE International Conference on Robot and Human Interac- tive Communication (RO-MAN). IEEE, 2020,...
2020
-
[15]
LEMMA: Learning language-conditioned multi- robot manipulation,
R. Gong, X. Gao, Q. Gao, S. Shakiah, G. Thattai, and G. S. Sukhatme, “LEMMA: Learning language-conditioned multi- robot manipulation,” IEEE Robotics and Automation Letters,
-
[16]
ARNOLD: A benchmark for language-grounded task learn- ing with continuous states in realistic 3D scenes,
R. Gong, J. Huang, Y . Zhao, H. Geng, X. Gao, Q. Wu, W. Ai, Z. Zhou, D. Terzopoulos, S.-C. Zhu, B. Jia, and S. Huang, “ARNOLD: A benchmark for language-grounded task learn- ing with continuous states in realistic 3D scenes,” in Proceed- ings of the IEEE/CVF International Confe...
2023
-
[17]
MindAgent: Emergent gaming interaction,
R. Gong, Q. Huang, X. Ma, H. V o, Z. Durante, Y . Noda, Z. Zheng, S.-C. Zhu, D. Terzopoulos, L. Fei-Fei, and J. Gao, “MindAgent: Emergent gaming interaction,” 2023. 1, 2
2023
-
[18]
IQA: Visual question answering in interac- tive environments,
D. Gordon, A. Kembhavi, M. Rastegari, J. Redmon, D. Fox, and A. Farhadi, “IQA: Visual question answering in interac- tive environments,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 4089–
2018
-
[19]
MineRL: A large-scale dataset of Minecraft demonstrations,
W. H. Guss, B. Houghton, N. Topin, P. Wang, C. Codel, M. Veloso, and R. Salakhutdinov, “MineRL: A large-scale dataset of Minecraft demonstrations,” 2019. 2
2019
-
[20]
Human-level performance in 3D multiplayer games with population-based reinforcement learning,
M. Jaderberg, W. M. Czarnecki, I. Dunning, L. Marris, G. Lever, A. G. Castaneda, C. Beattie, N. C. Rabinowitz, A. S. Morcos, A. Ruderman et al., “Human-level performance in 3D multiplayer games with population-based reinforcement learning,” Science, vol. 364, no. 6443, pp. 859...
2019
-
[21]
A cordial sync: Going beyond marginal policies for multi-agent embodied tasks,
U. Jain, L. Weihs, E. Kolve, A. Farhadi, S. Lazebnik, A. Kem- bhavi, and A. Schwing, “A cordial sync: Going beyond marginal policies for multi-agent embodied tasks,” 2020. 1, 2
2020
-
[22]
Two body prob- lem: Collaborative visual task completion,
U. Jain, L. Weihs, E. Kolve, M. Rastegari, S. Lazebnik, A. Farhadi, A. Schwing, and A. Kembhavi, “Two body prob- lem: Collaborative visual task completion,” 2019. 1, 2, 9
2019
-
[23]
Learn- ing to execute instructions in a Minecraft dialogue,
P. Jayannavar, A. Narayan-Chen, and J. Hockenmaier, “Learn- ing to execute instructions in a Minecraft dialogue,” in Pro- ceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020, pp. 2589–2602. 2, 9
2020
-
[24]
VIMA: General robot manipulation with multimodal prompts,
Y . Jiang, A. Gupta, Z. Zhang, G. Wang, Y . Dou, Y . Chen, L. Fei-Fei, A. Anandkumar, Y . Zhu, and L. Fan, “VIMA: General robot manipulation with multimodal prompts,” arXiv preprint arXiv:2210.03094, vol. 2, no. 3, p. 6, 2022. 1, 2
-
[25]
The malmo platform for artificial intelligence experimentation,
M. Johnson, K. Hofmann, T. Hutton, and D. Bignell, “The malmo platform for artificial intelligence experimentation,” in Ijcai, 2016, pp. 4246–4247. 2
2016
-
[26]
Scalable evaluation of multi-agent reinforcement learning with melting pot,
J. Z. Leibo, E. Duéñez-Guzmán, A. S. Vezhnevets, J. P. Aga- piou, P. Sunehag, R. Koster, J. Matyas, C. Beattie, I. Mor- datch, and T. Graepel, “Scalable evaluation of multi-agent reinforcement learning with melting pot,” 2021. 1, 2 48
2021
-
[27]
CAMEL: Communicative agents for “mind
G. Li, H. Hammoud, H. Itani, D. Khizbullin, and B. Ghanem, “CAMEL: Communicative agents for “mind” exploration of large language model society,”Advances in Neural Informa- tion Processing Systems, vol. 36, pp. 51 991–52 008, 2023. 2
2023
-
[28]
Visual instruction tuning,
H. Liu, C. Li, Q. Wu, and Y . J. Lee, “Visual instruction tuning,” Advances in neural information processing systems, vol. 36,
-
[29]
Multi-agent embodied visual semantic navigation with scene prior knowledge,
X. Liu, D. Guo, H. Liu, and F. Sun, “Multi-agent embodied visual semantic navigation with scene prior knowledge,”IEEE Robotics and Automation Letters, vol. 7, no. 2, pp. 3154–3161,
-
[30]
Embodied multi-agent task planning from ambiguous instruction,
X. Liu, X. Li, D. Guo, S. Tan, H. Liu, and F. Sun, “Embodied multi-agent task planning from ambiguous instruction,” Pro- ceedings of Robotics: Science and Systems, New York City, NY, USA, pp. 1–14, 2022. 1, 2
2022
-
[31]
Inverse attention agent for multi-agent system,
Q. Long, R. Li, M. Zhao, T. Gao, and D. Terzopoulos, “Inverse attention agent for multi-agent system,” 2024. [Online]. Available: https://arxiv.org/abs/2410.21794 1
2024 arXiv
-
[32]
So- cialGFs: Learning social gradient fields for multi-agent rein- forcement learning,
Q. Long, F. Zhong, M. Wu, Y . Wang, and S.-C. Zhu, “So- cialGFs: Learning social gradient fields for multi-agent rein- forcement learning,” arXiv preprint arXiv:2405.01839, 2024. 2
2024 arXiv
-
[33]
Evolutionary population curriculum for scaling multi-agent reinforcement learning,
Q. Long, Z. Zhou, A. Gupta, F. Fang, Y . Wu, and X. Wang, “Evolutionary population curriculum for scaling multi-agent reinforcement learning,” arXiv preprint arXiv:2003.10423, 2020
2003 arXiv
-
[34]
Multi-agent actor-critic for mixed cooperative- competitive environments,
R. Lowe, Y . Wu, A. Tamar, J. Harb, P. Abbeel, and I. Mordatch, “Multi-agent actor-critic for mixed cooperative- competitive environments,” 2020. 2
2020
-
[35]
SQA3D: Situated question answering in 3D scenes,
X. Ma, S. Yong, Z. Zheng, Q. Li, Y . Liang, S.-C. Zhu, and S. Huang, “SQA3D: Situated question answering in 3D scenes,” 2023. 2
2023
-
[36]
OpenEQA: Embodied question answering in the era of foundation models,
A. Majumdar, A. Ajay, X. Zhang, P. Putta, S. Yenamandra, M. Henaff, S. Silwal, P. Mcvay, O. Maksymets, S. Arnaud et al., “OpenEQA: Embodied question answering in the era of foundation models,” in Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recogniti...
2024
-
[37]
RoCo: Dialectic multi-robot collaboration with large language models,
Z. Mandi, S. Jain, and S. Song, “RoCo: Dialectic multi-robot collaboration with large language models,” in Proceedings of the IEEE International Conference on Robotics and Automa- tion (ICRA). IEEE, 2024, pp. 286–299. 1, 2, 9
2024
-
[38]
CALVIN: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks,
O. Mees, L. Hermann, E. Rosete-Beas, and W. Burgard, “CALVIN: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks,” IEEE Robotics and Automation Letters, vol. 7, no. 3, pp. 7327–7334,
-
[39]
Emergence of grounded compo- sitional language in multi-agent populations,
I. Mordatch and P. Abbeel, “Emergence of grounded compo- sitional language in multi-agent populations,” arXiv preprint arXiv:1703.04908, 2017. 2
2017 arXiv
-
[40]
Col- laborative dialogue in Minecraft,
A. Narayan-Chen, P. Jayannavar, and J. Hockenmaier, “Col- laborative dialogue in Minecraft,” in Proceedings of the 57th Annual Meeting of the Association for Computational Lin- guistics, 2019, pp. 5405–5415. 2, 9
2019
-
[41]
TEACh: Task-driven embodied agents that chat,
A. Padmakumar, J. Thomason, A. Shrivastava, P. Lange, A. Narayan-Chen, S. Gella, R. Piramuthu, G. Tur, and D. Hakkani-Tur, “TEACh: Task-driven embodied agents that chat,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 2, 2022, pp. 2017–2025. 2
2022
-
[42]
Generative agents: Interactive simulacra of human behavior,
J. S. Park, J. O’Brien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein, “Generative agents: Interactive simulacra of human behavior,” in Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, 2023, pp. 1–22. 1, 2
2023
-
[43]
The multi- agent reinforcement learning in MalmÖ (MARLÖ) competi- tion,
D. Perez-Liebana, K. Hofmann, S. P. Mohanty, N. Kuno, A. Kramer, S. Devlin, R. D. Gaina, and D. Ionita, “The multi- agent reinforcement learning in MalmÖ (MARLÖ) competi- tion,” arXiv preprint arXiv:1901.08129, 2019. 1, 2
1901 arXiv
-
[44]
Watch-And-Help: A challenge for social perception and human-ai collaboration,
X. Puig, T. Shu, S. Li, Z. Wang, Y .-H. Liao, J. B. Tenenbaum, S. Fidler, and A. Torralba, “Watch-And-Help: A challenge for social perception and human-ai collaboration,” 2021. 1, 2
2021
-
[45]
NOPA: Neurally-guided online probabilistic assistance for building socially intelligent home assistants,
X. Puig, T. Shu, J. B. Tenenbaum, and A. Torralba, “NOPA: Neurally-guided online probabilistic assistance for building socially intelligent home assistants,” in Proceedings of the IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2023, pp. 7628–7634. 2
2023
-
[46]
CLIPort: What and where pathways for robotic manipulation,
M. Shridhar, L. Manuelli, and D. Fox, “CLIPort: What and where pathways for robotic manipulation,” in Proceedings of the 5th Conference on Robot Learning (CoRL), 2021. 2
2021
-
[47]
ALFRED: A benchmark for interpreting grounded instructions for every- day tasks,
M. Shridhar, J. Thomason, D. Gordon, Y . Bisk, W. Han, R. Mottaghi, L. Zettlemoyer, and D. Fox, “ALFRED: A benchmark for interpreting grounded instructions for every- day tasks,” 2020. 2
2020
-
[48]
ALFWorld: Aligning text and embod- ied environments for interactive learning,
M. Shridhar, X. Yuan, M.-A. Côté, Y . Bisk, A. Trischler, and M. Hausknecht, “ALFWorld: Aligning text and embod- ied environments for interactive learning,” arXiv preprint arXiv:2010.03768, 2020. 2
2010 arXiv
-
[49]
The development of embodied cognition: Six lessons from babies,
L. Smith and M. Gasser, “The development of embodied cognition: Six lessons from babies,” Artificial life, vol. 11, no. 1-2, pp. 13–29, 2005. 1
2005
-
[50]
Multiagent systems: A survey from a machine learning perspective,
P. Stone and M. Veloso, “Multiagent systems: A survey from a machine learning perspective,” Autonomous Robots, vol. 8, pp. 345–383, 2000. 1
2000
-
[51]
Neural MMO 2.0: A massively multi-task addition to massively multi-agent learning,
J. Suárez, D. Bloomin, K. W. Choe, H. X. Li, R. Sullivan, N. Kanna, D. Scott, R. Shuman, H. Bradley, L. Castricato et al., “Neural MMO 2.0: A massively multi-task addition to massively multi-agent learning,” Advances in Neural Infor- mation Processing Systems, vol. 36, 2024. 2
2024
-
[52]
The neural MMO platform for massively multiagent research,
J. Suarez, Y . Du, C. Zhu, I. Mordatch, and P. Isola, “The neural MMO platform for massively multiagent research,” in Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, J. Vanschoren and S. Ye- ung, Eds., vol. 1, 2021. 1, 2
2021
-
[53]
Habitat 2.0: Training home assistants to rearrange their habitat,
A. Szot, A. Clegg, E. Undersander, E. Wijmans, Y . Zhao, J. Turner, N. Maestre, M. Mukadam, D. S. Chaplot, O. Maksymets et al., “Habitat 2.0: Training home assistants to rearrange their habitat,” Advances in neural information processing systems, vol. 34, pp. 251–266, 2021. 2
2021
-
[54]
Multi-agent embodied question answering in interactive environments,
S. Tan, W. Xiang, H. Liu, D. Guo, and F. Sun, “Multi-agent embodied question answering in interactive environments,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XIII 16. Springer, 2020, pp. 663–678. 2 49
2020
-
[55]
Grandmaster level in StarCraft II using multi-agent reinforcement learning,
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev et al., “Grandmaster level in StarCraft II using multi-agent reinforcement learning,” Nature, vol. 575, no. 7782, pp. 350–354, 2019. 2
2019
-
[56]
HandMeThat: Human- robot communication in physical and social environments,
Y . Wan, J. Mao, and J. Tenenbaum, “HandMeThat: Human- robot communication in physical and social environments,” Advances in Neural Information Processing Systems, vol. 35, pp. 12 014–12 026, 2022. 2
2022
-
[57]
V oyager: An open-ended em- bodied agent with large language models,
G. Wang, Y . Xie, Y . Jiang, A. Mandlekar, C. Xiao, Y . Zhu, L. Fan, and A. Anandkumar, “V oyager: An open-ended em- bodied agent with large language models,” 2023. 1, 2
2023
-
[58]
Jarvis-1: Open-world multi-task agents with memory-augmented mul- timodal language models,
Z. Wang, S. Cai, A. Liu, Y . Jin, J. Hou, B. Zhang, H. Lin, Z. He, Z. Zheng, Y . Yang, X. Ma, and Y . Liang, “Jarvis-1: Open-world multi-task agents with memory-augmented mul- timodal language models,” arXiv preprint arXiv:2311.05997,
-
[59]
Describe, explain, plan and select: Interactive planning with large lan- guage models enables open-world multi-task agents,
Z. Wang, S. Cai, A. Liu, X. Ma, and Y . Liang, “Describe, explain, plan and select: Interactive planning with large lan- guage models enables open-world multi-task agents,” 2023. 9
2023
-
[60]
Too many cooks: Bayesian inference for coordinating multi-agent collaboration,
S. A. Wu, R. E. Wang, J. A. Evans, J. B. Tenenbaum, D. C. Parkes, and M. Kleiman-Weiner, “Too many cooks: Bayesian inference for coordinating multi-agent collaboration,” Topics in Cognitive Science, vol. 13, no. 2, pp. 414–432, 2021. 1
2021
-
[61]
The surprising effectiveness of PPO in cooperative multi-agent games,
C. Yu, A. Velu, E. Vinitsky, J. Gao, Y . Wang, A. Bayen, and Y . Wu, “The surprising effectiveness of PPO in cooperative multi-agent games,” 2021. 2
2021
-
[62]
MineLand: Simulating large-scale multi-agent interactions with limited multimodal senses and physical needs,
X. Yu, J. Fu, R. Deng, and W. Han, “MineLand: Simulating large-scale multi-agent interactions with limited multimodal senses and physical needs,” 2024. 1
2024
-
[63]
Transporter networks: Rearranging the visual world for robotic manipulation,
A. Zeng, P. Florence, J. Tompson, S. Welker, J. Chien, M. At- tarian, T. Armstrong, I. Krasin, D. Duong, V . Sindhwani, and J. Lee, “Transporter networks: Rearranging the visual world for robotic manipulation,” Conference on Robot Learning (CoRL), 2020. 2
2020
-
[64]
ProAgent: Building proactive cooper- ative agents with large language models,
C. Zhang, K. Yang, S. Hu, Z. Wang, G. Li, Y . Sun, C. Zhang, Z. Zhang, A. Liu, S.-C. Zhu, X. Chang, J. Zhang, F. Yin, Y . Liang, and Y . Yang, “ProAgent: Building proactive cooper- ative agents with large language models,” inProceedings of the AAAI Conference on Artificial Int...
2024
-
[65]
Building cooperative embodied agents modularly with large language models,
H. Zhang, W. Du, J. Shan, Q. Zhou, Y . Du, J. B. Tenenbaum, T. Shu, and C. Gan, “Building cooperative embodied agents modularly with large language models,” in Proceedings of the Twelfth International Conference on Learning Representations , 2024. [Online]. Available: https://...
2024
-
[66]
COMBO: Compositional world mod- els for embodied multi-agent cooperation,
H. Zhang, Z. Wang, Q. Lyu, Z. Zhang, S. Chen, T. Shu, Y . Du, and C. Gan, “COMBO: Compositional world mod- els for embodied multi-agent cooperation,” arXiv preprint arXiv:2404.10775, 2024. 1
2024 arXiv
-
[67]
VLMbench: A compositional benchmark for vision-and-language manipu- lation,
K. Zheng, X. Chen, O. C. Jenkins, and X. Wang, “VLMbench: A compositional benchmark for vision-and-language manipu- lation,” Advances in Neural Information Processing Systems, vol. 35, pp. 665–678, 2022. 2 50
2022
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.