REVIEW 7 cited by
Watch-And-Help: A Challenge for Social Perception and Human-AI Collaboration
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In this paper, we introduce Watch-And-Help (WAH), a challenge for testing social intelligence in agents. In WAH, an AI agent needs to help a human-like agent perform a complex household task efficiently. To succeed, the AI agent needs to i) understand the underlying goal of the task by watching a single demonstration of the human-like agent performing the same task (social perception), and ii) coordinate with the human-like agent to solve the task in an unseen environment as fast as possible (human-AI collaboration). For this challenge, we build VirtualHome-Social, a multi-agent household environment, and provide a benchmark including both planning and learning based baselines. We evaluate the performance of AI agents with the human-like agent as well as with real humans using objective metrics and subjective user ratings. Experimental results demonstrate that the proposed challenge and virtual environment enable a systematic evaluation on the important aspects of machine social intelligence at scale.
Forward citations
Cited by 7 Pith papers
-
Embodied Multi-Agent Coordination by Aligning World Models Through Dialogue
Dialogue between partially-observing LLM agents cuts action conflicts by 40-83 points but lowers task success versus silent coordination, with new metrics exposing limited genuine world-model alignment.
-
R^3-VQA: "Read the Room" by Video Social Reasoning
R3-VQA is a new real-world video benchmark on which the best tested model, GPT-4o, scores 83% on generated questions but only 54% on human-written ones, while humans score 91% and 80%.
-
CoMind: Understanding Collaborative Human Activity from Multiple Minds and Views
CoMind releases 41 h of synchronized multi-view cooking collaboration with social-cue annotations and three ToM-oriented benchmarks on which current VLMs score poorly until fine-tuned.
-
CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models
Topology-aware attention over hierarchical scene graphs lets a 3D-LLM ground, caption, and answer questions across multi-room homes, with large gains on a new HM3D benchmark.
-
Human-AI Collaboration for Estimating Scientific Replicability
Hybrid human-AI prediction markets match or slightly outperform AI-only markets at forecasting scientific replication outcomes across six social science disciplines.
-
Auto-SLURP: A Benchmark Dataset for Evaluating Multi-Agent Frameworks in Smart Personal Assistant
Auto-SLURP is a new end-to-end benchmark for LLM multi-agent personal assistants, and the best current framework still fails a majority of user requests.
-
Large Language Models for Planning: A Comprehensive and Systematic Survey
A structured survey of LLM planning methods, benchmarks, and interpretability work, organized around a three-way taxonomy.
Discussion (0). Continue with ORCID to comment.