Pith. sign in

REVIEW 2 cited by

SIMMC 2.0: A Task-oriented Dialog Dataset for Immersive Multimodal Conversations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.08667 v2 pith:HVH3MUN2 submitted 2021-04-18 cs.CL cs.AI

classification cs.CLcs.AI
keywords dialogmultimodaltask-orienteddatasetcollectedconversationsdialogsimmersive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Next generation task-oriented dialog systems need to understand conversational contexts with their perceived surroundings, to effectively help users in the real-world multimodal environment. Existing task-oriented dialog datasets aimed towards virtual assistance fall short and do not situate the dialog in the user's multimodal context. To overcome, we present a new dataset for Situated and Interactive Multimodal Conversations, SIMMC 2.0, which includes 11K task-oriented user<->assistant dialogs (117K utterances) in the shopping domain, grounded in immersive and photo-realistic scenes. The dialogs are collected using a two-phase pipeline: (1) A novel multimodal dialog simulator generates simulated dialog flows, with an emphasis on diversity and richness of interactions, (2) Manual paraphrasing of the generated utterances to collect diverse referring expressions. We provide an in-depth analysis of the collected dataset, and describe in detail the four main benchmark tasks we propose. Our baseline model, powered by the state-of-the-art language model, shows promising results, and highlights new challenges and directions for the community to study.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Augmenting speech transcripts of VR recordings with gaze, pointing, and visual context for multimodal coreference resolution

    cs.HC 2025-09 conditional novelty 6.0 of 10

    Augmenting VR speech transcripts with gaze and pointing cues improved GPT-4 coreference resolution from 40.6% to 67.1% accuracy.

  2. Muse: A Multimodal Conversational Recommendation Dataset with Scenario-Grounded User Profiles

    cs.MM 2024-12 conditional novelty 6.0 of 10

    MUSE is a 7,000-conversation multimodal conversational recommendation dataset synthesized by MLLM agents with scenario-grounded user profiles, including 83,148 utterances and product images.

Pith tools