Adding coreference resolution supervision to a phrase-grounding model improves pronoun-to-object grounding in Japanese dialogue images and enables handling of zero references, outperforming MDETR and GLIP on pronouns in J-CRe3.
Moura, Devi Parikh, and Dhruv Batra
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Disambiguating Reference in Visually Grounded Dialogues through Joint Modeling of Textual and Multimodal Semantic Structures
Adding coreference resolution supervision to a phrase-grounding model improves pronoun-to-object grounding in Japanese dialogue images and enables handling of zero references, outperforming MDETR and GLIP on pronouns in J-CRe3.