Pith. sign in

REVIEW 1 cited by

Using Game Play to Investigate Multimodal and Conversational Grounding in Large Multimodal Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.14035 v3 pith:LZA5T32B submitted 2024-06-20 cs.CL cs.AI

classification cs.CLcs.AI
keywords modelsmultimodalevaluationdefinefindgamegameslargest
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While the situation has improved for text-only models, it again seems to be the case currently that multimodal (text and image) models develop faster than ways to evaluate them. In this paper, we bring a recently developed evaluation paradigm from text models to multimodal models, namely evaluation through the goal-oriented game (self) play, complementing reference-based and preference-based evaluation. Specifically, we define games that challenge a model's capability to represent a situation from visual information and align such representations through dialogue. We find that the largest closed models perform rather well on the games that we define, while even the best open-weight models struggle with them. On further analysis, we find that the exceptional deep captioning capabilities of the largest models drive some of the performance. There is still room to grow for both kinds of models, ensuring the continued relevance of the benchmark.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Vision-Language Model Dialog Games for Self-Improvement

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Self-play dialog games between two VLMs generate filtered synthetic data that, when fine-tuned on, improves VQA and robotics success detection.

Pith tools