Pith. sign in

REVIEW 4 cited by

NExT-Chat: An LMM for Chat, Detection and Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.04498 v4 pith:HEYNZHWO submitted 2023-11-08 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords next-chatlocationmultimodalboundingdifferentlargelmmsmethod
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The development of large language models (LLMs) has greatly advanced the field of multimodal understanding, leading to the emergence of large multimodal models (LMMs). In order to enhance the level of visual comprehension, recent studies have equipped LMMs with region-level understanding capabilities by representing object bounding box coordinates as a series of text sequences (pix2seq). In this paper, we introduce a novel paradigm for object location modeling called pix2emb method, where we ask the LMM to output the location embeddings and then decode them with different decoders. This paradigm allows us to use different location formats (such as bounding boxes and masks) in multimodal conversations. Leveraging the proposed pix2emb method, we train an LMM named NExT-Chat and demonstrate its capability of handling multiple tasks like visual grounding, region captioning, and grounded reasoning. Comprehensive experiments show the effectiveness of our NExT-Chat on various tasks, e.g., NExT-Chat (87.7) vs. Shikra (86.9) on POPE-Random, NExT-Chat (68.9) vs. LISA (67.9) on referring expression segmentation task, and NExT-Chat (79.6) vs. Kosmos-2 (62.3) on region caption task. The code and model are released at https://github.com/NExT-ChatV/NExT-Chat.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Under one zero-shot protocol, general-purpose MLLMs match or outperform remote-sensing-specific MLLMs on several RS benchmarks, while RS-MLLMs keep advantages in visual grounding, RS-VQA, and ultra-high-resolution und...

  2. ReMeREC: Relation-aware and Multi-entity Referring Expression Comprehension

    cs.CV 2025-07 conditional novelty 6.0 of 10

    ReMeREC introduces a relation-aware multi-entity referring expression comprehension framework and the ReMeX dataset, reporting state-of-the-art grounding and relation prediction, with some evaluation caveats.

  3. HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    HRSeg combines high-resolution image crops, region attention, and cross-attention mask enhancement, improving reasoning segmentation over LLM-Seg by up to 13.2 gIoU points on LLM-Seg40K.

  4. MediRound: Multi-Round Entity-Level Reasoning Segmentation in Medical Images

    cs.CV 2025-11 conditional novelty 5.0 of 10

    MediRound introduces a multi-round, entity-level medical segmentation task, a 177K-dialogue dataset built from SA-Med2D-20M with GPT-5, and a LLaVA-Med/MedSAM baseline whose inference-time judgment-and-correction modu...

Pith tools