Pith. sign in

REVIEW 3 cited by

GPT4SGG: Synthesizing Scene Graphs from Holistic and Region-specific Narratives

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.04314 v2 pith:FRDTGG2Y submitted 2023-12-07 cs.CV

classification cs.CV
keywords sceneaccuratecomprehensivedatagraphcaptiongpt4sgggraphs
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Training Scene Graph Generation (SGG) models with natural language captions has become increasingly popular due to the abundant, cost-effective, and open-world generalization supervision signals that natural language offers. However, such unstructured caption data and its processing pose significant challenges in learning accurate and comprehensive scene graphs. The challenges can be summarized as three aspects: 1) traditional scene graph parsers based on linguistic representation often fail to extract meaningful relationship triplets from caption data. 2) grounding unlocalized objects of parsed triplets will meet ambiguity issues in visual-language alignment. 3) caption data typically are sparse and exhibit bias to partial observations of image content. Aiming to address these problems, we propose a divide-and-conquer strategy with a novel framework named \textit{GPT4SGG}, to obtain more accurate and comprehensive scene graph signals. This framework decomposes a complex scene into a bunch of simple regions, resulting in a set of region-specific narratives. With these region-specific narratives (partial observations) and a holistic narrative (global observation) for an image, a large language model (LLM) performs the relationship reasoning to synthesize an accurate and comprehensive scene graph. Experimental results demonstrate \textit{GPT4SGG} significantly improves the performance of SGG models trained on image-caption data, in which the ambiguity issue and long-tail bias have been well-handled with more accurate and comprehensive scene graphs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question Answering

    cs.CV 2026-01 unverdicted novelty 6.0 of 10

    KG-ViP fuses scene graphs and commonsense graphs via a query-based retrieval-and-fusion pipeline to improve multi-modal LLM performance on visual question answering.

  2. Open World Scene Graph Generation using Vision Language Models

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A zero-shot VLM pipeline can generate scene graphs with unseen objects and relations, and a new open-world evaluation setting exposes how much capacity remains untapped.

  3. From Data to Modeling: Fully Open-vocabulary Scene Graph Generation

    cs.CV 2025-05 conditional novelty 4.0 of 10

    OvSGTR jointly predicts unseen objects and relationships in scene graphs using a DETR-like transformer, relation-aware pre-training, and knowledge distillation, achieving state-of-the-art results on VG150 and GQA200.

Pith tools