REVIEW 3 cited by
AUG: A New Dataset and An Efficient Model for Aerial Image Urban Scene Graph Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Scene graph generation (SGG) aims to understand the visual objects and their semantic relationships from one given image. Until now, lots of SGG datasets with the eyelevel view are released but the SGG dataset with the overhead view is scarcely studied. By contrast to the object occlusion problem in the eyelevel view, which impedes the SGG, the overhead view provides a new perspective that helps to promote the SGG by providing a clear perception of the spatial relationships of objects in the ground scene. To fill in the gap of the overhead view dataset, this paper constructs and releases an aerial image urban scene graph generation (AUG) dataset. Images from the AUG dataset are captured with the low-attitude overhead view. In the AUG dataset, 25,594 objects, 16,970 relationships, and 27,175 attributes are manually annotated. To avoid the local context being overwhelmed in the complex aerial urban scene, this paper proposes one new locality-preserving graph convolutional network (LPG). Different from the traditional graph convolutional network, which has the natural advantage of capturing the global context for SGG, the convolutional layer in the LPG integrates the non-destructive initial features of the objects with dynamically updated neighborhood information to preserve the local context under the premise of mining the global context. To address the problem that there exists an extra-large number of potential object relationship pairs but only a small part of them is meaningful in AUG, we propose the adaptive bounding box scaling factor for potential relationship detection (ABS-PRD) to intelligently prune the meaningless relationship pairs. Extensive experiments on the AUG dataset show that our LPG can significantly outperform the state-of-the-art methods and the effectiveness of the proposed locality-preserving strategy.
Forward citations
Cited by 3 Pith papers
-
GeoSelect: Spatial-Program Execution for Training-Free Referring Remote Sensing Image Segmentation
A training-free pipeline synthesises referring expressions into a typed geometric DSL, executes them over scored candidate boxes, and reaches 58.86 mIoU on RRSIS-D—over twice the previous training-free best.
-
Scene Understanding Enabled Semantic Communication with Open Channel Coding
OpenSC transmits structured scene graphs produced by a scene-graph generator, maps their text tokens to QAM symbols via BERT token IDs, and uses an LLM with retrieval-augmented generation to answer visual questions.
-
Enhanced Multimodal RAG-LLM for Accurate Visual Question Answering
A pipeline that converts an image into a scene graph, embeds the graph chunks, retrieves the most relevant chunks, and prompts an LLM with them reports high VQA accuracy, but the comparison to MLLMs is not credible.
Discussion (0). Continue with ORCID to comment.