Pith. sign in

REVIEW 2 cited by

What is Right for Me is Not Yet Right for You: A Dataset for Grounding Relative Directions via Multi-Task Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.02671 v1 pith:52LSBESW submitted 2022-05-05 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords directionsrelativegroundingdatasetobjectsspatialabsoluteanswering
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Understanding spatial relations is essential for intelligent agents to act and communicate in the physical world. Relative directions are spatial relations that describe the relative positions of target objects with regard to the intrinsic orientation of reference objects. Grounding relative directions is more difficult than grounding absolute directions because it not only requires a model to detect objects in the image and to identify spatial relation based on this information, but it also needs to recognize the orientation of objects and integrate this information into the reasoning process. We investigate the challenging problem of grounding relative directions with end-to-end neural networks. To this end, we provide GRiD-3D, a novel dataset that features relative directions and complements existing visual question answering (VQA) datasets, such as CLEVR, that involve only absolute directions. We also provide baselines for the dataset with two established end-to-end VQA models. Experimental evaluations show that answering questions on relative directions is feasible when questions in the dataset simulate the necessary subtasks for grounding relative directions. We discover that those subtasks are learned in an order that reflects the steps of an intuitive pipeline for processing relative directions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RoboVista: Evaluating Vision Language Models for Diverse Robot Applications

    cs.RO 2026-07 accept novelty 6.5 of 10

    Expert-curated modular Robot-VQA benchmark of 474 questions across 39 robot tasks shows SOTA VLMs have large gaps that correlate with physical robot execution.

  2. Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

    cs.RO 2025-05 conditional novelty 6.0 of 10

    Robot sensor logs can be converted automatically into 684,710 labeled visual QA questions, and fine-tuning on them improves VLM answers on the same test distribution, though models still trail humans on fine spatial r...

Pith tools