Pith. sign in

REVIEW 2 cited by

Interpretable Counting for Visual Question Answering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1712.08697 v2 pith:P6LKJ6K4 submitted 2017-12-23 cs.AI cs.CLcs.CV

classification cs.AIcs.CLcs.CV
keywords countingimageobjectsquestionansweringcountsdiscreteinterpretable
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Questions that require counting a variety of objects in images remain a major challenge in visual question answering (VQA). The most common approaches to VQA involve either classifying answers based on fixed length representations of both the image and question or summing fractional counts estimated from each section of the image. In contrast, we treat counting as a sequential decision process and force our model to make discrete choices of what to count. Specifically, the model sequentially selects from detected objects and learns interactions between objects that influence subsequent selections. A distinction of our approach is its intuitive and interpretable output, as discrete counts are automatically grounded in the image. Furthermore, our method outperforms the state of the art architecture for VQA on multiple metrics that evaluate counting.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A rule-guided spatial-aware network that localizes all mentioned entities in a 3D scene and uses target-position weak supervision raises ScanRefer 3D-RES mIoU from 39.5 to 44.6.

  2. A Comprehensive Survey on Visual Question Answering Datasets and Algorithms

    cs.CV 2024-11 unverdicted

    A broad but dated survey of VQA datasets and algorithms that organizes the pre-2021 literature into four dataset categories and six model paradigms.

Pith tools