Pith. sign in

REVIEW 3 cited by

Generate then Select: Open-ended Visual Question Answering Guided by World Knowledge

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.18842 v1 pith:J666W3XE submitted 2023-05-30 cs.CL cs.AIcs.CV

classification cs.CLcs.AIcs.CV
keywords knowledgemodelsworldanswergeneraterasovisualanswering
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The open-ended Visual Question Answering (VQA) task requires AI models to jointly reason over visual and natural language inputs using world knowledge. Recently, pre-trained Language Models (PLM) such as GPT-3 have been applied to the task and shown to be powerful world knowledge sources. However, these methods suffer from low knowledge coverage caused by PLM bias -- the tendency to generate certain tokens over other tokens regardless of prompt changes, and high dependency on the PLM quality -- only models using GPT-3 can achieve the best result. To address the aforementioned challenges, we propose RASO: a new VQA pipeline that deploys a generate-then-select strategy guided by world knowledge for the first time. Rather than following the de facto standard to train a multi-modal model that directly generates the VQA answer, RASO first adopts PLM to generate all the possible answers, and then trains a lightweight answer selection model for the correct answer. As proved in our analysis, RASO expands the knowledge coverage from in-domain training data by a large margin. We provide extensive experimentation and show the effectiveness of our pipeline by advancing the state-of-the-art by 4.1% on OK-VQA, without additional computation cost. Code and models are released at http://cogcomp.org/page/publication_view/1010

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Court of LLMs: Evidence-Augmented Generation via Multi-LLM Collaboration for Text-Attributed Graph Anomaly Detection

    cs.LG 2025-08 conditional novelty 6.0 of 10

    CoLL uses two specialized LLM 'prosecutors' and an LLM 'judge' to generate textual anomaly evidence, which a gated GNN then fuses with graph structure for state-of-the-art text-attributed graph anomaly detection.

  2. GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A multi-agent vision-language framework that extracts event, time, and location from public event images, evaluated with a new soft metric on VLM-augmented datasets.

  3. Augmented Vision-Language Models: A Systematic Review

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A structured taxonomy of inference-time augmentation techniques that connect vision-language models to external symbolic systems, tools, and knowledge sources.

Pith tools