REVIEW 5 cited by
Visual question answering: from early developments to recent advances -- a survey
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Visual Question Answering (VQA) is an evolving research field aimed at enabling machines to answer questions about visual content by integrating image and language processing techniques such as feature extraction, object detection, text embedding, natural language understanding, and language generation. With the growth of multimodal data research, VQA has gained significant attention due to its broad applications, including interactive educational tools, medical image diagnosis, customer service, entertainment, and social media captioning. Additionally, VQA plays a vital role in assisting visually impaired individuals by generating descriptive content from images. This survey introduces a taxonomy of VQA architectures, categorizing them based on design choices and key components to facilitate comparative analysis and evaluation. We review major VQA approaches, focusing on deep learning-based methods, and explore the emerging field of Large Visual Language Models (LVLMs) that have demonstrated success in multimodal tasks like VQA. The paper further examines available datasets and evaluation metrics essential for measuring VQA system performance, followed by an exploration of real-world VQA applications. Finally, we highlight ongoing challenges and future directions in VQA research, presenting open questions and potential areas for further development. This survey serves as a comprehensive resource for researchers and practitioners interested in the latest advancements and future
Forward citations
Cited by 5 Pith papers
-
Spectral Heat Flow for Conservative Token Condensation in Vision-Language Models
Training-free SpecFlow condenses VLM visual tokens via kNN heat diffusion, adaptive quadtree budgets, and coreset sinks, retaining 95.6% LLaVA-1.5 performance after pruning 88.9% of tokens.
-
Affordance Benchmark for MLLMs
A new 2,000-question benchmark finds multimodal AI models recognize object affordances far worse than humans, with top model Gemini-2.0-Pro at 18.05% versus 85.34% human best.
-
Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method
OmniVQA is a first open-source dataset and benchmark for 360-degree visual question answering, and 360-R1 uses GRPO with three LLM-based rewards to improve an existing multimodal model on it.
-
Exploring the Application of Visual Question Answering (VQA) for Classroom Activity Monitoring
Four open-source VQA models reach moderate accuracy on a new classroom video dataset, with yes/no questions easiest and counting/reasoning hardest.
-
From Reasoning to Generalization: Knowledge-Augmented LLMs for ARC Benchmark
A staged knowledge-prompting method (KAAR) improves LLM test accuracy on ARC by about 5 absolute points over repeated-sampling plan-guided code generation, reaching 35% with GPT-o3-mini.
Discussion (0). Continue with ORCID to comment.