{"id":"c6b148e5-6a7d-4333-83e8-057d3f544fbd","arxiv_id":"2501.03939","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey that classifies VQA architectures by encoder, fusion, and decoder, reviews datasets and metrics, and discusses applications and future directions.","lead":"This paper surveys Visual Question Answering, organizing models, datasets, metrics, and applications into a taxonomy of encoders, fusion methods, and decoders. It is a reference review, not a new method, so its value depends on how accurately it summarizes the prior work it cites.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 2's RSVQA rows cite Gurari et al. instead of Lobry et al., Section 7.4 cites a non-existent 'Bongini Brown et al.' for the GPT-3 study, and Table 3's Qwen-VL row cites the ChartQA paper instead of Qwen-VL — a reference survey's core value is traceability, and these breaks undermine it.","rationale":"The paper is explicitly a survey, so the evidentiary standard is compilation accuracy and traceability, not novel experimental results. I assessed the taxonomy, the dataset tables, the model tables, and the application sections. The taxonomy is conventional and overlaps prior surveys the authors cite; that is a novelty limitation, not a soundness failure. The load-bearing concern is traceability: the reference value of the survey depends on every entry in Tables 2-3 and every in-text attribution pointing to its correct source. The reader identified two unambiguous failures, and the full text confirms them plus a third: RSVQA is misattributed in Table 2 while Section 7.3 is correct; Section 7.4 cites a non-existent author and conflates the GPT-3 paper with an artwork-description study; and Table 3's Qwen-VL row cites the ChartQA paper rather than the Qwen-VL source. Three independent citation breaks across different sections indicate a systematic verification gap rather than isolated typos. The reader's analysis is accurate, and I find no reason to move away from CONDITIONAL: the survey is usable as a broad orientation once the citation-to-source mapping and accuracy transcriptions are corrected. I set agreement_with_reader to 'agree' because the reader's weakest assumption — that citations and accuracy numbers are correctly transcribed — is exactly the assumption that fails, and it is the load-bearing one for a survey claiming to be a comprehensive resource.","tokens_in":42564,"tokens_out":2318,"duration_ms":19548,"concrete_test":"Audit every row of Tables 2 and 3 by resolving each cited reference to its actual title and venue, then checking that the dataset/model name and reported accuracy figure match the cited paper. Concretely: (1) For the Qwen-VL 7B row, open Bai et al. 2023 and verify the 78.80 VQA v2.0 test-dev accuracy appears there and not in Masry et al. 2022 (ChartQA); (2) For Table 2 rows 37-38, verify the RSVQA entries against Lobry et al. 2020 rather than Gurari et al. 2018; (3) Search for any publication by 'Bongini Brown' and check whether Section 7.4's GPT-3 artwork-description claim is actually supported by Brown et al. 2020. If more than 10% of audited rows fail, the survey cannot serve its stated reference function.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's stated purpose is to serve as a comprehensive reference resource, with Table 2 compiling datasets and Table 3 compiling model accuracies. That purpose depends on every entry being traceable to the correct primary source. The reader identified several breaks, and the full text confirms: (1) Table 2 rows 37-38 attribute RSVQA to Gurari et al. 2018, but Section 7.3 correctly names Lobry et al. 2019/2020 as the RSVQA authors; readers cannot reconcile the table with the text or trace the dataset to its source. (2) Section 7.4 cites 'Bongini Brown et al. [2020]' for a GPT-3 artwork-description study, while the bibliography contains only Brown et al. 2020, 'Language models are few-shot learners,' which describes no such study. (3) Table 3's Qwen-VL 7B row cites Masry et al. 2022 (ChartQA) as both the model and language-encoder source; the correct source for Qwen-VL is Bai et al. 2023 (Qwen Technical Report), and the 78.80 figure must be verified against that paper. These are not stylistic slips: a survey that cannot be used for lookups fails its central function. The missing systematic search protocol further weakens the 'comprehensive' claim, but the factual citation errors alone justify a conditional verdict.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey proposes a taxonomy of VQA architectures organized by vision encoder, language encoder, fusion method, and answer decoder, and uses it to structure a review of deep-learning VQA methods, LVLMs, datasets, metrics, applications, and future directions. The paper includes a large table of datasets (Table 2) and a large table of model accuracies (Table 3), and it describes applications in medicine, accessibility, remote sensing, cultural heritage, advertising, and education. The stated contribution is to serve as a comprehensive, up-to-date reference resource for VQA researchers and practitioners as of early 2025.","tokens_in":42752,"tokens_out":5150,"duration_ms":42826,"significance":"If the factual content were reliable, this would be a genuinely useful synthesis: the taxonomy in Figure 3 is reasonable, the coverage of historical periods from 2015 to 2024 is broad, and the treatment of LVLMs as the current dominant paradigm is appropriate. The paper also covers multiple applied domains and gives a clear account of the standard VQA pipeline. However, the survey's main value as a reference depends on the accuracy of its tables and citations, and the manuscript currently contains several concrete traceability failures. Those failures are not cosmetic; they directly undermine the paper's central claim to be a comprehensive resource for lookup and comparison. I therefore evaluate the manuscript as promising but not yet publishable in its present form.","major_comments":[{"comment":"RSVQA-low and RSVQA-high are attributed to Gurari et al. [2018a], but Section 7.3 correctly attributes RSVQA to Lobry et al. [2019] and Lobry et al. [2020]. A reader using Table 2 as a lookup cannot trace the dataset to its primary source, so the table fails its stated reference function.","section":"Table 2, rows 37–38"},{"comment":"The Qwen-VL 7B row cites Masry et al. [2022b] for both the image encoder and the language encoder, but Qwen-VL is introduced in Bai et al. [2023]; Masry et al. is the ChartQA paper. The 78.80 accuracy should be verified against the Qwen-VL report, and the row should cite the correct source.","section":"Table 3, Qwen-VL 7B row"},{"comment":"The sentence 'To overcome the need for expert-generated descriptions, Bongini Brown et al. [2020] used GPT-3 to automatically generate detailed descriptions of artworks' cites a work that does not appear in the bibliography. The closest entry, Brown et al. [2020] 'Language models are few-shot learners', describes no such artwork study, so this claim is currently unsupported and needs a correct reference or removal.","section":"Section 7.4"},{"comment":"AI2D is cited to Sheng et al. [2016], the same source used for Art-VQA in row 9, but the actual AI2D source is Kembhavi et al. [2016], which is cited in the text. This is another instance in which the dataset table provides incorrect provenance for an entry.","section":"Table 2, row 10"},{"comment":"The text states that VQA v2.0 accuracy during Period III ranged from 71% to 78%, but Table 3 lists Florence (80.16), SimVLM (80.34), and VLMo (82.78) with 2021 dates inside that period. The text and table contradict each other; the range should be corrected or the period boundaries clarified.","section":"Section 8.3, Period III"}],"minor_comments":[{"comment":"The example question is 'What is the object on the refrigerator?', but the described reasoning step says 'identifying the object on the table'; this internal inconsistency should be fixed.","section":"Section 3.3.3"},{"comment":"The sentence 'simple fusion techniques need to provide insight into how the visual and textual inputs are combined' appears to mean the opposite; it should say 'fail to provide insight' or 'provide no insight'.","section":"Section 3.3.1"},{"comment":"The axis labels are garbled, for example 'VI515/0(20  - 08)017/2 7/08(201  - 8)19/020 /089201(  - )1/08202 1/08(202  - )222/120'; the figure should be regenerated with legible period labels.","section":"Figure 4"},{"comment":"The text says questions 'must begin with one of six letters: what, where, how, why, who, and when', but these are interrogative words rather than letters; the wording should be corrected to 'six question categories' or similar.","section":"Section 4.2"},{"comment":"The bibliography contains duplicate entries (e.g., Antol et al. 2015a/2015b, Devlin et al. 2018/2019, Liu et al. 2019a/2019b) and a malformed author string in the Masry et al. [2022b] entry ('andbai2023qwen Villavicencio'); these need cleanup.","section":"References"},{"comment":"The final sentence about the SimpsonsVQA dataset appears without citation or connection to the surrounding discussion of question relevance; it should be integrated or removed.","section":"Section 5"}],"recommendation":"major_revision","confidential_remarks":"The paper has the right scope and structure for a survey, and with a careful correction pass it could become a useful reference. My main concern is that the citation and transcription errors are numerous enough that the authors should audit every row of Tables 2 and 3 against the primary sources and verify the accuracy numbers in Table 3 before resubmission. I would not require a new methodology or a formal meta-analysis, but the fact-check burden is substantial and should be part of the revision, not deferred."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a survey of VQA that would like to be the reference table you hand a new student. It organizes the field by vision encoder, language encoder, fusion, and answer decoder, compiles 61 datasets and a large model comparison table, and adds sections on LVLMs, medical VQA, accessibility, remote sensing, ads, education, and cultural heritage. The structure is sensible and the coverage of recent LVLMs is a genuine update over the earlier surveys it cites. For a newcomer wanting a bird's-eye view, it mostly does the job.\n\nThe problem is that the job of a reference survey is traceability, and the traceability is broken in several places you can see without even trying hard. Table 2 credits RSVQA to Gurari et al. (the VizWiz paper) when Section 7.3 correctly credits Lobry et al. Section 7.4 cites a non-existent 'Bongini Brown et al.' for a GPT-3 artwork-description study; the bibliography only has Brown et al.'s 'Language models are few-shot learners', which describes no such study. Table 3's Qwen-VL row cites the ChartQA paper (Masry et al. 2022) as both the model and the language encoder source, when the correct source is Bai et al.'s Qwen Technical Report. There are also duplicated rows in Table 3 (MLIN-BERT, mPLUG) and at least one entry with a year that does not match the cited reference. These are not stylistic slips. A reader who wants to look up RSVQA or Qwen-VL will be directed to the wrong paper, and that defeats the purpose of the survey.\n\nThe stress-test note got this right. I checked the full text and the errors are real and load-bearing for the paper's stated purpose. The taxonomy itself is conventional, overlapping with Manmadhan and Kovoor and others the authors themselves cite, so the survey's added value is precisely the compilation — which is exactly what can't be trusted as-is.\n\nThat said, the errors look fixable rather than fatal. The authors clearly have read the literature; Section 7.3 names Lobry correctly, so the table mistake is a transcription slip, and the bibliography includes the Qwen report. A careful revision with independent verification of every row in Tables 2 and 3, a corrected Section 7.4 citation, and an explicit statement of the search/selection methodology would make this a usable resource. Without that, I would not want to cite it.\n\nVerdict: worth sending to a serious referee, but the referee should be asked to check the tables line by line. I would not desk-reject; I also would not accept as-is. Conditional on fixing the citations and being honest about the method.","headline":"A useful but sloppy VQA survey: structure and LVLM coverage are fine, but citation and transcription errors in the dataset and model tables break its value as a reference until fixed.","tokens_in":43403,"tokens_out":2696,"would_cite":false,"duration_ms":23925,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper's central claim is that a single taxonomy of VQA architectures—organized by vision encoder, language encoder, fusion machine, and answer decoder—can structure the whole field and expose its shift from simple fusion to large…","keywords":["visual question answering","multimodal learning","vision-language pretraining","large visual language models","attention mechanisms","benchmark datasets","taxonomy","survey"],"falsifier":"Select a random sample of rows from Table 3, locate the cited original papers, and verify the reported accuracy values and dataset names; if a substantial share of entries differ from the originals, the paper's claim to be a reliable VQA resource fails.","tokens_in":42264,"feed_emoji":"🧩","tokens_out":10394,"duration_ms":83354,"temperature":0.7,"pith_summary":"This survey sets out to give the visual question answering field a structured map: a taxonomy that sorts VQA architectures by their vision encoder, language encoder, fusion mechanism, and answer decoder. The authors argue that earlier surveys covered only pieces, such as fusion techniques or datasets, whereas an updated survey is needed because large visual language models (LVLMs) have reshaped the field since 2019. If the taxonomy is right, researchers can compare any VQA model along the same four axes and see how the field moved from simple fusion, through attention and bilinear pooling, to pretrained LVLMs. The paper also assembles the main datasets, evaluation metrics, and application domains as part of the same organizing scheme.","feed_headline":"One taxonomy sorts a decade of VQA models","feed_subtitle":"It groups every major VQA model by its four building blocks: encoders, fusion, and decoder.","key_machinery":"The load-bearing object is the taxonomy shown in Figure 3, which decomposes every VQA system into the same four functional blocks and then subdivides each block into the technique families used in the literature. It does the organizing work of the survey: every model reviewed in the paper is placed in this grid, and Table 3 and Figure 4 translate the taxonomy into quantitative comparisons by reporting each model's accuracy on benchmarks such as VQA v1.0, VQA v2.0, DAQUAR, Visual7W, COCO-QA, CLEVR, and GQA. The taxonomy is what allows the paper to draw timeline conclusions, such as object-based vision encoders giving way to ViT patch encoders and transformer-based fusion displacing earlier attention designs.","core_discovery":"The paper's central claim is that the whole of VQA architecture design can be organized by four components: vision encoder (grid-based, object-based, or ViT patch-based), language encoder (bag-of-words, RNN, CNN, or transformer), fusion machine (simple fusion, attention, bilinear pooling, relation networks, neural module networks, or large visual-language models), and answer decoder (closed-vocabulary or open-vocabulary). Within this taxonomy, the survey charts a historical progression: early models used CNNs and LSTMs with simple fusion, attention-based and bilinear pooling methods then raised accuracy, and after 2019 LVLMs pretrained on image-text data became the dominant approach, with accuracy on VQA v2.0 climbing from about 52 percent in 2015 to above 84 percent by 2022. The authors' stated intent is that this structured framework makes comparative analysis and evaluation straightforward and gives researchers and practitioners a single resource covering methods, datasets, metrics, and applications.","pith_inferences":["If the taxonomy is taken as the field's organizing scheme, an implicit testable prediction is that future state-of-the-art VQA models will be describable as LVLMs with ViT encoders and LLM backbones, making the older fusion categories mostly historical.","The survey's own section on question relevance reports that a large share of human questions are unrelated to their images; a production VQA system could therefore benefit from an explicit relevance-detection gate before answering, a design the paper does not develop.","A consequence the authors leave implicit is that Table 3's benchmark rankings conflate architecture choice with model scale; separating those two effects would require controlled comparisons the table does not provide.","One testable extension would be to use the taxonomy's four axes as features for predicting which new VQA datasets a model will transfer to, treating the survey's classification as input to a prediction study."],"forward_implications":["Researchers can locate any VQA model's design choices on the taxonomy's four axes and compare models that differ in only one component, isolating which part drives accuracy.","The four-period timeline implies that future advances will come from LVLMs, with ViT-based patch encoders and pretrained large language model backbones dominating the field.","The accuracy tables imply that transformer-based fusion, especially self-attention and cross-attention, was the strongest fusion family before the LVLM era, with MCAN and its extensions setting the pre-LVLM ceiling on VQA v2.0.","Because the survey maps datasets by domain, it implies that progress in medical, remote sensing, education, cultural heritage, and advertising VQA tracks dataset creation as much as model design.","For closed-vocabulary VQA tasks, the decoder taxonomy suggests that open-ended and multiple-choice setups should be treated as distinct problems with distinct evaluation metrics."],"supporting_citations":[{"why":"Defines the VQA task and the original VQA v1.0 dataset, the starting point for the survey's accuracy timeline.","marker":"Antol et al. [2015b]"},{"why":"Introduces VQA v2.0, the main benchmark used in the survey's model comparison table.","marker":"Goyal et al. [2017b]"},{"why":"Supplies the bottom-up top-down attention and object-based encoder family that anchors one branch of the taxonomy.","marker":"Anderson et al. [2018]"},{"why":"Provides MCAN, the transformer-based fusion model that anchors the survey's attention and transformer branch.","marker":"Yu et al. [2019]"},{"why":"Introduces ViLT, used as evidence that ViT patch encoders cut model size and runtime compared with grid and object encoders.","marker":"Kim et al. [2021]"},{"why":"Supplies CLIP, the contrastive vision-language pretraining model the taxonomy lists under ViT-based encoders.","marker":"Radford et al. [2021]"},{"why":"Supplies BERT, the transformer language encoder on which many LVLM entries in the survey are built.","marker":"Devlin et al. [2018]"},{"why":"Provides Flamingo, the pretrained-backbone LVLM example supporting the survey's claims about large-scale few-shot VQA.","marker":"Alayrac et al. [2022]"},{"why":"Presents BLIP-2, the Q-Former adapter example used to describe fine-tuning a frozen LLM backbone for VQA.","marker":"Li et al. [2023a]"},{"why":"Introduces LLaVA, the visual instruction tuning work cited as a leading adapter-based LVLM for VQA.","marker":"Liu et al. [2023a]"}],"fun_headline_variants":["VQA survey: 4 components, 52% to 84% accuracy","Four-part taxonomy maps VQA's rise to 84% accuracy","From CNN-LSTM to LVLM: a VQA survey's roadmap","How VQA accuracy doubled: 52% to 84% in a decade","Survey: VQA's four components and its leap to 84%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's usefulness as a reference depends on the accuracy numbers in Table 3 and the dataset attributions in Table 2 being transcribed correctly from the cited papers; if those transcriptions are unreliable, the comparison tables cannot be trusted.","fun_headline_variants_meta":{"raw":{"variants":["VQA survey: 4 components, 52% to 84% accuracy","Four-part taxonomy maps VQA's rise to 84% accuracy","From CNN-LSTM to LVLM: a VQA survey's roadmap","How VQA accuracy doubled: 52% to 84% in a decade","Survey: VQA's four components and its leap to 84%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000748,"raw_usage":{"total_tokens":3339,"prompt_tokens":961,"completion_tokens":2378,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":577,"completion_tokens_details":{"reasoning_tokens":2279}},"tokens_in":577,"tokens_out":2378,"duration_ms":14920,"temperature":1.0,"reasoning_tokens":2279,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:43:06.905260+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Select a random sample of rows from Table 3, locate the cited original papers, and verify the reported accuracy values and dataset names; if a substantial share of entries differ from the originals, the paper's claim to be a reliable VQA resource fails.","supporting_citations":[],"review_version":1}