Pith. sign in

REVIEW 3 cited by

FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.19237 v2 pith:5IIVDVA3 submitted 2024-06-27 cs.CL cs.CVcs.IRcs.LG

classification cs.CLcs.CVcs.IRcs.LG
keywords visualmultimodalreasoningflowvqaansweringbenchmarkflowchartslanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Existing benchmarks for visual question answering lack in visual grounding and complexity, particularly in evaluating spatial reasoning skills. We introduce FlowVQA, a novel benchmark aimed at assessing the capabilities of visual question-answering multimodal language models in reasoning with flowcharts as visual contexts. FlowVQA comprises 2,272 carefully generated and human-verified flowchart images from three distinct content sources, along with 22,413 diverse question-answer pairs, to test a spectrum of reasoning tasks, including information localization, decision-making, and logical progression. We conduct a thorough baseline evaluation on a suite of both open-source and proprietary multimodal language models using various strategies, followed by an analysis of directional bias. The results underscore the benchmark's potential as a vital tool for advancing the field of multimodal modeling, providing a focused and challenging environment for enhancing model performance in visual and logical reasoning tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multilingual Multimodal Software Developer for Code Generation

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A 7B vision-language model trained on synthetic diagram-to-code data outperforms several larger open-weight models on a new 10-language UML/flowchart code-generation benchmark.

  2. Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Fine-tuning vision-language models with sentence-formatted outputs instead of tuple outputs improves shape attribute and coordinate prediction for larger models, and scaling the loss on numeric tokens sharpens numeric...

  3. Rethinking Comprehensive Benchmark for Chart Understanding: A Perspective from Scientific Literature

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A new benchmark built from real scientific paper charts, including flowcharts and context-dependent questions, shows large multimodal models perform far below human level on chart understanding.

Pith tools