Pith. sign in

REVIEW 11 cited by

GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2105.14517 v3 pith:TK4RPPYI submitted 2021-05-30 cs.AI

classification cs.AI
keywords geometricgeoqaproblemsbenchmarksolvingansweringauxiliarydataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Automatic math problem solving has recently attracted increasing attention as a long-standing AI benchmark. In this paper, we focus on solving geometric problems, which requires a comprehensive understanding of textual descriptions, visual diagrams, and theorem knowledge. However, the existing methods were highly dependent on handcraft rules and were merely evaluated on small-scale datasets. Therefore, we propose a Geometric Question Answering dataset GeoQA, containing 4,998 geometric problems with corresponding annotated programs, which illustrate the solving process of the given problems. Compared with another publicly available dataset GeoS, GeoQA is 25 times larger, in which the program annotations can provide a practical testbed for future research on explicit and explainable numerical reasoning. Moreover, we introduce a Neural Geometric Solver (NGS) to address geometric problems by comprehensively parsing multimodal information and generating interpretable programs. We further add multiple self-supervised auxiliary tasks on NGS to enhance cross-modal semantic representation. Extensive experiments on GeoQA validate the effectiveness of our proposed NGS and auxiliary tasks. However, the results are still significantly lower than human performance, which leaves large room for future research. Our benchmark and code are released at https://github.com/chen-judge/GeoQA .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LaRe: Latent Refocusing for Multimodal Reasoning

    cs.CV 2025-11 reject novelty 6.0 of 10

    LaRe performs iterative visual refocusing in latent space and reports accuracy gains with fewer tokens, but its main experiments compare against baselines trained with less data.

  2. EasyARC: Evaluating Vision Language Models on True Visual Reasoning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    EasyARC is a new procedurally generated visual reasoning benchmark where state-of-the-art vision-language models score below 20%, despite tasks designed to be easy.

  3. Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Mixed-R1 uses four reward types under GRPO, including a new bidirectional max-average token similarity (BMAS) reward, and lifts MLLM reasoning benchmarks by 2-5%.

  4. Decomposing Elements of Problem Solving: What "Math" Does RL Teach?

    cs.AI 2025-05 conditional novelty 6.0 of 10

    Reinforcement learning (GRPO) on math LLMs primarily increases execution robustness on already-solvable problems, not planning or coverage of new problems.

  5. Can Visual Encoder Learn to See Arrows?

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Contrastive training on randomly laid-out, alphabet-labeled synthetic diagrams makes CLIP encoders detect edge direction and caption diagrams, beating pretrained CLIP and zero-shot GPT-4o on the synthetic test set.

  6. LaViDa: A Large Diffusion Language Model for Multimodal Understanding

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A diffusion-based vision-language model matches several autoregressive baselines on multimodal benchmarks while enabling controllable generation and faster decoding at reduced quality.

  7. Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery

    cs.AI 2026-06 unverdicted novelty 5.0 of 10

    An integrated survey organizing AI mathematical reasoning into informal, formal, discovery, and technique axes while cataloging benchmarks and assessing failure modes.

  8. MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning

    cs.CV 2025-06 conditional novelty 5.0 of 10

    MINT-CoT-7B interleaves fine-grained visual tokens into each math reasoning step and reports 73.70 on MathVista-Math, 64.72 on GeoQA, and 69.6 on MMStar-Math.

  9. MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A multimodal LLM with a learnable alignment projector and a combined cross-entropy and mean-squared-error loss reports state-of-the-art scores on several vision-language benchmarks using only 144 visual tokens.

  10. Towards Geometry Problem Solving in the Large Model Era: A Survey

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A survey that organizes geometry problem-solving research into benchmark construction, parsing, and reasoning, and proposes a unified parse-then-reason paradigm for the large-model era.

  11. NAN: A Training-Free Solution to Coefficient Estimation in Model Merging

    cs.LG 2025-05 reject novelty 4.0 of 10

    NAN sets merging coefficients inversely proportional to each model's parameter norm and claims a least-squares justification, but the derivation yields a different formula and performance gains are inconsistent.

Pith tools