Pith. sign in

REVIEW 5 cited by

On Pre-training of Multimodal Language Models Customized for Chart Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.14506 v3 pith:4U5QL2PM submitted 2024-07-19 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords chartcomprehensiondatachartschopinllmlanguagemllmsmodels
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recent studies customizing Multimodal Large Language Models (MLLMs) for domain-specific tasks have yielded promising results, especially in the field of scientific chart comprehension. These studies generally utilize visual instruction tuning with specialized datasets to enhance question and answer (QA) accuracy within the chart domain. However, they often neglect the fundamental discrepancy between natural image-caption pre-training data and digital chart image-QA data, particularly in the models' capacity to extract underlying numeric values from charts. This paper tackles this oversight by exploring the training processes necessary to improve MLLMs' comprehension of charts. We present three key findings: (1) Incorporating raw data values in alignment pre-training markedly improves comprehension of chart data. (2) Replacing images with their textual representation randomly during end-to-end fine-tuning transfer the language reasoning capability to chart interpretation skills. (3) Requiring the model to first extract the underlying chart data and then answer the question in the fine-tuning can further improve the accuracy. Consequently, we introduce CHOPINLLM, an MLLM tailored for in-depth chart comprehension. CHOPINLLM effectively interprets various types of charts, including unannotated ones, while maintaining robust reasoning abilities. Furthermore, we establish a new benchmark to evaluate MLLMs' understanding of different chart types across various comprehension levels. Experimental results show that CHOPINLLM exhibits strong performance in understanding both annotated and unannotated charts across a wide range of types.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities

    cs.LG 2025-07 conditional novelty 6.0 of 10

    By probing visual, projection, and response representations, the authors find that most VLM visual knowledge loss for recognition and counting occurs in the language decoder, while spatial understanding is lost in the...

  2. Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Training VLMs to explicitly convert images to text before reasoning transfers simple-to-hard generalization from text to image, and this conversion skill can be internalized to keep inference cheap.

  3. MMFactory: A Universal Solution Search Engine for Vision-Language Tasks

    cs.CV 2024-12 conditional novelty 6.0 of 10

    MMFactory automatically generates and benchmarks a pool of reusable programmatic vision-language solutions from a few examples, letting users pick one that fits their accuracy and speed constraints.

  4. Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes?

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A new 1,430-item multimodal benchmark shows that leading multimodal LLMs rarely notice small visual traps needed for commonsense safety reasoning, with top scores near 0.46 on a scale whose maximum is about 0.97.

  5. Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training

    cs.CV 2025-05 conditional novelty 5.0 of 10

    PRIOR reweights the next-token prediction loss in vision-language pretraining by 1 minus the probability assigned by a text-only reference LLM, and reports consistent benchmark improvements over standard NTP.

Pith tools