Pith. sign in

Paper Citation Record · LEDGER

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations

As of 14 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 1 inbound Pith citation observation for arXiv:2501.04675.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.04675 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:30:50.269302Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-07T13:43:30.511187Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T08:51:23.968119Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact1
  • verified fuzzy8
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cc8690f2-ac2b-4c1a-a9c9-54e16d2953be · outbound

This paper cites DePlot: One-shot visual language reasoning by plot-to-table translation.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations DePlot: One-shot visual language reasoning by plot-to-table translation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:50.137954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:50.137954Z digest=sha256:d4201a0e89f10d2dc46e189f13d2c6c136079a27e3a953f36fc7e9d54523f6cf

Observation c065acc8-9347-4b1d-91de-a03a54e68b49 · outbound

This paper cites Language models are few-shot learners,.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations Language models are few-shot learners,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:50.144427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:50.144427Z digest=sha256:5673e3863a37c87eeffa89cc234cb970e61248b69fcff64cd2221bef23fd2b00

Observation 2cabc3fe-8bdc-4867-a1a7-fb08ecacc421 · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning,.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations ChartQA: A benchmark for question answering about charts with visual and logical reasoning,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:50.769783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:30:50.156783Z digest=sha256:2de52db8f7201082ee041a2ba7af513ddbe3d1463c850d4ea94c70e4b16c5cea

Observation ed2c8629-63cb-4eb8-99fa-719f34fd5a59 · outbound

This paper cites Chartocr: Data extraction from charts images via a deep hybrid framework,.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations Chartocr: Data extraction from charts images via a deep hybrid framework,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:50.750520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:30:50.163197Z digest=sha256:23d45d0b6e3d899fa3bf48359fabb929557433573b0764cb936d6710bf8ed65f

Observation 1f04fdea-0a22-45d2-8f09-dc1056214568 · outbound

This paper cites Figureseer: Parsing result-figures in research papers,.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations Figureseer: Parsing result-figures in research papers,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:50.732270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:30:50.169804Z digest=sha256:53694747b7875d67f6908987b13d765369d4e74e91b5d6611829e17d0e051630

Observation 31dd110e-10a2-4816-8b41-b976d911a9cd · outbound

This paper cites MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart Derendering.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart Derendering

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:50.176774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:50.176774Z digest=sha256:a4a099f224c60e32ed3227fafb50b2c65fb808ee9ab300a493da7e4fa05a9b88

Observation a97f4161-3543-47fd-b634-f125db8ec85d · outbound

This paper cites Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:50.183976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:50.183976Z digest=sha256:2ec3aeb6849418bdca8044dd9ffd3e68835772c0d59eaf01d303ebc870efc7d9

Observation 1e4e07ce-9a85-40cd-8800-549027aac684 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:50.190566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:50.190566Z digest=sha256:13b1c27865d27a498c76e9bc35e0de39f4ea2282d9989407e95d3f8fd1fc3f92

Observation d72d906a-735f-404f-b5c8-511c5ec8be8a · outbound

This paper cites From Data Quality to Model Quality: an Exploratory Study on Deep Learning.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations From Data Quality to Model Quality: an Exploratory Study on Deep Learning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:30:50.484052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:30:50.197079Z digest=sha256:0bb1d65bb1e119cdf9fb93ded4f043bdaf62b378af700ae92ff71e613315ce01

Observation 1dd429fd-4f64-4059-8f7d-32a4f78c4977 · outbound

This paper cites The Effects of Data Quality on Machine Learning Performance on Tabular Data.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations The Effects of Data Quality on Machine Learning Performance on Tabular Data

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:50.203758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:50.203758Z digest=sha256:ea26a3edfa122177af6305d5816d2809767ab15efbc90f8a879a8152ca20277f

Observation 31036630-9623-4890-adf0-5163f1ba81eb · outbound

This paper cites Matplotlib: A 2d graphics environment,.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations Matplotlib: A 2d graphics environment,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:50.210416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:50.210416Z digest=sha256:5b6b5519eba218934eeb007212d18f006850651bb77b01fccc85ea8b99aecc7b

Observation b7769dce-25e2-40ba-a27e-7164297e5b0d · outbound

This paper cites seaborn: statistical data visualization,.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations seaborn: statistical data visualization,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:50.698321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:30:50.216592Z digest=sha256:164a473ac10a090b66fba46c00e78720e2e2ada9dc0387c862e08f701142447a

Observation 3262c74d-2ae0-4cfc-901c-20b2cb7b9a05 · outbound

This paper cites Icdar 2019 competition on scene text visual question answering,.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations Icdar 2019 competition on scene text visual question answering,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:50.679744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:30:50.222759Z digest=sha256:8cd38a1c232c2042380f469246b3f9ef140a75438c2b7af9da35a7434c62c42f

Observation 52233d0a-fc1d-4624-bafc-5c65cb234b5b · outbound

This paper cites Enhancing large vision language models with self-training on image comprehension,.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations Enhancing large vision language models with self-training on image comprehension,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:50.662218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:30:50.228260Z digest=sha256:b978e0cdb2c2ec26b0c7e828be6dbbb53a2972dc8ac8ca64afc62d03fab1e7c5

Observation 938b0371-4e26-4bdd-a144-bd1136ac483e · outbound

This paper cites Fine-tuning Smaller Language Models for Question Answering over Financial Documents.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations Fine-tuning Smaller Language Models for Question Answering over Financial Documents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:50.234543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:50.234543Z digest=sha256:9cb06f64cf8ad7f981a88ae2c0679f3c99d76220fed24305b68ef2c7dc87ccf4

Observation f55ea03d-7b43-4398-9fc3-be2cc900ec57 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations LLaMA: Open and Efficient Foundation Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:50.240903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:50.240903Z digest=sha256:8b076bf7bd6a661d33ea558d572139ce004d9359d76136b6bc4273af01a5c2e4

Observation 1856b9db-2a53-4c4f-bdc9-144f50168c51 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:50.246368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:50.246368Z digest=sha256:4e841bbceab108808adf0f6085244510c35fb7269a5fb2c21ad6766850e31b50

Observation cc435a35-b579-454f-ae40-23e05aca991c · outbound

This paper cites Pixtral 12B.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations Pixtral 12B

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:50.252461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:50.252461Z digest=sha256:38fb4c61e8434caa4302a76223bc2575bd243a9dd0ba4aeadc1d12510c890600

Observation 6196ccf5-b5b3-4cd2-8c06-24c9e8ffa41c · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:50.258087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:50.258087Z digest=sha256:bfc85e98af3c0ec437c4d8df0056c2d7af7bc69c1d8e7e3683fabeeba112f3ef

Observation 2aef1462-7321-45d8-b103-96c755c070fe · outbound

This paper cites You are a helpful assistant. Help me with my math homework!.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations You are a helpful assistant. Help me with my math homework!

Reference 150

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:50.622858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:30:50.269302Z digest=sha256:25acd7a771861dff525313eaf3b6ddd8fa16f44cdf0a6190ed8e3a0848d64bc8

Observation aa605e70-1963-45d1-9af4-87e472b36bc0 · outbound

This paper cites The difference in value between Reserves and Cash is: -30 - (-20) = -10 Therefore, the difference in Value between Reserves and Cash is -10.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations The difference in value between Reserves and Cash is: -30 - (-20) = -10 Therefore, the difference in Value between Reserves and Cash is -10

Reference 800

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:50.642916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:30:50.263957Z digest=sha256:b667f1dcdc51db0fb8b16415b2600bb2b795b4f2ddc74a98b4f0bd87fd4d5ca8

Observation b2704af8-aea8-4c37-9f35-450423c5d384 · outbound

This paper cites Language Models are Few-Shot Learners.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations Language Models are Few-Shot Learners

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:50.150645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:50.150645Z digest=sha256:b96c725aaf589e455eac9a18ceac7fbb4f6244a8c2d8c8049a1432192420af78

Pith citing papers

Observation 44f8cb2f-8c35-4433-9921-7b254445a636 · inbound

A Multistage Extraction Pipeline for Long Scanned Financial Documents: An Empirical Study in Industrial KYC Workflows cites this paper.

A Multistage Extraction Pipeline for Long Scanned Financial Documents: An Empirical Study in Industrial KYC Workflows Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:51:23.970512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-07T13:43:30.511187Z digest=sha256:edce4c59309ee8c641578766bc9a8159617aa600c0122223e5bd714a0b1d7942