Pith. sign in

Paper Citation Record · LEDGER

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations

As of 14 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 1 inbound Pith citation observation for arXiv:2501.04675.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.04675 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:30:50.269302Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-07T13:43:30.511187Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T08:51:23.968119Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact1
  • verified fuzzy8
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cc8690f2-ac2b-4c1a-a9c9-54e16d2953be · outbound

This paper cites DePlot: One-shot visual language reasoning by plot-to-table translation.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations DePlot: One-shot visual language reasoning by plot-to-table translation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:50.137954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:50.137954Z digest=sha256:d4201a0e89f10d2dc46e189f13d2c6c136079a27e3a953f36fc7e9d54523f6cf

Observation c065acc8-9347-4b1d-91de-a03a54e68b49 · outbound

This paper cites Language models are few-shot learners,.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations Language models are few-shot learners,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:50.144427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:50.144427Z digest=sha256:5673e3863a37c87eeffa89cc234cb970e61248b69fcff64cd2221bef23fd2b00

Observation 2cabc3fe-8bdc-4867-a1a7-fb08ecacc421 · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning,.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations ChartQA: A benchmark for question answering about charts with visual and logical reasoning,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:50.769783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:50.156783Z digest=sha256:20876840965093fa000bf37053dd0fb89849c7ac27fc558f0c47c78e9aea4b04

Observation ed2c8629-63cb-4eb8-99fa-719f34fd5a59 · outbound

This paper cites Chartocr: Data extraction from charts images via a deep hybrid framework,.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations Chartocr: Data extraction from charts images via a deep hybrid framework,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:50.750520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:50.163197Z digest=sha256:5091c9dbbe1b1f15ef3dccd6c377479e418ea729b898ae10c42f0b65c2427ece

Observation 1f04fdea-0a22-45d2-8f09-dc1056214568 · outbound

This paper cites Figureseer: Parsing result-figures in research papers,.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations Figureseer: Parsing result-figures in research papers,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:50.732270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:50.169804Z digest=sha256:4863302a6b4aca522583352fbc06c052399bf5996dbb1f14a562fdbfb83d57b7

Observation 31dd110e-10a2-4816-8b41-b976d911a9cd · outbound

This paper cites MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart Derendering.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart Derendering

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:50.176774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:50.176774Z digest=sha256:a4a099f224c60e32ed3227fafb50b2c65fb808ee9ab300a493da7e4fa05a9b88

Observation a97f4161-3543-47fd-b634-f125db8ec85d · outbound

This paper cites Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:50.183976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:50.183976Z digest=sha256:2ec3aeb6849418bdca8044dd9ffd3e68835772c0d59eaf01d303ebc870efc7d9

Observation 1e4e07ce-9a85-40cd-8800-549027aac684 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:50.190566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:50.190566Z digest=sha256:13b1c27865d27a498c76e9bc35e0de39f4ea2282d9989407e95d3f8fd1fc3f92

Observation d72d906a-735f-404f-b5c8-511c5ec8be8a · outbound

This paper cites From Data Quality to Model Quality: an Exploratory Study on Deep Learning.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations From Data Quality to Model Quality: an Exploratory Study on Deep Learning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:30:50.484052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:50.197079Z digest=sha256:b18ba8542b1722e7485cd88a9dd675522c8e0ecfac11e2bfe353f9838bf0a647

Observation 1dd429fd-4f64-4059-8f7d-32a4f78c4977 · outbound

This paper cites The Effects of Data Quality on Machine Learning Performance on Tabular Data.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations The Effects of Data Quality on Machine Learning Performance on Tabular Data

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:50.203758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:50.203758Z digest=sha256:3ca3791179590c32100f3b34b704232932f683452031c570de59a82246e701d6

Observation 31036630-9623-4890-adf0-5163f1ba81eb · outbound

This paper cites Matplotlib: A 2d graphics environment,.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations Matplotlib: A 2d graphics environment,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:50.210416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:50.210416Z digest=sha256:5b6b5519eba218934eeb007212d18f006850651bb77b01fccc85ea8b99aecc7b

Observation b7769dce-25e2-40ba-a27e-7164297e5b0d · outbound

This paper cites seaborn: statistical data visualization,.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations seaborn: statistical data visualization,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:50.698321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:50.216592Z digest=sha256:7505c838bac464c3333170efe96bf7d4d42799459cce613d5d92eed2fe572b42

Observation 3262c74d-2ae0-4cfc-901c-20b2cb7b9a05 · outbound

This paper cites Icdar 2019 competition on scene text visual question answering,.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations Icdar 2019 competition on scene text visual question answering,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:50.679744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:50.222759Z digest=sha256:2b446ce7efa31f1168d01ecf3e8cdcc60d451f2fa0de0bf80d5df2756aeb4f45

Observation 52233d0a-fc1d-4624-bafc-5c65cb234b5b · outbound

This paper cites Enhancing large vision language models with self-training on image comprehension,.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations Enhancing large vision language models with self-training on image comprehension,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:50.662218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:50.228260Z digest=sha256:380ab5ae63fba107cb8259d1f56593086d6a806912246e30d5fadccc08b6dc5c

Observation 938b0371-4e26-4bdd-a144-bd1136ac483e · outbound

This paper cites Fine-tuning Smaller Language Models for Question Answering over Financial Documents.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations Fine-tuning Smaller Language Models for Question Answering over Financial Documents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:50.234543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:50.234543Z digest=sha256:9cb06f64cf8ad7f981a88ae2c0679f3c99d76220fed24305b68ef2c7dc87ccf4

Observation f55ea03d-7b43-4398-9fc3-be2cc900ec57 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations LLaMA: Open and Efficient Foundation Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:50.240903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:50.240903Z digest=sha256:8b076bf7bd6a661d33ea558d572139ce004d9359d76136b6bc4273af01a5c2e4

Observation 1856b9db-2a53-4c4f-bdc9-144f50168c51 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:50.246368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:50.246368Z digest=sha256:4e841bbceab108808adf0f6085244510c35fb7269a5fb2c21ad6766850e31b50

Observation cc435a35-b579-454f-ae40-23e05aca991c · outbound

This paper cites Pixtral 12B.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations Pixtral 12B

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:50.252461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:50.252461Z digest=sha256:38fb4c61e8434caa4302a76223bc2575bd243a9dd0ba4aeadc1d12510c890600

Observation 6196ccf5-b5b3-4cd2-8c06-24c9e8ffa41c · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:50.258087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:50.258087Z digest=sha256:bfc85e98af3c0ec437c4d8df0056c2d7af7bc69c1d8e7e3683fabeeba112f3ef

Observation 2aef1462-7321-45d8-b103-96c755c070fe · outbound

This paper cites You are a helpful assistant. Help me with my math homework!.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations You are a helpful assistant. Help me with my math homework!

Reference 150

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:50.622858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:50.269302Z digest=sha256:880d3c9e2e0d21843618e49b1b1957cf1de7f0b3c094675cbcdd650477424356

Observation aa605e70-1963-45d1-9af4-87e472b36bc0 · outbound

This paper cites The difference in value between Reserves and Cash is: -30 - (-20) = -10 Therefore, the difference in Value between Reserves and Cash is -10.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations The difference in value between Reserves and Cash is: -30 - (-20) = -10 Therefore, the difference in Value between Reserves and Cash is -10

Reference 800

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:50.642916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T21:30:50.263957Z digest=sha256:4cadb9034b32901ee5ad4e2f6aee8ec5e5eea7035ab5b2302a479ec94f3a897f

Observation b2704af8-aea8-4c37-9f35-450423c5d384 · outbound

This paper cites Language Models are Few-Shot Learners.

Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations Language Models are Few-Shot Learners

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:50.150645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:50.150645Z digest=sha256:b96c725aaf589e455eac9a18ceac7fbb4f6244a8c2d8c8049a1432192420af78

Pith citing papers

Observation 44f8cb2f-8c35-4433-9921-7b254445a636 · inbound

A Multistage Extraction Pipeline for Long Scanned Financial Documents: An Empirical Study in Industrial KYC Workflows cites this paper.

A Multistage Extraction Pipeline for Long Scanned Financial Documents: An Empirical Study in Industrial KYC Workflows Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:51:23.970512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-07T13:43:30.511187Z digest=sha256:fe7d5eb19ac77f2cac4f021037fba1aa024b6ddac778082aaa35a4eb05474b6e