Pith. sign in

Paper Citation Record · LEDGER

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing

As of 10 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2509.04469.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.04469 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:20:26.273267Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact4
  • verified fuzzy13
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a8c3042f-6252-44b7-8a55-71f5ec398a58 · outbound

This paper cites From Theory to Practice: Real-World Use Cases on Trustworthy LLM-Driven Process Modeling, Prediction and Automation.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing From Theory to Practice: Real-World Use Cases on Trustworthy LLM-Driven Process Modeling, Prediction and Automation

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:20:28.336247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T14:20:23.118070Z digest=sha256:a4c8806cc15008e47005d842475f7f4fa206392749bd266a96562bb9129d89d4

Observation 5a4f03fa-640c-45c2-9896-a7fbdbc26a68 · outbound

This paper cites Towards automated auditing with machine learning,.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Towards automated auditing with machine learning,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:20:31.555716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T14:20:23.226947Z digest=sha256:621535446bd9986ce3c79d6cef5071cff4c1bb42ac94f9586e27d392ec40bd53

Observation c881b657-495f-4979-ab75-5612ad701b3c · outbound

This paper cites Advancing risk and quality assurance: A rag chatbot for improved regulatory compliance,.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Advancing risk and quality assurance: A rag chatbot for improved regulatory compliance,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:20:31.274844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T14:20:23.386519Z digest=sha256:70ae10bd9b7b96bd87f997c5f0bb5c5cd62d7bbbc364a2a63cb90370385c1f99

Observation d3fd2493-b4f4-4333-bbf0-38a0d62f60d7 · outbound

This paper cites Fine-tuning large language models for compliance checks,.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Fine-tuning large language models for compliance checks,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:20:31.033404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T14:20:23.486011Z digest=sha256:a9bd6f697763d84d42a507924d14e4aa69150fce9174e8bb5f5151d2c95b35c8

Observation bb6c00df-3c0b-4bb5-9eda-4d0ee44c3118 · outbound

This paper cites An overview of data extraction from invoices,.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing An overview of data extraction from invoices,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:20:30.695448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T14:20:23.624840Z digest=sha256:a9b2e97db38a060ad4e45c10235d2ac6ffebb9614671d25919e594019a15a4bc

Observation a5f28a53-32a9-4176-afea-91452a2ba6fd · outbound

This paper cites A Survey of Deep Learning Approaches for OCR and Document Understanding.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing A Survey of Deep Learning Approaches for OCR and Document Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T14:20:23.707156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:20:23.707156Z digest=sha256:c3703648c6a6103c630fc8ddf353bd45de5350165c13b2e149237c282de91fa0

Observation a577dc6c-28e4-4526-bf59-61180c98fc14 · outbound

This paper cites Deep Learning based Visually Rich Document Content Understanding: A Survey.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Deep Learning based Visually Rich Document Content Understanding: A Survey

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T14:20:23.881572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:20:23.881572Z digest=sha256:4d03f5a9bd9300e066b99811c27701c3af1617ca2242ed95b682b1e9e8b8e8d2

Observation 3f2573d4-a449-446b-831a-a0bee46287a9 · outbound

This paper cites Memory-augmented agent training for business document understanding,.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Memory-augmented agent training for business document understanding,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:20:30.461236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T14:20:24.050023Z digest=sha256:137c42775a9618f8e0ebee347b0374d2e2730f55fdce55d2aba4fd1df891e7fb

Observation a025c8de-59a3-43eb-bcae-a02bba2427cf · outbound

This paper cites Wonderbread: A benchmark for evaluating multimodal foundation models on business process management tasks,.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Wonderbread: A benchmark for evaluating multimodal foundation models on business process management tasks,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:20:30.301934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T14:20:24.224802Z digest=sha256:940fe603101552595ee58c9cf2d03ddda70d5e63a9ec14c183226a76ed89622a

Observation 8f82b2bb-18c4-4e96-9a2a-b11a4d3fad16 · outbound

This paper cites MMLONGBENCH-DOC: Benchmarking long-context document understanding with visualizations,.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing MMLONGBENCH-DOC: Benchmarking long-context document understanding with visualizations,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:20:30.148048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T14:20:24.436236Z digest=sha256:686cd0907e1fc491aa85090368f4ac45c7943269ca5bddda5bc3b2673cb7f582

Observation 5b09afa9-9cd9-4903-9a5a-e813cffb0543 · outbound

This paper cites M-longdoc: A benchmark for multimodal super- long document understanding and a retrieval-aware tuning framework,.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing M-longdoc: A benchmark for multimodal super- long document understanding and a retrieval-aware tuning framework,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:20:29.998179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T14:20:24.581017Z digest=sha256:134f6a551cca7737dd22aae37c23b3133efde02c4de6db1be3afbb1ecdfeb93c

Observation a7068012-6c16-455a-ba9a-a227db41d7bb · outbound

This paper cites DOCBENCH: A Benchmark for Evaluating LLM-based Document Reading Systems.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing DOCBENCH: A Benchmark for Evaluating LLM-based Document Reading Systems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T14:20:24.762308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:20:24.762308Z digest=sha256:f3088c78e320c1a713911582d5454370dc6231017e168b76372f37a6887d2ba2

Observation 463c174f-422f-4bf7-8a24-ff1346cb8e08 · outbound

This paper cites Gemma 3 Technical Report.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Gemma 3 Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T14:20:24.883307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:20:24.883307Z digest=sha256:deb0dc4377cbb1a049938236473807826f1dd9a5660969782d4c7b95479cf407

Observation 63cf2eeb-754b-40e3-829e-c38033740bdd · outbound

This paper cites M-Longdoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing M-Longdoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T14:20:24.672980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:20:24.672980Z digest=sha256:0a36eb798f066732c46d3edefe5dde7efa572c6ab6bae11c31c0fdc1d3924951

Observation 0b8afe44-8a1a-46f0-bb21-6192ce1ac89a · outbound

This paper cites SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T14:20:25.170342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:20:25.170342Z digest=sha256:6bf53f1f6935ec90a76e80b87edbb968baadea505908506f18d3bb9dd5789e70

Observation 63da7747-6544-42c2-b6f3-28763bbb382e · outbound

This paper cites Pdf data extraction benchmark 2025: Comparing docling, unstructured, and llamaparse for document processing pipelines,.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Pdf data extraction benchmark 2025: Comparing docling, unstructured, and llamaparse for document processing pipelines,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:20:29.714799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T14:20:25.310390Z digest=sha256:e6c26d69ba996888ef75f404cde951c7da21d734532588ccfe50c36ef6caf357

Observation d89c6c31-da03-41d0-9765-b874d46061af · outbound

This paper cites Docling Technical Report.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Docling Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T14:20:25.043109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:20:25.043109Z digest=sha256:a53ee72b0e49470e977fe189f1055e334d54e8a22354bb6e4955d2955571a03a

Observation 28516716-8c73-4d53-8e84-be9d4e3feef8 · outbound

This paper cites Icdar 2019 robust reading challenge on scanned receipts ocr and information extraction,.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Icdar 2019 robust reading challenge on scanned receipts ocr and information extraction,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:20:29.494652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T14:20:25.574862Z digest=sha256:7dd91829185887a03972b343221d78b7b40836d6e9d0754eecb5e3eb1257a517

Observation d495aad1-eb43-4047-83c5-ac40093a04a4 · outbound

This paper cites Field extraction from forms with unlabeled data,.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Field extraction from forms with unlabeled data,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:20:29.253485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T14:20:25.713563Z digest=sha256:cecb8c5822e3686298ba6ef81f15466c83369f67bdccfbe834a98206838db489

Observation 420f57b5-134c-4055-91b6-f8da0a8e243e · outbound

This paper cites Ocr-free document understanding transformer,.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Ocr-free document understanding transformer,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T14:20:25.424347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:20:25.424347Z digest=sha256:23f459da3cecc119f7e35e8bdb55ce2926f0f95ed7b60301debabe054778e420

Observation 8f899c6e-b581-4a55-bb0e-b5dc7e5d3b86 · outbound

This paper cites Layoutlm: Pre-training of text and layout for document image understanding,.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Layoutlm: Pre-training of text and layout for document image understanding,

Reference 21

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-05T14:20:27.117317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T14:20:25.924804Z digest=sha256:24646ca0b1da29e87e246f7c148687c325a12bca9da5f7ac2fb2bf99ba78f6bf

Observation cd29b14b-566c-4a13-8616-c2885530940f · outbound

This paper cites Layoutlmv3: Pre-training for document ai with unified text and image masking,.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Layoutlmv3: Pre-training for document ai with unified text and image masking,

Reference 22

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-05T14:20:26.734768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T14:20:26.035589Z digest=sha256:4f29b26eb8d066d822f73eb3eaf781ee5ba6f5ccb27e32075c07f915fbd1084c

Observation 8f6fbf9d-d63a-42fb-8116-d30b7f2a2fe9 · outbound

This paper cites Truth tobacco industry documents (formerly legacy tobacco documents library),.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Truth tobacco industry documents (formerly legacy tobacco documents library),

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:20:29.089939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T14:20:25.840205Z digest=sha256:066fb1fe5b0edf93fe702558c036a7800e26709ed7c5194e8af0dd72fdb86999

Observation 1453c3ad-843d-464a-9243-9ea03ac71c51 · outbound

This paper cites Lilt: A simple yet effective language- independent layout transformer for structured document understanding,.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Lilt: A simple yet effective language- independent layout transformer for structured document understanding,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:20:28.848561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T14:20:26.134753Z digest=sha256:143cb9450a9bcf54839fc2c7391ee16c4e5f85189d7fed43401e775edff95330

Observation 4c7db77f-c7e1-43c0-9a40-0aa40ab351d7 · outbound

This paper cites Available: https://doi.org/10.1145/3342558.3345421.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Available: https://doi.org/10.1145/3342558.3345421

Reference 2019

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-05T14:20:28.028635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T14:20:23.281575Z digest=sha256:a817b8f8a19219ca208fbdf7c379a1b1b7a0524e88ef6a56b8a63bfcaae38d16

Observation 9fb6ba9c-b98d-43be-a961-0e4f1390640f · outbound

This paper cites an unresolved cited work.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:20:28.634779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T14:20:26.273267Z digest=sha256:ce4e414072a46875f62e4187031dd8789a28c910c323562de3f5e5a828f04499

Observation 21ee41fd-fcae-48f7-a742-f5ac024f7e33 · outbound

This paper cites WONDERBREAD: A Benchmark for Evaluating Multimodal Foundation Models on Business Process Management Tasks.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing WONDERBREAD: A Benchmark for Evaluating Multimodal Foundation Models on Business Process Management Tasks

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T14:20:24.320091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:20:24.320091Z digest=sha256:b2b49fda1d59b3f91bb7f5f36f2a9b5bc2584416e945b51b2a9a8859a55abbd2

Pith citing papers

No inbound Pith citation observations are available.