Pith. sign in

Paper Citation Record · LEDGER

DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model

As of 17 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2504.17315.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.17315 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:47:47.562309Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact2
  • verified fuzzy13
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f45c9075-195d-4c54-89cf-42b4ed02a5af · outbound

This paper cites LayoutDIT: Layout-aware end-to-end document image translation with multi-step conductive decoder[C]//Findings of the Associa- tion for Computational Linguistics: EMNLP 2023.

DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model LayoutDIT: Layout-aware end-to-end document image translation with multi-step conductive decoder[C]//Findings of the Associa- tion for Computational Linguistics: EMNLP 2023

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:47.980404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:47.447418Z digest=sha256:5e479867d04a0552f1f96ca823f8859d740dff47e6ea40f64a2dda863202a431

Observation e4155984-a997-4684-979b-d3dad6273107 · outbound

This paper cites an unresolved cited work.

DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:47:47.965980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:47.452703Z digest=sha256:6b8f65375a794e69ec30e02aad5fc98dff562652949b8db765b2b8fd833e16bb

Observation 825a9231-da36-49e8-a801-4d6f47f70022 · outbound

This paper cites Document Image Analysis for Text Extraction and Trans- lation.

DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model Document Image Analysis for Text Extraction and Trans- lation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:47.951691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:47.456905Z digest=sha256:878724da41a9323c1d3e24d85bfcc1404ac2cea0101605c24f1c91f37a5304de

Observation 403d075b-e023-4843-9ea0-0f5399dc2fd7 · outbound

This paper cites Deep learning.

DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model Deep learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:47.461312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:47.461312Z digest=sha256:56e35ee2ab7364ed374a231a30604f5699c99b6ffe03f4387f75585fba04f0f3

Observation fcee8566-22ae-4eed-80c1-6355d9b72821 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:47.465781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:47.465781Z digest=sha256:57870346916f144aef9feab0d69577a48f3a803acdfcaddc65348846093451b6

Observation 46af64b2-b8e3-4125-917a-c519bb5b47df · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:47.470788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:47.470788Z digest=sha256:b9245134dc93dea32ba24459edf5c0101523b13c558fb6dca6edaf1460531104

Observation 97bbf57a-3532-4290-8a5d-9a83750c5f29 · outbound

This paper cites A survey on multi-task learning.

DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model A survey on multi-task learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:47.927089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:47.476033Z digest=sha256:223525055c0a58a9738a8f30d3083835759f74daafca884ef9d24c7f27a7cd7d

Observation 6935d653-19f3-46aa-a7e2-69f1ea800afa · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model Chain-of-thought prompting elicits reasoning in large language models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:47.480190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:47.480190Z digest=sha256:7f57ad8bf511cac2dd8615ba5e19265b6ba6abea79099ede34756af0d5d93f3b

Observation 2a25e677-5e52-4639-a666-8b5e28b382ce · outbound

This paper cites An Empirical Study of Translation Hypothesis Ensembling with Large Language Models.

DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model An Empirical Study of Translation Hypothesis Ensembling with Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:47.484285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:47.484285Z digest=sha256:428e3faaaf981691dc5745a72ffd8406d151474b773a9a2b45e830697bb534ef

Observation 604153f1-8fca-492e-99b6-6e792bc5b711 · outbound

This paper cites Optical character recogni- tion.

DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model Optical character recogni- tion

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:47.904388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:47.488580Z digest=sha256:cddbba6c058947e486889d5f43d69bf1c4be0280189f848e0f803cc1351382d2

Observation 3605f8bb-c856-495f-8e40-a0513dfb2c8c · outbound

This paper cites A survey on evaluation of large language models.

DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model A survey on evaluation of large language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:47.890502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:47.492420Z digest=sha256:642874b57f155b7f1e3e5beeb208e0e6950d6e203d1fcefd0ca2810097770b59

Observation 7219d817-b958-442c-91da-defe2cc6d403 · outbound

This paper cites ChatGPT for good? On opportunities and challenges of large language models for education.

DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model ChatGPT for good? On opportunities and challenges of large language models for education

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:47.876496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:47.496652Z digest=sha256:5e6eb777b0fd4ca751602884098c49cd872d64c3e977a726e8fae0b7ccfc89a6

Observation 9e75d4c4-e974-4390-91f0-802c0112428e · outbound

This paper cites Continual Pre-Training of Large Language Models: How to (re)warm your model?.

DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model Continual Pre-Training of Large Language Models: How to (re)warm your model?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:47.500876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:47.500876Z digest=sha256:50590aa2cc6858ec00f38f2db85e4dd683181bd40d99307242f166ad928fa18a

Observation a82b8e21-60a5-4ff1-880f-1df157a127d0 · outbound

This paper cites How Abilities in Large Language Models are Affected by Supervised Fine-tuning Data Composition.

DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model How Abilities in Large Language Models are Affected by Supervised Fine-tuning Data Composition

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:47.505387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:47.505387Z digest=sha256:a6770fad6840f575bcc3cb007c01f3948af67d77da3089ebcdcf8cf9b7b80d3b

Observation 5db83e95-d9e0-4a17-823f-44afd0d807b7 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:47.510015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:47.510015Z digest=sha256:aaffc7cfc112f257b7f8b10598fc23f7d33b7be81a677b242d1fa066bafcdee9

Observation 455e59ca-f62a-49eb-bf65-92a0ee4aa882 · outbound

This paper cites Qwen2.5-VL Technical Report.

DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model Qwen2.5-VL Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:47.514354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:47.514354Z digest=sha256:2d566604cbd7c64e56024ab8f4208b14c6daf0f9956c11f762e4a7ef34f52f89

Observation ed1e2735-4177-4506-8001-bc0f25518fe7 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model Gemini: A Family of Highly Capable Multimodal Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:47.518908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:47.518908Z digest=sha256:07a94ea16887ac0adab43fedaf81b75fdfd6c831aaca2c5f5c3919791e67c91d

Observation 42ea20de-f230-48a4-98e1-9f7793733e77 · outbound

This paper cites Perspectives and prospects on transformer architecture for cross-modal tasks with language and vision.

DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model Perspectives and prospects on transformer architecture for cross-modal tasks with language and vision

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:47.860813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:47.523484Z digest=sha256:3b5524ff19b610c1c804d542351740a4e9af0fac01112f2b35f7870db7022626

Observation 8c609b1d-26df-44bd-b872-b7cd8935c92b · outbound

This paper cites Machine learning paradigms.

DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model Machine learning paradigms

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:47.845366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:47.527662Z digest=sha256:bcfc7f00f0378f1b3b47d97db48575f9fe278318ff971e41dafec6011abdc55e

Observation 7ebbbc20-7f37-42dd-bf89-8e0205be7685 · outbound

This paper cites A comparison of multi-task learning and single-task learning approaches.

DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model A comparison of multi-task learning and single-task learning approaches

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:47.829178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:47.531891Z digest=sha256:10f9c65242c0f301e37ec4b7393bd1bd438f025a239d8b90d973b799e3f44fc5

Observation 591fc156-4013-4f24-a126-e614b9de4bf4 · outbound

This paper cites HW-TSC’s Submission to the CCMT 2024 Machine Transla- tion Tasks.

DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model HW-TSC’s Submission to the CCMT 2024 Machine Transla- tion Tasks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:47.814913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:47.536110Z digest=sha256:bf3d74f592e712e17fc1414a763eee186a43ec9af9d27ba6c4fedac038938dbb

Observation 4aaa53a8-1cb9-4ade-a0e3-830027787d9e · outbound

This paper cites Choose the Final Translation from NMT and LLM hypotheses Using MBR Decoding: HW-TSC's Submission to the WMT24 General MT Shared Task.

DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model Choose the Final Translation from NMT and LLM hypotheses Using MBR Decoding: HW-TSC's Submission to the WMT24 General MT Shared Task

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-16T10:47:47.625335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:47.540546Z digest=sha256:129372cdfe5a674fb723afe29aa82a4cb7ce0bea6ca1b5903df420f64536bab5

Observation 26624a39-6366-49fb-9bca-76a5b8a4d557 · outbound

This paper cites ImprovingtheQualityofIWLST2024CascadeOfflineSpeech Translation and Speech-to-Speech Translation via Translation Hypothesis Ensem- bling with NMT models and Large Language Models.

DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model ImprovingtheQualityofIWLST2024CascadeOfflineSpeech Translation and Speech-to-Speech Translation via Translation Hypothesis Ensem- bling with NMT models and Large Language Models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:47.800247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:47.544972Z digest=sha256:c2ea2592c98925befbc809c18dc133863bce1954d12adf92bfb21e6da075e3b7

Observation 707e9544-7332-4711-b299-ca1099d5b060 · outbound

This paper cites COMET: A Neural Framework for MT Evaluation.

DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model COMET: A Neural Framework for MT Evaluation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:47.784757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:47.549272Z digest=sha256:ae11b4589f725749d085c026c7b965976643d1ba545b7791ef499d51e8f791a3

Observation c1609d2a-32e7-4371-8a1f-424cf017e222 · outbound

This paper cites COMET-22: Unbabel-IST 2022 submission for the met- rics shared task.

DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model COMET-22: Unbabel-IST 2022 submission for the met- rics shared task

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:47.769912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:47.553556Z digest=sha256:ec62efce541ab32b66378e5a7ff8abe1d8573c60acdc763bbad1a2646e7062e9

Observation 0a1690c7-7106-444c-911d-b7a9bf81133a · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model Bleu: a method for automatic evaluation of machine translation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:47.558093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:47.558093Z digest=sha256:e0a6b6ba9b4eaefe0a96f826de5675f7161ad07698bd798c446f986d5aab8afd

Observation 9a02fca6-6b1b-44a6-b0a7-992664cae950 · outbound

This paper cites Evaluating Menu OCR and Translation: A Benchmark for Aligning Human and Automated Evaluations in Large Vision-Language Models.

DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model Evaluating Menu OCR and Translation: A Benchmark for Aligning Human and Automated Evaluations in Large Vision-Language Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-16T10:47:47.604611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:47.562309Z digest=sha256:f4d9732fd7183b9e89ed066bce35fca4fe430ca350e7493ba6100f8b37d4b659

Pith citing papers

No inbound Pith citation observations are available.