Pith. sign in

Paper Citation Record · LEDGER

MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2405.11985.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.11985 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T18:14:20.479979Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

5
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2c0cffd4-9796-468c-9105-a91c41822bbc · inbound

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? cites this paper.

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:59:32.703331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T07:59:32.638758Z digest=sha256:c8007167d7e665dab40a5e8b21e5aa5a4ba41ea94aa94c3e440953625103c18e

Observation 0bc929f0-7cd8-4d2c-a166-8b4f02b75fa7 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 228

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.036152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:47487d557ff5614e8d2591c2f0bdb7374857a89ccec218220cfa1485401ec5b6

Observation d65ed6ca-8fed-4637-9dbd-badccf0998b8 · inbound

Cross-Modal Synergies: Unveiling the Potential of Motion-Aware Fusion Networks in Handling Dynamic and Static ReID Scenarios cites this paper.

Cross-Modal Synergies: Unveiling the Potential of Motion-Aware Fusion Networks in Handling Dynamic and Static ReID Scenarios MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T18:14:20.479979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:14:20.479979Z digest=sha256:d1cbd88955e3aea28f2668222a97248f9883ac4d80d62b4714d18fbed226ce36

Observation ebb1b6e8-f7fb-468b-9459-b4bfb34acc37 · inbound

Trajectory Map-Matching in Urban Road Networks Based on RSS Measurements cites this paper.

Trajectory Map-Matching in Urban Road Networks Based on RSS Measurements MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-09T15:56:19.626881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:56:19.626881Z digest=sha256:27716c21d79cd7f7cafe8004cc11da0300e58a064d5789b49864f4b7aeb74ccc

Observation 38269d09-3bca-4cbd-8bb8-73503b83363f · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:53:26.122658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:ccffd776850b7099fdd450b63a4b7b63c1033164f7ca44039c5ffef3ea551c5e

Observation 61cd1887-8c50-4ebe-85e3-3ea3c9b7988c · inbound

Qwen2.5-VL Technical Report cites this paper.

Qwen2.5-VL Technical Report MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:25:19.101037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T02:25:04.405036Z digest=sha256:b1f0c531ffe064e2241df4b7b285464c0adad5be25142a6d6681563abe5fdc43

Observation b33f1460-4f7d-43ab-8b0d-7e224e99433d · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.096968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:0fa4fbdf6a6138cebdf010998c1d84da737d189bd651d44a280ff74fc27b96b5

Observation 5c42ea2b-b3de-4070-8fb4-ef412deb4b9d · inbound

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting cites this paper.

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:27.648631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:27.648631Z digest=sha256:d1465697e3b5ead0dad45ba286e9f758daf2fa07d8fc378aac8e421e65db0fbe

Observation 8f0e8020-f461-475d-83dd-13ddd401c8bc · inbound

Chart-to-Experience: Benchmarking Multimodal LLMs for Predicting Experiential Impact of Charts cites this paper.

Chart-to-Experience: Benchmarking Multimodal LLMs for Predicting Experiential Impact of Charts MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:09.826103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:09.826103Z digest=sha256:67b01197a3e3d413a9f7f49505bc20756bf32c4ef6e5e9e2945534ae096df46a

Observation 03f0d185-6e7d-4775-a7bf-4f27119e219c · inbound

Multilingual Multimodal Software Developer for Code Generation cites this paper.

Multilingual Multimodal Software Developer for Code Generation MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:58.185565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:58.185565Z digest=sha256:025708db3a10ccd58916580f413466ff15642de887200a33101219fcf7a13749

Observation 4ddae192-bd41-49eb-aed9-21f4a513d996 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 119

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:58.775085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:ac72353c643aede371cdb1e247ef3440cea622f0b2c13b92af8c413ddb88b111

Observation fa77b7a4-1134-4c0f-9ab0-63b9d36b2162 · inbound

LaV-CoT: Language-Aware Visual CoT with Multi-Aspect Reward Optimization for Real-World Multilingual VQA cites this paper.

LaV-CoT: Language-Aware Visual CoT with Multi-Aspect Reward Optimization for Real-World Multilingual VQA MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:56:42.310528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T17:53:08.136677Z digest=sha256:27a9b13e756e84b70c5f7d739bffd7c52a06f654f3cb96e18a8ddb00f9d7011f

Observation 34fb4af9-5580-4886-ad1e-d1a63464aaad · inbound

Multilingual Vision-Language Models, A Survey cites this paper.

Multilingual Vision-Language Models, A Survey MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 138

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:02:37.275213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T13:02:08.000814Z digest=sha256:7349cfcbc184f883288624a07d319bbe4266349f40060e2fb46ecc38a2ed7501

Observation 9f3b2e45-ec2a-407a-b6f0-eaddc4eb220f · inbound

LPCAN: Lightweight Pyramid Cross-Attention Network for Rail Surface Defect Detection Using RGB-D Data cites this paper.

LPCAN: Lightweight Pyramid Cross-Attention Network for Rail Surface Defect Detection Using RGB-D Data MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T10:43:50.130815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:43:50.130815Z digest=sha256:d54f5559855e7b1de13e08f7ec3afb961dda1c508fd20dfcaa4187a90a39a322

Observation 111207c7-2583-48df-907a-babeb134fe9c · inbound

Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method cites this paper.

Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T10:43:42.114265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:43:42.114265Z digest=sha256:984be6b941e1e7bb5685435c0f3ca8e43a15c54f13a134e07f05435391e2ed4e

Observation cf8efb61-fa51-4edd-884f-b7ca1d56c4e0 · inbound

Hierarchical Awareness Adapters with Hybrid Pyramid Feature Fusion for Dense Depth Prediction cites this paper.

Hierarchical Awareness Adapters with Hybrid Pyramid Feature Fusion for Dense Depth Prediction MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:03:12.256459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T20:01:57.029143Z digest=sha256:2b21f39c4a8467401118b90662b887d3eb0dcc285ffdd6d418375afb0aef7b7b

Observation cfb4d057-5666-4f78-a646-a46b32205dc3 · inbound

Multilingual Training and Evaluation Resources for Vision-Language Models cites this paper.

Multilingual Training and Evaluation Resources for Vision-Language Models MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:38:43.165665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T05:11:00.461167Z digest=sha256:ee7828ef40e4580661df10173aa0dc8683f9c328ceab840d6f61a3a039abe516

Observation d132e84b-6c28-4475-8e41-eb619ea4a887 · inbound

Multilingual Training and Evaluation Resources for Vision-Language Models cites this paper.

Multilingual Training and Evaluation Resources for Vision-Language Models MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-07-05T12:30:59.869841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:09cd89d9268ea54778fe21988a45ea63d71fbcce7895d89d52c5d34be4421e99

Observation ed168971-597e-4e1d-b4c1-736851e055d9 · inbound

Do You Need Text Rectification? Soft Attention Mask Embedding for Rectification-Free Scene Text Spotting cites this paper.

Do You Need Text Rectification? Soft Attention Mask Embedding for Rectification-Free Scene Text Spotting MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:33:14.576013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T11:28:51.711892Z digest=sha256:91adbcf9999a3427f39cb158803a88db61e163945d16b1ddef0398c234717216

Observation f45e8b3e-a9b3-4e73-9033-66b690564a28 · inbound

DAIN: Dynamic Agent-Based Interaction Network for Efficient and Collaborative Multimodal Reasoning cites this paper.

DAIN: Dynamic Agent-Based Interaction Network for Efficient and Collaborative Multimodal Reasoning MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:04:21.359361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T06:01:20.803078Z digest=sha256:f7f983ec2e9816f7660b2380f56493e30a02e188c399f89d3d57ab3e2a21567c