Pith. sign in

Paper Citation Record · LEDGER

Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution

As of 22 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 1 inbound Pith citation observation for arXiv:2505.10921.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.10921 v2

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:04:46.282592Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T14:55:22.057166Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T15:03:31.569140Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ed1fdcbe-0717-44a3-bff5-911cbcc95db2 · outbound

This paper cites In: Proceedings of the IEEE international confer- ence on computer vision.

Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution In: Proceedings of the IEEE international confer- ence on computer vision

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:04:46.174700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:04:46.174700Z digest=sha256:414c3d29f162b04e9520aa5b6fd9731429e2c1e6a064ea86db485e4951ad91fc

Observation c568d6a4-687d-43a5-a764-8d3378579d01 · outbound

This paper cites ACM Journal on Computing and Cultural Heritage17(1), 1–20 (2024).

Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution ACM Journal on Computing and Cultural Heritage17(1), 1–20 (2024)

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:04:46.719915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:04:46.181205Z digest=sha256:2de9a5c931904cd9618e1fc44c4b25caa44e41c725d577e3c1da589f7d63aa69

Observation d2ceae6d-4425-4b9f-a81b-91a5b6d656dd · outbound

This paper cites In: Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval.

Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution In: Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:04:46.703613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:04:46.187517Z digest=sha256:0874c69985c8710b716bd6c5cef1f4267df1f62dd73b75f29b57ee401ec337b2

Observation b004b617-831b-42dd-ae05-4b021818db7d · outbound

This paper cites AltCLIP: Altering the Language Encoder in CLIP for Extended Language Capabilities.

Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution AltCLIP: Altering the Language Encoder in CLIP for Extended Language Capabilities

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:04:46.192504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:04:46.192504Z digest=sha256:5df999bd2c6f924113d68ce27736dcb69469ad1ef08ca26179d6abde3f691122

Observation 5461cb29-daaf-4d55-98c1-88cca2c35658 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:04:46.197820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:04:46.197820Z digest=sha256:122e564598630b993f4fa08ffd57117e7ee3e7b613cdd56d2f57988cf7c399ef

Observation 1d66c9d2-d027-4738-8302-cb93984e3664 · outbound

This paper cites Advances in Neural Information Processing Systems35, 26418–26431 (2022).

Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution Advances in Neural Information Processing Systems35, 26418–26431 (2022)

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:04:46.686136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:04:46.203373Z digest=sha256:2817d29f7c280c2dd60d7981e11e7f86e43e8a9bf3019eac4f3a1fca45edd705

Observation 503c4793-71eb-47e1-8424-5e1147d58485 · outbound

This paper cites GPT-4o System Card.

Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution GPT-4o System Card

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:04:46.209026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:04:46.209026Z digest=sha256:c7eb193d55b34f85134595904df1d07ef73f41c828664311df707327b4e19737

Observation ce23eee7-3b19-4225-a78f-a1bc5e230b4a · outbound

This paper cites In: International conference on document analysis and recognition.

Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution In: International conference on document analysis and recognition

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:04:46.666776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:04:46.214233Z digest=sha256:b7399d3d32b81246c8240143d1bbeef5c42295b46eea038bf4985fa05b8b1a56

Observation 20fb0435-3751-4ec4-8c65-e6cb8b70773d · outbound

This paper cites In: Pro- ceedings of the 25th ACM international conference on Multimedia.

Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution In: Pro- ceedings of the 25th ACM international conference on Multimedia

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:04:46.650239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:04:46.218560Z digest=sha256:65a0ce62d2e7a8580341efb698c4253a15c2570b91accb636b98dc04cdeb06e4

Observation 21974551-5459-4aca-aaa7-1ed61b4f807f · outbound

This paper cites IEEE Transactions on Multimedia 21(9), 2347–2360 (2019).

Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution IEEE Transactions on Multimedia 21(9), 2347–2360 (2019)

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:04:46.633889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:04:46.222332Z digest=sha256:2ea837e4874f11da3eb01db7a7aeef2d5b7c5b3bec5bb9c58734db5ffef6c76d

Observation e21b32e2-0300-45ff-a60c-1410e23b82ef · outbound

This paper cites Advances in Neural Information Processing Systems 37, 1141–1161 (2024).

Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution Advances in Neural Information Processing Systems 37, 1141–1161 (2024)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:04:46.615142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:04:46.226150Z digest=sha256:7b765b91ae24abda6048a2bf228e16723c26fb2f7115723bc7345af074225bf9

Observation bd7dc1dd-7867-4a2e-bd95-1a42d5f66fed · outbound

This paper cites International Circular of Graphic Education and Research pp.

Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution International Circular of Graphic Education and Research pp

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:04:46.598743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:04:46.230092Z digest=sha256:8a7679c584bd937a2c63b86b4e577cf31bebe39b7e5a44c38e9a5899d181c0f0

Observation 14795443-05a6-4931-9f2d-9363148f4639 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T21:04:46.234254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:04:46.234254Z digest=sha256:873becdd45c977815dfd0778c60d8bdea974a195de02df123ae769667861c611

Observation 49062ea1-062f-45a8-b2e4-c80f0e117ad8 · outbound

This paper cites ACM Computing Surveys (CSUR)54(6), 1–37 (2021).

Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution ACM Computing Surveys (CSUR)54(6), 1–37 (2021)

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:04:46.584172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:04:46.238729Z digest=sha256:a00509167531311d00560fc3768cef6efb08e53f4d657cbad611a3b713326b21

Observation 853eb5ce-1ae3-4e96-8f9b-caf8a3dea925 · outbound

This paper cites Halmstad University Press (2018).

Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution Halmstad University Press (2018)

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:04:46.569987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:04:46.242518Z digest=sha256:75743ae7d1f8531ed435c7eeed647b089679e01690b364963b32c633a1b9e936

Observation cba4acab-9bf2-4b85-8d57-c4f6e20912cc · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.

Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:04:46.554606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:04:46.246407Z digest=sha256:d7568f2ec912201ac55d5de8f9351dc4ffb608ce2769282366e5973cf9840aed

Observation d3758079-26d9-4508-b9f5-67a0b04aa978 · outbound

This paper cites Gondwana Research26(3-4), 1216–1221 (2014).

Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution Gondwana Research26(3-4), 1216–1221 (2014)

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:04:46.537010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:04:46.250137Z digest=sha256:43422c61d6a229ed821553ecab3d80599975e8b4c4b819a8838acbb1de582120

Observation 7d3fd0fb-cd45-4fc5-a23d-4b5cbd52440a · outbound

This paper cites In: International conference on machine learning.

Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution In: International conference on machine learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T21:04:46.253799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:04:46.253799Z digest=sha256:3b76b0ef9a662033d11a1a8ad79e98d7b57ab1cf03e756b3a79345da099c5a39

Observation 1a9044bd-6ae0-4eae-ab6a-ebd733fe9e25 · outbound

This paper cites Expert Systems with Applications255, 124811 (2024).

Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution Expert Systems with Applications255, 124811 (2024)

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:04:46.500634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:04:46.257621Z digest=sha256:dd186eaf8e1763cc60e91030ae6bf60beb78bae167b27fa1f5398a650f7cbac6

Observation e169781d-40ff-4310-9888-f0caa0281bb7 · outbound

This paper cites In: Proceedings of the 31st ACM International Conference on Multimedia.

Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution In: Proceedings of the 31st ACM International Conference on Multimedia

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:04:46.481278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:04:46.261426Z digest=sha256:a01e05c7e646cc8cf3627e862f51c49da978289654a82b3ad49d201d452a013e

Observation 810a1e1e-88ba-404a-8d59-e2005a2c0637 · outbound

This paper cites Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese.

Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:04:46.265412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:04:46.265412Z digest=sha256:d1d47968e00a75d3a244788302a6bc7b8091c9735d24f27c5ccffc8de41cc37a

Observation fa666afd-4efd-4822-8234-b45c11ee35b8 · outbound

This paper cites SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs.

Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:04:46.269774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:04:46.269774Z digest=sha256:503672a137c56c9c1bbc226524081eb73e1d44b14b030ff8f754246aad23c488

Observation d82c41a2-b1e2-458c-b523-4123c738f925 · outbound

This paper cites Dunhuang Grottoes Painting Dataset and Benchmark.

Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution Dunhuang Grottoes Painting Dataset and Benchmark

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:04:46.274045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:04:46.274045Z digest=sha256:8485bed1d2537a00c9da8fb6424832599d26efd26d11705243e0c31b881eab58

Observation 44ae1fe7-a01c-4f03-b1fb-1c9875004e01 · outbound

This paper cites IEEE Transactions on Pattern Analysis and Machine Intelligence (2024).

Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution IEEE Transactions on Pattern Analysis and Machine Intelligence (2024)

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:04:46.278282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:04:46.278282Z digest=sha256:cad31351b1140a62b321c60ffae102e9a666dafb549d04dce97a641543d0df66

Observation 1a77e20e-5edd-4fd7-9140-1843d143022f · outbound

This paper cites Światowit11(52), 42–57 (2013).

Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution Światowit11(52), 42–57 (2013)

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:04:46.454747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:04:46.282592Z digest=sha256:21625fab0ef8525fdd2ea89f451e81229cfffa7d018bbdff5321117924d17229

Pith citing papers

Observation a8e703d7-3471-4fb5-9b8d-4d68a8fbdd5d · inbound

MinhwaNet: Faithful but Insufficient Object Grounding in Korean Folk Painting cites this paper.

MinhwaNet: Faithful but Insufficient Object Grounding in Korean Folk Painting Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-29T15:03:31.570418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T14:55:22.057166Z digest=sha256:601e4dcbaa8f644cf0b6d44809219d8a44c0823e11140dfe05f6c50c1b75a196