Pith. sign in

Paper Citation Record · LEDGER

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval

As of 19 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 4 inbound Pith citation observations for arXiv:2507.21917.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21917 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:21:15.520982Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:13:58.950328Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T18:23:51.303795Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact1
  • verified fuzzy35
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e58c71a6-42b4-4ffa-b0b1-921d762e9479 · outbound

This paper cites an unresolved cited work.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T12:21:15.946480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.376287Z digest=sha256:7c41866e259e417432a4af923a2a6d6995ca63c4d011d32e69eb55100d331294

Observation bb92f3e5-ee8b-4339-9da7-04b63b97b603 · outbound

This paper cites Leveraging knowledge graphs and deep learning for automatic art analysis,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Leveraging knowledge graphs and deep learning for automatic art analysis,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.937813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.379980Z digest=sha256:7931175239dc8d0286be3c0de829c0b0cc69ce9c9ad8d0e30b72776d46604b9f

Observation 85581239-a6e7-49b6-b7fe-088a5ce8c4e9 · outbound

This paper cites GraphCLIP: Image-graph contrastive learning for multimodal artwork classification,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval GraphCLIP: Image-graph contrastive learning for multimodal artwork classification,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.929473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.383516Z digest=sha256:f785191eb26eb14156184659d1fbc42405390849a48015d0fbb5cb14fd34889c

Observation 09f67382-5b4e-41f5-b76b-b8608701fff6 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Learning transferable visual models from natural language supervision,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.921080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.386819Z digest=sha256:c83072bb80196a11edd7753ff2b65e2df49b360bacfed6851500b3ed4e2dad1c

Observation 89069abc-1406-4f9d-9b92-372b40e2c1cf · outbound

This paper cites GPT-4 Technical Report.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval GPT-4 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:15.389939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:15.389939Z digest=sha256:72d901b73c5f8fd80db50342bca7ac7185f536ccc20ce64c5f272d088cce6a87

Observation 82ddbe2f-079b-4080-8e8e-f674a9773db3 · outbound

This paper cites ArtGPT-4: Towards Artistic-understanding Large Vision-Language Models with Enhanced Adapter.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval ArtGPT-4: Towards Artistic-understanding Large Vision-Language Models with Enhanced Adapter

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:15.393243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:15.393243Z digest=sha256:23ad2a5a5820ac612c07b660c0fd9b18ca700bc9ab3fca5ae961afd998c5b4bf

Observation 303579f4-8f4d-4066-ae43-bd80e62d7ad1 · outbound

This paper cites Gallerygpt: Analyzing paintings with large multimodal models,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Gallerygpt: Analyzing paintings with large multimodal models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.912785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.396782Z digest=sha256:972fd113b99604dd3be1831ab3f2ff055322a39afbb60030df3cf47bdab0c711

Observation 96df4660-a14e-427b-b6f7-17c4164ac575 · outbound

This paper cites KALE: An Artwork Image Captioning System Augmented with Heterogeneous Graph.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval KALE: An Artwork Image Captioning System Augmented with Heterogeneous Graph

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T12:21:15.597754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.399409Z digest=sha256:30b6ec7b2b8ca15b65cdb125aa65ccefc9d2174777a37645b0a2e3d31f82c641

Observation d17427ac-7cd6-4fd3-a2c1-15891a390e1a · outbound

This paper cites Colbert: Efficient and effective passage search via contextualized late interaction over bert,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Colbert: Efficient and effective passage search via contextualized late interaction over bert,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.903528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.402733Z digest=sha256:3c1f0177ddd4893e6c70b603a3e861247b44919a2063657f1b0cfd0428a2cb2d

Observation 5218236b-b32b-4c5f-ae03-183c3e8f9dfb · outbound

This paper cites Artpedia: A new visual-semantic dataset with visual and contextual sentences in the artistic domain,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Artpedia: A new visual-semantic dataset with visual and contextual sentences in the artistic domain,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.895336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.405602Z digest=sha256:314e50b50f5d053e67c2a746c6f42d264a3ae667ab867421cd2623725647bd0f

Observation cb8a1aa8-1494-4cab-973d-395f0f91ca35 · outbound

This paper cites Deep learning approaches to pattern extraction and recognition in paintings and drawings: An overview,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Deep learning approaches to pattern extraction and recognition in paintings and drawings: An overview,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.886284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.408392Z digest=sha256:faa72d35759216891d224d041b027ea78f9e6c780892882f6a02f88a21cff412

Observation bf5868df-3321-4f3c-991a-9a2cd3b0770f · outbound

This paper cites Machine learning for cultural heritage: A survey,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Machine learning for cultural heritage: A survey,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.877773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.411512Z digest=sha256:0f3213f6e37872628ffaa903dc390af3d27ae67390cb6c6f785dca8c4f524bb6

Observation fd1a1c4a-be3f-4538-8249-2b852a36360e · outbound

This paper cites Fine-tuning convolutional neural networks for fine art classification,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Fine-tuning convolutional neural networks for fine art classification,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.869536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.415129Z digest=sha256:69d4374414bfe8ec8feabe7f22f170dc6fca7c21944046d6420d63aee99cd0ae

Observation 92586de4-cc32-4b0c-957d-e42eef8eea45 · outbound

This paper cites Recognizing Image Style.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Recognizing Image Style

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:15.417940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:15.417940Z digest=sha256:da6c07b67ec2d26d22323e801c0c313d5618c09289fe0d73dee2c3a078cdd8e3

Observation f80128db-6fba-4cf4-a641-c745b61c6bf9 · outbound

This paper cites Large-scale Classification of Fine-Art Paintings: Learning The Right Metric on The Right Feature.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Large-scale Classification of Fine-Art Paintings: Learning The Right Metric on The Right Feature

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:15.421188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:15.421188Z digest=sha256:0f97f5fe54cd741d7cbb39f65fde00cd3ef8dc999afba8001d99d0c93f7ece86

Observation 1ff222b3-d6e5-4844-9d06-8e57cf17d884 · outbound

This paper cites Toward Discovery of the Artist’s Style: Learning to recognize artists by their artworks,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Toward Discovery of the Artist’s Style: Learning to recognize artists by their artworks,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.861235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.424471Z digest=sha256:c9df1b8d8b3714c227333f8000679789122424c5454676510d39e937930f01ce

Observation 1f9b41be-672e-4017-8e92-b83db479d8b6 · outbound

This paper cites A deep learning approach to clustering visual arts,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval A deep learning approach to clustering visual arts,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.852876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.428015Z digest=sha256:d2ab4147e15ec6c96cb3dfc3ceea4d43b83797fefcfe716072203a7a497e1297

Observation f4e985b6-a115-4f2d-8495-4c3b7d2d6ce4 · outbound

This paper cites Toward automated discovery of artistic influence,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Toward automated discovery of artistic influence,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.844686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.431094Z digest=sha256:7bb22ddac5656ed06bd8a16f9a8f6ab8c6859036d8cb91b9a559b5b17a40178d

Observation 21280f40-d304-4669-b5fa-44b502e038a3 · outbound

This paper cites Wasielewski, Computational formalism: Art history and machine learning.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Wasielewski, Computational formalism: Art history and machine learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.835852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.433853Z digest=sha256:3aacb1146ab563f447007259105c71f3b733f2540f029a4ad514f10c3eb95d4f

Observation cb2c3577-f29e-4ff7-ba22-6848451ddf9e · outbound

This paper cites ContextNet: representation and exploration for painting classification and retrieval in context,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval ContextNet: representation and exploration for painting classification and retrieval in context,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.827653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.437235Z digest=sha256:9895d1353b40e207704e16680d84fc58c9d7c82c19398697dfdd1619f2964291

Observation 227ce487-ca6a-4f85-b30b-e37473388c43 · outbound

This paper cites How to read paintings: semantic art understanding with multi-modal retrieval,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval How to read paintings: semantic art understanding with multi-modal retrieval,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.818505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.440087Z digest=sha256:ec3c85e69912654922fbdc3cf83fc6c2c61f4d1e29f3f6115308ca2ed14d6527

Observation b8fde18c-aff2-42f4-b86d-791f7c01a484 · outbound

This paper cites Generating captions for images of ancient artworks,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Generating captions for images of ancient artworks,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.810228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.442949Z digest=sha256:32fe4b17e9b2a45377561471ee5b79c3c612ce553f20419ce55dccda492fb0a8

Observation 803fa69d-df7a-40d3-95bb-2278452f278f · outbound

This paper cites A dataset and baselines for visual question answering on art,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval A dataset and baselines for visual question answering on art,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.801853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.445701Z digest=sha256:987890785ecd27773c61caa72879c03ce7bc6dc1b5d704194ce4337af7dd32a4

Observation 2528bd1d-25c3-4ea5-97de-16a86f7f7764 · outbound

This paper cites Iconographic image captioning for artworks,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Iconographic image captioning for artworks,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.793638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.448635Z digest=sha256:eb3269b8f009bedb4d545e12137f9022999192314d6adb0e3030a44b3c614710

Observation f6a71688-807f-4219-9c59-382d2bb1c4fe · outbound

This paper cites Iconclass: an iconographic classification system,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Iconclass: an iconographic classification system,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.785017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.451450Z digest=sha256:d8b9be8ef10fd6c58c7d8c1a5467ecbbf6724cdc5d19b4321db8b1e7f14aeb52

Observation 634affea-ec92-4ce7-874b-3a3714ce6a93 · outbound

This paper cites Explain me the painting: Multi-topic knowledgeable art description generation,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Explain me the painting: Multi-topic knowledgeable art description generation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.775091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.454256Z digest=sha256:797d04074d9ef6f3573f9a95a40bb71ad46f391fae75f72cdf0765c1233c36bb

Observation b6e747ae-a0df-4e17-a0f2-a48a7065f2b0 · outbound

This paper cites Reading Wikipedia to Answer Open-Domain Questions.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Reading Wikipedia to Answer Open-Domain Questions

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:15.457109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:15.457109Z digest=sha256:3b6b1a447bdf8a97fda19622c5991ffe4d2a5f3b8597cf7ad53156d439a3b15d

Observation 0582f23c-164b-4137-8317-8aa4fe6e72a9 · outbound

This paper cites Is GPT-3 all you need for visual question answering in cultural heritage?.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Is GPT-3 all you need for visual question answering in cultural heritage?

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.765957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.460567Z digest=sha256:0207278a2693d3f68294464b7638bc2043fa1c6e2a6001ab647bea63a6dfaee6

Observation cb6d3a00-b468-43c0-9457-842b4bf6981c · outbound

This paper cites Exploring the Synergy Between Vision-Language Pretrain- ing and ChatGPT for Artwork Captioning: A Preliminary Study,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Exploring the Synergy Between Vision-Language Pretrain- ing and ChatGPT for Artwork Captioning: A Preliminary Study,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.756816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.463245Z digest=sha256:1e4504e6bb92f3c9af08ce0c159c1f0c13f04e0e6357147f8cf99300b47710f3

Observation 854e98c0-f9c8-42c8-af33-75ac959ec5d1 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.747936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.466019Z digest=sha256:4e4794dfd863eeb0b78ab4b9a790afb72096763dc091844712e39c4364196095

Observation 150cff47-c133-499e-b3d4-4e4f410708fe · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Flamingo: a visual language model for few-shot learning,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.739098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.468917Z digest=sha256:2f034617d7540545da8c228a8c2d4a55ff7eda2aff7b04e91c1f874a95735986

Observation c9a50e0a-707c-40ec-aa9b-29fad0d34079 · outbound

This paper cites Visual instruction tuning,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Visual instruction tuning,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:15.471718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:15.471718Z digest=sha256:4fe1284903604ac85daf2e3615a92a5743adca4c8604748f6de6806fa86ebadd

Observation 6d0664b4-9995-4025-96fc-bc41d9ab6673 · outbound

This paper cites InstructBLIP: Towards General- purpose Vision-Language Models with Instruction Tuning,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval InstructBLIP: Towards General- purpose Vision-Language Models with Instruction Tuning,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.725332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.474586Z digest=sha256:d26bdf80f9d0d0bc304c0af322dc43af0fb9cad00cf78c18df03cb4893a5d1ca

Observation 796d7235-184e-4229-9233-793acc8f5ea4 · outbound

This paper cites Reveal: Retrieval- augmented visual-language pre-training with multi-source multimodal knowledge memory,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Reveal: Retrieval- augmented visual-language pre-training with multi-source multimodal knowledge memory,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.717080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.477353Z digest=sha256:4fb3601e1fa1e47e2bafc925a78b302b687ddbb8f6a22e414849616c8703b175

Observation e0c4461a-5074-4713-aac2-d48110b8557c · outbound

This paper cites EchoSight: Advancing Visual-Language Models with Wiki Knowledge,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval EchoSight: Advancing Visual-Language Models with Wiki Knowledge,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.708566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.480040Z digest=sha256:0c9b515fb7acb22ad02d67659d875564374f6fe61e6cdc4d1ec90d27eb643048

Observation d4765d3a-bea7-46bb-823d-108c18c0938c · outbound

This paper cites Wiki-llava: Hierarchical retrieval-augmented generation for multimodal llms,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Wiki-llava: Hierarchical retrieval-augmented generation for multimodal llms,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.699418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.482995Z digest=sha256:121c926524893315c51760ddf9445b78e83ea86d8352cf458bb8179570568caa

Observation 60a1f0ce-cbdd-4d57-87b8-6a4c87f7f94b · outbound

This paper cites Qwen2.5-VL Technical Report.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Qwen2.5-VL Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:15.486031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:15.486031Z digest=sha256:c6b0f0135b5a3ea3d0ab7249fba50f6db34dbefb6aad5ef86c7786d5084e300f

Observation b1ff6eda-cc9f-404b-bbe8-e9a45320b9d4 · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Toolformer: Language Models Can Teach Themselves to Use Tools,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.690765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.489064Z digest=sha256:f5fb27098a282c76122a41af37efbfa671992771b4c5adbe9072ef8b9867f41f

Observation fcc00fb5-76f2-47b9-bf60-adc2a0cb7931 · outbound

This paper cites React: Synergizing reasoning and acting in language models,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval React: Synergizing reasoning and acting in language models,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:15.491934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:15.491934Z digest=sha256:9b618e960538ed6ba1d7e15b4a3b542d7819606ff99ece8a0ff2c29cfd7221cd

Observation 6b1053e5-fd02-4592-9b53-b6675d9f4ca3 · outbound

This paper cites Colpali: Efficient document retrieval with vision language models,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Colpali: Efficient document retrieval with vision language models,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.676205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.494631Z digest=sha256:2e9aaa4e0eebde7dfa09b50ed64624272f5cbaeaf75b5d95a73f2773f87c672a

Observation ad6f203a-8928-413e-a0fb-bbc6b3324ca3 · outbound

This paper cites Wikiextractor,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Wikiextractor,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.667440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.497512Z digest=sha256:8a8772cb8342b2f3d4a22fd0542a15b2fb6d9e00bbeed8095e0061fb8a0eb51d

Observation a4726a7f-1eef-41f6-918b-9b5744df52f2 · outbound

This paper cites Reducing the Footprint of Multi-Vector Retrieval with Minimal Performance Impact via Token Pooling.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Reducing the Footprint of Multi-Vector Retrieval with Minimal Performance Impact via Token Pooling

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:15.500544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:15.500544Z digest=sha256:ea9b9474e2fc6c939b4032eeafc3c9d4a3e61fb739d162fa2c10e50a72572b61

Observation 26aa50d7-e60d-4a0c-b09f-33329d07c8c3 · outbound

This paper cites Sigmoid loss for language image pre-training,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Sigmoid loss for language image pre-training,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:15.503536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:15.503536Z digest=sha256:993c9df34f18eed14179eb7bd9db859703cac3030299e61eddb366bfda3271be

Observation 8e9eafa2-5e31-410c-a71a-abea7a1ca900 · outbound

This paper cites Multi-task learning using uncertainty to weigh losses for scene geometry and semantics,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Multi-task learning using uncertainty to weigh losses for scene geometry and semantics,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:15.506829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:15.506829Z digest=sha256:016a7fad78904954fcbe0661b6ebdb15d77e13009414995e4413a34539f02064

Observation c7b3448b-451d-4bf3-9902-a94a24d34c47 · outbound

This paper cites Art History: A Preliminary Handbook (1996).

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Art History: A Preliminary Handbook (1996)

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.648818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.509740Z digest=sha256:4b3b2fd1123af28b158bcd67210c7c9a844afe6d1a0b3d4ce0e5d640e6fc431a

Observation 659a6a7a-cd51-4859-9372-5b7321fd9989 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:15.512431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:15.512431Z digest=sha256:a3e40d61d7ec075f4a106ce268c1952c2bd58bd7a5426572d3690b5053de92b4

Observation e1c506db-022a-43b2-bcb1-58e6f8c08359 · outbound

This paper cites Composed image retrieval using contrastive learning and task-oriented CLIP-based features,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Composed image retrieval using contrastive learning and task-oriented CLIP-based features,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.640222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.515457Z digest=sha256:53b69a93b16c95faee12d8bdfa96767b2b2281811ad02496436790c291d15fda

Observation baa7618f-90fc-4fe1-b022-3a3206541662 · outbound

This paper cites Artwork interpretation,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Artwork interpretation,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.631502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.518248Z digest=sha256:80ef647ed0ca4928861c88b5560ad9f8f07b3a621e70d27f1c6850796eccdf42

Observation c841696b-c349-4518-a218-08b8bee8ffa7 · outbound

This paper cites Artquest: Countering hidden language biases in artvqa,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Artquest: Countering hidden language biases in artvqa,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.623100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:21:15.520982Z digest=sha256:44819adf60ce63eff9fe648a5513271258f755daabab68dc75f21286f0494e54

Pith citing papers

Observation 40f2e3ed-7984-40f9-bb3f-3843d57451db · inbound

Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding cites this paper.

Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:15:21.084154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T09:11:31.870441Z digest=sha256:f577dcebd73f7e1e019c3a69951bbf688734826997b69514bdaf2fc4eff94f9e

Observation e6393261-fd1e-4261-a288-b4a7d816a19c · inbound

Appear2Meaning: A Cross-Cultural Benchmark for Structured Cultural Metadata Inference from Images cites this paper.

Appear2Meaning: A Cross-Cultural Benchmark for Structured Cultural Metadata Inference from Images ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:05:55.467704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T17:49:25.227568Z digest=sha256:561871763b005aecc0da96c5e40df5cab3179ae94f22a27eb0ae305b44115d6e

Observation 710e3f64-12c0-4227-a319-519d0174c297 · inbound

Understanding How MLLMs Describe Artworks Using Token Activation Maps cites this paper.

Understanding How MLLMs Describe Artworks Using Token Activation Maps ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:23:51.305082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T05:06:35.085802Z digest=sha256:5c450bb3408dfa7b861ee537c01f88d46b495a2e451dda415d026e0eec0b5acc

Observation 6879c8b8-a8b5-4248-bb49-3e95baadafc6 · inbound

ArtAnno: Annotating Implicit Semantics in Artworks through LLM Agent-Driven Bidirectional Human-AI Augmentation cites this paper.

ArtAnno: Annotating Implicit Semantics in Artworks through LLM Agent-Driven Bidirectional Human-AI Augmentation ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T11:13:58.950328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:13:58.950328Z digest=sha256:4c1f467a9943ea5f54e295ff491553ae71fade6c9b86ff41a8c660ccbdc96dd5