Pith. sign in

Paper Citation Record · LEDGER

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval

As of 8 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 4 inbound Pith citation observations for arXiv:2507.21917.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21917 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:21:15.520982Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:13:58.950328Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T18:23:51.303795Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact1
  • verified fuzzy35
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e58c71a6-42b4-4ffa-b0b1-921d762e9479 · outbound

This paper cites an unresolved cited work.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T12:21:15.946480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.376287Z digest=sha256:3b426d8e1cbbf079f63c64fdbaffec94f9b4644890c3495e23cc193a1b71e1b7

Observation bb92f3e5-ee8b-4339-9da7-04b63b97b603 · outbound

This paper cites Leveraging knowledge graphs and deep learning for automatic art analysis,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Leveraging knowledge graphs and deep learning for automatic art analysis,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.937813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.379980Z digest=sha256:358ead09883b3f5adf6bd9132f7264e177ba7cccb6ba20e1d6db2e6aa5514ddc

Observation 85581239-a6e7-49b6-b7fe-088a5ce8c4e9 · outbound

This paper cites GraphCLIP: Image-graph contrastive learning for multimodal artwork classification,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval GraphCLIP: Image-graph contrastive learning for multimodal artwork classification,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.929473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.383516Z digest=sha256:ece39415b175575707ee9c34597ee0d29521ab05351a75c263a270fe3dc3550e

Observation 09f67382-5b4e-41f5-b76b-b8608701fff6 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Learning transferable visual models from natural language supervision,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.921080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.386819Z digest=sha256:d654a5423cddec0afc7963d9672dbc2ca46390835d3a213620f6378c00ae268f

Observation 89069abc-1406-4f9d-9b92-372b40e2c1cf · outbound

This paper cites GPT-4 Technical Report.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval GPT-4 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:15.389939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:15.389939Z digest=sha256:975e863724498ea4764d9eb498c414d707119916462760385ec630244923ad33

Observation 82ddbe2f-079b-4080-8e8e-f674a9773db3 · outbound

This paper cites ArtGPT-4: Towards Artistic-understanding Large Vision-Language Models with Enhanced Adapter.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval ArtGPT-4: Towards Artistic-understanding Large Vision-Language Models with Enhanced Adapter

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:15.393243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:15.393243Z digest=sha256:9c26d9396d9e1d67ca3f628a36a6933bdd787b679847a882d7d96cefedc0b6fe

Observation 303579f4-8f4d-4066-ae43-bd80e62d7ad1 · outbound

This paper cites Gallerygpt: Analyzing paintings with large multimodal models,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Gallerygpt: Analyzing paintings with large multimodal models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.912785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.396782Z digest=sha256:302c593019110d7b9c4b872e5d33fa73374c1a46ba0176919afd03886cd96da6

Observation 96df4660-a14e-427b-b6f7-17c4164ac575 · outbound

This paper cites KALE: An Artwork Image Captioning System Augmented with Heterogeneous Graph.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval KALE: An Artwork Image Captioning System Augmented with Heterogeneous Graph

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T12:21:15.597754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.399409Z digest=sha256:fa6345ec75bba739fbee555d75ec15011cb6f9b04755628bf73b5e1608572cd3

Observation d17427ac-7cd6-4fd3-a2c1-15891a390e1a · outbound

This paper cites Colbert: Efficient and effective passage search via contextualized late interaction over bert,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Colbert: Efficient and effective passage search via contextualized late interaction over bert,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.903528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.402733Z digest=sha256:e02ed6df706317b1bab83ddb1a625c81a4e2d3b1f22fb0af034966338b0946ea

Observation 5218236b-b32b-4c5f-ae03-183c3e8f9dfb · outbound

This paper cites Artpedia: A new visual-semantic dataset with visual and contextual sentences in the artistic domain,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Artpedia: A new visual-semantic dataset with visual and contextual sentences in the artistic domain,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.895336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.405602Z digest=sha256:409cc0ed7f490b1da55f199c429f7a5bd1c0a63021557428717b021c1572e9a5

Observation cb8a1aa8-1494-4cab-973d-395f0f91ca35 · outbound

This paper cites Deep learning approaches to pattern extraction and recognition in paintings and drawings: An overview,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Deep learning approaches to pattern extraction and recognition in paintings and drawings: An overview,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.886284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.408392Z digest=sha256:7f609fb3d45ee5d6be5f439e163c367bcba665b1163c5f7566eb7d5790cea803

Observation bf5868df-3321-4f3c-991a-9a2cd3b0770f · outbound

This paper cites Machine learning for cultural heritage: A survey,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Machine learning for cultural heritage: A survey,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.877773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.411512Z digest=sha256:79ce21b8059fc58e9158728bebd803764ade507c8def62250f5b51383d658824

Observation fd1a1c4a-be3f-4538-8249-2b852a36360e · outbound

This paper cites Fine-tuning convolutional neural networks for fine art classification,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Fine-tuning convolutional neural networks for fine art classification,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.869536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.415129Z digest=sha256:d42eb0bc98336145ca48d2c8c9aeb5b270144cd1b631c4aca46cf6474f920078

Observation 92586de4-cc32-4b0c-957d-e42eef8eea45 · outbound

This paper cites Recognizing Image Style.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Recognizing Image Style

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:15.417940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:15.417940Z digest=sha256:71131c91fbcad6a86998fd97a361767a8ea3566591e8d017d4f24629c25873b2

Observation f80128db-6fba-4cf4-a641-c745b61c6bf9 · outbound

This paper cites Large-scale Classification of Fine-Art Paintings: Learning The Right Metric on The Right Feature.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Large-scale Classification of Fine-Art Paintings: Learning The Right Metric on The Right Feature

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:15.421188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:15.421188Z digest=sha256:125e7de7682ce9d42b7d6af0f0e3c0bbba803494ed33864842e9a22138202948

Observation 1ff222b3-d6e5-4844-9d06-8e57cf17d884 · outbound

This paper cites Toward Discovery of the Artist’s Style: Learning to recognize artists by their artworks,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Toward Discovery of the Artist’s Style: Learning to recognize artists by their artworks,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.861235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.424471Z digest=sha256:34f26505f8f32c3453379fd15c10b34e73151bdd56677eb496b42ac808e5e052

Observation 1f9b41be-672e-4017-8e92-b83db479d8b6 · outbound

This paper cites A deep learning approach to clustering visual arts,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval A deep learning approach to clustering visual arts,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.852876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.428015Z digest=sha256:c7deed258d110438068d7ead646f893d84a999ff98cfdb70d9beb3cb3edb3d22

Observation f4e985b6-a115-4f2d-8495-4c3b7d2d6ce4 · outbound

This paper cites Toward automated discovery of artistic influence,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Toward automated discovery of artistic influence,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.844686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.431094Z digest=sha256:5c6e8f4545f2a31f86ef79902431641886758efa94705709d50a0c3001fb1755

Observation 21280f40-d304-4669-b5fa-44b502e038a3 · outbound

This paper cites Wasielewski, Computational formalism: Art history and machine learning.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Wasielewski, Computational formalism: Art history and machine learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.835852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.433853Z digest=sha256:2b1f35e641e646cb6183af750a3181252dc577833df0c36d709a40c97825564a

Observation cb2c3577-f29e-4ff7-ba22-6848451ddf9e · outbound

This paper cites ContextNet: representation and exploration for painting classification and retrieval in context,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval ContextNet: representation and exploration for painting classification and retrieval in context,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.827653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.437235Z digest=sha256:a94e8c8dc6ae9b19d16e8c8796a1ef3813d92d9ea4eec0f8a2b36450de827b86

Observation 227ce487-ca6a-4f85-b30b-e37473388c43 · outbound

This paper cites How to read paintings: semantic art understanding with multi-modal retrieval,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval How to read paintings: semantic art understanding with multi-modal retrieval,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.818505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.440087Z digest=sha256:d247d43a8ec62880dd2714396ab59a1033275611bcfbbbedb129be4a1852c6fc

Observation b8fde18c-aff2-42f4-b86d-791f7c01a484 · outbound

This paper cites Generating captions for images of ancient artworks,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Generating captions for images of ancient artworks,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.810228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.442949Z digest=sha256:15ed74f971fafe36a9d9965630eaa625da69f228db2b7e0f4a56ddaa8d1a84d5

Observation 803fa69d-df7a-40d3-95bb-2278452f278f · outbound

This paper cites A dataset and baselines for visual question answering on art,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval A dataset and baselines for visual question answering on art,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.801853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.445701Z digest=sha256:c7ced670adaf6b0efbd1ca96870fa56ad017374af22d53963b63f31db5fbd2cc

Observation 2528bd1d-25c3-4ea5-97de-16a86f7f7764 · outbound

This paper cites Iconographic image captioning for artworks,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Iconographic image captioning for artworks,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.793638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.448635Z digest=sha256:be2ad8c8992c2089cb102f6a668583c5ae2b2b2d78b85601d96258495a23514d

Observation f6a71688-807f-4219-9c59-382d2bb1c4fe · outbound

This paper cites Iconclass: an iconographic classification system,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Iconclass: an iconographic classification system,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.785017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.451450Z digest=sha256:061e9ff36d9aaee7d6e9982688253a3ee74dca1e9d802b6126d8c232a5663b69

Observation 634affea-ec92-4ce7-874b-3a3714ce6a93 · outbound

This paper cites Explain me the painting: Multi-topic knowledgeable art description generation,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Explain me the painting: Multi-topic knowledgeable art description generation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.775091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.454256Z digest=sha256:8570c7e2376a3fae09429a48b06809ec9fb4c04edc2d350765e561a7bfd66149

Observation b6e747ae-a0df-4e17-a0f2-a48a7065f2b0 · outbound

This paper cites Reading Wikipedia to Answer Open-Domain Questions.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Reading Wikipedia to Answer Open-Domain Questions

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:15.457109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:15.457109Z digest=sha256:c3851858bbd79190ce7db2b454875c0a0c7f30b436611cc741ab9b3fd4950eb9

Observation 0582f23c-164b-4137-8317-8aa4fe6e72a9 · outbound

This paper cites Is GPT-3 all you need for visual question answering in cultural heritage?.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Is GPT-3 all you need for visual question answering in cultural heritage?

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.765957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.460567Z digest=sha256:bcd49da8d24a2e89ab49dd3b45a9b9825887d4c02f4fc2addb3c996401959a7c

Observation cb6d3a00-b468-43c0-9457-842b4bf6981c · outbound

This paper cites Exploring the Synergy Between Vision-Language Pretrain- ing and ChatGPT for Artwork Captioning: A Preliminary Study,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Exploring the Synergy Between Vision-Language Pretrain- ing and ChatGPT for Artwork Captioning: A Preliminary Study,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.756816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.463245Z digest=sha256:d8d99aa84fe55f0189f81190ff642b473386fc1cb93e142f806116d8c67aaf27

Observation 854e98c0-f9c8-42c8-af33-75ac959ec5d1 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.747936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.466019Z digest=sha256:bf2a943ff1c281f36259f9c155bffdfdd720c655b406be7e86e059ce027bf1c3

Observation 150cff47-c133-499e-b3d4-4e4f410708fe · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Flamingo: a visual language model for few-shot learning,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.739098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.468917Z digest=sha256:e88273d376a15e2fb63da851c66563e786197183749b64308f06be376f116bdc

Observation c9a50e0a-707c-40ec-aa9b-29fad0d34079 · outbound

This paper cites Visual instruction tuning,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Visual instruction tuning,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:15.471718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:15.471718Z digest=sha256:52bfa4e4419abcc8db1393d8a907cfb752ba67d87ec2f921dbd051b75fcf75ca

Observation 6d0664b4-9995-4025-96fc-bc41d9ab6673 · outbound

This paper cites InstructBLIP: Towards General- purpose Vision-Language Models with Instruction Tuning,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval InstructBLIP: Towards General- purpose Vision-Language Models with Instruction Tuning,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.725332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.474586Z digest=sha256:295a6cdac0e5e0c8d00d1fc328fd295902e1a8a2e677fa4c0c963b9211bac4d4

Observation 796d7235-184e-4229-9233-793acc8f5ea4 · outbound

This paper cites Reveal: Retrieval- augmented visual-language pre-training with multi-source multimodal knowledge memory,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Reveal: Retrieval- augmented visual-language pre-training with multi-source multimodal knowledge memory,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.717080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.477353Z digest=sha256:7b3bfa3bc5afa5a0a7f8d202bf6fcccf66e4cb5c267ba5b1fce4ba87cd2ffff0

Observation e0c4461a-5074-4713-aac2-d48110b8557c · outbound

This paper cites EchoSight: Advancing Visual-Language Models with Wiki Knowledge,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval EchoSight: Advancing Visual-Language Models with Wiki Knowledge,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.708566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.480040Z digest=sha256:c2fa8ab1345fdfc9b8ee273b44ffb545aa4868ff9bec34af2009a72adb0f88cf

Observation d4765d3a-bea7-46bb-823d-108c18c0938c · outbound

This paper cites Wiki-llava: Hierarchical retrieval-augmented generation for multimodal llms,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Wiki-llava: Hierarchical retrieval-augmented generation for multimodal llms,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.699418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.482995Z digest=sha256:8f392784dd296b2956cc6e5b4eb0339dccc856ea3b3ea87001a98d2be36d96dd

Observation 60a1f0ce-cbdd-4d57-87b8-6a4c87f7f94b · outbound

This paper cites Qwen2.5-VL Technical Report.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Qwen2.5-VL Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:15.486031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:15.486031Z digest=sha256:d206b4534f95d11d70f734c8cca30c45465c8c403350ae5c9768d3fce7485730

Observation b1ff6eda-cc9f-404b-bbe8-e9a45320b9d4 · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Toolformer: Language Models Can Teach Themselves to Use Tools,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.690765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.489064Z digest=sha256:e520d504268f54dc1b4ffeced69d988225cb78eaec6db97558d0dbe3950ebd75

Observation fcc00fb5-76f2-47b9-bf60-adc2a0cb7931 · outbound

This paper cites React: Synergizing reasoning and acting in language models,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval React: Synergizing reasoning and acting in language models,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:15.491934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:15.491934Z digest=sha256:1767fc7e5e58883eb5d0c5190045766954603f18bbf578f1b28d24ae2a0d51d9

Observation 6b1053e5-fd02-4592-9b53-b6675d9f4ca3 · outbound

This paper cites Colpali: Efficient document retrieval with vision language models,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Colpali: Efficient document retrieval with vision language models,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.676205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.494631Z digest=sha256:6f635b1cbb0b3652baa052825132c9e46c9bd5444c675c05e20d51bc4a707448

Observation ad6f203a-8928-413e-a0fb-bbc6b3324ca3 · outbound

This paper cites Wikiextractor,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Wikiextractor,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.667440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.497512Z digest=sha256:d0738bb78bce1c7215d2bb5f6c4a333b645f9807ca9471eed79804f047f789f3

Observation a4726a7f-1eef-41f6-918b-9b5744df52f2 · outbound

This paper cites Reducing the Footprint of Multi-Vector Retrieval with Minimal Performance Impact via Token Pooling.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Reducing the Footprint of Multi-Vector Retrieval with Minimal Performance Impact via Token Pooling

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:15.500544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:15.500544Z digest=sha256:f201b6154494b53580da1452ae0161549197007dd790e8bde9bc00df12d62216

Observation 26aa50d7-e60d-4a0c-b09f-33329d07c8c3 · outbound

This paper cites Sigmoid loss for language image pre-training,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Sigmoid loss for language image pre-training,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:15.503536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:15.503536Z digest=sha256:05339f8e265f10c819b4bb200ca1b574f932d927170565d95e7c30130447ab9c

Observation 8e9eafa2-5e31-410c-a71a-abea7a1ca900 · outbound

This paper cites Multi-task learning using uncertainty to weigh losses for scene geometry and semantics,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Multi-task learning using uncertainty to weigh losses for scene geometry and semantics,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:15.506829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:15.506829Z digest=sha256:289ae1ed5be2b53333adf4055c782ab7be4c3f0c3ca2db03caf0bd3e89245a03

Observation c7b3448b-451d-4bf3-9902-a94a24d34c47 · outbound

This paper cites Art History: A Preliminary Handbook (1996).

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Art History: A Preliminary Handbook (1996)

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.648818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.509740Z digest=sha256:b10c7665938f2033fc65a298308c5b2777fda163a5f11505ad62ad5131f44542

Observation 659a6a7a-cd51-4859-9372-5b7321fd9989 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:15.512431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:15.512431Z digest=sha256:13c4f99d8afc8780f4dd98106c1d98e6faf44e82c953ec9c1432c2b176c04db9

Observation e1c506db-022a-43b2-bcb1-58e6f8c08359 · outbound

This paper cites Composed image retrieval using contrastive learning and task-oriented CLIP-based features,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Composed image retrieval using contrastive learning and task-oriented CLIP-based features,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.640222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.515457Z digest=sha256:ecd1a4d22b93c0af7286d3b27a3d21cded5820953ea79cd1296606f6137b822b

Observation baa7618f-90fc-4fe1-b022-3a3206541662 · outbound

This paper cites Artwork interpretation,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Artwork interpretation,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.631502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.518248Z digest=sha256:af427a220c919e721c21d5bf2e50bf49a6771dd23ad4b1a412573ebbdc63c55f

Observation c841696b-c349-4518-a218-08b8bee8ffa7 · outbound

This paper cites Artquest: Countering hidden language biases in artvqa,.

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval Artquest: Countering hidden language biases in artvqa,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:15.623100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:21:15.520982Z digest=sha256:ebb6151e240f4ce0c6b1082a13fb27ccbb05cc7d4ee51f91a88fe9966e269443

Pith citing papers

Observation 40f2e3ed-7984-40f9-bb3f-3843d57451db · inbound

Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding cites this paper.

Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:15:21.084154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T09:11:31.870441Z digest=sha256:05d180a9ec4d735e77005e298c8e00a0424fc98f6a774dd23195a7c6b2efa877

Observation e6393261-fd1e-4261-a288-b4a7d816a19c · inbound

Appear2Meaning: A Cross-Cultural Benchmark for Structured Cultural Metadata Inference from Images cites this paper.

Appear2Meaning: A Cross-Cultural Benchmark for Structured Cultural Metadata Inference from Images ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:05:55.467704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:49:25.227568Z digest=sha256:c370b92a014a6e55f135b65186cc70b6684ace8a08d2d177c02be1f595f58392

Observation 710e3f64-12c0-4227-a319-519d0174c297 · inbound

Understanding How MLLMs Describe Artworks Using Token Activation Maps cites this paper.

Understanding How MLLMs Describe Artworks Using Token Activation Maps ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:23:51.305082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T05:06:35.085802Z digest=sha256:bfe458e3a7a1d9c3bbfec26075da6fafa9a5fca241ff8f82dd945c1a9ed75a2c

Observation 6879c8b8-a8b5-4248-bb49-3e95baadafc6 · inbound

ArtAnno: Annotating Implicit Semantics in Artworks through LLM Agent-Driven Bidirectional Human-AI Augmentation cites this paper.

ArtAnno: Annotating Implicit Semantics in Artworks through LLM Agent-Driven Bidirectional Human-AI Augmentation ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T11:13:58.950328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:13:58.950328Z digest=sha256:92dacfd83b94680606a320df7eee21b1ece638eb5a96066a42057bcd26840ce1