Pith. sign in

Paper Citation Record · LEDGER

Understanding Museum Exhibits using Vision-Language Reasoning

As of 21 August 2026, this Paper Citation Record lists 100 of 100 outbound references and 0 inbound Pith citation observations for arXiv:2412.01370.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01370 v2

Coverage vector

measured 100 of 100 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:28:39.810901Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 100 outbound references displayed

  • verified exact4
  • verified fuzzy49
  • unresolved46
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7ab34dd0-f126-446a-9a2e-6666b0b96ed3 · outbound

This paper cites Artemis: Affective language for visual art.

Understanding Museum Exhibits using Vision-Language Reasoning Artemis: Affective language for visual art

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.003322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.003322Z digest=sha256:536fb4b6dd1d9142bc37c52220bbc917de764d9112423d0c8aad91faced32602

Observation a559219f-0e2e-4a95-92b1-c3bc742964ff · outbound

This paper cites Feelingblue: A corpus for understanding the emotional con- notation of color in context.Transactions of the Association for Computational Linguistics, 11:176–190, 2023.

Understanding Museum Exhibits using Vision-Language Reasoning Feelingblue: A corpus for understanding the emotional con- notation of color in context.Transactions of the Association for Computational Linguistics, 11:176–190, 2023

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.054864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.054864Z digest=sha256:ead5495598e02956925fc8e7e096dc439f81a54f056b9f664eeb64c538a6c437

Observation a32c6281-9c4d-4ce6-9854-f0f7ed3dc9aa · outbound

This paper cites Vqa: Visual question answering.

Understanding Museum Exhibits using Vision-Language Reasoning Vqa: Visual question answering

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.145741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.145741Z digest=sha256:e50bfd65c5cc89e1dc7923b3ca1ca355c34ce96b8f2b9426626e9ae51813cee5

Observation 74832a63-777f-4fca-b59a-95b253eaf969 · outbound

This paper cites Explain me the painting: Multi-topic knowledgeable art description gen- eration.

Understanding Museum Exhibits using Vision-Language Reasoning Explain me the painting: Multi-topic knowledgeable art description gen- eration

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.150738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.150738Z digest=sha256:f5be80a4d0ed01fb4f85f92bfb3ff13a521bcbfe6261f078e17fae5934ff0539

Observation 19fa2347-3eef-48e7-a2c2-1f85b6db6be8 · outbound

This paper cites Taming CLIP for Fine-grained and Structured Visual Understanding of Museum Exhibits.

Understanding Museum Exhibits using Vision-Language Reasoning Taming CLIP for Fine-grained and Structured Visual Understanding of Museum Exhibits

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-12T04:28:40.472696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:38.155517Z digest=sha256:129e4046eda2831353e3682c0599af98a2baaa4390a648b393b84db14310ef02

Observation b79c97ce-f162-480e-8e25-1f1a355f36a8 · outbound

This paper cites Bridg- ing the gap between object and image-level representations for open-vocabulary detection.Advances in Neural Informa- tion Processing Systems, 35:33781–33794, 2022.

Understanding Museum Exhibits using Vision-Language Reasoning Bridg- ing the gap between object and image-level representations for open-vocabulary detection.Advances in Neural Informa- tion Processing Systems, 35:33781–33794, 2022

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.161415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.161415Z digest=sha256:2be06ae216d447596bb3c08f0495f8cdd6396a72c8338c868d2a37c197b1f8e4

Observation d2a3c8a7-0c14-412d-8e15-50deef6273e0 · outbound

This paper cites Clip retrieval: Easily compute clip embeddings and build a clip retrieval system with them.https : / / github.

Understanding Museum Exhibits using Vision-Language Reasoning Clip retrieval: Easily compute clip embeddings and build a clip retrieval system with them.https : / / github

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.166687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.166687Z digest=sha256:422bd3cecac622ba9163cc8417093c8cf7cdd6b81b372c33896bb58c8dce7c45

Observation 217bea4a-3900-4979-b31d-fe583452871a · outbound

This paper cites Viscounth: A large-scale multilin- gual visual question answering dataset for cultural heritage.

Understanding Museum Exhibits using Vision-Language Reasoning Viscounth: A large-scale multilin- gual visual question answering dataset for cultural heritage

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.171052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.171052Z digest=sha256:236afdc1c403c7647ffaed57952e8a7f3f015c0d2a63bdeecfd2a68be14364c4

Observation 968a67bb-7162-4c82-855b-8b34d7154f9c · outbound

This paper cites Predicting image aesthetics with deep learning.

Understanding Museum Exhibits using Vision-Language Reasoning Predicting image aesthetics with deep learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.176728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.176728Z digest=sha256:02d07a3c39bea25a19fa0d8421f422e1e5d68e4496ff6e03d2fc809b816a2b64

Observation 686e0b9d-99aa-4b79-8619-5920307db9ed · outbound

This paper cites Vizwiz: nearly real-time answers to visual questions.

Understanding Museum Exhibits using Vision-Language Reasoning Vizwiz: nearly real-time answers to visual questions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.181314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.181314Z digest=sha256:3544c4c5ec340e615608692f1d06c96e1f0d9e74993c2c65b40dd68db0967a2a

Observation e287ec9c-024c-42f7-bbe0-239a78dbecc4 · outbound

This paper cites Visual question answering for cul- tural heritage.

Understanding Museum Exhibits using Vision-Language Reasoning Visual question answering for cul- tural heritage

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.233064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.233064Z digest=sha256:b8c19daa638608d9861ed166792283873c2a09cd5d4764d2b4811bec495a13f2

Observation a943caab-bd28-4bf6-b97e-9e6b98f24f67 · outbound

This paper cites Fine-tuning convolutional neural networks for fine art classification.Ex- pert Systems with Applications, 114:107–118, 2018.

Understanding Museum Exhibits using Vision-Language Reasoning Fine-tuning convolutional neural networks for fine art classification.Ex- pert Systems with Applications, 114:107–118, 2018

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.318618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.318618Z digest=sha256:4fd803ef0020214e839c5b222427b96bd2d04955c0128ee7484598bb0b74520c

Observation 90f98f85-fceb-4b7b-8d52-d5e9a64d7d93 · outbound

This paper cites Uniter: Universal image-text representation learning.

Understanding Museum Exhibits using Vision-Language Reasoning Uniter: Universal image-text representation learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.402518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.402518Z digest=sha256:61851f8f7d89c75531cfc6553ca46fc24f5a5879291539e5a6a5470f7263637e

Observation dd62fb5a-e778-4519-af9e-5bc7a3bf1f7d · outbound

This paper cites Clip-art: Contrastive pre-training for fine-grained art classification.

Understanding Museum Exhibits using Vision-Language Reasoning Clip-art: Contrastive pre-training for fine-grained art classification

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.423884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.423884Z digest=sha256:cbd208737624ccb048f0712ca36eec48c57b9998ba8345961b6563c346dcfc9a

Observation 993ed531-a51d-4f56-b838-f2b835c78e60 · outbound

This paper cites Learning sample difficulty from pre-trained models for reliable prediction.Advances in Neural Information Process- ing Systems, 36, 2024.

Understanding Museum Exhibits using Vision-Language Reasoning Learning sample difficulty from pre-trained models for reliable prediction.Advances in Neural Information Process- ing Systems, 36, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.428640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.428640Z digest=sha256:f68e84d8208f91ae899ef11ac76d6178b435b3e5e2826e8730cab2e3acbf7997

Observation 45e79656-ab91-4148-8fd7-ecace7141253 · outbound

This paper cites Novel datasets for fine-grained image categoriza- tion.

Understanding Museum Exhibits using Vision-Language Reasoning Novel datasets for fine-grained image categoriza- tion

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.433206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.433206Z digest=sha256:ca88e726a32a653e9b0717acb4567c16e03e94ad53209a096cf2ab2a4a226d4b

Observation dc951444-420d-4949-908f-6d3910591cc0 · outbound

This paper cites Noisyart: A dataset for webly-supervised art- work recognition.

Understanding Museum Exhibits using Vision-Language Reasoning Noisyart: A dataset for webly-supervised art- work recognition

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.438200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.438200Z digest=sha256:d04f61039aae4d6545316a699bf6f0abdec80a8496e7b6abe6ee5eafbd09697a

Observation 3e23ee05-b3d7-4684-a88d-373dbc491836 · outbound

This paper cites Webly-supervised zero-shot learning for artwork instance recognition.Pattern Recognition Letters, 128:420– 426, 2019.

Understanding Museum Exhibits using Vision-Language Reasoning Webly-supervised zero-shot learning for artwork instance recognition.Pattern Recognition Letters, 128:420– 426, 2019

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.443109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.443109Z digest=sha256:464959cc8a8f4bdb2e55afd73f4e9d8deecf8faa5b72ee92d543b85bf05a417e

Observation c87627a0-b92f-44e6-a460-ca8043706bd8 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Understanding Museum Exhibits using Vision-Language Reasoning Imagenet: A large-scale hierarchical image database

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.448126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.448126Z digest=sha256:043518955af0bd8cf4d170dd66541b39cccd8874619809bbf2360ed2923ebaa0

Observation 39835090-6393-42d0-a6f4-58b124e00f7b · outbound

This paper cites Stytr2: Im- age style transfer with transformers.

Understanding Museum Exhibits using Vision-Language Reasoning Stytr2: Im- age style transfer with transformers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.452791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.452791Z digest=sha256:8212707d9e7dd76cf841131c99fd24ff1a11cb15a1fc33e9300814048c93e47d

Observation a5dd0a3b-ad59-40fd-869c-379dad4de782 · outbound

This paper cites De- coupling zero-shot semantic segmentation.

Understanding Museum Exhibits using Vision-Language Reasoning De- coupling zero-shot semantic segmentation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.458317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.458317Z digest=sha256:cca11207a8c9fa37b35830bf28589c790026cd0eb5b9adb7176002b18392a9f7

Observation ba65d43f-f6d0-4044-a369-8bacfcb9f522 · outbound

This paper cites A survey on bias in visual datasets.Computer Vision and Image Understanding, 223: 103552, 2022.

Understanding Museum Exhibits using Vision-Language Reasoning A survey on bias in visual datasets.Computer Vision and Image Understanding, 223: 103552, 2022

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:43.808497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:38.462791Z digest=sha256:ffde141ef2b969cdca8027b441b39ce3a5e638fae1a29d3cf5b76546e9512b21

Observation d704fb76-e5b4-48f7-95a1-e172a16b8c4b · outbound

This paper cites Are you talking to a machine? dataset and methods for multilingual image question.Advances in neural information processing systems, 28, 2015.

Understanding Museum Exhibits using Vision-Language Reasoning Are you talking to a machine? dataset and methods for multilingual image question.Advances in neural information processing systems, 28, 2015

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:43.625655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:38.525414Z digest=sha256:79198d22fcc3cce4cfa1e560a3ebd5e70bdf3c357833722b97235e2388215b27

Observation 9dff17b3-abb8-4c4b-9996-d23ab488b45c · outbound

This paper cites How to read paintings: semantic art understanding with multi-modal retrieval.

Understanding Museum Exhibits using Vision-Language Reasoning How to read paintings: semantic art understanding with multi-modal retrieval

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:43.611046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:38.585444Z digest=sha256:5026e0b163030a59507463a4ae8195f2617922881ec203069ad206c445bf8372

Observation 37e29f7e-c594-4e44-b004-7ff0a3c51795 · outbound

This paper cites Knowit vqa: Answering knowledge-based ques- tions about videos.

Understanding Museum Exhibits using Vision-Language Reasoning Knowit vqa: Answering knowledge-based ques- tions about videos

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:43.594896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:38.642759Z digest=sha256:e06a5e86880b2cb6bd31e402d4dd0d98f0e26db5e7d3f1dd85737c4334110acb

Observation 93fc2cf7-a3cf-4e9e-9e18-10fcfd7a18e8 · outbound

This paper cites A dataset and baselines for visual question answering on art.

Understanding Museum Exhibits using Vision-Language Reasoning A dataset and baselines for visual question answering on art

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:43.547921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:38.647582Z digest=sha256:84156cfbbd84e23a4b90263694e2cfb9b2dc8c91031e18415d4829c67d79d879

Observation 0aa7eadb-4c48-4b01-8426-e79ec7ed97f1 · outbound

This paper cites A dataset and baselines for visual question answering on art.

Understanding Museum Exhibits using Vision-Language Reasoning A dataset and baselines for visual question answering on art

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:43.421716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:38.652677Z digest=sha256:81066ea796d2a5c9b1ed080f6fa824e0a5bbc551dd7d4048b72a8b3ff5c1b1b1

Observation ba39c09d-9e5d-43fb-92dd-f939e800650e · outbound

This paper cites Im- age style transfer using convolutional neural networks.

Understanding Museum Exhibits using Vision-Language Reasoning Im- age style transfer using convolutional neural networks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.656505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.656505Z digest=sha256:814c9c46f80b09c713b2a3398e5aaffe4ae0d6cfe03ea554c9afcc0de8b53575

Observation bd7ee42a-0928-46e3-976c-8a75fbd95d3d · outbound

This paper cites Aes- thetic image captioning from weakly-labelled photographs.

Understanding Museum Exhibits using Vision-Language Reasoning Aes- thetic image captioning from weakly-labelled photographs

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:43.394999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:38.660234Z digest=sha256:fbdb0c7cf26c58e9afb38571c86af6de09c1d4acfb482bba37c144b801a02128

Observation 73516bb5-921b-4289-be5d-a6e62b090371 · outbound

This paper cites Beyond language bias: Over- coming multimodal shortcut and distribution biases for ro- bust visual question answering.

Understanding Museum Exhibits using Vision-Language Reasoning Beyond language bias: Over- coming multimodal shortcut and distribution biases for ro- bust visual question answering

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:43.378533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:38.664520Z digest=sha256:c18b0981bc43a8616ad3a8199981b48c28ff0821a8a5e4bd459b62ea445ecb97

Observation b4a5b2ae-22b6-481f-a7b0-5a8031636e2c · outbound

This paper cites Open-vocabulary Object Detection via Vision and Language Knowledge Distillation.

Understanding Museum Exhibits using Vision-Language Reasoning Open-vocabulary Object Detection via Vision and Language Knowledge Distillation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.668359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.668359Z digest=sha256:0026d1523c169dc870ae39d8b365ab328ee87ac79bc48547b9a1ccd3b5b2483f

Observation f4d8dea0-41d4-4287-a82d-315618f7e9ca · outbound

This paper cites Many- modalqa: Modality disambiguation and qa over diverse in- puts.

Understanding Museum Exhibits using Vision-Language Reasoning Many- modalqa: Modality disambiguation and qa over diverse in- puts

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:43.247869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:38.673304Z digest=sha256:4f46dc9e9e5fe5af98cb429e7c02553091cc8b8b45aeea9ae527c903cb1e3cbb

Observation 1e2af5e0-222f-4f25-bbba-534691a917ab · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

Understanding Museum Exhibits using Vision-Language Reasoning Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.677467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.677467Z digest=sha256:a82a48bcc294e4d91403a68c718c9f64a90433037b93ac2e29ef690872aafd1a

Observation f9ac5b94-5043-4e0c-81d4-d2127b34d9d7 · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elementary visual reasoning.

Understanding Museum Exhibits using Vision-Language Reasoning Clevr: A diagnostic dataset for compositional language and elementary visual reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.681754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.681754Z digest=sha256:10fe8d4d702c1c751bcca64dc286083d2a3814f32cdf468785b6af92371090e5

Observation 014e6e6c-03ae-4007-be21-1b55ee5b5f38 · outbound

This paper cites Prompting visual-language models for efficient video understanding.

Understanding Museum Exhibits using Vision-Language Reasoning Prompting visual-language models for efficient video understanding

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:43.186286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:38.689495Z digest=sha256:35f63d84defd585bafb1be83da875bc525c7f0b54cc9342573ac9bde3c5e76e9

Observation 56a5d640-b34e-4d6f-9d64-48d150e4c2b6 · outbound

This paper cites FigureQA: An Annotated Figure Dataset for Visual Reasoning.

Understanding Museum Exhibits using Vision-Language Reasoning FigureQA: An Annotated Figure Dataset for Visual Reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.693551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.693551Z digest=sha256:eabdbc7d50f02176c5f1c521010116ab0262bda99b256ee69a9b8c7f5b0a8df5

Observation 91a5111b-56c1-408a-b851-afc328280b05 · outbound

This paper cites Mdetr- modulated detection for end-to-end multi-modal understand- ing.

Understanding Museum Exhibits using Vision-Language Reasoning Mdetr- modulated detection for end-to-end multi-modal understand- ing

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:43.169881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:38.751828Z digest=sha256:90b7c1b9ee651c534a9103e7ed89e0d19115834e4759f03bddc95935315b645f

Observation e9ad3cc9-660a-4927-9144-b1305de58ffc · outbound

This paper cites From word embeddings to document distances.

Understanding Museum Exhibits using Vision-Language Reasoning From word embeddings to document distances

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:43.066016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:38.844763Z digest=sha256:5bff9adbdbb90183334c570bd724284e3c463fc9856f3e266a51184122490397

Observation f243837c-0604-4e8e-8bc2-d85da5c838bb · outbound

This paper cites Clipstyler: Image style transfer with a single text condition.

Understanding Museum Exhibits using Vision-Language Reasoning Clipstyler: Image style transfer with a single text condition

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:42.938721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:38.883392Z digest=sha256:875e927b37aeffed7802fc763a257fb80aefb0a400f5be9c19ce765fbb1ad778

Observation a4071822-4b8d-42fe-8d76-44e138dbacfa · outbound

This paper cites Language-driven semantic seg- mentation.

Understanding Museum Exhibits using Vision-Language Reasoning Language-driven semantic seg- mentation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:42.899404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:38.888272Z digest=sha256:76e028ce2231960efeb58c7fb1d1dca1892fac438f4334397b3620d50a45c927

Observation 354d88fa-3e2b-4e3a-b2e8-9ba7caf022bc · outbound

This paper cites Language-driven Semantic Segmentation.

Understanding Museum Exhibits using Vision-Language Reasoning Language-driven Semantic Segmentation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.892694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.892694Z digest=sha256:3d870511352124c1bc7a8305cb7f8241fdcf01ccd5101d128bf7dcba761039c9

Observation 6c046909-4675-4dbd-aa31-1c2964475703 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Understanding Museum Exhibits using Vision-Language Reasoning Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:42.883668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:38.897377Z digest=sha256:ef79375dd40ed0b79d79262d4cda4708017dfac9dd6480231a0975b69fd3e60f

Observation 9753162f-58c1-49f5-b48e-d55cf528f2d4 · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

Understanding Museum Exhibits using Vision-Language Reasoning VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.901539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.901539Z digest=sha256:40d7fbbd5aa3b288bb59453ef8168a40786b177bee0de48002b567f98de61a46

Observation e9c55e5e-e360-4186-b272-66c623722a76 · outbound

This paper cites Oscar: Object-semantics aligned pre-training for vision-language tasks.

Understanding Museum Exhibits using Vision-Language Reasoning Oscar: Object-semantics aligned pre-training for vision-language tasks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.906247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.906247Z digest=sha256:bb143b1deb3ff1adf11e7285f6510c65f7873e405e68c21c564953117bbd8eba

Observation 668da33f-32df-46d9-9636-97f15eca9a50 · outbound

This paper cites Microsoft coco: Common objects in context.

Understanding Museum Exhibits using Vision-Language Reasoning Microsoft coco: Common objects in context

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.910224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.910224Z digest=sha256:5d325bee737cb8df5bd793923a4f080beac084712bfb601ea033e6bb2944fa97

Observation acf6eea0-eca4-4a60-8565-670eac19241c · outbound

This paper cites Fine-grained late-interaction multi-modal retrieval for retrieval augmented visual question answering.

Understanding Museum Exhibits using Vision-Language Reasoning Fine-grained late-interaction multi-modal retrieval for retrieval augmented visual question answering

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:42.844970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:38.914415Z digest=sha256:1f94bf265cad4a09079f9942ecd1c80032f3e9ba2be072b91828e27ded446b87

Observation 098a63f3-b619-4544-9194-ea597e1a9687 · outbound

This paper cites Visual instruction tuning, 2023.

Understanding Museum Exhibits using Vision-Language Reasoning Visual instruction tuning, 2023

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:42.717687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:38.918582Z digest=sha256:5df2d675e5bae9401cd7cb2be77b8eec87503266164830f19ae69ecb73a8234e

Observation c81f9129-ed1e-4b7d-a392-f2eea8294171 · outbound

This paper cites Decoupled Weight Decay Regularization.

Understanding Museum Exhibits using Vision-Language Reasoning Decoupled Weight Decay Regularization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.956848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.956848Z digest=sha256:0356742e526fd26f5713bdf74947f03241d3c4838f5e1a366ebaa74bf84eec4a

Observation 003cb9a8-bf1f-47de-a4f8-cd0c19128723 · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.Advances in neural information processing systems, 32, 2019.

Understanding Museum Exhibits using Vision-Language Reasoning Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.Advances in neural information processing systems, 32, 2019

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:39.001582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:39.001582Z digest=sha256:dc5ee44e91ede16a5ffdc819feee4e603c730a730fd488ae8fe25c93e6e0fa32

Observation 8cbcad21-eaf7-44e9-8c4d-a5ae0d4fd08e · outbound

This paper cites Data- efficient image captioning of fine art paintings via virtual- real semantic alignment training.Neurocomputing, 490:163– 180, 2022.

Understanding Museum Exhibits using Vision-Language Reasoning Data- efficient image captioning of fine art paintings via virtual- real semantic alignment training.Neurocomputing, 490:163– 180, 2022

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:42.635288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.068238Z digest=sha256:586658a27a13fef8354868dcd2e90c48863d95282941f32e8b543ab41d111751

Observation 6cdcf8fc-e786-4aae-83d4-5b6c3ff86082 · outbound

This paper cites Class-agnostic object detection with multi- modal transformer.

Understanding Museum Exhibits using Vision-Language Reasoning Class-agnostic object detection with multi- modal transformer

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:42.618986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.090766Z digest=sha256:a52ab2c43ab90b80efd539a8483f05bf6b08c67382754273dd8b4ba274ddfa4a

Observation cbfab324-7c10-42f4-b27a-7158d197c08a · outbound

This paper cites Fine-Grained Visual Classification of Aircraft.

Understanding Museum Exhibits using Vision-Language Reasoning Fine-Grained Visual Classification of Aircraft

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:39.096188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:39.096188Z digest=sha256:9c43d9f55f687d792c6324b51ae3c9d8821f069a0ed145ba8a7fe9f20f71d73d

Observation 0748b918-e7fc-4108-9cd0-4ce397e63939 · outbound

This paper cites A multi-world ap- proach to question answering about real-world scenes based on uncertain input.Advances in neural information process- ing systems, 27, 2014.

Understanding Museum Exhibits using Vision-Language Reasoning A multi-world ap- proach to question answering about real-world scenes based on uncertain input.Advances in neural information process- ing systems, 27, 2014

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:42.603957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.100756Z digest=sha256:e08a99cf16e20ee6551dd5e2e68c41b73ed6e95d00aa5cb04223f974c91f05bb

Observation e6246572-4564-4010-9ec5-c189df868763 · outbound

This paper cites Ask your neurons: A neural-based approach to answering questions about images.

Understanding Museum Exhibits using Vision-Language Reasoning Ask your neurons: A neural-based approach to answering questions about images

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:42.588756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.105087Z digest=sha256:cc36f8afd313a649679016ea473d7c59e80fada76454e3bb5a42af1bc0701e5b

Observation 3c9a29dc-8cc1-4c7a-81d1-7531986fefe4 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Understanding Museum Exhibits using Vision-Language Reasoning Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:42.573678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.109382Z digest=sha256:933ebf1f6d97d2f45f5bc0f0d374023e5a031d61b8f1387070b3104a84524d5b

Observation eb64fad0-9d73-4b02-86f3-debbafd07243 · outbound

This paper cites Taylor & Francis, 2008.

Understanding Museum Exhibits using Vision-Language Reasoning Taylor & Francis, 2008

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:42.425263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.113376Z digest=sha256:a51d224509478cf1b3a8b53b9531ffba88617db2233c35dec1d4afe5eb92b34c

Observation a19689f0-cd54-4a7e-abc3-667faf400d9a · outbound

This paper cites Foundation Model is Efficient Multimodal Multitask Model Selector.

Understanding Museum Exhibits using Vision-Language Reasoning Foundation Model is Efficient Multimodal Multitask Model Selector

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-12T04:28:40.216299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.117862Z digest=sha256:e641e04e62cf2694041571df3d2ba6e452800dbac53a4ffe64c33d595945542c

Observation f9aafe1e-cd7c-4ad7-b6d2-8e9df7233284 · outbound

This paper cites The rijksmuseum challenge: Museum-centered visual recognition.

Understanding Museum Exhibits using Vision-Language Reasoning The rijksmuseum challenge: Museum-centered visual recognition

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:42.318571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.123393Z digest=sha256:936271545cb97a2889f23c47401cefdf1494301cb66b1ddb54f1eea18770644a

Observation b784edc1-0dd0-4ff1-a425-551c1f242782 · outbound

This paper cites Encyclopedic vqa: Visual questions about detailed properties of fine-grained categories.

Understanding Museum Exhibits using Vision-Language Reasoning Encyclopedic vqa: Visual questions about detailed properties of fine-grained categories

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:42.302337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.127890Z digest=sha256:07b46d3d0300aff89a97b9f5352986eda488230503a323f4f9e702e8a88d4e6e

Observation 26c6844f-aa70-497a-9df4-85d53dd1cb30 · outbound

This paper cites A dataset and a con- volutional model for iconography classification in paintings.

Understanding Museum Exhibits using Vision-Language Reasoning A dataset and a con- volutional model for iconography classification in paintings

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:42.285300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.132310Z digest=sha256:3516676bd57fbcbdb7e2c6a6f726fbf71ef35cd41a4984ab37c1a46742398bc8

Observation 51c3790b-0b65-4663-b441-bc3ac2f5e2e5 · outbound

This paper cites Expanding language-image pretrained models for gen- eral video recognition.

Understanding Museum Exhibits using Vision-Language Reasoning Expanding language-image pretrained models for gen- eral video recognition

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:42.267591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.167324Z digest=sha256:acf885a0ac44327c025a5a0d5d37064c62d36c55d127bae2e8fac508d2dad9d9

Observation 8b3dc484-7cf8-4189-9d71-02bb7a0dd433 · outbound

This paper cites A survey of geospatial semantic web for cultural heritage.

Understanding Museum Exhibits using Vision-Language Reasoning A survey of geospatial semantic web for cultural heritage

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:42.072177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.222939Z digest=sha256:d9238667862d91b0b4591ad0cd3d3a0ab7e475ddc3c5712c10220b59579889b3

Observation 7482a97c-7bc9-4211-b6ac-3a571966d964 · outbound

This paper cites Suppressing biased samples for robust vqa.IEEE Transactions on Multimedia, 24:3405– 3415, 2021.

Understanding Museum Exhibits using Vision-Language Reasoning Suppressing biased samples for robust vqa.IEEE Transactions on Multimedia, 24:3405– 3415, 2021

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:41.957607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.285954Z digest=sha256:ff0e5a2601d4ca36f5fad896ac91e8450aa92b7ea6a35e476c49696b65bbb49d

Observation dc5d31ef-e9c6-4d4b-9511-18ae44cab577 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

Understanding Museum Exhibits using Vision-Language Reasoning Bleu: a method for automatic evaluation of machine translation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:41.941560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.314365Z digest=sha256:3a87d44ba7e79a24f20bc602a1f6c04452c09c852ed7f12e53b4d08fb86898d9

Observation b795c7f2-ee04-4d8e-a045-5abfd8ee07a3 · outbound

This paper cites Combined scal- ing for zero-shot transfer learning.Neurocomputing, 555: 126658, 2023.

Understanding Museum Exhibits using Vision-Language Reasoning Combined scal- ing for zero-shot transfer learning.Neurocomputing, 555: 126658, 2023

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:41.925823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.319230Z digest=sha256:1af27e8c2d1a37b347526d292ae84b87d1cf633720e510e154edb4680d6cf82f

Observation b3ce181e-f90d-4a03-993e-6941ebef5c9b · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Understanding Museum Exhibits using Vision-Language Reasoning Learning transferable visual models from natural language supervi- sion

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:39.324548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:39.324548Z digest=sha256:523f8165f7b8f098c3e0b14a46ef1dbdcd8984b3078a9e8181627942bf682331

Observation 3393842a-f175-4673-bf18-b1edf5d1071e · outbound

This paper cites Zero-shot text-to-image generation.

Understanding Museum Exhibits using Vision-Language Reasoning Zero-shot text-to-image generation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:39.330273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:39.330273Z digest=sha256:a87735ce85be8d0b8afe733a725223c2606c540ad5a5a9a33e7bf254188a4501

Observation 74699bd3-f761-4f52-9766-8131ff3b5a6e · outbound

This paper cites Fine-tuned clip models are efficient video learners.

Understanding Museum Exhibits using Vision-Language Reasoning Fine-tuned clip models are efficient video learners

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:41.730928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.334548Z digest=sha256:41611c066115738a435742506a8f7397850c22f6bd3720a4d099e09f673e0ca7

Observation dee15f97-4987-4f51-a781-fd26ff89b7d8 · outbound

This paper cites Stylebabel: Artistic style tag- ging and captioning.

Understanding Museum Exhibits using Vision-Language Reasoning Stylebabel: Artistic style tag- ging and captioning

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:41.713807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.339012Z digest=sha256:65aba08fd766ec87c3fd9c3e058c01af4a33123e2798de1307fde00d9cf5c1e8

Observation c979b1ca-624e-4080-8416-cca8ccfa4758 · outbound

This paper cites Viske: Visual knowledge extraction and question answering by visual verification of relation phrases.

Understanding Museum Exhibits using Vision-Language Reasoning Viske: Visual knowledge extraction and question answering by visual verification of relation phrases

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:41.698238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.343060Z digest=sha256:5bae33ffeb9b790b717ab4d9c5da244d652c0109e02d2738a9a23a9aca8db029

Observation c7124083-46a9-4827-bd99-77b995396922 · outbound

This paper cites A dataset for multimodal question answering in the cultural heritage domain.

Understanding Museum Exhibits using Vision-Language Reasoning A dataset for multimodal question answering in the cultural heritage domain

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:41.559034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.347348Z digest=sha256:a7ca8b67e8b3b3c471f3187f4caf23d9393c4e459d2c937bdfe440bca8848916

Observation e259545d-07a7-4e8d-b1c1-de2dc9a6d980 · outbound

This paper cites Towards vqa models that can read.

Understanding Museum Exhibits using Vision-Language Reasoning Towards vqa models that can read

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:39.351442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:39.351442Z digest=sha256:59d296bd8e5a93dcd8468cade4291298f55225392b2a4f7aafe0a5a468d97c6e

Observation d8483041-520a-4f3e-835b-c6508cc17e72 · outbound

This paper cites Bioclip: A vision foundation model for the tree of life.

Understanding Museum Exhibits using Vision-Language Reasoning Bioclip: A vision foundation model for the tree of life

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:39.355265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:39.355265Z digest=sha256:03381eaacc4b637765e88574596a6914c87a335233196206c376438e7bcfc91e

Observation bba0832d-4e52-4df3-b4e9-8017004c2042 · outbound

This paper cites Omniart: a large- scale artistic benchmark.ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), 14 (4):1–21, 2018.

Understanding Museum Exhibits using Vision-Language Reasoning Omniart: a large- scale artistic benchmark.ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), 14 (4):1–21, 2018

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:41.384474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.360221Z digest=sha256:a9fff305bbfd352c3afeb30176a69eb25b59d76782f0379871651c4f0db00886

Observation a9614f6d-52ae-4f24-bfa0-b516f6e0cca1 · outbound

This paper cites MultiModalQA: Complex Question Answering over Text, Tables and Images.

Understanding Museum Exhibits using Vision-Language Reasoning MultiModalQA: Complex Question Answering over Text, Tables and Images

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:39.364514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:39.364514Z digest=sha256:081771ae147b8d78d5cf3f8d6eb863f4143cddd942e49599774346a1387374c1

Observation c06e7cdc-e446-442e-b6e8-a092dee5fe18 · outbound

This paper cites Ceci n’est pas une pipe: A deep convo- lutional network for fine-art paintings classification.

Understanding Museum Exhibits using Vision-Language Reasoning Ceci n’est pas une pipe: A deep convo- lutional network for fine-art paintings classification

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:41.291712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.369035Z digest=sha256:88c46bd6f7bdc2478586fea19f7fbe08cfd73ca77fd94a1fbc0ed48948a3afa6

Observation cd9533cc-20c9-4381-b957-369923c57141 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Understanding Museum Exhibits using Vision-Language Reasoning Gemini: A Family of Highly Capable Multimodal Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:39.373214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:39.373214Z digest=sha256:2927f8717a84267e534901ef9188897df95d655203c969a4967ce7db94a5c1cd

Observation f73830ac-b551-4adf-b5fa-53ed64e0ebd8 · outbound

This paper cites The caltech-ucsd birds-200–2011 dataset.

Understanding Museum Exhibits using Vision-Language Reasoning The caltech-ucsd birds-200–2011 dataset

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:41.178391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.377956Z digest=sha256:5f56676119cb11ea31903826bb33407c9458ddfc6e7215f270eb5cbb5b199527

Observation 9dbba161-2b99-46c5-a7e2-96ae1cca3792 · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

Understanding Museum Exhibits using Vision-Language Reasoning ActionCLIP: A New Paradigm for Video Action Recognition

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:39.410964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:39.410964Z digest=sha256:d7394bdd7816ef030e482764da9be91474e7dc4195bc303437d82f22014bda5a

Observation 3abc7c45-00ef-4af2-a55a-71ade113b9d1 · outbound

This paper cites Explicit Knowledge-based Reasoning for Visual Question Answering.

Understanding Museum Exhibits using Vision-Language Reasoning Explicit Knowledge-based Reasoning for Visual Question Answering

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:39.481628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:39.481628Z digest=sha256:2343ba13813aed484c458f157368e756e030c7076774a65988ac37f7a77bdf8e

Observation d5deeac1-c870-458c-aa33-2d7f6d327429 · outbound

This paper cites Fvqa: Fact-based visual question an- swering.IEEE transactions on pattern analysis and machine intelligence, 40(10):2413–2427, 2017.

Understanding Museum Exhibits using Vision-Language Reasoning Fvqa: Fact-based visual question an- swering.IEEE transactions on pattern analysis and machine intelligence, 40(10):2413–2427, 2017

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:41.162644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.508493Z digest=sha256:e78797ddacda6016a2dfd06f1ad5a4122b1169eec2f829ef1d3e6ade9359e0e3

Observation d8ad7152-db21-408c-8a55-867df9072f73 · outbound

This paper cites MedCLIP: Contrastive Learning from Unpaired Medical Images and Text.

Understanding Museum Exhibits using Vision-Language Reasoning MedCLIP: Contrastive Learning from Unpaired Medical Images and Text

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:39.513298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:39.513298Z digest=sha256:f535e48da73f32f593deeb46d2f9aba4994d75297dd8a4d8944eb9e63d370857

Observation bde93852-87da-444a-9af4-350c3fbebaab · outbound

This paper cites Im- proving clip fine-tuning performance.

Understanding Museum Exhibits using Vision-Language Reasoning Im- proving clip fine-tuning performance

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:41.145829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.519200Z digest=sha256:33911c5f8cc7d62aa31e289a3291113a09930c627f01fc8ff471183d8d72281a

Observation 3e1bf36c-3379-4ffa-80aa-57ff33191ce3 · outbound

This paper cites Bam! the behance artistic media dataset for recognition beyond photography.

Understanding Museum Exhibits using Vision-Language Reasoning Bam! the behance artistic media dataset for recognition beyond photography

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:41.130513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.524116Z digest=sha256:602bb362ba0eb40ea0eb5c99658399a1c75c2faee2cfa5ef419844d24b2c50eb

Observation ad4bd166-652f-48e3-a8b3-01743e97da80 · outbound

This paper cites Ask me anything: Free-form vi- sual question answering based on knowledge from external sources.

Understanding Museum Exhibits using Vision-Language Reasoning Ask me anything: Free-form vi- sual question answering based on knowledge from external sources

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:41.115755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.528325Z digest=sha256:951d4fd942fc36949012205601923764f2906455d9447f627854555e3acc94ba

Observation c61d608d-9033-4403-bef2-4b8ff3d2fab9 · outbound

This paper cites Language bias in Visual Question Answering: A Survey and Taxonomy.

Understanding Museum Exhibits using Vision-Language Reasoning Language bias in Visual Question Answering: A Survey and Taxonomy

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-08-12T04:28:40.036189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.532262Z digest=sha256:0da18802c6b5044dae53cf526db652c4b6fae365d2dea413073c36f123257834

Observation 9481fe34-7874-4896-803e-f8985656b4b9 · outbound

This paper cites Lit: Zero-shot transfer with locked-image text tuning.

Understanding Museum Exhibits using Vision-Language Reasoning Lit: Zero-shot transfer with locked-image text tuning

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:41.049754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.537077Z digest=sha256:f585246fd32696d059ea1ccb750b69ca146fd3d9c8df0c120ac38fc42536474c

Observation 6333fb85-8626-4a9a-b17e-2f6f341026e1 · outbound

This paper cites The iMet Collection 2019 Challenge Dataset.

Understanding Museum Exhibits using Vision-Language Reasoning The iMet Collection 2019 Challenge Dataset

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-08-12T04:28:39.948829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.541446Z digest=sha256:ed27d97a1b5085244d9054bcee1a3b2820bed2474f4b9c65283e05f30258ef7f

Observation 6f322d1f-f461-4a98-a33e-59c7ec488409 · outbound

This paper cites Tip-Adapter: Training-free CLIP-Adapter for Better Vision-Language Modeling.

Understanding Museum Exhibits using Vision-Language Reasoning Tip-Adapter: Training-free CLIP-Adapter for Better Vision-Language Modeling

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:39.545705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:39.545705Z digest=sha256:6ea9179d424a36e2db2a6bc52094afc25385ae9d81bb24989fb8d64fec7883e6

Observation 0795a025-0709-4b0d-a524-20aa8616c062 · outbound

This paper cites Extract free dense labels from clip.

Understanding Museum Exhibits using Vision-Language Reasoning Extract free dense labels from clip

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:40.902216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.622118Z digest=sha256:12e866827a2c7836d25fcc06e39be44894da8069a91b7924b1a6d80e34f62c01

Observation f0398813-e35d-4c93-a692-fdfb94ee4b49 · outbound

This paper cites Conditional prompt learning for vision-language mod- els.

Understanding Museum Exhibits using Vision-Language Reasoning Conditional prompt learning for vision-language mod- els

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:39.722009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:39.722009Z digest=sha256:5c031bf693b8d9f3cad0f6c715e567038f535a60f853c7e43cd10fd9f084ba84

Observation 80673caf-d90d-4d30-a895-4da7219100f5 · outbound

This paper cites Learning to prompt for vision-language models.In- ternational Journal of Computer Vision, 130(9):2337–2348,.

Understanding Museum Exhibits using Vision-Language Reasoning Learning to prompt for vision-language models.In- ternational Journal of Computer Vision, 130(9):2337–2348,

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:39.774502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:39.774502Z digest=sha256:df2effcaa06fbc721bc1e14a55d3d68cf82f43cdf7fd3afca1d9f4478558c014

Observation 7af21584-2ef2-4c9b-88e2-43a4b8a837ac · outbound

This paper cites Detecting twenty-thousand classes using image-level supervision.

Understanding Museum Exhibits using Vision-Language Reasoning Detecting twenty-thousand classes using image-level supervision

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:40.867945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.778885Z digest=sha256:3d1c610ef491f4b550fb46330bc98a2ac43bcfdcf524ab702371abe8daf1a7df

Observation 76d54b7f-72ef-45a1-a906-56d441fc0fe2 · outbound

This paper cites Visual7w: Grounded question answering in images.

Understanding Museum Exhibits using Vision-Language Reasoning Visual7w: Grounded question answering in images

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:40.853371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.783992Z digest=sha256:618f0b8e82483497dde9b26e744bef6b0e8b925185402bbeb37187e25f28469b

Observation 8433939e-def4-4455-b368-f57453b87aea · outbound

This paper cites LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model.

Understanding Museum Exhibits using Vision-Language Reasoning LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:39.788524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:39.788524Z digest=sha256:2d98c03b01e285e719b3bafd87cdab7d04bbaaf16aeab68ae49b6a7a69649037

Observation 9e1d174e-af6e-49ce-ba92-ccce016abe46 · outbound

This paper cites • These aggregators provide access to extensive digitized collections from major museums across Europe and America and offer structured data through platform- specific APIs.

Understanding Museum Exhibits using Vision-Language Reasoning • These aggregators provide access to extensive digitized collections from major museums across Europe and America and offer structured data through platform- specific APIs

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:40.751013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.793770Z digest=sha256:7a27d4622f2a7bfc9cf795c25f5ab86733aaf7633c4aeffa2e3d91129cf1ad63

Observation ebe8df6b-8dac-4910-a943-ebc0d528d8f1 · outbound

This paper cites Curation in- volved minimal edits: removing redundant attributes (in- ventory numbers, bibliographic info); extraneous symbols and numbers.

Understanding Museum Exhibits using Vision-Language Reasoning Curation in- volved minimal edits: removing redundant attributes (in- ventory numbers, bibliographic info); extraneous symbols and numbers

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:40.619143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.798428Z digest=sha256:cddc6882a2df2017c414becb8c52cd166242e2d80ad4fe6acb4ad4570b58fc7d

Observation 91b95f06-7661-4bdb-bbf8-2a5fae826243 · outbound

This paper cites an unresolved cited work.

Understanding Museum Exhibits using Vision-Language Reasoning Unresolved cited work

Reference 98

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:28:40.604256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.802733Z digest=sha256:270845789e75411feba1fd8201eacc1fc3bcaa84f9569de381b1300a04c20400

Observation 68c51512-01ec-4ae6-99e5-3cc2bb9dce3b · outbound

This paper cites Which primary material is the object made of?.

Understanding Museum Exhibits using Vision-Language Reasoning Which primary material is the object made of?

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:40.587709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.806871Z digest=sha256:77aec053116f59099a6e5991c7d125445abe72f7e522f34b8ca360f0ba328700

Observation fc021a6f-b5f5-4beb-af3d-2e790aefa241 · outbound

This paper cites For each object, we now have a list of images and a set of question-answer pairs, omitting the answers for which the value is not known.

Understanding Museum Exhibits using Vision-Language Reasoning For each object, we now have a list of images and a set of question-answer pairs, omitting the answers for which the value is not known

Reference 100

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T04:28:40.572404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T04:28:39.810901Z digest=sha256:02af2352250b103ed0ea8b461105ada8f10857637d0d0e4111e8d2fc3abeae07

Pith citing papers

No inbound Pith citation observations are available.