Pith. sign in

Paper Citation Record · LEDGER

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models?

As of 12 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2502.06600.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06600 v2

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:58:48.251828Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 83f9ae22-6664-4e3e-9ae7-facd58bf7ce4 · outbound

This paper cites From Images to Sentences through Scene Description Graphs using Commonsense Reasoning and Knowledge.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? From Images to Sentences through Scene Description Graphs using Commonsense Reasoning and Knowledge

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T14:58:48.107848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:58:48.107848Z digest=sha256:90b10f49f7fd544d4533ac44eeaaa357df85258ecab777dce7ddf991f69e3e6e

Observation 8d4fc923-288d-469e-8913-2e82bf926c33 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.703743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.112701Z digest=sha256:4702d4758a001bf9cb6c97378f69fa8bb782da3ab617775c2d9e07775e7ad729

Observation 6b138ef1-d3f6-4927-8fbb-ec2a0c8e06e4 · outbound

This paper cites Tower: An Open Multilingual Large Language Model for Translation-Related Tasks.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Tower: An Open Multilingual Large Language Model for Translation-Related Tasks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T14:58:48.116638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:58:48.116638Z digest=sha256:b01e2373d352ca31ed980f3a43cd96ec912ac352d7f284abf6c1316cf770a566

Observation 4e69acfb-352d-46d3-80bc-8845f2411477 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.693345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.120954Z digest=sha256:bfbd29872ae128399f70326aad818b95753ef1c56067cc9c6eba63186b40b8d0

Observation ab46c83d-b001-4dd1-8c64-eeadb081bf10 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.683627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.124736Z digest=sha256:a8069e4ebaa032c5a48ea5ae58da9b68e9409d9daca7a9d4618c01e1f9131aca

Observation 58f9782a-809f-4fa1-933c-d12dff79cff4 · outbound

This paper cites Evaluating Image Caption via Cycle-consistent Text-to-Image Generation.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Evaluating Image Caption via Cycle-consistent Text-to-Image Generation

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-08T14:58:48.351746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.128379Z digest=sha256:f7e76b878249f5bbb64812a9029af24ce718d77477d796b9a735628da44bc012

Observation 1c9b6edf-f166-431e-9e0e-68708a488229 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.673976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.132436Z digest=sha256:5c28b0eb5ce2b6a712b831b58c8cac22a545995d9d5ac85d191211a4e1d28df0

Observation 7d6d3166-20b1-43c9-ad06-3b9f14f4478f · outbound

This paper cites mBLIP: Efficient Bootstrapping of Multilingual Vision-LLMs.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? mBLIP: Efficient Bootstrapping of Multilingual Vision-LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T14:58:48.135855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:58:48.135855Z digest=sha256:110840e3730d4546a524e63d797b68f462d2e10d5129854397a2ef4d1d129af7

Observation 054bdc68-448b-4a6b-900d-c853408c8bd1 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.664055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.139766Z digest=sha256:abfcfc17a5aa59533c16ac51f3e33efa798b6711a163f68ca17b9a622f8a6df4

Observation dcd3ac4a-2f04-42d5-86ba-ed55c2335154 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.653900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.143162Z digest=sha256:599f1c9bf643435e665197b748583afdd96263be0e2d203a3f5267cb98b8a939

Observation e8b7e6cb-d43e-4f60-bd21-d1e3534c6313 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.643350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.146463Z digest=sha256:8b62537dd3796eeadbf588b54af4a51350be4ae4e3a8dc12913cf600f7dea7f5

Observation f4cd076b-50e7-4c45-88e4-127b8d59cb26 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.632497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.149925Z digest=sha256:e8d3437562e725cf3bcab6cd0253496793714d6b85384791577eaf41708640e2

Observation f86b738c-e88b-462f-84b7-616c8fbf42da · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.622305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.152564Z digest=sha256:06d19dc1c5e2f806b94297e497a066f172e333e8cebc97f78aa5f31d6df89a1f

Observation 0509c59d-0ce0-4f90-b056-583aad488a4d · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.611322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.155242Z digest=sha256:cc21accfb2fabe0fb1159764338a6ff7a0c1716247d8d060574752004b8ef8cc

Observation 262e7b40-e123-4ea2-a546-18028f092554 · outbound

This paper cites FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T14:58:48.157795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:58:48.157795Z digest=sha256:dd677b7565d850937252620665c36e75401b4e3d13e45a6f21a2ba1e7d4d0b04

Observation e94ee4ea-f377-4b0f-8732-a7ba9c3bb5ab · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.600724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.160659Z digest=sha256:9bd3da14a8b919f21587bb3117dffe62b3c7cbcd06c3e412cfa138ca73ddd88c

Observation 5dd98dfd-ec45-4b48-b16e-190a8b5ee264 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.589632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.163384Z digest=sha256:fca378069cd7a533e0e1af90cd247fa7a2598347eb44eea3a107ae494a56fb24

Observation 2a499ef0-c798-4326-958d-05772ebe347c · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.578526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.166210Z digest=sha256:a3621cd14dc952c3ca4091ed329a589dcb36c488db44636247f7756ad20f9c06

Observation dcabe155-8dab-4e76-91c0-868983d29dd0 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.567322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.168871Z digest=sha256:fee5491858dde48458c69d678a4f886b6e72c47feb635206b1b9dd57302c55de

Observation 52e30fe1-49a0-4c7a-a590-48eef4db58df · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.556817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.171937Z digest=sha256:d8fc873d6858d26fd6b4cc5aec8d9f61ce6336a50e9d1ade46da6549acdcaad1

Observation e827993b-975c-49bd-b84a-0efffdec7d0e · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.545931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.175223Z digest=sha256:cfde53689356436db9c0d3fce5b1df1d522df48056500cb1605ed98a69178953

Observation cd7d2f62-94a8-4d85-a4a6-1f32bfd8f481 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T14:58:48.178520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:58:48.178520Z digest=sha256:315c31ab947e8cd5840c674b569ecd6db85fcce30ee015fbc97b273e7ae15b08

Observation 720e87c9-ba5b-4ca1-9252-aee9ddfa6290 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.529185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.181934Z digest=sha256:ae7eb82f6d0ce3268252a1e64628ea6b33ae1a8e367fc6f3634f3a72e3ff7862

Observation 2c560504-c6a2-4af7-a220-b923ba8bf3ad · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.520419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.185359Z digest=sha256:54fcd9c158398ccb55747069ffc35f75ae5e24b5a38a6a3d92cf111895993462

Observation 91ae5dad-3086-4095-902b-c279a8d85fa4 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.511567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.188730Z digest=sha256:363dd8b2b4294c94ee432ac8dc2a671ddc6a067f2d391f18bfdd41a8a4e0c140

Observation 93864a7b-81d6-4f6b-9c5d-fe781abf2d33 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.502102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.192388Z digest=sha256:1657fa90c11fea6f8b66f1ed2637e391b2f437e927a5cde6773ec44e461105bc

Observation 3ff3180b-db37-4e66-a805-e6ba78042af0 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.492448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.195667Z digest=sha256:146b2111ebaa36d0f752bee1b197708425a9339322765c46a2f8d874183323b6

Observation 0128b44b-877f-49ea-b537-4d2c28710619 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.482354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.199242Z digest=sha256:0cc79930a4c3b9a664db7d3f8c3434b572b75c13c2a69c96cfb47cd4879927f4

Observation f4db2265-c362-4c07-bb56-8c486bd1e2b1 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.472530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.202653Z digest=sha256:2c79c87369d90204571f4850bd56f61dd2b2b0fa1e3adac670a74ef23a82cc14

Observation 88661f33-1a84-4df6-8210-b8a6f49236f4 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.462397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.206031Z digest=sha256:24076704c331ba50f0c4a01e5a00c499ec49c2b834c65cfbb021f685b9c4ff64

Observation 70bab504-fbe1-4dcd-94b0-de0d35278fbd · outbound

This paper cites Positive-Augmented Contrastive Learning for Vision-and-Language Evaluation and Training.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Positive-Augmented Contrastive Learning for Vision-and-Language Evaluation and Training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T14:58:48.209337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:58:48.209337Z digest=sha256:5aa2b298435ed282ee18b4dc6e01d05b191e0da5d26e1bcb38e103bdcd96f133

Observation 30ead303-1f8d-4752-8115-bc8cd0becdc4 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.452291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.213093Z digest=sha256:a1c0233de53684e69d62001dfdaf188ad8d66c562209be5d974e5857befbcf38

Observation a4d74aca-a523-4c46-b549-d18939fdc182 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.442298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.216626Z digest=sha256:9d0f1c0a09f0b570044c4dceaf1d254615090c9fdc5cbc3fb035ea0594a3fd3e

Observation 91a1aeb2-7791-4bb0-806e-49a13e7c416f · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.431991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.219896Z digest=sha256:ca492d5fe9df6b0a102d3d779e488b3fe6fc45a0e8ce7e1d69976cf39f0c37ca

Observation 5444d6e4-1a51-4718-8f98-e699647f718d · outbound

This paper cites G-VEval: A Versatile Metric for Evaluating Image and Video Captions Using GPT-4o.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? G-VEval: A Versatile Metric for Evaluating Image and Video Captions Using GPT-4o

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-08T14:58:48.309144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.223313Z digest=sha256:0afa27333045b37d84b6e4605030fe3921d94215332e79eb80156f0a54aed64f

Observation 80b7e105-15a2-437c-83b2-7746e24bfdb2 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.421535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.227281Z digest=sha256:2f8e8f2f8623185f947d0a80ec047afa1d8bc454342ec337c10dd8a4c5920656

Observation 8e4b798c-3d69-4815-acbf-c5c5db183cf4 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.411384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.230725Z digest=sha256:1b7d756ce7c180e686761e6bd6442ceca8591660d5b32f3c8de4232370fb9df7

Observation 9fede902-8fa7-4281-a23a-e75100d5d95e · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.401688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.234045Z digest=sha256:f7caa2518946e3113e7a188fd38b84a878b7cef1e798b748f45e57192550dc0c

Observation e0a843d9-4994-4c1b-8a96-0b8dd8c8be83 · outbound

This paper cites Visual Entailment: A Novel Task for Fine-Grained Image Understanding.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Visual Entailment: A Novel Task for Fine-Grained Image Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T14:58:48.237478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:58:48.237478Z digest=sha256:006e7cfecf6af6fed0dfcc2aaeaed4b17eb0ccc88e2d333a9c34f8f626ad293d

Observation 3e996417-493e-4955-a628-82ecb0dba27f · outbound

This paper cites When are Lemons Purple? The Concept Association Bias of Vision-Language Models.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? When are Lemons Purple? The Concept Association Bias of Vision-Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T14:58:48.240959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:58:48.240959Z digest=sha256:f85bd7747b1542154ada9cd880abb8167796cff9683b8ce27f09cb2f6208994f

Observation 55d6fdf1-22a0-49e7-b462-f7684214929e · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.392417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.244606Z digest=sha256:5fa1fee40000d72b09b42ecd9d04bae6cb48890eeb3acf1bdd69a5029e72c8d5

Observation d455749d-bc74-4cf7-b9b3-4288580f2fab · outbound

This paper cites online" 'onlinestring :=.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? online" 'onlinestring :=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T14:58:48.248032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:58:48.248032Z digest=sha256:7f347d016e0cc69199d683baf895b4f106153112ceb266af7af91edb4e2d96ae

Observation ecbced9a-afae-4de3-ac64-1f0c38fbdb2e · outbound

This paper cites write newline.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? write newline

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T14:58:48.251828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:58:48.251828Z digest=sha256:cf70be1d249c1b041bf60b7c3b4f4c7bdf063df8c856d63dfb8255682c05afbe

Pith citing papers

No inbound Pith citation observations are available.