Pith. sign in

Paper Citation Record · LEDGER

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning

As of 18 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2507.20411.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20411 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:51:16.765188Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact2
  • verified fuzzy19
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 85bc558f-938b-4dd0-b042-62e02b334afc · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.353819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.576514Z digest=sha256:2c45ad163687f0a4b86eb0ce71fdd5dad1f857c2c6919e5bb649acc69a30b03c

Observation 47ea6d02-c789-4414-a71a-eb1e62ded684 · outbound

This paper cites mCLIP : Multilingual clip via cross-lingual transfer.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning mCLIP : Multilingual clip via cross-lingual transfer

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.342951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.581430Z digest=sha256:74a63b73f3fe329ce284202a20dd8da16e6c3d476b0a2b65e5e9fefd0c0b6682

Observation 6a0dbf99-f686-4062-af51-915912b086e7 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.585591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.585591Z digest=sha256:5e124f6f6628743da29e574e9684fee8a2368e5d52e7624052097dddfe340190

Observation c314bc97-07df-4dc5-a2ff-61a0536bd5a7 · outbound

This paper cites Vicuna: An open-source chatbot impressing GPT-4 with 90\ See https://vicuna.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Vicuna: An open-source chatbot impressing GPT-4 with 90\ See https://vicuna

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.330972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.590190Z digest=sha256:bb5b58a4dffb54ab72f7ee26d676df29a8d073b2439cb06cb4b3cdb2c65e146e

Observation ecfe5081-d108-4786-af45-3b16dd096d42 · outbound

This paper cites The Faiss library.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning The Faiss library

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.319457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.594311Z digest=sha256:6778a62406e509d62dc43b2e20a731c82190898c8f75b5082a9ccde9c2a7440d

Observation 846e8690-b5db-4738-954d-6530ab540ff4 · outbound

This paper cites A survey on rag meeting llms: Towards retrieval-augmented large language models.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning A survey on rag meeting llms: Towards retrieval-augmented large language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.309117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.598553Z digest=sha256:671534f0e9837abfdc2a690854bac16b364385acc45ffdf3e1afaa9a82115c65

Observation 17276bd8-4f06-4674-b3f3-105646f266da · outbound

This paper cites m BLIP : Efficient bootstrapping of multilingual vision- LLM s.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning m BLIP : Efficient bootstrapping of multilingual vision- LLM s

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.607728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.607728Z digest=sha256:11530c09d1400cdf1a1ac7582da5444f78070cd2a1a6df676235f4ab8f781d8c

Observation 09c2b0af-e638-4a32-88da-12c53adbc512 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning LoRA: Low-Rank Adaptation of Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.612200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.612200Z digest=sha256:3644c3feb44f4f7e5e2ba5d58db10e3cc667027592b45cc2ef3e88f4ba371f98

Observation fe708efe-721b-44df-a740-84d2c62d4573 · outbound

This paper cites A Survey on Retrieval-Augmented Text Generation for Large Language Models.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning A Survey on Retrieval-Augmented Text Generation for Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.616754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.616754Z digest=sha256:ba4181cdca379758843153c63a42565fef1b733a20bc98e11d678ecdd616322a

Observation 8cd63b7e-5e07-44f3-8f5b-77e4685657c2 · outbound

This paper cites Jieba Chinese text segmentation.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Jieba Chinese text segmentation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.299339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.621083Z digest=sha256:60cc1fbfa897e13a6c27a22e68af81244c574134d6a81b2c2d3389e9f08890ff

Observation e3736bd0-a3b5-4a49-80ed-ab15cf9f287b · outbound

This paper cites JEEM: Vision-Language Understanding in Four Arabic Dialects.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning JEEM: Vision-Language Understanding in Four Arabic Dialects

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.624934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.624934Z digest=sha256:09b9b634443829bd0ac439cd5b9282b5f8931ff04b3cf7f0d53fdc27323a96ee

Observation 33061be7-7d06-4e84-8641-a6c5f7f813a7 · outbound

This paper cites Jina CLIP: Your CLIP Model Is Also Your Text Retriever.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.629099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.629099Z digest=sha256:4e151c2bd81d3a9b3b8faf711cc889c99eabc070cf21a17d77dff22c55747a18

Observation fd59b8a9-4a33-46f3-b852-e12dc5597075 · outbound

This paper cites MeCab: Yet Another Part-of-Speech and Morphological Analyzer.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning MeCab: Yet Another Part-of-Speech and Morphological Analyzer

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.289332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.632935Z digest=sha256:ca163cf9bc8c5934efcca83116c7269baf11011411f89b13164b33bc6a217a79

Observation 0dcfd769-727d-4592-9c0e-202d19b25e9c · outbound

This paper cites The IndicNLP Library.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning The IndicNLP Library

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.279721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.636690Z digest=sha256:a1e65974358e1ccba4cc5cd1fd724f71f5ebdfc41867a198b4090ca48864752b

Observation c7a22470-775d-43e5-9b8d-5d0631d8897b · outbound

This paper cites What matters when building vision-language models? Advances in Neural Information Processing Systems, 37: 0 87874--87907, 2024.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning What matters when building vision-language models? Advances in Neural Information Processing Systems, 37: 0 87874--87907, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.640675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.640675Z digest=sha256:99a2cdf0f8b35db85b0e28c3f36b0b76355fd2464ada528a726fde87aa5abf07

Observation faa9f74a-0fa5-4724-a853-9082954a96ac · outbound

This paper cites u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt\.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt\

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.263507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.644569Z digest=sha256:3111d21fa16e7b969467d0fa9290be65acabeeea5c17b23fd662299015760f3b

Observation 517d34f0-010d-46e9-8aaf-1d6de556f505 · outbound

This paper cites Evcap: Retrieval-augmented image captioning with external visual-name memory for open-world comprehension.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Evcap: Retrieval-augmented image captioning with external visual-name memory for open-world comprehension

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.253454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.648440Z digest=sha256:df185a3cea88bbb2637826dbd513b76224270ee2a5bfda997e75be029c916f80

Observation 74afd818-a77f-403c-8703-2959bce25ead · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.652000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.652000Z digest=sha256:475b2d65b0ff38c7004ea65b395fa67e9d6abedbf5fe28017db112a781dfbb02

Observation e7874f0f-5989-4824-b5ef-fc0ce10a10bf · outbound

This paper cites Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.656228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.656228Z digest=sha256:6b67cab9e070ab0ba3a7ccf081bc78e525bd1cd9c34479f4142d44342a17390a

Observation 81a63aee-6015-44d3-8eac-4a2bdcc64d73 · outbound

This paper cites Few-shot Learning with Multilingual Language Models.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Few-shot Learning with Multilingual Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.660499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.660499Z digest=sha256:1be78a0a3db394d3359a30f0cc0541daa84fde6b1bc0f77efc1ea0edcd03f8f0

Observation f8bb8130-295a-4539-86fb-3a3060895571 · outbound

This paper cites Visual Instruction Tuning.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Visual Instruction Tuning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.242606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.664943Z digest=sha256:492ed64ece78ca2ef03a2d1d02ad3ec1813f3b5734469da1d9970ef66956e1dc

Observation fadedef2-a003-4e9f-9924-bdd11e7842db · outbound

This paper cites LLaVA-NeXT: Improved reasoning, OCR, and world knowledge , January 2024.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning LLaVA-NeXT: Improved reasoning, OCR, and world knowledge , January 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.232573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.668652Z digest=sha256:20f50548114b7ebf460e04aa128189b15f38912845f0e83fee9d164ac4cb8c23

Observation 14378ac8-bae8-4aae-8198-fc83ac36b32a · outbound

This paper cites fugashi, a tool for tokenizing J apanese in python.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning fugashi, a tool for tokenizing J apanese in python

Reference 23

Resolution
verified exact
doi, observed 2026-08-15T17:51:16.819538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.672148Z digest=sha256:fa78b5c3c6b9e5b661a76edf5bcbe6b78dd6a0ba850880f95a3f20c746ae862f

Observation 5f30e3c7-ca49-4e1e-b8d7-3dbba3b0fcc1 · outbound

This paper cites ClipCap: CLIP Prefix for Image Captioning.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning ClipCap: CLIP Prefix for Image Captioning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.675792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.675792Z digest=sha256:ede5a1b7dc4744a2648edb7b8193c095b3d9b9ae99b359c52fb5f1f95beb6239

Observation 51b04839-ec5b-4803-9217-1de436dbe949 · outbound

This paper cites Crosslingual Generalization through Multitask Finetuning.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Crosslingual Generalization through Multitask Finetuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.679464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.679464Z digest=sha256:3fcaf78e16b076a43a1900570fab1507865a12405040d5344096859b4d974362

Observation e1ea91a9-9c40-41a8-953c-93c3e9a78fd3 · outbound

This paper cites GPT-4o System Card.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning GPT-4o System Card

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.683396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.683396Z digest=sha256:7e3a7722ca44296f2173838f2665842db75ed1da74eb6a51f5e1910d95e784a9

Observation 74b74614-b132-4d43-ad93-f59553198c90 · outbound

This paper cites X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.687278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.687278Z digest=sha256:73f04dc488e530d3ced310eaac33f165959b345d902023fffb0007f410315f08

Observation 4afd289a-bc32-4e9b-9e69-9ba034fdd7c7 · outbound

This paper cites BLEU: a method for automatic evaluation of machine translation.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning BLEU: a method for automatic evaluation of machine translation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.690863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.690863Z digest=sha256:f7cc62ad87595ce881c5f05ee4670c1c70ca518dc4d5ece3c5ad29996163f585

Observation 8f4cda7b-4054-4873-aff6-45af173add3a · outbound

This paper cites P y T hai NLP : T hai natural language processing in P ython, June 2024.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning P y T hai NLP : T hai natural language processing in P ython, June 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.221933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.694081Z digest=sha256:3e0d520f5f6a5fae7046689a34a9c7f02bed39038b84d7871a0c25ed02108bcd

Observation 910c2351-826b-4f7b-9d9a-dfe03e42fa35 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Learning Transferable Visual Models From Natural Language Supervision

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.698048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.698048Z digest=sha256:8404e9fa77b8fa781d8394cf949f29e7e73b6d6ffe5f2ae1e6124b09eadc4c41

Observation 3d508bb6-c1f8-4088-8ddc-3708d4d07367 · outbound

This paper cites LMC ap: Few-shot multilingual image captioning by retrieval augmented language model prompting.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning LMC ap: Few-shot multilingual image captioning by retrieval augmented language model prompting

Reference 31

Resolution
verified exact
doi, observed 2026-08-15T17:51:16.807453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.701713Z digest=sha256:b3e9d4cba8b04aa7eb355806996cb09791838c1d539332da1b71cab867469a07

Observation 1417952e-efa7-4b22-af52-96deeef79851 · outbound

This paper cites SmallCap : Lightweight image captioning prompted with retrieval augmentation.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning SmallCap : Lightweight image captioning prompted with retrieval augmentation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.208661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.705022Z digest=sha256:868485d795b6fc996c3065c4ce3b564d9c08ad9f196e48587ce294d3e56b9da2

Observation ee0ff7e6-5416-4f46-b12b-be8e792b8ad5 · outbound

This paper cites PAELLA : Parameter-efficient lightweight language-agnostic captioning model.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning PAELLA : Parameter-efficient lightweight language-agnostic captioning model

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.197085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.708324Z digest=sha256:1474671ff2103d2954ab836e2a002fad280617cc6f4e1de8b2c7d0a6d75b3ae4

Observation 348fd5ae-21c5-4771-bf9e-add921b3b67b · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.186408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.711713Z digest=sha256:9e67759b8d786ccf88e8329b6c6b32eb2d2282df9d54e428905a5f231170a47c

Observation 0b5cf44d-1522-4951-a8d2-79e1cdaf2a25 · outbound

This paper cites Stanford Alpaca: An instruction-following LLaMA model , 2023.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Stanford Alpaca: An instruction-following LLaMA model , 2023

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.174255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.715464Z digest=sha256:aeea8b0b0db3193e41c2ea4529e649644a19d4fc71b1b234e8eac597df8e4761

Observation ebd4d8dd-71aa-4cc6-b4ca-184996c1f9ce · outbound

This paper cites Thapliyal, Jordi Pont Tuset, Xi Chen, and Radu Soricut.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Thapliyal, Jordi Pont Tuset, Xi Chen, and Radu Soricut

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.719471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.719471Z digest=sha256:d6c1de2408f1a87ded65837766da18c3a368229ceac2b407831b4b04b90de50b

Observation 3daa7291-0601-4c6a-87b9-5db7461ab864 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning LLaMA: Open and Efficient Foundation Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.722756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.722756Z digest=sha256:b3c529a735527b873c6d06422e1150ffd3ffe40003c3736cf109e310d2ebbb3c

Observation f6a6cb77-0ca8-4be7-956e-aebe5d0cdd46 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.726343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.726343Z digest=sha256:a83e8a22502f5137696b9fcc862eca2358f6817a1253e26a1fd8e56d843dd858

Observation afa3a0b6-b36e-42fa-9d3e-fa14adb62fdd · outbound

This paper cites CIDEr: Consensus-based Image Description Evaluation.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning CIDEr: Consensus-based Image Description Evaluation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.729976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.729976Z digest=sha256:93efb497e93d9fe75daf15362e70e984f33f0adb69b0b472673d01ef8007fe8c

Observation b1049281-ed97-4bca-9353-afbc0cf6d975 · outbound

This paper cites mT5: A massively multilingual pre-trained text-to-text transformer.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning mT5: A massively multilingual pre-trained text-to-text transformer

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.733490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.733490Z digest=sha256:6f9d75e410619692b077f0e69ccf63892dd51a7ba9545c0191c29aa5778194ca

Observation bc8b0d94-0737-4c43-a615-beb1341521ce · outbound

This paper cites Retrieval-Augmented Multimodal Language Modeling.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Retrieval-Augmented Multimodal Language Modeling

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.736921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.736921Z digest=sha256:eb12b1cd6ef46403e9d3ed29139b09ea4090845f91448d79ed4e70cd78facb44

Observation b0f06c3a-7cae-4d47-8a28-c71914589cdc · outbound

This paper cites Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.740304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.740304Z digest=sha256:c63160574e25d52c0f8ccc6621168bfd856250922e1017a87c1d674167ac9456

Observation 61bf576e-1c58-47fb-964e-63ef406e5c32 · outbound

This paper cites MeaCap: Memory-augmented zero-shot image captioning.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning MeaCap: Memory-augmented zero-shot image captioning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.162716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.744297Z digest=sha256:10165611d15f02b0c211248435c0be38fe758090ef4e89a7ffe75ac4351b2f54

Observation a668557f-71c6-444b-9443-b16a5e1139cc · outbound

This paper cites Scaling vision transformers.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Scaling vision transformers

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.151494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.747605Z digest=sha256:1355bae3ed3222e23400db14494ee08fafcc53c655f33a4936856543e4564e4c

Observation 222e58a0-d79d-4940-8502-f02eb0b2b441 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Sigmoid Loss for Language Image Pre-Training

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.750717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.750717Z digest=sha256:190c2b0709fe29737aa3e9d5c96ef0461abe2c11af752319e885c16fec4c983b

Observation 02e3af54-5b06-47c0-aa73-30d05a6e3708 · outbound

This paper cites write newline.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning write newline

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.754102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.754102Z digest=sha256:1eee06c448930fcff773e217dcba5f1f01dfc39d9bdbfbff10d36f49b8eecd95

Observation 3a7dcc8f-acfc-4615-a000-b301174eea7b · outbound

This paper cites @esa (Ref.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning @esa (Ref

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.758121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.758121Z digest=sha256:dece3854e18d3b74a10552fca095cce022cd06f93eb05b4906ded0831cb55956

Observation f4ab3af3-559a-407a-85ce-f38e069fd2e4 · outbound

This paper cites an unresolved cited work.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.761894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.761894Z digest=sha256:8c2be418b6c494fc936ea90fdeaeb7c54c84a87d7108d111e7b5b0991593874c

Observation fc570287-e508-4ad7-90fb-4fe2b8459c0c · outbound

This paper cites an unresolved cited work.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.765188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.765188Z digest=sha256:e1526bacfce9b261a7081aa807808e9e97a6c08e61f3deb4dc5ba93cbeb85f8b

Pith citing papers

No inbound Pith citation observations are available.