Pith. sign in

Paper Citation Record · LEDGER

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning

As of 16 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2507.20411.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20411 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:51:16.765188Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact2
  • verified fuzzy19
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 85bc558f-938b-4dd0-b042-62e02b334afc · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.353819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.576514Z digest=sha256:4ccadf318b1a6d0281edfdbb49a2475e179ad9fa5a059507187b84f6a9c03b91

Observation 47ea6d02-c789-4414-a71a-eb1e62ded684 · outbound

This paper cites mCLIP : Multilingual clip via cross-lingual transfer.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning mCLIP : Multilingual clip via cross-lingual transfer

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.342951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.581430Z digest=sha256:270fc86e04864f10fb563c2e433fac5627646e7a9b2801a4d3df5df5415b0d26

Observation 6a0dbf99-f686-4062-af51-915912b086e7 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.585591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.585591Z digest=sha256:4b20afb7bed6ff7dbc5fa14415d5c965aeb40c93ec0b9c0d73fb574b39d99254

Observation c314bc97-07df-4dc5-a2ff-61a0536bd5a7 · outbound

This paper cites Vicuna: An open-source chatbot impressing GPT-4 with 90\ See https://vicuna.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Vicuna: An open-source chatbot impressing GPT-4 with 90\ See https://vicuna

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.330972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.590190Z digest=sha256:87d50683f82670bf7602122505d35930781c1da7e41fade7626a35504e6010d0

Observation ecfe5081-d108-4786-af45-3b16dd096d42 · outbound

This paper cites The Faiss library.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning The Faiss library

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.319457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.594311Z digest=sha256:8aa6f89665253e7dd16a30a011587238c7600a9e44c415720a08120b5da35a4c

Observation 846e8690-b5db-4738-954d-6530ab540ff4 · outbound

This paper cites A survey on rag meeting llms: Towards retrieval-augmented large language models.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning A survey on rag meeting llms: Towards retrieval-augmented large language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.309117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.598553Z digest=sha256:678d02dea51501ad4449a795b642c97df83a1917bafc87bd465243b7585d0bc4

Observation 17276bd8-4f06-4674-b3f3-105646f266da · outbound

This paper cites m BLIP : Efficient bootstrapping of multilingual vision- LLM s.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning m BLIP : Efficient bootstrapping of multilingual vision- LLM s

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.607728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.607728Z digest=sha256:8a7a8dd42ad6c08e399bb4e47ddcdb1e0336cb221b9a3f4e63de02b1388f4d91

Observation 09c2b0af-e638-4a32-88da-12c53adbc512 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning LoRA: Low-Rank Adaptation of Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.612200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.612200Z digest=sha256:7653b986cd5d222707a099b5fca9ea7d6787dac19c953a6e0617d73a04d70eb8

Observation fe708efe-721b-44df-a740-84d2c62d4573 · outbound

This paper cites A Survey on Retrieval-Augmented Text Generation for Large Language Models.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning A Survey on Retrieval-Augmented Text Generation for Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.616754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.616754Z digest=sha256:bf5c830d7fe9188261776059383119bf7b61732594f9eab42fcc845f8bbbca69

Observation 8cd63b7e-5e07-44f3-8f5b-77e4685657c2 · outbound

This paper cites Jieba Chinese text segmentation.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Jieba Chinese text segmentation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.299339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.621083Z digest=sha256:3125ce74da41662532e69c3a346830d33838bd30f1f601989f276b22847d9668

Observation e3736bd0-a3b5-4a49-80ed-ab15cf9f287b · outbound

This paper cites JEEM: Vision-Language Understanding in Four Arabic Dialects.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning JEEM: Vision-Language Understanding in Four Arabic Dialects

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.624934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.624934Z digest=sha256:eb47f37d374aafd3a26b3efe8348d9505f537e2af0feb5c21990e23b99020e13

Observation 33061be7-7d06-4e84-8641-a6c5f7f813a7 · outbound

This paper cites Jina CLIP: Your CLIP Model Is Also Your Text Retriever.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.629099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.629099Z digest=sha256:ab3c46e0d562d4ae09695e3ccc34a5dbc2a443f1fcf457bdc99df7c25a884d0a

Observation fd59b8a9-4a33-46f3-b852-e12dc5597075 · outbound

This paper cites MeCab: Yet Another Part-of-Speech and Morphological Analyzer.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning MeCab: Yet Another Part-of-Speech and Morphological Analyzer

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.289332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.632935Z digest=sha256:d9fed2eefdf1c11c2f3d26cae0677a816644c6c7547d4492af086ea95aea5c94

Observation 0dcfd769-727d-4592-9c0e-202d19b25e9c · outbound

This paper cites The IndicNLP Library.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning The IndicNLP Library

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.279721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.636690Z digest=sha256:30eea6fdbb9e79e60333a5c1054234b8897d12bd7f595196628559baed97735b

Observation c7a22470-775d-43e5-9b8d-5d0631d8897b · outbound

This paper cites What matters when building vision-language models? Advances in Neural Information Processing Systems, 37: 0 87874--87907, 2024.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning What matters when building vision-language models? Advances in Neural Information Processing Systems, 37: 0 87874--87907, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.640675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.640675Z digest=sha256:eafb1a6076835c708cb934ce9dd5df999c30396f5279f9f31502cc54a991c9b0

Observation faa9f74a-0fa5-4724-a853-9082954a96ac · outbound

This paper cites u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt\.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt\

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.263507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.644569Z digest=sha256:d898542c173371067c7840e03c24148916f45c143d445d2f6d556cfafcff0d04

Observation 517d34f0-010d-46e9-8aaf-1d6de556f505 · outbound

This paper cites Evcap: Retrieval-augmented image captioning with external visual-name memory for open-world comprehension.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Evcap: Retrieval-augmented image captioning with external visual-name memory for open-world comprehension

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.253454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.648440Z digest=sha256:d1da498012ce1881c1079866abfa9e25e9034b3c49b773912e5bffe2028e68cd

Observation 74afd818-a77f-403c-8703-2959bce25ead · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.652000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.652000Z digest=sha256:a8d92f047ecabd5197f143afd3c0e29d73c58520140be5b2ec8ab17ddd5cb48c

Observation e7874f0f-5989-4824-b5ef-fc0ce10a10bf · outbound

This paper cites Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.656228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.656228Z digest=sha256:98c424fb58829e1ad40952e616074e2418bab3c92494cacd54b8d472743efe02

Observation 81a63aee-6015-44d3-8eac-4a2bdcc64d73 · outbound

This paper cites Few-shot Learning with Multilingual Language Models.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Few-shot Learning with Multilingual Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.660499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.660499Z digest=sha256:6a5949657d78620c7e29dc99fcce106e461d1be73e63d639daae81b75b3bc33f

Observation f8bb8130-295a-4539-86fb-3a3060895571 · outbound

This paper cites Visual Instruction Tuning.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Visual Instruction Tuning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.242606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.664943Z digest=sha256:6717f3e5154572bc625bf664d47caddbe5d1bb93c5e4c41516771087e92acba2

Observation fadedef2-a003-4e9f-9924-bdd11e7842db · outbound

This paper cites LLaVA-NeXT: Improved reasoning, OCR, and world knowledge , January 2024.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning LLaVA-NeXT: Improved reasoning, OCR, and world knowledge , January 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.232573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.668652Z digest=sha256:721cadeee9369033028043e23b7441b8fd06597782ddd9dcc56a96a1288fdf9f

Observation 14378ac8-bae8-4aae-8198-fc83ac36b32a · outbound

This paper cites fugashi, a tool for tokenizing J apanese in python.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning fugashi, a tool for tokenizing J apanese in python

Reference 23

Resolution
verified exact
doi, observed 2026-08-15T17:51:16.819538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.672148Z digest=sha256:7f14e88bb0bb6ad55caaeed49f9126f092d8c2c70eb1215b1ef8f5d4f5a98e3a

Observation 5f30e3c7-ca49-4e1e-b8d7-3dbba3b0fcc1 · outbound

This paper cites ClipCap: CLIP Prefix for Image Captioning.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning ClipCap: CLIP Prefix for Image Captioning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.675792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.675792Z digest=sha256:b3488b6e3632d18031ee801b60bc5dc94e351e10ceb4bfd83b8455df8a50a1cb

Observation 51b04839-ec5b-4803-9217-1de436dbe949 · outbound

This paper cites Crosslingual Generalization through Multitask Finetuning.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Crosslingual Generalization through Multitask Finetuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.679464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.679464Z digest=sha256:69cfef5ce0962dc219eddd4996d2feb7a016e723a0264ba626f764da65b956f2

Observation e1ea91a9-9c40-41a8-953c-93c3e9a78fd3 · outbound

This paper cites GPT-4o System Card.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning GPT-4o System Card

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.683396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.683396Z digest=sha256:cdf0ec35ae0aa036bbc76b7aa46411bdbd765cbf2c3bdbfa71faa1b9d1c1c07b

Observation 74b74614-b132-4d43-ad93-f59553198c90 · outbound

This paper cites X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.687278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.687278Z digest=sha256:b248daab32b4603a563ed61532c801c427857802ba597277d60a523c0ee66351

Observation 4afd289a-bc32-4e9b-9e69-9ba034fdd7c7 · outbound

This paper cites BLEU: a method for automatic evaluation of machine translation.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning BLEU: a method for automatic evaluation of machine translation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.690863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.690863Z digest=sha256:d50d87b93d97921127d802cf0c54856879de6821e604cb99e28cf54354c022b9

Observation 8f4cda7b-4054-4873-aff6-45af173add3a · outbound

This paper cites P y T hai NLP : T hai natural language processing in P ython, June 2024.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning P y T hai NLP : T hai natural language processing in P ython, June 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.221933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.694081Z digest=sha256:7e53fb91ce39ce2a36e4f22ced4fefee0c139a4d374695d9097d635c94ac00f8

Observation 910c2351-826b-4f7b-9d9a-dfe03e42fa35 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Learning Transferable Visual Models From Natural Language Supervision

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.698048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.698048Z digest=sha256:8cacd1ccee0598f5a7696ef77b016fbbe55b241d090c9390a759115baf97f762

Observation 3d508bb6-c1f8-4088-8ddc-3708d4d07367 · outbound

This paper cites LMC ap: Few-shot multilingual image captioning by retrieval augmented language model prompting.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning LMC ap: Few-shot multilingual image captioning by retrieval augmented language model prompting

Reference 31

Resolution
verified exact
doi, observed 2026-08-15T17:51:16.807453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.701713Z digest=sha256:60f9cd8c0cf8a95d6b2c25ca2d8d254d7eec5aa6278007d6a513ddb43c560d15

Observation 1417952e-efa7-4b22-af52-96deeef79851 · outbound

This paper cites SmallCap : Lightweight image captioning prompted with retrieval augmentation.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning SmallCap : Lightweight image captioning prompted with retrieval augmentation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.208661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.705022Z digest=sha256:8ec93dc180fa939d6caa4465b9ccbf692ca665f578b59f62cf29f9e97c8bca0c

Observation ee0ff7e6-5416-4f46-b12b-be8e792b8ad5 · outbound

This paper cites PAELLA : Parameter-efficient lightweight language-agnostic captioning model.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning PAELLA : Parameter-efficient lightweight language-agnostic captioning model

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.197085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.708324Z digest=sha256:58aef2f38e5e7bd57820fa64a722f77f6ea17929dafb6544ddfb33f0d90dc4a4

Observation 348fd5ae-21c5-4771-bf9e-add921b3b67b · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.186408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.711713Z digest=sha256:d9aab8cdbe060ef3215288669371073f63e9e7c0a928ca961f30e352fb8fdb75

Observation 0b5cf44d-1522-4951-a8d2-79e1cdaf2a25 · outbound

This paper cites Stanford Alpaca: An instruction-following LLaMA model , 2023.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Stanford Alpaca: An instruction-following LLaMA model , 2023

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.174255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.715464Z digest=sha256:b728f29fff6c902cbf42b248dfd83de5964f17637db13ff3aa9ddd59f448e3b9

Observation ebd4d8dd-71aa-4cc6-b4ca-184996c1f9ce · outbound

This paper cites Thapliyal, Jordi Pont Tuset, Xi Chen, and Radu Soricut.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Thapliyal, Jordi Pont Tuset, Xi Chen, and Radu Soricut

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.719471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.719471Z digest=sha256:99a731783638864ebb0fa7bd15c94c00067909b94226dc3cb05cd868c5b357f3

Observation 3daa7291-0601-4c6a-87b9-5db7461ab864 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning LLaMA: Open and Efficient Foundation Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.722756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.722756Z digest=sha256:d066ef972bb336578a30aa55108f7069be53e0091a2a5523e719c92939e34560

Observation f6a6cb77-0ca8-4be7-956e-aebe5d0cdd46 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.726343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.726343Z digest=sha256:45b4f461b6a0b892e11c68a93aedbc155d6adb5910e6bded5a96512a3f1d99b9

Observation afa3a0b6-b36e-42fa-9d3e-fa14adb62fdd · outbound

This paper cites CIDEr: Consensus-based Image Description Evaluation.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning CIDEr: Consensus-based Image Description Evaluation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.729976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.729976Z digest=sha256:bbbfc3923c9dfd33b2793ed2a18e0b324c9ebf8ac6b5a1d4f9a3bdc306812178

Observation b1049281-ed97-4bca-9353-afbc0cf6d975 · outbound

This paper cites mT5: A massively multilingual pre-trained text-to-text transformer.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning mT5: A massively multilingual pre-trained text-to-text transformer

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.733490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.733490Z digest=sha256:f1f0d0c5f9061b2f4d1793af9d8346047d1155e37ea164b65a90a0b50181d5b9

Observation bc8b0d94-0737-4c43-a615-beb1341521ce · outbound

This paper cites Retrieval-Augmented Multimodal Language Modeling.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Retrieval-Augmented Multimodal Language Modeling

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.736921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.736921Z digest=sha256:dbae517656ab79bca1da56c31020b11c3ea4d62ddb166640c53672b6f595726a

Observation b0f06c3a-7cae-4d47-8a28-c71914589cdc · outbound

This paper cites Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.740304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.740304Z digest=sha256:93e497ef3e66bab35f698b7fcc0d5033f6274381095697c7230fa697a062ed49

Observation 61bf576e-1c58-47fb-964e-63ef406e5c32 · outbound

This paper cites MeaCap: Memory-augmented zero-shot image captioning.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning MeaCap: Memory-augmented zero-shot image captioning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.162716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.744297Z digest=sha256:6b8374da156ea90f5b7b75daafd8bb5f1d38b37423814c629e09dcf9cbebb290

Observation a668557f-71c6-444b-9443-b16a5e1139cc · outbound

This paper cites Scaling vision transformers.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Scaling vision transformers

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:17.151494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:51:16.747605Z digest=sha256:e640dbc619fd23fd3c2aaa60559340e0a0f8db822935f32caaf2a3b12bb2d4df

Observation 222e58a0-d79d-4940-8502-f02eb0b2b441 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Sigmoid Loss for Language Image Pre-Training

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.750717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.750717Z digest=sha256:49c9c120d4c94c542b616400d639cf6323d832dc236ccaa95441c708e4cced36

Observation 02e3af54-5b06-47c0-aa73-30d05a6e3708 · outbound

This paper cites write newline.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning write newline

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.754102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.754102Z digest=sha256:0fe603453c7d5358441fd75caaf0552c969af7f01c75a641312f89075f078a59

Observation 3a7dcc8f-acfc-4615-a000-b301174eea7b · outbound

This paper cites @esa (Ref.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning @esa (Ref

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.758121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.758121Z digest=sha256:9254338d4cc5b8aad914588c832dc3b203011f89b0f5e34dad3fd47ba06aec0f

Observation f4ab3af3-559a-407a-85ce-f38e069fd2e4 · outbound

This paper cites an unresolved cited work.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.761894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.761894Z digest=sha256:289d4528b3b84323ab62944534ffc71cd580a35c6a5449aa7a016954f6523e21

Observation fc570287-e508-4ad7-90fb-4fe2b8459c0c · outbound

This paper cites an unresolved cited work.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.765188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.765188Z digest=sha256:81a0b2f713199c660b9e27c3d24648ae1a5bb64d926148c1ecddbea788057e35

Pith citing papers

No inbound Pith citation observations are available.