Pith. sign in

Paper Citation Record · LEDGER

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training

As of 18 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2507.08710.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.08710 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:18:51.876640Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

68 of 68 outbound references displayed

  • verified exact0
  • verified fuzzy42
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation af9e9b67-ed51-43f0-a289-8f9970342520 · outbound

This paper cites Learning cnn-lstm architectures for image caption generation,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Learning cnn-lstm architectures for image caption generation,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.538709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:44.496453Z digest=sha256:1bfac999b3588753e926ead358a87b30daa965602caec19cb2a0d38857a6a566

Observation 98920df9-8da1-4c92-b48d-4e976d93e013 · outbound

This paper cites Image captioning through image transformer,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Image captioning through image transformer,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.528365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:44.581263Z digest=sha256:d2f72812b84e3190809af6c8f09e4ef86fd7a7b25059d70a958f6232e56f8bf9

Observation 2b78848c-363e-4dbb-818d-7779191acc0b · outbound

This paper cites Meshed-memory transformer for image captioning,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Meshed-memory transformer for image captioning,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.517298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:44.763308Z digest=sha256:096f07eb2408e4e0cf1b88616dec4df5f21ca111f29a03f90732d4d638f38712

Observation 724a695d-4006-446b-b5d5-c6b7680e669c · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Bottom-up and top-down attention for image captioning and visual question answering,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.505011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:44.883631Z digest=sha256:284805b73ef23c1f5d3bf1e4e28f7d8d5c184f159a4e68b927472a56c36984b2

Observation 3a749b07-904b-437b-a0b3-6be7fa4d1177 · outbound

This paper cites Self- critical sequence training for image captioning,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Self- critical sequence training for image captioning,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:45.008929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:45.008929Z digest=sha256:aa6e783e925e6ed40c1231b56692c3b14ff99141cc4bc50c557ba6439daab0f1

Observation 21ca654c-8798-4748-ae07-cac4c4654afa · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image cap- tioning,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image cap- tioning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.482559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:45.151174Z digest=sha256:da2178848b81a0444c8f5f9dafd39bd6a6fef1c9883dc9e96127b0f7c4d8036b

Observation 6233736e-00d0-45bb-a043-a7cba9c56470 · outbound

This paper cites Microsoft coco: Common objects in context,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Microsoft coco: Common objects in context,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:45.215989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:45.215989Z digest=sha256:40a0be3b5f7b7b0e975f04862152f91f89d0ffb1cbe31426497cec170ba72c30

Observation 8fe3c809-e63e-4f17-ae2a-7ed19de952a0 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Bleu: a method for automatic evaluation of machine translation,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:45.325805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:45.325805Z digest=sha256:196853e7e9131f4072b6364175c404430d20751bb6841ba2412da673e9a3fdb8

Observation c7189918-21d6-4020-a0d7-b828d7412d6d · outbound

This paper cites Cider: Consensus- based image description evaluation,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Cider: Consensus- based image description evaluation,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:45.401247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:45.401247Z digest=sha256:16851a0af4ed3c423e94d71601c4d68e8aa1b4dcbe60e550b4c50baf086f2d27

Observation 1a1376e0-81c8-435b-9c09-1c69de0c4497 · outbound

This paper cites Bertscore: Evaluating text generation with bert,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Bertscore: Evaluating text generation with bert,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.451864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:45.492571Z digest=sha256:3c2c8bd05f5ba3e47b9c25253b78765cd5278b75a67c14e37cafc775e11a7738

Observation 196d3fa2-6bb0-4897-a325-0db460aa2a3f · outbound

This paper cites What you see is what you read? improv- ing text-image alignment evaluation,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training What you see is what you read? improv- ing text-image alignment evaluation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.442462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:45.585156Z digest=sha256:ee2553cad4597df27931f8437f0b5a9eae1deba94ff048b67d96ff95b7577b47

Observation 738d87d5-2d8d-40ca-a429-d349426f4e66 · outbound

This paper cites Fast, accurate, and lightweight memory-enhanced embedding learning framework for image-text retrieval,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Fast, accurate, and lightweight memory-enhanced embedding learning framework for image-text retrieval,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.432955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:45.683674Z digest=sha256:ebaaafbf7af9ec20bfc5c7e214294256db6b0f5a3b50fe313fb2936851b94d39

Observation 34fd9183-3893-4f64-bb97-884567a1abe8 · outbound

This paper cites Vilbertscore: Evaluating image caption using vision-and-language bert,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Vilbertscore: Evaluating image caption using vision-and-language bert,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.421682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:45.769679Z digest=sha256:d2d380b0622422a21dad52f98ff295e5eba1843c8da17b5aad0aef9b7fcf8686

Observation 16462428-c0fa-474d-a16c-51ad33ce8d5e · outbound

This paper cites Image-text alignment and retrieval using light-weight transformer,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Image-text alignment and retrieval using light-weight transformer,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.396031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:45.864014Z digest=sha256:72a5ba8e69f7d23253a8203bf50152f77ad008d7efe8cdcedcd5cc60b86656eb

Observation b86d9278-ae41-4bd3-929a-1f025224a9f0 · outbound

This paper cites UMIC: An Unreferenced Metric for Image Captioning via Contrastive Learning.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training UMIC: An Unreferenced Metric for Image Captioning via Contrastive Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:45.987661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:45.987661Z digest=sha256:00703bc833df83d4f17a33e418937538019be5aec4721c45e6a6e19b55930cc8

Observation b11e06f7-676b-404a-b3d5-0d11a036e1a1 · outbound

This paper cites Less is more: Clipbert for video-and-language learning via sparse sampling,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Less is more: Clipbert for video-and-language learning via sparse sampling,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.382031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:46.060303Z digest=sha256:e0a0fbe071536da0930d477cf604b45f7f5c4a8d86a335cc1ddded877cd275ab

Observation 523b857d-58f0-4d9e-94a8-81e3b919172a · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.369989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:46.129387Z digest=sha256:a7e8f654cf0f59e4fd9570c8af0396df779ed91adbcd52c7c51b233861f573d5

Observation 10672b9f-a868-466f-8a6c-b1c479216c38 · outbound

This paper cites Quality estimation for image captions based on large-scale human evaluations,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Quality estimation for image captions based on large-scale human evaluations,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.346489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:46.244803Z digest=sha256:84cc02c4e9bebff203523a2d1201b3c0ae5e18f5aee2a1b0ac8158e7f5402b73

Observation 067683fa-cef1-4c83-811c-27ee9be55865 · outbound

This paper cites Revealing the dark secrets of bert,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Revealing the dark secrets of bert,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.328710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:46.370999Z digest=sha256:f220f864912c4d9d303f52afb213b7a66e145e34a4339fb70f706d033149a5af

Observation 104ac830-39ae-4fed-9f7f-ce993210402d · outbound

This paper cites Minivit: Compressing vision transformers with weight multiplexing,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Minivit: Compressing vision transformers with weight multiplexing,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.318090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:46.521919Z digest=sha256:953f2fb9d7573d034616dd91a7e58b3a9ff941501ddc10aed6ebeaa60431d23f

Observation 8ad54b67-70ec-4ebe-9c0d-6cd9ca908c74 · outbound

This paper cites ALBERT: A Lite BERT for Self-supervised Learning of Language Representations.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training ALBERT: A Lite BERT for Self-supervised Learning of Language Representations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:46.607900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:46.607900Z digest=sha256:651a24fa68155493d7d2b9b3ae0f2387c2ebdbd9806b5ee8546277078db266e8

Observation ee9aa492-63a3-4d64-b84b-55e70a95c154 · outbound

This paper cites Enabling multimodal generation on clip via vision-language knowledge distilla- tion,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Enabling multimodal generation on clip via vision-language knowledge distilla- tion,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.291873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:46.699499Z digest=sha256:8be6642d6d7f5456df83c42ab84ca9937f059e5320fc635f9f1c5503b72e7507

Observation 65ec115c-e93a-478d-a9bb-c43a8a8ca597 · outbound

This paper cites CLIP-TD: CLIP Targeted Distillation for Vision-Language Tasks.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training CLIP-TD: CLIP Targeted Distillation for Vision-Language Tasks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:46.816291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:46.816291Z digest=sha256:21fa9dfabce9220d1e3b2ebbb24fefa2506ada4f4a6130a7f95beb3bf14ab5ec

Observation 60ba08d4-fc9b-4506-a040-d906a169dae3 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Learning transferable visual models from natural language supervision,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:46.902472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:46.902472Z digest=sha256:b5889cd360989dcb3ed33ad2c9f23f6907d38bba2fb4e8578b5e738e280b00f7

Observation c0f769db-72fa-4407-b8cd-bcbe1e6ffbd3 · outbound

This paper cites Sca-cnn: Spatial and channel-wise attention in convolutional networks for image captioning,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Sca-cnn: Spatial and channel-wise attention in convolutional networks for image captioning,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.226808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:47.017461Z digest=sha256:a4df7c65efc7e8c5ec1df6aaa7665d23cbb74d0dcffd0dfbcb03791ff2094f24

Observation 571e6b59-0e11-4213-bfca-acffcc3b3899 · outbound

This paper cites Knowing when to look: Adaptive attention via a visual sentinel for image captioning,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Knowing when to look: Adaptive attention via a visual sentinel for image captioning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.182859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:47.108684Z digest=sha256:277bd63b061f45276fa7830424f0e24a89199f0699e7db6c5c6d48812f043d91

Observation c435e904-4efc-468e-90e1-77a151c923e6 · outbound

This paper cites Know more say less: Image captioning based on scene graphs,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Know more say less: Image captioning based on scene graphs,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:57.441150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:47.214435Z digest=sha256:7802115c5db49a5560a4191cb8502fe21a905c56cdf40c28c963c92b13579485

Observation 9192a4e7-23b4-4c75-aad3-0e2a41fc051a · outbound

This paper cites High-quality image cap- tioning with fine-grained and semantic-guided visual attention,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training High-quality image cap- tioning with fine-grained and semantic-guided visual attention,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:57.142395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:47.379950Z digest=sha256:31a7cc42cf4507d944a7236fd57dd46da305ed31acb01421a54f8d4e50b72c67

Observation 45cf7118-9da8-4cb3-b021-9ce38f332a1e · outbound

This paper cites Multimodal transformer with multi- view visual representation for image captioning,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Multimodal transformer with multi- view visual representation for image captioning,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:56.954306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:47.450394Z digest=sha256:8e0c4132cea4ee0cb54158cb0f8da87279b6853f81bd036402c33935e56c895f

Observation 9dbb5512-162e-4619-b482-2081258a674a · outbound

This paper cites Compact bidirectional transformer for image captioning,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Compact bidirectional transformer for image captioning,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:56.756704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:47.571278Z digest=sha256:f32d7af29edee4a6432f3f957154a25f42cdf0d76d4eca8db10c6c2d93d20990

Observation 58c688ad-e4fc-45fb-a40f-f41e3ad78348 · outbound

This paper cites Task-adaptive attention for image captioning,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Task-adaptive attention for image captioning,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:56.571741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:47.661175Z digest=sha256:5cd240090b3e841615a2740c29b2c46ffc2a0d7167f6a7e8428fa1b16c4237f6

Observation b8faab56-4325-4121-a7fa-e0fc61873d7c · outbound

This paper cites Textual context-aware dense captioning with diverse words,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Textual context-aware dense captioning with diverse words,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:56.262369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:47.728248Z digest=sha256:91987626ecc596e7095e9e77ca0693fc9a402849f647357fe5949919f35780f8

Observation 46457da1-e0aa-4070-8373-e2adfac91838 · outbound

This paper cites Meteor: An automatic metric for mt evalua- tion with improved correlation with human judgments,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Meteor: An automatic metric for mt evalua- tion with improved correlation with human judgments,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:47.800623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:47.800623Z digest=sha256:ad3d7f75a7b93d8598c6c93b3d2549757b38b9838f52d488d1c04536c6903828

Observation a5722b42-588b-466a-be6c-900b74189f3e · outbound

This paper cites Rouge: A package for automatic evaluation of summaries,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Rouge: A package for automatic evaluation of summaries,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:47.891704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:47.891704Z digest=sha256:2a820d775f66eb90c5c4b53f1fefb99c80d462c19698fde3c8eb5718d4791a5c

Observation b440dc69-0d02-42c1-a27d-3735867b32a8 · outbound

This paper cites Spice: Semantic propositional image caption evaluation,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Spice: Semantic propositional image caption evaluation,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:47.956604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:47.956604Z digest=sha256:23209a6e4310a5b4408bac60338897ac07881c6ac40a5ed813d6f40a26f7aec4

Observation 9266a018-e0eb-40f0-a5ec-d9b3eb341cf1 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:48.077301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:48.077301Z digest=sha256:d665a347ee2e2af9c4636747360a79d92f27bbbf3df2cb841449381b682ec6c2

Observation ac99a651-3121-4d31-8668-a42d84296387 · outbound

This paper cites Clipscore: A reference-free evaluation metric for image captioning,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Clipscore: A reference-free evaluation metric for image captioning,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:55.989883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:48.194780Z digest=sha256:7cffe2d691ff34a64cae87a99306571bcc819338df5a39ea81c9b304578eaf99

Observation 9e56f2de-0a44-46f0-8363-b3366ecad143 · outbound

This paper cites Tiger: Text-to-image grounding for image caption evalua- tion,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Tiger: Text-to-image grounding for image caption evalua- tion,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:55.753289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:48.326568Z digest=sha256:96038dfc2d2a7d0463830e57066819c6fc6b9cfd6a92a47b75db15a9a20451c1

Observation 473024b0-5264-48ed-91fa-011520487af3 · outbound

This paper cites Em- score: Evaluating video captioning via coarse-grained and fine-grained embedding matching,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Em- score: Evaluating video captioning via coarse-grained and fine-grained embedding matching,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:55.505044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:48.435085Z digest=sha256:920689d9983a361ea4d958783b5b96acdff810461f78518a47b99ffedd6ccfa2

Observation 2bebff7b-075d-4dc3-9477-81a543bc9c20 · outbound

This paper cites Gpt-3: Its nature, scope, limits, and consequences,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Gpt-3: Its nature, scope, limits, and consequences,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:48.555233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:48.555233Z digest=sha256:45bb1ed4935565e31394424a3d878a8bf675f099313f10d0af9d292a794ecb13

Observation 0e0d5cf6-ec6d-4216-9234-d04b0fb41152 · outbound

This paper cites Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehen- sion,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehen- sion,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:55.274386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:48.636684Z digest=sha256:3a5e8206cae4a02f33a08454caa3a3da6e3587b859e622ad584a5ff2491046b5

Observation 1b91fbc6-9386-456e-9543-9953ab43f225 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Swin transformer: Hierarchical vision transformer using shifted windows,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:48.724702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:48.724702Z digest=sha256:50b602e7a8531657780b8dfb8a20ca49d8f09cae9161ffa382f1c9e2bb2d5286

Observation 2f07502a-6d12-430e-8cbe-bf26a807362b · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:48.992310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:48.992310Z digest=sha256:c31e30ed8c1b116ba400533d22aed5fbc05fa390050cf6f46e8b3432bd367817

Observation e4c820e8-97fb-4d58-8dae-ad870c5dcfc1 · outbound

This paper cites VL-BERT: Pre-training of Generic Visual-Linguistic Representations.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training VL-BERT: Pre-training of Generic Visual-Linguistic Representations

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:49.108576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:49.108576Z digest=sha256:e6b642d31076ef6307f7cd25d58a7945c16ebaef037a394a1aa86bb5179489fa

Observation 671a853a-d350-4185-a7bd-c000fcfee516 · outbound

This paper cites Slip: Self-supervision meets language-image pre-training,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Slip: Self-supervision meets language-image pre-training,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:54.996841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:49.294938Z digest=sha256:6a8b034d83146b361fc388813b65374372ae70d1ea3f753bf019ece5a1bfff7c

Observation d5a17254-2f11-44e3-be44-f90bd97b8975 · outbound

This paper cites Pruning Filters for Efficient ConvNets.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Pruning Filters for Efficient ConvNets

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:49.439407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:49.439407Z digest=sha256:c83f831cee7ce123cc6ce5e8c346b9f6b4ddf9eaec25e9765c40203a4887abc8

Observation ede5327c-922a-4ed8-9555-cb7fb26c0b0c · outbound

This paper cites Vision Transformer Pruning.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Vision Transformer Pruning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:49.586224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:49.586224Z digest=sha256:b6a27e45fe40bbdfc955a8bcd65c2c71a792e889bb7264cc2cd6846b37f73d31

Observation df4b184b-157f-4be5-a4cf-3c00dbcc0f2f · outbound

This paper cites Post-training quantization for vision transformer,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Post-training quantization for vision transformer,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:54.652256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:49.787638Z digest=sha256:26320db5a36bb7fce41c46eda221ce697efc1a85830af2f629a888df09012672

Observation 0505ab65-2ef6-496c-98ae-c4ec38866818 · outbound

This paper cites Tinyvit: Fast pretraining distillation for small vision transformers,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Tinyvit: Fast pretraining distillation for small vision transformers,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:54.444617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:49.905742Z digest=sha256:2d5b5992a314c30c5046b243594620dac99c9560f9267072a6006d4846dd1a89

Observation 0b0011c6-d04c-4753-abb2-3a5bdf78d27f · outbound

This paper cites Tinybert: Distilling bert for natural language understanding,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Tinybert: Distilling bert for natural language understanding,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:54.259692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:50.076329Z digest=sha256:6b696b44710269bac5b601197089ca9e41371581d79acb8aefffa464417d166a

Observation 2b89c3d1-403f-4093-b9d5-8456c1de0eb0 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:50.162524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:50.162524Z digest=sha256:6a15d6609e9a2b3850517a6d55e5b30f5d8c617e795300fb3b8d9f43ef52dada

Observation b39e71e5-76e7-4887-a9a7-cdab83f87c18 · outbound

This paper cites Training data-efficient image transformers & distillation through attention,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Training data-efficient image transformers & distillation through attention,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:50.303885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:50.303885Z digest=sha256:3c5f6e026e75dc165c09195967516c9ecd30a5e7481b0dbe4e45a7425efc1433

Observation 57e846a7-95fc-4f08-b549-f0006eaa6703 · outbound

This paper cites Attention is all you need,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Attention is all you need,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:50.460291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:50.460291Z digest=sha256:531cbc4e478d53863d2ef6d8b3d5acb35e42ae7ae98259fbd269a44b8f3ecc01

Observation 507da7cf-40af-4f7e-b50a-a3421316f702 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:50.608287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:50.608287Z digest=sha256:e9d3652862a4255bc249ddc9bdc835d77a11c99be7aa16e95e6219d5b464574c

Observation 6b05c5f1-1756-47bf-91a4-56717f2c1ae2 · outbound

This paper cites Bag of tricks for image classification with convolutional neural networks,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Bag of tricks for image classification with convolutional neural networks,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:54.104734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:50.761420Z digest=sha256:247e75d070d10c1d4fbda7bdad14bb71063049b572263ce8acd7e2fb42b2ea43

Observation 4af3bb92-6134-416d-b4aa-62c6937a5550 · outbound

This paper cites Making convolutional networks shift-invariant again,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Making convolutional networks shift-invariant again,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:53.800069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:50.886873Z digest=sha256:8c7b58b0ab0ea08366e54e5d32a97d1d09fb3bc9cdfc3511da0fbd4118cf7c20

Observation 20d691ff-b298-4771-935b-70f04e205385 · outbound

This paper cites Discriminability objective for training descriptive captions,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Discriminability objective for training descriptive captions,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:53.570868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:50.986021Z digest=sha256:1ab60931ca336078b2c82de03a41d8a867d243321297bfb1e2048ea602928401

Observation 07f51bdc-c850-49b2-81e8-fb98b53f422f · outbound

This paper cites ImageNet Large Scale Visual Recognition Challenge,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training ImageNet Large Scale Visual Recognition Challenge,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:51.056134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:51.056134Z digest=sha256:0b01db373ce0da4dc2291032648622cfb385c5c4ef7f0483c1672ad1968126b0

Observation 2ba5db4d-d95c-41f0-a492-b0611a25abff · outbound

This paper cites Framing image description as a ranking task: Data, models and evaluation metrics,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Framing image description as a ranking task: Data, models and evaluation metrics,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:53.266631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:51.171017Z digest=sha256:da984f631d2b2de81568568633bcf942aa91303873c435b50585ee647b64ce90

Observation fa8cb4f5-4a9b-4e46-b2d8-af3aa3151a8b · outbound

This paper cites From Images to Sentences through Scene Description Graphs using Commonsense Reasoning and Knowledge.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training From Images to Sentences through Scene Description Graphs using Commonsense Reasoning and Knowledge

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:51.241514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:51.241514Z digest=sha256:4492827b599f2fff0c4c4c3ea0403a00c34e8d2a127a27aaa7ac221e1f788156

Observation b5ca7f10-23d0-4445-9220-275089ddd46e · outbound

This paper cites Consensus-based image description evaluation,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Consensus-based image description evaluation,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:53.042948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:51.305407Z digest=sha256:11d892830c59534da54789325fe7c1d9e513dc36e1056f75692a3bd30c2d2822

Observation 0433243b-fd85-4fcc-af24-11e652ff5144 · outbound

This paper cites Foil it! find one mismatch between image and language caption,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Foil it! find one mismatch between image and language caption,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:52.700175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:51.395791Z digest=sha256:78e6bcc0b5db9c72d5f170bff4952654175b1ac53c57c5e72d4c5e10cc8d17f9

Observation 6705d649-e7d7-4e3f-9634-a4129066491c · outbound

This paper cites Deep visual-semantic alignments for gen- erating image descriptions,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Deep visual-semantic alignments for gen- erating image descriptions,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:52.538573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:51.477940Z digest=sha256:8148c17a45071cef93eb2b51553854ce7e928e0accead59d6fb4739803a5fb47

Observation 24f8f060-4a72-4d69-a4fb-4b7d196df7b4 · outbound

This paper cites Vinvl: Revisiting visual representations in vision-language models,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Vinvl: Revisiting visual representations in vision-language models,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:52.392262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:51.554693Z digest=sha256:8eb6c258e0e26387e080917f5357a54433f2e41be8199474e5dd7e1179cc110f

Observation 8dd056fb-e566-42b9-a6c3-2a426bbf0ca9 · outbound

This paper cites Adam: A method for stochastic optimization,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Adam: A method for stochastic optimization,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:52.258827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:51.640789Z digest=sha256:ece13ffb443cf863c7ecb3a83f48eaa6028020edf59799041410a0e139703c70

Observation fc4aa0b8-2168-462a-8dd9-58111a0fc563 · outbound

This paper cites Improving image captioning evaluation by considering inter references variance,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Improving image captioning evaluation by considering inter references variance,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:52.129394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:18:51.705906Z digest=sha256:6da99a7dd590e090dfe7d950f907986b48c14fd7f64b6ac28935020b50b25f94

Observation 380b0f08-6362-4ce7-8203-5776db952ae8 · outbound

This paper cites Concrete Problems in AI Safety.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Concrete Problems in AI Safety

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:51.809319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:51.809319Z digest=sha256:ce5b3c44aafe5a7dea29e84ffd9349f378b86fd7a7f11e04225b10bbcd33cc6c

Observation 61a01979-bf04-494a-9ff6-eac046908ce5 · outbound

This paper cites Spontaneous Reward Hacking in Iterative Self-Refinement.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Spontaneous Reward Hacking in Iterative Self-Refinement

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:51.876640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:51.876640Z digest=sha256:a427173124296013b924d7d66ea6084016b04741e875a6c813b5f79da4d876e0

Pith citing papers

No inbound Pith citation observations are available.