Pith. sign in

Paper Citation Record · LEDGER

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding

As of 14 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 1 inbound Pith citation observation for arXiv:2412.01115.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01115 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:43:43.115768Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T21:35:11.654114Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T23:59:06.852265Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact2
  • verified fuzzy23
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9f629a68-5eaf-4176-9847-00c3feca6a4f · outbound

This paper cites nocaps: novel object caption- ing at scale.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding nocaps: novel object caption- ing at scale

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:43:43.571606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:43:42.969444Z digest=sha256:f58dc937d04d9c8bc5cfaad8d09141beed22bf05d19357d55a1613543097ccc1

Observation 9f9f82b4-414e-4f83-b87e-562ca87c7941 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:43:43.561109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:43:42.973506Z digest=sha256:9fa6db3de245fc08da95962961406aaffbbcbf2826dd5a900cd1268c11317c26

Observation 57d7ef26-e5f3-44a2-b7b0-7e8fc9d1ff88 · outbound

This paper cites Spice: Semantic propositional image cap- tion evaluation.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Spice: Semantic propositional image cap- tion evaluation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T04:43:42.977330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:43:42.977330Z digest=sha256:be061b02597f57284f964f7efd59d7034ec368c79459ff0f182bec3925e13682

Observation 1f7c5749-991c-47f2-9ceb-fc3d0c0c42b7 · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Bottom-up and top-down attention for image captioning and visual question answering

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:43:43.544373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:43:42.981044Z digest=sha256:e7476a7fe30db3f007a165a21de9b5755ec45991d0385a2fbe4020273f1d6c81

Observation 30a10b3f-81e1-4cba-8d5d-d70e843ade48 · outbound

This paper cites Self-RAG: Learning to retrieve, generate, and critique through self-reflection.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Self-RAG: Learning to retrieve, generate, and critique through self-reflection

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:43:43.533738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:43:42.984854Z digest=sha256:718b27ee51d08ae6134c3912d62b11e95ba3497304404391f10cb26e93e8b634

Observation ecb6e715-0921-4fd6-b787-51d1bfb5ca57 · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T04:43:42.988626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:43:42.988626Z digest=sha256:db5d09ba17aeab6a8f85e8653f471104232d8de2c06a3bf153e258f60fdc2106

Observation ecef8cc7-b673-4a9d-8e07-6a4f5e973bd6 · outbound

This paper cites Label-efficient se- mantic segmentation with diffusion models.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Label-efficient se- mantic segmentation with diffusion models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T04:43:42.992524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:43:42.992524Z digest=sha256:fe0660b8078b513fc118ffb685cf9304ccb9eb7b489905b6889b7fad8a5d9225

Observation 1e84edba-05ad-4dfe-a235-6f95d7f901fb · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T04:43:42.996694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:43:42.996694Z digest=sha256:10c113842307fdc21408590b417c44fe20a28eed84c4808f02fd4a560518d112

Observation bc5d7ae1-ded4-44ad-a8e3-f1e9c200255c · outbound

This paper cites Bert: Pre-training of deep bidirectional trans- formers for language understanding.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Bert: Pre-training of deep bidirectional trans- formers for language understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T04:43:43.000178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:43:43.000178Z digest=sha256:7d25fa4780389c4115d60c75d83fb6a5753b65116eb3cabef422c4e581770656

Observation 01a069d3-cc53-4cba-9933-d41d3c84939f · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T04:43:43.003606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:43:43.003606Z digest=sha256:bc0b35e752cc02700705ccc8ec3935fb695504e19720f94e46c23cd451255966

Observation f371f76e-ab92-458d-88a9-3b40d79e0f92 · outbound

This paper cites Eva: Exploring the limits of masked visual representation learning at scale.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Eva: Exploring the limits of masked visual representation learning at scale

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:43:43.500292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:43:43.007173Z digest=sha256:1bd35dbb1cd061d93de078d695b9446aa271ef25ec9991c2090ffc6d3d846eb6

Observation 46911fbe-c54e-4344-8cec-a9485f6ff6a2 · outbound

This paper cites Transferable decoding with visual entities for zero-shot image captioning.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Transferable decoding with visual entities for zero-shot image captioning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:43:43.490657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:43:43.010597Z digest=sha256:17361a8793dfb87cd599861c80ffb642947300ac6d5cd97bc90a52da88eb2c9a

Observation df0c99a1-a8ec-433c-a900-bf03106f9c36 · outbound

This paper cites Retrieval augmented language model pre- training.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Retrieval augmented language model pre- training

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:43:43.480829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:43:43.013976Z digest=sha256:833ac3c4f32a72b61c459f9cba65a5e1f7dd63291cece166c5c8ff0eb3a9422d

Observation 68be8fe1-8976-446d-ad40-1bac41046606 · outbound

This paper cites Deep visual-semantic align- ments for generating image descriptions.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Deep visual-semantic align- ments for generating image descriptions

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:43:43.470359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:43:43.017333Z digest=sha256:9efd7e526d842d4cd5e02ae3af11380fc63830ce293310d59f2edb7a1f15a873

Observation 990f8361-e9b3-431c-9c96-a3b7cd12cf5e · outbound

This paper cites Do you remember? dense video captioning with cross-modal memory retrieval.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Do you remember? dense video captioning with cross-modal memory retrieval

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:43:43.459814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:43:43.020118Z digest=sha256:0fb1ba999907fa4792b29c6730e6f437337ca2c1f91855ffa5ef8f907505b309

Observation 060a5ffb-ae1f-4927-aa25-9e78b8e0d6b5 · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive nlp tasks.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Retrieval-augmented generation for knowledge-intensive nlp tasks

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:43:43.448584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:43:43.022706Z digest=sha256:090f8286a7fd682e4ae68d6508a7077d32f235837bc9a1d40bdfb6de19ca5c74

Observation d3069805-a769-4317-a2b8-317d9d593d0f · outbound

This paper cites an unresolved cited work.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:43:43.437301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:43:43.025459Z digest=sha256:ababe10dbf5b080e047fe639bf337872e451592cd3c07da01cfbc0474c19b358

Observation 8fc193d5-96b3-4f32-a045-2baea68f0b4b · outbound

This paper cites an unresolved cited work.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:43:43.426406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:43:43.028165Z digest=sha256:d690cbc49f4a8c2767e5e1be027c67351bbd536b52a8dfe5ea61a7adfcecaa5a

Observation f5ab841b-bc38-4beb-92ea-b6fb5748a5ff · outbound

This paper cites Evcap: Retrieval-augmented image captioning 9 with external visual-name memory for open-world compre- hension.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Evcap: Retrieval-augmented image captioning 9 with external visual-name memory for open-world compre- hension

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:43:43.415912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:43:43.030855Z digest=sha256:83491a9ec2dc88efaf018a0f273a5c7da47b3ba9dda25e40bdf247389daa31ab

Observation f13e7679-7e75-4b19-a8cc-bc9b645dce1f · outbound

This paper cites Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:43:43.405063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:43:43.033636Z digest=sha256:8b615134d695f8eab72957e89611706c5ba7b1642a464499427104cec46f92e8

Observation b1c281d9-abeb-4985-a3b9-559479573ca0 · outbound

This paper cites Prefix-diffusion: A Lightweight Diffusion Model for Diverse Image Captioning.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Prefix-diffusion: A Lightweight Diffusion Model for Diverse Image Captioning

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-12T04:43:43.218984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:43:43.036329Z digest=sha256:b29a3b9cee5f5f654df5ae4fee31384346f71484a97409f9b17574310313bd0b

Observation 87e8123a-ffa7-4a86-bc32-a52cfb937224 · outbound

This paper cites Diffusion hyperfeatures: Searching through time and space for semantic correspondence.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Diffusion hyperfeatures: Searching through time and space for semantic correspondence

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T04:43:43.039463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:43:43.039463Z digest=sha256:ba23a379b591a92e30a7f399607ab6d2848faf4027f1e96389b4888a6250d3c9

Observation 37b683c8-40c7-4fa8-a99d-f42add8cdb4c · outbound

This paper cites Semantic-conditional dif- fusion networks for image captioning*.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Semantic-conditional dif- fusion networks for image captioning*

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:43:43.387994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:43:43.042841Z digest=sha256:4c005072cdbc77ecbd156f66cff5e37cebbd80613593f9434a5105b28f42fd31

Observation b02820e3-6aaa-4279-9d01-d0fb14dfdd77 · outbound

This paper cites ClipCap: CLIP Prefix for Image Captioning.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding ClipCap: CLIP Prefix for Image Captioning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T04:43:43.046108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:43:43.046108Z digest=sha256:6f6bcaa88bea9a631617849d5e540b319f7b54ace746ed5db37d64dbd80f4766

Observation 77d9f822-b9da-4672-bc68-362f4d09f598 · outbound

This paper cites an unresolved cited work.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:43:43.377272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:43:43.049896Z digest=sha256:7784f7a1492c3287bba955f466b76e0ed254e92b2bb876501dffe6348cec26cc

Observation 78176fae-007c-4638-9723-60ddb8d68179 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Bleu: a method for automatic evaluation of machine translation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:43:43.368069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:43:43.053181Z digest=sha256:1252b34a93466021b39ed2556a42962c01b3567a96a5217693cba282ab6425f6

Observation bacbc150-f363-4b77-b62f-9772b5d8ae9c · outbound

This paper cites Plummer, Liwei Wang, Christopher M.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Plummer, Liwei Wang, Christopher M

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T04:43:43.056659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:43:43.056659Z digest=sha256:8f6b3c6324701e8c4b83c212a141ba00bb6a27b83b6396ad4e733d662a462008

Observation e05529d1-9cac-4d33-ad06-481455ac4446 · outbound

This paper cites Diffusion-TTA: Test-time Adaptation of Discriminative Models via Generative Feedback.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Diffusion-TTA: Test-time Adaptation of Discriminative Models via Generative Feedback

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T04:43:43.193514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:43:43.060391Z digest=sha256:c1bc1369b311e2a26965cae76840135f9d5450adc572ab33761437b6ca20d4e9

Observation 5d5504fc-006a-4c5f-9704-c28e55b879ad · outbound

This paper cites Smallcap: Lightweight image captioning prompted with retrieval augmentation.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Smallcap: Lightweight image captioning prompted with retrieval augmentation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:43:43.351786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:43:43.064268Z digest=sha256:0a27e413220c7c0abd5a06df27776f9817a0ebb0b4232554aa50afbec7936bbb

Observation c8d15c66-6239-4ae0-87ac-bf3fec9ec12b · outbound

This paper cites Retrieval-augmented image captioning.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Retrieval-augmented image captioning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:43:43.342752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:43:43.067476Z digest=sha256:e71c88905fb2636956955dc3ca683e668d5f042811f401b8a1639d57099c0c8d

Observation e3fa4a6b-4244-4dea-881c-8092ac8c8d57 · outbound

This paper cites Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:43:43.332357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:43:43.070857Z digest=sha256:76e8d43721944a81793a48b2de31cc2c0c44045121ab6f84906aed8aeb8b8a95

Observation bb2682d7-8688-4452-983d-7a195fbb5094 · outbound

This paper cites an unresolved cited work.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:43:43.321676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:43:43.074187Z digest=sha256:8a6ced942165dc6b2fbe1499a06e327438b6f9e09d64f77a22f96f265a905de0

Observation 21eeccbb-b9d7-4fa4-947a-e2117a12272c · outbound

This paper cites Attention is all you need.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Attention is all you need

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T04:43:43.077492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:43:43.077492Z digest=sha256:599aea39648f9c07c25613ba960df19fb98a2eb3764e0483008422551a83ed79

Observation cbef04e2-0d9d-4e0c-953d-8c56f6d35f8b · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Lawrence Zitnick, and Devi Parikh

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:43:43.304232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:43:43.080756Z digest=sha256:8764a27fe7296fe98445916467e15384b0b01ea66bc836748a908735628d5823

Observation 4caaeb35-510d-492e-bdbf-e610e37fbe19 · outbound

This paper cites Diffusion Feedback Helps CLIP See Better.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Diffusion Feedback Helps CLIP See Better

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T04:43:43.084063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:43:43.084063Z digest=sha256:958247bd18f928dd09dc21bbe8110013386c6a82ab7378247efa05af3b6de3e4

Observation 7512b51a-6713-43c6-b7af-c16252e6c7c7 · outbound

This paper cites Language Models with Image Descriptors are Strong Few-Shot Video-Language Learners.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Language Models with Image Descriptors are Strong Few-Shot Video-Language Learners

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-12T04:43:43.169825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:43:43.087911Z digest=sha256:42bf896d5f34d5f7c14542da97d19d4c2ab6ed424130190e901a84a221add6d7

Observation 7a0fe769-fa90-434a-b4e1-470feae52d14 · outbound

This paper cites Re-vilm: Retrieval-augmented visual lan- guage model for zero and few-shot image captioning.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Re-vilm: Retrieval-augmented visual lan- guage model for zero and few-shot image captioning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:43:43.293734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:43:43.091386Z digest=sha256:cf7c9c5cf67122f866eafdaa14a81484da3ef8665dadb36e147135da9221eb6d

Observation 4bd4f742-490c-4586-ad62-57946d644bb2 · outbound

This paper cites Retrieval-Augmented Multimodal Language Modeling.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Retrieval-Augmented Multimodal Language Modeling

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T04:43:43.094591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:43:43.094591Z digest=sha256:02dd754f5d8fbb8350ec260354b46abdab204e94c03faba1ff72d257704eef9f

Observation 8784c95e-fd2f-40c9-ac84-bbc64046fefa · outbound

This paper cites Meacap: Memory-augmented zero- shot image captioning.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Meacap: Memory-augmented zero- shot image captioning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:43:43.282125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:43:43.098031Z digest=sha256:641a6ac6926724f881e44f33c91ab110a34a0ffeb5d2e6ead3d30aec74508b4b

Observation d024bdf6-d798-44ee-83c2-2af253e33ae1 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T04:43:43.101124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:43:43.101124Z digest=sha256:aad2020afa9280a9f6d4714e7df44c295c28000aea300aa64258b1f3987e663f

Observation 1b6e84c1-05cd-4793-8439-9237cb32bf36 · outbound

This paper cites This approach significantly reduces memory and processing requirements while preserving critical visual information.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding This approach significantly reduces memory and processing requirements while preserving critical visual information

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:43:43.272484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:43:43.104667Z digest=sha256:ffe56099be11803bc83a82be50ca9f39a4a3b3baaf857360e4d21e84a7e7dfad

Observation 7a14cb31-2e90-4280-a9c6-13e52123753a · outbound

This paper cites NOUN” or “PROPN.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding NOUN” or “PROPN

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:43:43.262215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:43:43.108226Z digest=sha256:7f17a14b833b91f00a4380bf601bf59c8dfc240cf7ea3369bfbf1e0fb63751e2

Observation 541524c1-893b-4906-a1f7-df5bdb574aad · outbound

This paper cites While the SPICE scores remain similar across different top- n values, the CIDEr scores exhibit more noticeable variation.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding While the SPICE scores remain similar across different top- n values, the CIDEr scores exhibit more noticeable variation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:43:43.251968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:43:43.111732Z digest=sha256:82cd06c1e6ff39ef939274ac865661bc5e934d0e94f51e0fa1f357988de0fbe8

Observation 593fb9af-9731-449e-85a6-ee2bda005b6c · outbound

This paper cites an unresolved cited work.

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:43:43.241682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:43:43.115768Z digest=sha256:15668366c377dad75d56555592ed351208d133bd5bd6720fed28e2a96e6bbdba

Pith citing papers

Observation d9f57b73-834e-4c94-ab03-233692c9035c · inbound

Hierarchical Multi-Modal Retrieval for Knowledge-Grounded News Image Captioning cites this paper.

Hierarchical Multi-Modal Retrieval for Knowledge-Grounded News Image Captioning DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:59:06.853907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T21:35:11.654114Z digest=sha256:7fcadd5425e305a9045fe01911b4ecacaef9895c6cc2203cfdf2216d289d308b