Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T22:16:58.754447Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 2 inbound Pith citation observations for arXiv:2412.03665.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T22:16:58.754447Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:46:34.401485Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T12:45:52.393487Z
70 of 70 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 88aff19c-47dd-4379-b481-3392b6cf10ce · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f5c1fa7-5043-4211-b46d-1a118c9f3d00 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ICCV (2019) Personalizing Multimodal Large Language Models for Image Captioning 15
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 910e4384-33dd-4150-8624-aacdd8ee9700 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: NeurIPS (2022)
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd854f97-112a-43e4-bc44-a205ad45c790 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ECCV (2016)
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3764e77d-439e-4868-a124-38ade9500983 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: CVPR (2018)
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4f7a5e1-d867-4d2d-8453-68ea91ec47f0 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis Gemini: A Family of Highly Capable Multimodal Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 995aa4ff-f4a5-4a0e-9190-2ec5bfeba01b · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e3cabb5-00f7-43b4-a5e7-175db1906366 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ebf49cd-7a18-4a18-8b92-a1d61af78b4b · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ACL Workshops (2005)
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5754b34e-a594-4334-81a8-3065d9f07345 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: CVPR Workshops (2022)
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 86e52088-6dd4-419b-aeb1-4ab6bba63b3e · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ICCV (2023)
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cc4d6944-e606-463d-a5e6-d9b7746791d3 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: NeurIPS (2020)
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 35e36477-59b0-4780-b077-34a7af105be3 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ACL Findings (2024)
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2138525e-80f8-4c82-8cb0-8042c9a11ded · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: CVPR Workshops (2024)
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8a1b7b9c-7d17-479a-af75-df22f09efc9a · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: CVPR (2024)
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a2bf7dea-865d-4ee1-9603-7fe25d30cfbc · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2151e4eb-b94b-46c7-befc-d70bd9392107 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88470436-b342-4621-82b2-bf6a7ebf575a · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: NAACL (2022)
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d2582b90-0d09-4e29-86c2-5f976391ca3f · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ICRA (2020)
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9abdeb89-358e-4b48-9f75-bd6887b71352 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis AI Communications35(2), 111–129 (2022) 16 D
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7ed477d6-b019-43b3-8bfb-aa39aab00757 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: CVPR (2023)
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e79dbcf7-d40c-47c4-b121-b135fe6c5c1b · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: NeurIPS (2023)
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 27d51029-679a-4be9-8767-3a5fe3267e12 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ea61bf5-f009-4a01-9953-6a6494dbdd55 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ECCV (2020)
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9238119e-4856-41f7-b97e-9c28fbcadbc1 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis WARP: Word-level Adversarial ReProgramming
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a53b8984-6834-444b-a697-644b86da5462 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: EMNLP (2021)
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7edc9b36-77b7-4f53-b0ea-dd5389aafefe · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ICLR (2021)
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ebe83922-685f-43e9-98bc-44c951c2eea6 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ICCV (2019)
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 206f6240-2aad-4a67-924f-f7c60e6fc1fb · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ICML (2021)
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c9d05278-9a27-4641-9021-f5f92fc830f6 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: CVPR (2015)
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a8bb973-d5a3-4de9-9ac8-31b0b0bb996a · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ICCV (2023)
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 391ecc13-847a-4d78-9b7d-b3fd59875024 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis IJCV128(7), 1956–1981 (2020)
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation aee69a7a-4216-4aec-871d-b8fb996824a2 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: EMNLP (2021)
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d6592922-bfbe-4c72-a979-c1baa20be5ce · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78590323-17ec-459b-aa7d-6c2fb6549981 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis Prefix-Tuning: Optimizing Continuous Prompts for Generation
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a019a3fd-76a0-4199-95cc-cf7e4430cf55 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: CVPR (2022)
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1f6004f9-19fe-4d08-89c9-a65cabb75b48 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ACL Workshops (2004)
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 27b982ec-3669-44fb-bede-41011964ecbf · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ECCV (2014)
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ef243cb5-94ad-40e3-8a7b-da1589174f68 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2861bc2-a02a-4b35-8f5f-68d680d25a0c · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ICLR (2024)
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5b8bb909-92ba-4f67-b342-362f2725b765 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: CVPR (2024)
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fe19fdf-4aa8-4602-b57f-84859f0c300d · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e530ccb5-fe98-46d8-8257-28e9122ed245 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: NeurIPS (2023)
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08ef1a79-1e93-4578-988a-e78a6eefeadd · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ICML (2024)
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8af3cdca-f7f8-4231-8e3c-5050f44c865f · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis AI Open (2023)
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d23ffe5e-79a5-40e2-83da-a342bb42bef7 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis ClipCap: CLIP Prefix for Image Captioning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f901596f-36d3-49be-96c8-b92059dd8d24 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis IEEE Intelligent Systems (2024)
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c3834755-49ea-4860-b4f3-7e9272b9d0bf · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: BMVC (2024)
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c83fb3cb-40e8-4d9e-8c93-dc2610e2dba9 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: NeurIPS (2022)
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c70fbae-54eb-4f02-926c-c904bf16c1be · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ACL (2002)
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a521a331-500c-417d-b8ba-dcd6158b7a0b · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ICML (2021)
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0df053e-ff07-4e7b-b4ea-aec08fd98c96 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: CVPR (2023)
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e97c0f19-db0a-4dfe-8f09-a6fa0876f523 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: CVPR (2017)
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 339dc2fa-210f-4bac-b13f-610edc63035b · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: CVPR (2023)
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 272e898e-2b01-4670-9723-48f8a79a94f0 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ECCV (2024)
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 17f48872-303c-46c8-80c0-10583998400f · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ACL (2018)
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8da7e177-7d18-444e-a67c-0e9d320bc116 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ECCV (2020)
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation adc2fd34-5175-46c3-b673-255457601811 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: CVPR (2010)
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 187fdc54-7bf0-4347-a9eb-796a5d6aa4cd · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 324496d1-b198-4755-9f67-3bd2ddf9775f · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: CVPR (2024)
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66ef6ac9-5a83-4d59-8b4c-de9e6c9a6a58 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis LLaMA: Open and Efficient Foundation Language Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7151421f-b469-4715-a8b2-1c2ff8e6e526 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: CVPR (2015)
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ad689c1a-694b-41d6-b2ef-346c4ead1263 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: CVPR (2015) 18 D
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2106f69c-6cf8-451a-9f0a-b415e65e0a57 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: AAAI (2024)
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6146e7bc-49c0-40c8-8f60-38f6987a6fd9 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis CogVLM: Visual Expert for Pretrained Language Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cd75297-620f-4325-8d9c-a94eaed1a975 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: CVPR (2019)
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2ad190ec-6c3a-4ad2-8675-2c52b3184f45 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis Proceedings of the IEEE98(8) (2010)
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7c5c315c-baa8-4461-890b-879b7c2e9386 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61a90ef5-6057-4d49-8292-acf4db629436 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis Woodpecker: Hallucination Correction for Multimodal Large Language Models
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 262c8a4c-f739-4d70-b9f4-bd242d2fa1d1 · outbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis SVIT: Scaling up Visual Instruction Tuning
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf8fdaf7-4fae-4a8c-bbdb-27597834ea0e · inbound
R-Genie: Reasoning-Guided Generative Image Editing Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 939add0e-cde7-4d97-a8f7-4d8a88594e64 · inbound
VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.