Pith. sign in

Paper Citation Record · LEDGER

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis

As of 14 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 2 inbound Pith citation observations for arXiv:2412.03665.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.03665 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:16:58.754447Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:46:34.401485Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T12:45:52.393487Z

Reference resolution

70 of 70 outbound references displayed

  • verified exact0
  • verified fuzzy43
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 88aff19c-47dd-4379-b481-3392b6cf10ce · outbound

This paper cites GPT-4 Technical Report.

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T22:16:58.573484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:16:58.573484Z digest=sha256:0effef65517c4bc1d8f063b66501b575fb0a0a7365c96cd139dd0d4004ecd139

Observation 0f5c1fa7-5043-4211-b46d-1a118c9f3d00 · outbound

This paper cites In: ICCV (2019) Personalizing Multimodal Large Language Models for Image Captioning 15.

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ICCV (2019) Personalizing Multimodal Large Language Models for Image Captioning 15

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:59.237009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.577275Z digest=sha256:ddba3bb661b24c00261b6ba2d4cd1352e894b1401960c2392de8f1d085db66d0

Observation 910e4384-33dd-4150-8624-aacdd8ee9700 · outbound

This paper cites In: NeurIPS (2022).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: NeurIPS (2022)

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T22:16:58.580202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:16:58.580202Z digest=sha256:f3d914b564b083ce9b3af619b4c5fa98cc82f8efd7c1c381e0b107d881a5be7c

Observation cd854f97-112a-43e4-bc44-a205ad45c790 · outbound

This paper cites In: ECCV (2016).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ECCV (2016)

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:59.225309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.583048Z digest=sha256:b2d0554841adc0b8f84465b513a15866e7e9a0cca6e81f73c88f68c5fb3d6d20

Observation 3764e77d-439e-4868-a124-38ade9500983 · outbound

This paper cites In: CVPR (2018).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: CVPR (2018)

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T22:16:58.586122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:16:58.586122Z digest=sha256:c5574ff6ecbec293f8662a29274b5da89de1e7bf73c58e77c8f124074697cd40

Observation e4f7a5e1-d867-4d2d-8453-68ea91ec47f0 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis Gemini: A Family of Highly Capable Multimodal Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T22:16:58.588905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:16:58.588905Z digest=sha256:fef90baa5a63ff9987a7bfe4773d75704b7b67f2f70dd38b418c6171fad9775f

Observation 995aa4ff-f4a5-4a0e-9190-2ec5bfeba01b · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T22:16:58.592925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:16:58.592925Z digest=sha256:f6dcf6b8ebc6791aadc10e478ba9a07fed176e1255c27d0ff5ffc1c3c121af5e

Observation 3e3cabb5-00f7-43b4-a5e7-175db1906366 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T22:16:58.595888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:16:58.595888Z digest=sha256:d4f934b2d0d3f8d8ea376596c48978bb24b25b70884e4a06ab3df9cc9e58d52e

Observation 6ebf49cd-7a18-4a18-8b92-a1d61af78b4b · outbound

This paper cites In: ACL Workshops (2005).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ACL Workshops (2005)

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:59.213666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.598756Z digest=sha256:1a0fea5d07d42ba301f73f679d2d830cd7ec7fd9da6693d68320b388df7e84bb

Observation 5754b34e-a594-4334-81a8-3065d9f07345 · outbound

This paper cites In: CVPR Workshops (2022).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: CVPR Workshops (2022)

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:59.205384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.601296Z digest=sha256:4e04c3620c42a7918c527e428db8950e554ec4339eb026126e865ea54d9ebd46

Observation 86e52088-6dd4-419b-aeb1-4ab6bba63b3e · outbound

This paper cites In: ICCV (2023).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ICCV (2023)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:59.197621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.603791Z digest=sha256:c8aa980850c521e0f31f64025edf0aa677bc4f704691bb9ba4beca114cc4f838

Observation cc4d6944-e606-463d-a5e6-d9b7746791d3 · outbound

This paper cites In: NeurIPS (2020).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: NeurIPS (2020)

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:59.190462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.606382Z digest=sha256:8be81e5a1d3764008c412dad42a290b70aa4650b08bbdcb84243dae990a4e34f

Observation 35e36477-59b0-4780-b077-34a7af105be3 · outbound

This paper cites In: ACL Findings (2024).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ACL Findings (2024)

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:59.183321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.609023Z digest=sha256:7798993f4a5df1b46c0a57e438f2a711fa285add0461e76c04185e1feef7bc94

Observation 2138525e-80f8-4c82-8cb0-8042c9a11ded · outbound

This paper cites In: CVPR Workshops (2024).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: CVPR Workshops (2024)

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:59.175666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.611471Z digest=sha256:4d4c69012d446aa583f00fe5788499b253e68a857a6d090d328ad0a4b8db47c6

Observation 8a1b7b9c-7d17-479a-af75-df22f09efc9a · outbound

This paper cites In: CVPR (2024).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: CVPR (2024)

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:59.168134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.613997Z digest=sha256:df351c905255678e8d055f6ba5b686199c3cb7efed82bdb7885ef5af7a0300f1

Observation a2bf7dea-865d-4ee1-9603-7fe25d30cfbc · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T22:16:58.616899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:16:58.616899Z digest=sha256:9a426616468b90d5850e5abaed43137223630e20133a39fb8300ec3d5adc65e4

Observation 2151e4eb-b94b-46c7-befc-d70bd9392107 · outbound

This paper cites an unresolved cited work.

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T22:16:58.619616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:16:58.619616Z digest=sha256:cc6b070ecade09fc1e842110b683319c0156737820ae7d50490ab42fd3924866

Observation 88470436-b342-4621-82b2-bf6a7ebf575a · outbound

This paper cites In: NAACL (2022).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: NAACL (2022)

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:59.157081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.621926Z digest=sha256:3d61a88de013b7d0e9bad0b93e09ad1620b5a8fedf6d4cab7cb555b8f71b5235

Observation d2582b90-0d09-4e29-86c2-5f976391ca3f · outbound

This paper cites In: ICRA (2020).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ICRA (2020)

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:59.150266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.624419Z digest=sha256:6a44437a7dc3e8ee8a8711b945eb4bc86412643310e6b245a11c3fbcfadc463e

Observation 9abdeb89-358e-4b48-9f75-bd6887b71352 · outbound

This paper cites AI Communications35(2), 111–129 (2022) 16 D.

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis AI Communications35(2), 111–129 (2022) 16 D

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:59.142930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.626806Z digest=sha256:a36da12ae7644d5bac877c9b20c6e9974aa3ed98c38fdaba2f1a93326176f152

Observation 7ed477d6-b019-43b3-8bfb-aa39aab00757 · outbound

This paper cites In: CVPR (2023).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: CVPR (2023)

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:59.135861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.629271Z digest=sha256:e433cdeab6ef2d41dc1c9cfa1a1ec5c6980d55fca31dc2ea5a0829e022c45e25

Observation e79dbcf7-d40c-47c4-b121-b135fe6c5c1b · outbound

This paper cites In: NeurIPS (2023).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: NeurIPS (2023)

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:59.128855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.631690Z digest=sha256:5ce702b812d1a3141dbb72f100c120ef0872cdff4d36ebfbf417e3f7742aeb55

Observation 27d51029-679a-4be9-8767-3a5fe3267e12 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T22:16:58.634119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:16:58.634119Z digest=sha256:93c7c82ce85a8887ae4898b1b4d86dc566c72db8e13e8dc4dcbbfe42534e8f61

Observation 0ea61bf5-f009-4a01-9953-6a6494dbdd55 · outbound

This paper cites In: ECCV (2020).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ECCV (2020)

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:59.121559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.637157Z digest=sha256:2b13919b33ff8a696ce07b25c7ddb6cedab85a71d63c2e267e23cafb62883c21

Observation 9238119e-4856-41f7-b97e-9c28fbcadbc1 · outbound

This paper cites WARP: Word-level Adversarial ReProgramming.

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis WARP: Word-level Adversarial ReProgramming

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T22:16:58.639576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:16:58.639576Z digest=sha256:943a616e1d019ab8d9cc87848c801257a682aa3c967489a1ad3be4f39f4e90d9

Observation a53b8984-6834-444b-a697-644b86da5462 · outbound

This paper cites In: EMNLP (2021).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: EMNLP (2021)

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:59.114545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.643290Z digest=sha256:dc78d202bbcc2181f9c24570fb6738d25d74f5b7c8c06445ab2b6b251e026c62

Observation 7edc9b36-77b7-4f53-b0ea-dd5389aafefe · outbound

This paper cites In: ICLR (2021).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ICLR (2021)

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:59.106866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.645828Z digest=sha256:cf05ca3f1650b15a98c2d00c594bb058c01c259830c05f890d2b8fad35cce631

Observation ebe83922-685f-43e9-98bc-44c951c2eea6 · outbound

This paper cites In: ICCV (2019).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ICCV (2019)

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:59.099681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.648340Z digest=sha256:ec0530e5568f537637bdc9f3a8cc38d6e13b72811fef9ca9ce073ff61aa9cd43

Observation 206f6240-2aad-4a67-924f-f7c60e6fc1fb · outbound

This paper cites In: ICML (2021).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ICML (2021)

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:59.092595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.650869Z digest=sha256:e6c17ada8ab1791764d5e635d93a4038a45ae37583c8970876d761eab1abd175

Observation c9d05278-9a27-4641-9021-f5f92fc830f6 · outbound

This paper cites In: CVPR (2015).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: CVPR (2015)

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T22:16:58.653354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:16:58.653354Z digest=sha256:15c6111c7ca93797cab5e558fbf0dd3b40c6844835702a402471ece4f4e4fc5c

Observation 7a8bb973-d5a3-4de9-9ac8-31b0b0bb996a · outbound

This paper cites In: ICCV (2023).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ICCV (2023)

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:59.081314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.655889Z digest=sha256:231cac220e191b2b2a5fa78382a0fcdd416ae1d0ccac86e5c60c5a4556c35a69

Observation 391ecc13-847a-4d78-9b7d-b3fd59875024 · outbound

This paper cites IJCV128(7), 1956–1981 (2020).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis IJCV128(7), 1956–1981 (2020)

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:59.073698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.658296Z digest=sha256:f9579b48e8f53f0c1a3f470ba2f19e891bfe78e984f9cf510ec201d4f0f0a09a

Observation aee69a7a-4216-4aec-871d-b8fb996824a2 · outbound

This paper cites In: EMNLP (2021).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: EMNLP (2021)

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:59.066077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.660800Z digest=sha256:b855e9f09649b020f9777d6adb1d07b940bc31f39dff6d8a4be1f87e453d8795

Observation d6592922-bfbe-4c72-a979-c1baa20be5ce · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T22:16:58.663763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:16:58.663763Z digest=sha256:6d2944aff0a4fed829761c7561046203679cc8ffbf8c6c087b532c3a02d4c877

Observation 78590323-17ec-459b-aa7d-6c2fb6549981 · outbound

This paper cites Prefix-Tuning: Optimizing Continuous Prompts for Generation.

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis Prefix-Tuning: Optimizing Continuous Prompts for Generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T22:16:58.666398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:16:58.666398Z digest=sha256:16c2f59d479cbab1d3d589cfcb2328f0d50ec79784f85be5102e919f167423fe

Observation a019a3fd-76a0-4199-95cc-cf7e4430cf55 · outbound

This paper cites In: CVPR (2022).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: CVPR (2022)

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:59.058485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.669126Z digest=sha256:9f11bb65ceee04ae9f4a656acbabc318eea4cd019ec7c3fdacb4a940372852d3

Observation 1f6004f9-19fe-4d08-89c9-a65cabb75b48 · outbound

This paper cites In: ACL Workshops (2004).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ACL Workshops (2004)

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:59.051156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.671741Z digest=sha256:518718c35c04578b257784a36953f23606eca9a414d540b42956890fc601c0d4

Observation 27b982ec-3669-44fb-bede-41011964ecbf · outbound

This paper cites In: ECCV (2014).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ECCV (2014)

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:59.044016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.674167Z digest=sha256:c06fe21d5273ce84d175454302a07d3d40df43bdf4f5ba6ae69b673bbddac39f

Observation ef243cb5-94ad-40e3-8a7b-da1589174f68 · outbound

This paper cites SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models.

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T22:16:58.676660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:16:58.676660Z digest=sha256:360904e3b5b6fe0bb4280fa9594c38244b09e131e7f2b38f13ff3f88c791be6c

Observation b2861bc2-a02a-4b35-8f5f-68d680d25a0c · outbound

This paper cites In: ICLR (2024).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ICLR (2024)

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:59.036899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.679335Z digest=sha256:23f374080ad95ab08146690f82200235b5df37402fe9ced187ee09afe32755fe

Observation 5b8bb909-92ba-4f67-b342-362f2725b765 · outbound

This paper cites In: CVPR (2024).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: CVPR (2024)

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T22:16:58.681591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:16:58.681591Z digest=sha256:3bc487fd95a21f5d7e4b197205cf6d3e7438d016401b073b5a57376347be4539

Observation 0fe19fdf-4aa8-4602-b57f-84859f0c300d · outbound

This paper cites an unresolved cited work.

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:16:59.025883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.684009Z digest=sha256:e472fa6b930d5d13172fecbbc2d43c0a520bdc6e04e41973a9e1820ec16fb151

Observation e530ccb5-fe98-46d8-8257-28e9122ed245 · outbound

This paper cites In: NeurIPS (2023).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: NeurIPS (2023)

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T22:16:58.686457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:16:58.686457Z digest=sha256:57e1a9a474cd1e6f8ea47f408bfc2d9923ad85d1faf1b937982f17afbab157a7

Observation 08ef1a79-1e93-4578-988a-e78a6eefeadd · outbound

This paper cites In: ICML (2024).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ICML (2024)

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:59.014669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.688999Z digest=sha256:29c28496e8e70c78a8ee43afb54956144d3f040df89b6561cd126162f309187c

Observation 8af3cdca-f7f8-4231-8e3c-5050f44c865f · outbound

This paper cites AI Open (2023).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis AI Open (2023)

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:59.007392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.691389Z digest=sha256:1a594d6a30246818e863a78a66d8744a02822477173b68a26e1bc1857f756118

Observation d23ffe5e-79a5-40e2-83da-a342bb42bef7 · outbound

This paper cites ClipCap: CLIP Prefix for Image Captioning.

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis ClipCap: CLIP Prefix for Image Captioning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T22:16:58.693911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:16:58.693911Z digest=sha256:19bf9732425ed48fbbf2eb45295195fb52029f015e771a1465200acec83a2bbc

Observation f901596f-36d3-49be-96c8-b92059dd8d24 · outbound

This paper cites IEEE Intelligent Systems (2024).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis IEEE Intelligent Systems (2024)

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:59.000372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.696458Z digest=sha256:39f6e8ddd7b364d393dd0ecfb739f9b785e873d47867e4b80517a3647c270496

Observation c3834755-49ea-4860-b4f3-7e9272b9d0bf · outbound

This paper cites In: BMVC (2024).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: BMVC (2024)

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:58.992730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.698802Z digest=sha256:5602d3488c2f27bf9551a48a6ff3cf4a6d8f8fa7df25960d46455bcc910efb63

Observation c83fb3cb-40e8-4d9e-8c93-dc2610e2dba9 · outbound

This paper cites In: NeurIPS (2022).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: NeurIPS (2022)

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T22:16:58.701174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:16:58.701174Z digest=sha256:28cf464dc906a09d3fa5d5c94f30d52d8687a13eca5b8812bddfefe4b3cf8b8c

Observation 5c70fbae-54eb-4f02-926c-c904bf16c1be · outbound

This paper cites In: ACL (2002).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ACL (2002)

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:58.980571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.703631Z digest=sha256:e6ea3e1f4fd3335d37780e8baf5f350fe8a61027521a3b1af3796d9159f9477d

Observation a521a331-500c-417d-b8ba-dcd6158b7a0b · outbound

This paper cites In: ICML (2021).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ICML (2021)

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T22:16:58.706026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:16:58.706026Z digest=sha256:d1ae7e1132cc67897f122eabd0ebe77d922ab71892ae8e1109dcd41e8cc0ab73

Observation f0df053e-ff07-4e7b-b4ea-aec08fd98c96 · outbound

This paper cites In: CVPR (2023).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: CVPR (2023)

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:58.968770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.708368Z digest=sha256:9c4145edf0fb1494841c0c20b5acca033415d2f8cd35702b4ddcb993ab784a53

Observation e97c0f19-db0a-4dfe-8f09-a6fa0876f523 · outbound

This paper cites In: CVPR (2017).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: CVPR (2017)

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:58.960932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.710872Z digest=sha256:198a4d07c398550fc7e57341cded29e6b087ffa862aee0003e1365fb15a42f35

Observation 339dc2fa-210f-4bac-b13f-610edc63035b · outbound

This paper cites In: CVPR (2023).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: CVPR (2023)

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:58.953356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.713286Z digest=sha256:9ba1d1f395e36d66fbe5ee2fe924e9d9527484271bd7a9fb7369ad364394b165

Observation 272e898e-2b01-4670-9723-48f8a79a94f0 · outbound

This paper cites In: ECCV (2024).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ECCV (2024)

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:58.946238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.715603Z digest=sha256:3207ddafbbae27e9552063c78acd613b2a51bd2937ec358b2fc7a93fc7ad3475

Observation 17f48872-303c-46c8-80c0-10583998400f · outbound

This paper cites In: ACL (2018).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ACL (2018)

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:58.938784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.718001Z digest=sha256:43b7781da9bbcf859320b9bde699c65efd1464fa49727a3fbb736ae0ae7f9956

Observation 8da7e177-7d18-444e-a67c-0e9d320bc116 · outbound

This paper cites In: ECCV (2020).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: ECCV (2020)

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:58.931253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.720369Z digest=sha256:9fd48369436eaca2b74242033c1e17333917591e0dba54f9deaa9e2a9aa82d13

Observation adc2fd34-5175-46c3-b673-255457601811 · outbound

This paper cites In: CVPR (2010).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: CVPR (2010)

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:58.923858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.722862Z digest=sha256:b548753af639f3341968bb8d2666ebd4080a0d0da22dfeef6775dda3e8569bc2

Observation 187fdc54-7bf0-4347-a9eb-796a5d6aa4cd · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T22:16:58.725267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:16:58.725267Z digest=sha256:8edc7494cf231a7069ce536b16ea02e5e0702d6cf32774d7dbdc1180d58a6afc

Observation 324496d1-b198-4755-9f67-3bd2ddf9775f · outbound

This paper cites In: CVPR (2024).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: CVPR (2024)

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T22:16:58.727872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:16:58.727872Z digest=sha256:9a6f9f8c914e87ff233bb98c748b3923fc60a45f836a55333d407ca8f2135661

Observation 66ef6ac9-5a83-4d59-8b4c-de9e6c9a6a58 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis LLaMA: Open and Efficient Foundation Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T22:16:58.731179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:16:58.731179Z digest=sha256:b9da0672563945c02d716bc8f393834266c9604c64259f5a3a81cd449dc130fc

Observation 7151421f-b469-4715-a8b2-1c2ff8e6e526 · outbound

This paper cites In: CVPR (2015).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: CVPR (2015)

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:58.911614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.734254Z digest=sha256:d39baf8f62b3f9a4831a9c96332ac2e9d17ab33b1ae74443c7da7bb1d48fec08

Observation ad689c1a-694b-41d6-b2ef-346c4ead1263 · outbound

This paper cites In: CVPR (2015) 18 D.

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: CVPR (2015) 18 D

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:58.904323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.736584Z digest=sha256:8cc01c4d844cbedb5e8afdf2bb2a26cfef884aaf9a5979894fd46d440e6500c1

Observation 2106f69c-6cf8-451a-9f0a-b415e65e0a57 · outbound

This paper cites In: AAAI (2024).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: AAAI (2024)

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:58.897106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.738957Z digest=sha256:204ecc9e440ce9b01c9f7ce3da90dff445c2c6a9ce322f47a15e237417ec77a8

Observation 6146e7bc-49c0-40c8-8f60-38f6987a6fd9 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis CogVLM: Visual Expert for Pretrained Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T22:16:58.741333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:16:58.741333Z digest=sha256:1febfef650dd2c00f4eb1151d2ed7bf4395cd8328d76856d09c1881ed3134047

Observation 7cd75297-620f-4325-8d9c-a94eaed1a975 · outbound

This paper cites In: CVPR (2019).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis In: CVPR (2019)

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:58.889719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.743980Z digest=sha256:c4a1e68234811c0ff3489d5ee31024605137ecfaa7850072ea47b216f4e6bbbd

Observation 2ad190ec-6c3a-4ad2-8675-2c52b3184f45 · outbound

This paper cites Proceedings of the IEEE98(8) (2010).

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis Proceedings of the IEEE98(8) (2010)

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:16:58.881946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:16:58.746357Z digest=sha256:a7b5f772070f4dfaaedd58308283cb8216e391f80885bba78e784933e71593b8

Observation 7c5c315c-baa8-4461-890b-879b7c2e9386 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T22:16:58.748692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:16:58.748692Z digest=sha256:8b0e399c0d9e3672bfb2a43158972a4cd9b404f625ae03580186bfaea282e2c1

Observation 61a90ef5-6057-4d49-8292-acf4db629436 · outbound

This paper cites Woodpecker: Hallucination Correction for Multimodal Large Language Models.

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis Woodpecker: Hallucination Correction for Multimodal Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T22:16:58.751793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:16:58.751793Z digest=sha256:6957a98a561dc269b2d0e85eab2ab8550a83dcc99ea5171b77e3982ac8b404ef

Observation 262c8a4c-f739-4d70-b9f4-bd242d2fa1d1 · outbound

This paper cites SVIT: Scaling up Visual Instruction Tuning.

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis SVIT: Scaling up Visual Instruction Tuning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T22:16:58.754447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:16:58.754447Z digest=sha256:f9130c4eb200b838f3ff8b823e43918920458a0fde60c83c859a9f8006b035ed

Pith citing papers

Observation cf8fdaf7-4fae-4a8c-bbdb-27597834ea0e · inbound

R-Genie: Reasoning-Guided Generative Image Editing cites this paper.

R-Genie: Reasoning-Guided Generative Image Editing Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:34.401485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:34.401485Z digest=sha256:c42d71e154ba15c93cbbcf265703ca0359d570fbb9840db2177a949e651f613b

Observation 939add0e-cde7-4d97-a8f7-4d8a88594e64 · inbound

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos cites this paper.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:45:52.476729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:45:44.316344Z digest=sha256:e3a603824e6e563112ccc705e17b853832a811cecc3d8c3bac3b365a80a77c78