Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T10:35:48.056317Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 1 inbound Pith citation observation for arXiv:2411.19106.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T10:35:48.056317Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:26:15.313911Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T15:26:16.116871Z
70 of 70 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 01625707-a1fb-4b23-80fe-c7e910ae8ee4 · outbound
Detailed Object Description with Controllable Dimensions Describing objects by their attributes,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 98a7afe5-4f63-4af1-9b41-b71994d8ea98 · outbound
Detailed Object Description with Controllable Dimensions A new image captioning approach for visually impaired people,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f8072616-6ca4-43f8-aca7-d66831197f51 · outbound
Detailed Object Description with Controllable Dimensions Visual instruction tuning,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db1c39d6-f2ed-4bbf-9a2f-2d597d61f64d · outbound
Detailed Object Description with Controllable Dimensions InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3378b141-db95-40ab-b2a0-12a6ec8f5036 · outbound
Detailed Object Description with Controllable Dimensions Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d5890bf-6d27-48b3-ab1c-2e70abc6e4a9 · outbound
Detailed Object Description with Controllable Dimensions Ferret: Refer and ground anything anywhere at any granularity,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 56c4f11e-d411-4f34-93f5-f0d79587aaa6 · outbound
Detailed Object Description with Controllable Dimensions Osprey: Pixel understanding with visual instruction tuning,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation de5b4f10-c62b-4e9a-9b6b-ea1ea97dc266 · outbound
Detailed Object Description with Controllable Dimensions Gpt4roi: Instruction tuning large language model on region-of- interest,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a909808c-fc55-4a0d-90d8-cb50dcd1cae7 · outbound
Detailed Object Description with Controllable Dimensions Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ba6fc9cf-6e13-4474-8ffa-a4af28239ae7 · outbound
Detailed Object Description with Controllable Dimensions GPT-4 Technical Report
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40b2ef8c-f778-4c2e-9326-47cf79e61c84 · outbound
Detailed Object Description with Controllable Dimensions Say as you wish: Fine- grained control of image caption generation with abstract scene graphs,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation cb689d64-bb52-4b56-a096-6e6f5a6bb125 · outbound
Detailed Object Description with Controllable Dimensions Human-like controllable image captioning with verb-specific semantic roles,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 064b0a90-495b-4473-ba0e-3aeea78d620a · outbound
Detailed Object Description with Controllable Dimensions Length-Controllable Image Captioning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 843e0503-0ea8-4a99-9f54-0194369f9ddc · outbound
Detailed Object Description with Controllable Dimensions Image captioning with controllable and adaptive length levels,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 78a9ab64-72f7-4c6a-9eba-bd8a51811d3e · outbound
Detailed Object Description with Controllable Dimensions SentiCap: Generating Image Descriptions with Sentiments
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 353414c4-068d-43bb-90c6-dff0efe6be2d · outbound
Detailed Object Description with Controllable Dimensions Alpha-clip: A clip model focusing on wherever you want,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 17c5457e-12b3-41a9-999d-fe902ba830b8 · outbound
Detailed Object Description with Controllable Dimensions Kosmos-2: Grounding multimodal large language models to the world,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 82254e72-4d11-4a24-9f3b-a45684d2e4ae · outbound
Detailed Object Description with Controllable Dimensions Open-vocabulary attribute detection,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 28c062d1-fdf8-4f6f-ba7c-ca472fed1984 · outbound
Detailed Object Description with Controllable Dimensions OvarNet: Towards Open-vocabulary Object Attribute Recognition
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5e38b85b-607e-4f10-8b20-9054436ded70 · outbound
Detailed Object Description with Controllable Dimensions Hierarchical visual attribute learning in the wild,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc1fc82d-f697-458a-9b53-36ee71e43d57 · outbound
Detailed Object Description with Controllable Dimensions Coco attributes: Attributes for peo- ple, animals, and objects,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7563b39c-d52d-4f5d-843d-c298fc1ee45e · outbound
Detailed Object Description with Controllable Dimensions Reflective decod- ing network for image captioning,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation af887a95-cd45-4343-a65b-eec3ae886a29 · outbound
Detailed Object Description with Controllable Dimensions Show, control and tell: A framework for generating controllable and grounded captions,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d0997027-8721-4b6a-9a6f-8fd0da4cc149 · outbound
Detailed Object Description with Controllable Dimensions Open-set image tagging with multi-grained text supervision,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d5c422c-c1a1-43d1-917e-6618f65fb836 · outbound
Detailed Object Description with Controllable Dimensions Recognize Anything: A Strong Image Tagging Model
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cea303c-2bea-4d47-b397-696cac09c554 · outbound
Detailed Object Description with Controllable Dimensions Tag2Text: Guiding Vision-Language Model via Image Tagging
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9675df14-9443-47ca-9e2e-1783d99e4690 · outbound
Detailed Object Description with Controllable Dimensions Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image cap- tioning,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 594f5d7a-b24f-4360-8630-34a301c74082 · outbound
Detailed Object Description with Controllable Dimensions Laion- 5b: An open large-scale dataset for training next generation image-text models,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ad253e3f-9fce-4674-a0d4-ff678e1d337e · outbound
Detailed Object Description with Controllable Dimensions Docci: Descriptions of connected and contrasting images,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 19b32b7a-0943-4593-8318-435522822cea · outbound
Detailed Object Description with Controllable Dimensions Dense and aligned captions (dac) promote compositional reasoning in vl models,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a94cc381-02df-4ff5-b333-c36e5e55d52f · outbound
Detailed Object Description with Controllable Dimensions Pixlore: A dataset-driven approach to rich image caption- ing,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 13d4cd19-ccc3-41bf-88b0-4a282f552ba7 · outbound
Detailed Object Description with Controllable Dimensions SPICE: Semantic Propositional Image Caption Evaluation
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3d82ea8-504c-413c-ae82-8aeb7f70dc96 · outbound
Detailed Object Description with Controllable Dimensions Cdkm: Common and distinct knowledge mining network with content interaction for dense captioning,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8fdda0a4-519f-4c14-b4dc-2064fc8e02c3 · outbound
Detailed Object Description with Controllable Dimensions Icocap: Improving video captioning by compounding images,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5f221a8c-83e2-4ce4-8d64-a6894a515d8b · outbound
Detailed Object Description with Controllable Dimensions Fine-grained image captioning with global-local discriminative objective,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3a4e2ebd-d14d-477b-9362-7c56e870d7ad · outbound
Detailed Object Description with Controllable Dimensions Mul- titask learning for cross-domain image captioning,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4db71749-ac45-4575-9548-0e84ad3439c6 · outbound
Detailed Object Description with Controllable Dimensions Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40bc00fa-2b03-46d7-a8f4-5543bde6e78d · outbound
Detailed Object Description with Controllable Dimensions ImageInWords: Unlocking Hyper-Detailed Image Descriptions
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 952c4115-cedb-43a8-a491-a31df8724bf0 · outbound
Detailed Object Description with Controllable Dimensions Boosting entity-aware image captioning with multi-modal knowledge graph,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 47b3914d-8339-4831-905d-14e8ededf831 · outbound
Detailed Object Description with Controllable Dimensions Describing like humans: On diversity in image captioning,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9462c86b-b03c-4ed8-834c-07fdad930a34 · outbound
Detailed Object Description with Controllable Dimensions Benchmarking complex instruction-following with multiple constraints composition,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6c24d2b-b081-401b-961d-5806342bf03d · outbound
Detailed Object Description with Controllable Dimensions Densecap: Fully convolutional localization networks for dense captioning,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b26ba455-6353-4ab9-b9bb-8492f294d36b · outbound
Detailed Object Description with Controllable Dimensions Controlcap: Controllable region-level captioning,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7dc12e9f-7bb0-41b4-b2cb-279501b2d840 · outbound
Detailed Object Description with Controllable Dimensions OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d266fd7-6f17-43ec-aa48-8eafa2903e2a · outbound
Detailed Object Description with Controllable Dimensions Prompt Highlighter: Interactive Control for Multi-Modal LLMs
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 241b12de-afc9-4a69-8aa0-7c188425219a · outbound
Detailed Object Description with Controllable Dimensions FlexCap: Describe Anything in Images in Controllable Detail
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a17ba614-9c42-406a-be2a-490781254861 · outbound
Detailed Object Description with Controllable Dimensions Caption Anything: Interactive Image Description with Diverse Multimodal Controls
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b999e30-97e3-4414-ab62-820aaafc0919 · outbound
Detailed Object Description with Controllable Dimensions Debiasing pretrained generative models by uniformly sampling semantic attributes,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 971ced30-02e6-434f-98ec-559cd1f75ba9 · outbound
Detailed Object Description with Controllable Dimensions Unifying visual attribute learning with object recognition in a multiplicative framework,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f3c2e607-1ce7-49fd-825d-db348c506a6e · outbound
Detailed Object Description with Controllable Dimensions Learning to parameterize visual attributes for open-set fine-grained retrieval,
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2f5b3f3b-7c3f-4c2e-a7ed-d661bd9b06ba · outbound
Detailed Object Description with Controllable Dimensions Benchmarking Segmentation Models with Mask-Preserved Attribute Editing
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d0bf7c53-756a-426f-81a1-f75c003871b1 · outbound
Detailed Object Description with Controllable Dimensions Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering,
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5855857a-2d3e-48db-b118-5ffda950003e · outbound
Detailed Object Description with Controllable Dimensions MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b524447-7507-44f3-be44-13baf064773e · outbound
Detailed Object Description with Controllable Dimensions Mme: A comprehensive evaluation benchmark for multimodal large language models,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7ef05d19-4bdf-4ba8-a320-65cb317a3dd5 · outbound
Detailed Object Description with Controllable Dimensions SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a54a24cf-6522-4bd2-83e7-d1116c974f2b · outbound
Detailed Object Description with Controllable Dimensions SEED-Bench-2: Benchmarking Multimodal Large Language Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18e991f0-6e74-440d-ac84-0659174b7709 · outbound
Detailed Object Description with Controllable Dimensions SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f595b07-6d6b-471b-a02a-5b098af6f983 · outbound
Detailed Object Description with Controllable Dimensions Woodpecker: Hallucination Correction for Multimodal Large Language Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c992c0a2-7dc5-421f-ac95-62ad52703d06 · outbound
Detailed Object Description with Controllable Dimensions FaithScore: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 514c7eca-6f46-40c8-bfdc-8f0243fb3bfa · outbound
Detailed Object Description with Controllable Dimensions Evaluating Object Hallucination in Large Vision-Language Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73ef9f90-9d6e-4db9-99a7-1e95f222cbb7 · outbound
Detailed Object Description with Controllable Dimensions Davidsonian scene graph: Improving reliability in fine-grained evaluation for text-to-image generation,
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3cb156dc-3df5-43f7-993b-0464db26413a · outbound
Detailed Object Description with Controllable Dimensions Bleu: a method for automatic evaluation of machine translation,
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3e2b5125-5d48-40b7-8675-7ba40dd47cde · outbound
Detailed Object Description with Controllable Dimensions METEOR: An automatic metric for MT evaluation with improved correlation with human judgments,
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 20e61f48-e320-4dad-b4c3-44551232a430 · outbound
Detailed Object Description with Controllable Dimensions CIDEr: Consensus-based Image Description Evaluation
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3332ba3-5785-42de-b973-7e501621c279 · outbound
Detailed Object Description with Controllable Dimensions CLIPScore: A Reference-free Evaluation Metric for Image Captioning
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0577e74-aa53-4d04-8388-75a6a5aa47dc · outbound
Detailed Object Description with Controllable Dimensions Improved baselines with visual instruction tuning,
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1152873-3296-4039-8d8f-dd6aee64ba02 · outbound
Detailed Object Description with Controllable Dimensions Vicuna: An open-source chatbot impressing gpt- 4 with 90%* chatgpt quality,
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a779451d-4b33-46cf-8b22-2473c7e98cfa · outbound
Detailed Object Description with Controllable Dimensions Llama 3 model card,
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf3e4c25-36dc-4702-a277-c20a1ec407f8 · outbound
Detailed Object Description with Controllable Dimensions Microsoft coco: Common objects in context,
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64d73cef-163e-4ba0-ab26-dc194996979f · outbound
Detailed Object Description with Controllable Dimensions Benchmarking Complex Instruction-Following with Multiple Constraints Composition
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 144dfe14-17ee-4373-af5d-381632a1b5ca · inbound
Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation Detailed Object Description with Controllable Dimensions
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.