Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:16:58.956248Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 1 inbound Pith citation observation for arXiv:2505.13788.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:16:58.956248Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-18T15:35:30.549656Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-18T15:36:33.974443Z
65 of 65 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 09dd84bc-e99a-4a68-a73f-d66b0237e582 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d63c4fd-c8bb-4d54-9981-0daeaea3f22b · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Language models are few-shot learners.Advances in Neural Information Process- ing Systems, 2020
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 416fac4e-da57-43a9-a434-2ea8d99542e1 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Coco- stuff: Thing and stuff classes in context
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7f1fb598-e667-4965-96d5-90d1ea1ba6b6 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Masked-attention mask transformer for universal image segmentation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d15ef60e-8dbe-4eac-a076-3ddc0df88d7b · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Gonzalez, Ion Stoica, and Eric P
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e51618ff-1578-4644-8119-2a7b89ab50d2 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Scaling instruction- finetuned language models.Journal of Machine Learning Research, 25(70):1–53, 2024
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d92fd4bc-b7b4-4111-b98f-3405b708bcf1 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels InstructBLIP: Towards general-purpose vision-language models with instruction tuning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation eae0eda7-602a-4faa-b29d-c3e25a17db16 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Vision-language transformer and query generation for refer- ring segmentation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a74056c-a561-4845-b66e-2be5685b0bd3 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Segnext: Rethink- ing convolutional attention design for semantic segmenta- tion.Advances in Neural Information Processing Systems, 35:1140–1156, 2022
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 19ebaaa3-2ff0-4877-9359-db185a155d36 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Vizwiz grand challenge: Answering visual questions from blind people
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcf17a96-50ac-4f76-864f-71763da138fc · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Partimagenet: A large, high- quality dataset of parts
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7df6b6c8-d999-4f66-a59a-7899d8d1774f · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Segment any- thing
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 863d9aef-d4b2-41dc-8ace-95a32f96da2c · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Lisa: Reasoning segmentation via large language model
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4e75f98f-0b70-4022-bfb4-3e0cf84c6352 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73940d01-5787-4bff-85ea-6ef3212882e1 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Textbooks Are All You Need II: phi-1.5 technical report
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1633b25-78a9-4ad0-80da-509e450d59a6 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Evaluating Object Hallucination in Large Vision-Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67cc99a7-b38e-4e55-a6be-e8572b9dcf5c · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels A real-time cross-modality correlation fil- tering method for referring expression comprehension
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5ddddb08-c1a1-4a0f-bef9-93f07c7bac04 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Microsoft coco: Common objects in context
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9a237f9f-98aa-47a7-badc-941bc952a127 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Gres: Gen- eralized referring expression segmentation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1d092265-6276-44b7-b4fb-6e48a3f90470 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Visual instruction tuning.Advances in neural information processing systems, 36, 2023
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation eff0206e-393d-4f22-9c06-8c4144e9c8c2 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Improved baselines with visual instruction tuning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74ac52eb-a070-47db-b1c5-8beb11a367c3 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Poly- former: Referring image segmentation as sequential poly- gon generation
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 68b6d641-7668-4151-ad63-7a44f01197ef · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Grounding dino: Marry- ing dino with grounded pre-training for open-set object de- tection.ECCV, 2024
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation aa706191-db2f-45c3-9bbd-aa5100824b10 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vi- sion, pages 216–233
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 638043a9-af16-4de3-bc14-354e29fa3d0a · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Swin transformer: Hierarchical vision transformer using shifted windows
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0591ee7d-26c4-4864-a719-bf7a4e449334 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Decoupled weight de- cay regularization
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f545b85-dd49-4eb7-baa9-9a7fef71fbe3 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 495f148b-aaee-47c4-975a-2dc6b11ac49c · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Generation and comprehension of unambiguous object descriptions
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c378a0cd-837e-4a34-8ecf-503387121e27 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Mod- eling context between objects for referring expression under- standing
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d104872e-ec82-4836-94c7-a272f73ef646 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Ground- ing multimodal large language models to the world
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 92b54a11-894e-4f1e-8fb4-b30f3ffc4db1 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Perceptiongpt: Effectively fusing visual perception into llm
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 96ad3284-1722-4692-a652-f040e8780eb8 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Learning transferable visual models from natural language supervi- sion
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78df9661-6317-4dbe-a284-58cccc982189 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Vision language models are blind: Failing to translate detailed visual features into words
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17f95633-155f-4e49-ada9-1bd882e23a5d · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Paco: Parts and attributes of common objects
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 85722242-5d7a-4c35-80c8-adf3ab1399b0 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Learning to lo- calize objects improves spatial reasoning in visual-llms
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 40476545-e1db-4d61-892a-86128a5eefaa · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Glamm: Pixel grounding large multimodal model
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 632b7d50-0e95-40e9-a834-df1fcc34cd57 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Sam 2: Segment anything in images and videos,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b14ef63e-4cde-4dda-8c67-a8806a6a8a8f · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6239ae00-5d23-438b-9e9f-3772eb0f898a · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Pixellm: Pixel reasoning with large multimodal model
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9ef68c60-9c4e-4480-86bc-f34013e8b54c · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Eyes wide shut? exploring the visual shortcomings of multimodal llms
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9a2f73ce-5b07-436c-8fa0-a10b80332c85 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40b701da-635c-447a-8df8-cb0280a26aa2 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Smith, Daniel Khashabi, and Hannaneh Hajishirzi
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ca5d86b8-6261-4551-b0ac-0113d0e813f6 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Cris: Clip- driven referring image segmentation
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f65f6d2a-5017-4d8f-8ad6-a08bb3345e3c · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Gsva: Generalized segmentation via multimodal large language models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 49981b7a-c443-49b5-8d62-85ba87fb1b9f · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Described object detection: Liberating ob- ject detection with flexible expressions.Advances in Neural Information Processing Systems, 36, 2023
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bf6f09e9-e25a-4ad4-a635-392d2d4b58c6 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels u-LLaVA: Unifying Multi-Modal Tasks via Large Language Model
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a11fe01-d78f-439a-a348-5c2bcc1770d3 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Universal instance percep- tion as object discovery and retrieval
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c0f872fa-63cf-4518-87fe-a4e545761500 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Lavt: Language-aware vision transformer for referring image segmentation
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4f0b568c-c5c1-44a7-b430-f2f978e0c9a7 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Cross-modal self-attention network for referring image seg- mentation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation eb512e66-fed9-4fae-a4e3-ec12d96fd7a8 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Modeling context in referring ex- pressions
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1c129d7c-e875-416b-853c-170fdfce50a0 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Mattnet: Modular at- tention network for referring expression comprehension
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b05a6990-4bb2-41d3-8a60-0c3e709790b7 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ba8ef71-1fab-4795-afc3-334ac4aee03c · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Psalm: Pixelwise segmentation with large multi-modal model.ECCV, 2024
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 938ef0b2-7e23-4d15-ad74-a1309214d717 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Semantic under- standing of scenes through the ade20k dataset.International Journal of Computer Vision, 127:302–321, 2019
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 66b180ee-7c13-4264-8123-5b927ead9590 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Seqtr: A simple yet universal network for visual grounding
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2d64df02-9f33-4e03-9224-700845df341e · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels MiniGPT-4: Enhancing vision-language understanding with advanced large language models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a9afa6c4-e6a8-4cd6-8197-f5fbcc5d9e04 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Self-supervised multimodal learning: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation cc76aa36-84ac-4c55-8617-f4d0e434c4f5 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels African Bush Elephant
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0a367a30-0dd8-450b-9960-592763aefe07 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Example: ”Corgi,” ”Macbook,” ”SUV”
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f4753a50-ffbd-4c01-ab14-c275b8ae4886 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Example: ”Dog,” ”Computer,” ”Car”
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation fb9810c4-6277-45f4-b69a-2ea6d1c1eec5 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Unresolved cited work
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 16aa896e-2974-4248-845d-4ba7c0797225 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Unresolved cited work
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 725cc3dc-de9d-4460-9e94-aae57275f17c · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Unresolved cited work
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e5066cdb-9dd0-498d-950e-f4f2794c6898 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Unresolved cited work
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 58c074a2-72e6-4b46-9b35-4cceb24ce2c5 · outbound
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels (2) text answer: it should be one or a few coherent sentences connecting the objects in (1) that answer the question
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation dd59e59a-2026-4416-91d0-e3cdf7aa25c2 · inbound
MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.