Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:25:03.219873Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 90 of 90 outbound references and 0 inbound Pith citation observations for arXiv:2507.16213.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:25:03.219873Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
90 of 90 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 661f57e2-4e32-438c-b154-cecc90c08493 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6ac2b05-fc2e-46be-ae01-43ad91b33401 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception End-to- end object detection with transformers
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31b4a95e-7c6f-425c-b80d-b069b98b2a3c · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception See-through-text grouping for referring image segmentation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47df1633-dd0f-4905-95a1-a9e790ebee79 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Hybrid task cascade for instance seg- mentation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f727e24-feac-491c-87f6-c04008f95c50 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63d09218-62d6-4d70-9ae5-a041c6419513 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd29b6f5-3711-4d88-99ad-f6bbd9fa9895 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Masked-attention mask transformer for universal image segmentation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab19eb58-68c8-41ff-8140-995f6a9f6a23 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception The cityscapes dataset for semantic urban scene understanding
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 898316a0-dea6-4eea-9633-15ea60828160 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Phraseclick: toward achieving flexible interactive segmenta- tion by phrase and click
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1f8d708-1c0e-4e9b-9ad4-979bbb24ecc8 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Vision-language transformer and query generation for refer- ring segmentation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83ff79e8-4352-442e-be00-2935499567f7 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Open- vocabulary universal image segmentation with maskclip
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c7f3ab4-afa8-4929-a9b5-0f808f00106e · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception The pascal visual object classes (voc) challenge
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bcd79f2-9a5a-4a80-97f0-ca091d410fc5 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Scott, and Weilin Huang
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 663cab1d-8cc8-4886-9358-6fe0e6212210 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Prompt- det: Towards open-vocabulary detection using uncurated im- ages
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 44550596-7f95-4528-9499-cfb82c36b281 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Instagen: Enhancing object detection by training on syn- thetic dataset
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d86fcecc-3ec4-41d0-83dd-71f50e25360e · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Video-r1: Reinforcing video reasoning in mllms, 2025
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 76ea7b3a-3f72-4747-b11b-86cc895129d4 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Frozen-detr: Enhancing detr with image understanding from frozen foundation models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4a16c1d6-ff58-4e98-9d1d-9379d7b4e173 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Open-vocabulary object detection via vision and language knowledge distillation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 99d82295-8421-4552-897d-37e54cda6251 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Dataseg: Taming a universal multi-dataset multi-task segmentation model.Ad- vances in Neural Information Processing Systems, 36, 2024
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6d5f152c-c330-4b77-a032-e1b528425c83 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Textbooks Are All You Need
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a0f9581-8649-42bc-a3c2-150d50c2c733 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Llava-uhd: An lmm perceiving any aspect ratio and high- resolution images
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d7750ed9-fcd1-4a89-b2fe-202c81e2919c · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Open-vocabulary semantic segmentation with decou- pled one-pass network
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4b126936-388e-4bcb-9fb3-ce6583a596a5 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Mask r-cnn
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 532f1273-a3a6-4c33-8345-f58a67eb9341 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Bi-directional relationship inferring net- work for referring image segmentation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f170f6c4-87ad-442b-9138-18c255b56c3c · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Densely connected parameter- efficient tuning for referring image segmentation, 2025
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8937b392-bcdb-4806-998f-86ff53dc2c35 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Referring im- age segmentation via cross-modal progressive comprehen- sion
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4af64b3f-61c2-4889-81eb-fa0ee31fc2fe · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Linguistic structure guided context modeling for referring image segmentation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 742758de-c6f1-475c-9ffa-76cb73bafa49 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Oneformer: One transformer to rule universal image segmentation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 31ef1670-998f-420a-ba23-576d05d84415 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Mdetr- modulated detection for end-to-end multi-modal understand- ing
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d53df0f2-aa30-40d1-9b56-9037c47276d8 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception A spoken language dataset of descrip- tions for speech-based grounded language learning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0b4fbcd1-5e82-4d56-8527-7aebafa74693 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception F-vlm: Open-vocabulary object detection upon frozen vision and language models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8230abef-3c1f-45fa-a6dd-28c842cf370b · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception LISA: Reasoning Segmentation via Large Language Model
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3c4cae8-f981-4ecc-96ce-7fb7dd9ae564 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Lisa: Reasoning segmentation via large language model
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0c07afee-1d4f-4160-ae0d-529ef1a26bd5 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Discobox: Weakly supervised instance segmentation and semantic correspondence from box super- vision
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 63fea573-a68e-4e5b-8d5d-f9b7e661d10f · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Vision Transformers Are Good Mask Auto-Labelers
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b97777ef-b346-48bf-9c22-83ad347d63df · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Mask dino: Towards a unified transformer-based framework for object detection and segmentation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c765920b-5ec7-4141-bf16-ba1e6b80c835 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Distilling detr with visual-linguistic knowledge for open-vocabulary object detection
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f13d64f-ef6b-4553-9b6b-3ee39ac0a1b7 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Grounded language-image pre-training
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ee04b3ad-bf10-4c3d-9327-b6faafbf605d · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Box-supervised instance seg- mentation with level set evolution
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8aed0da2-f026-4012-b0c5-292b7ad09795 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Fully convolutional networks for panoptic segmentation with point-based supervision
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 00a7a282-fc22-4e44-a7d1-25e6ddf1e20e · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception A real-time cross-modality correlation fil- tering method for referring expression comprehension
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4a1657f9-3b3c-4191-a321-b25b7226597a · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2651bf09-4a8b-45ec-9acb-82ad86de3712 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Gres: Gener- alized referring expression segmentation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 516d1c06-fccd-472e-9311-c0d9b60eae89 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Visual instruction tuning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 396ea93a-94bd-43c9-bcfe-5e7eadd99cda · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Poly- former: Referring image segmentation as sequential poly- gon generation
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0b59ecd3-f4d4-4402-b6b1-8d487fe0098b · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4bc258ef-85e1-45bc-9f4e-ec3f03888c67 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Knowledge-guided pairwise recon- struction network for weakly supervised referring expression grounding
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b5080384-84b0-44e6-b417-7c5cd23d5ae8 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Learn- ing cross-modal context graph for visual grounding
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 11f99a97-e9db-4615-9b27-6c0bb48cc3bd · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Swin transformer: Hierarchical vision transformer using shifted windows
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90d19aa5-8690-4e28-99a6-18f1f3dec4e4 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Visual- rft: Visual reinforcement fine-tuning, 2025
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e9b21f22-e97f-48cb-9b98-250a81250791 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adf0279d-f848-4aa0-bb46-befbf3392261 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Scaling open-vocabulary object detection
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 01d20d78-7508-4559-a76f-22cb2c4321b9 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception The role of context for object detection and semantic segmentation in the wild
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 50d46028-436f-4fc3-a959-83ce2362efd2 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Mod- eling context between objects for referring expression under- standing
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c58cdb6b-764e-4a4f-bce1-383a37d9f6e5 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Vision-aware text features in referring image segmentation: From object understanding to context understanding, 2024
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 15adb626-08b6-4fb3-8aef-df41b9d5c851 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception GLaMM: Pixel Grounding Large Multimodal Model
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cf9db23-a1b3-4486-9bd3-6d6313aaeca9 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception PixelLM: Pixel Reasoning with Large Multimodal Model
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc0390e1-8c0b-464c-999f-62c436354f77 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Pixellm: Pixel reasoning with large multimodal model
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 13f2513a-d917-4376-bb03-c9e5c923fe19 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Grounding of textual phrases in images by reconstruction
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d5a2c00e-8dde-4ada-bc17-2b65970cbd35 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Objects365: A large-scale, high-quality dataset for object detection
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b2d08439-a2d8-4034-97ca-e4c024525b74 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Unresolved cited work
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef191971-6719-4529-8669-d85a212a1bb2 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cff9cb04-e1f7-441d-ad2e-147517293b21 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Boxinst: High-performance instance segmentation with box annotations
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 34336494-a6a6-4418-ab58-be450b0b032f · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Max-deeplab: End-to-end panoptic segmentation with mask transformers
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98f4faab-7e40-497d-b19a-127eb7ff4cb4 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb6d06a3-b170-4103-bd28-90dcb74f6005 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Neighbourhood watch: Refer- ring expression comprehension via language-guided graph attention networks
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cb440322-a61d-4a06-8915-060afcbbc754 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afadd51a-cd52-44c8-b374-d943f9aeb7d6 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Cris: Clip- driven referring image segmentation
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 25011881-7661-4cc4-ad98-2e6e512407ab · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception LaSagnA: Language-based Segmentation Assistant for Complex Queries
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c1f6087-bed1-425b-84d5-aa0f2c9c261e · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Hyperseg: Hybrid segmentation assistant with fine-grained visual perceiver
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 02f804a7-f5e7-4b3a-ace8-42250d64c4e5 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Chain-of-thought prompting elicits reasoning in large lan- guage models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6db5f32d-b051-46b5-9ea5-90634e417c42 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Aligning bag of regions for open- vocabulary object detection
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2c52f9d3-d099-4c71-a893-df38831c04ce · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception GSVA: Generalized Segmentation via Multimodal Large Language Models
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 724b5fbe-581d-4ebf-9e0c-5a5490c7dcda · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Upsnet: A unified panoptic segmentation network
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ba103859-1933-47a2-b932-916dd704c33b · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception A simple baseline for open- vocabulary semantic segmentation with pre-trained vision- language model
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ef29d085-269c-400e-9da7-81aeaa72966c · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Dynamic graph at- tention for referring expression comprehension
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0fa72d14-b0f8-446d-8546-d554b1f3a4a6 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception R1-onevision: Advancing generalized multimodal reasoning through cross- modal formalization, 2025
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c0a577bb-61fe-4fe8-ae83-69622b4af0f6 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Lavt: Language-aware vision transformer for referring image segmentation
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4c3979b2-5e45-45c8-9d4d-5aeb617b094f · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Modeling context in referring expres- sions
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b3d0bdee-5d38-4dca-beba-9523bd29be32 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Modeling context in referring expres- sions
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 55de4384-c3c4-4386-a750-18a9859a40bb · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception v-clr: View-consistent learning for open-world instance seg- mentation
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 025244fb-29af-46ee-accf-28ea3e981ce6 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Unresolved cited work
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 18bcd259-4f19-4ff8-b318-985f7371594f · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Grounding referring expressions in images by variational context
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 765e2d08-ab87-4cbf-bb9c-cbe49607ab6e · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception A simple framework for open-vocabulary segmentation and detection
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a3953e7e-616e-4d20-b302-4430a72a2996 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception R1-vl: Learning to reason with multimodal large language models via step- wise group relative policy optimization, 2025
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b42c9963-2ac2-42be-af19-2bc9d7ac31f2 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d326492-bdc2-4deb-9447-6726ca52a6e7 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Psalm: Pixelwise segmentation with large multi-modal model
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c5aaccc8-7ae1-40fb-a4bf-45c33e0a6902 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Semantic under- standing of scenes through the ade20k dataset
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 27c1b8fe-6856-40f3-812b-6a0db2ab204d · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Generalized decoding for pixel, image, and lan- guage
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ee94fd56-822c-4f04-8f19-475ccc1c2842 · outbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Segment everything everywhere all at once
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.