Pith. sign in

Paper Citation Record · LEDGER

Advancing Visual Large Language Model for Multi-granular Versatile Perception

As of 8 August 2026, this Paper Citation Record lists 90 of 90 outbound references and 0 inbound Pith citation observations for arXiv:2507.16213.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.16213 v1

Coverage vector

measured 90 of 90 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:25:03.219873Z

measured 90 of 90 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

90 of 90 outbound references displayed

  • verified exact1
  • verified fuzzy59
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 661f57e2-4e32-438c-b154-cecc90c08493 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:02.937567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:02.937567Z digest=sha256:1f162bafa04c1bacc8ac171f3822503970b9f00c434e57c69705e750692fd496

Observation a6ac2b05-fc2e-46be-ae01-43ad91b33401 · outbound

This paper cites End-to- end object detection with transformers.

Advancing Visual Large Language Model for Multi-granular Versatile Perception End-to- end object detection with transformers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:02.941989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:02.941989Z digest=sha256:e99ef3dd6ae5cba4f9ae57b1b5914e08d4e5b4924a25f1f06d6f07038c431ebf

Observation 31b4a95e-7c6f-425c-b80d-b069b98b2a3c · outbound

This paper cites See-through-text grouping for referring image segmentation.

Advancing Visual Large Language Model for Multi-granular Versatile Perception See-through-text grouping for referring image segmentation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:02.944924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:02.944924Z digest=sha256:b47ee67957f092f4039955b8a3a954a7e62e00a68f5fb781de1fd2e60fbe2854

Observation 47df1633-dd0f-4905-95a1-a9e790ebee79 · outbound

This paper cites Hybrid task cascade for instance seg- mentation.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Hybrid task cascade for instance seg- mentation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:02.948306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:02.948306Z digest=sha256:6ccdddb740ed877152c4b40d6716821b4a9b212f69bc91fdb53b421cda4b4a97

Observation 7f727e24-feac-491c-87f6-c04008f95c50 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:02.951332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:02.951332Z digest=sha256:181a31a2c574a18f648e449461d0fe9a4d4303e961419e4572a2e0f33776db6e

Observation 63d09218-62d6-4d70-9ae5-a041c6419513 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:02.954769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:02.954769Z digest=sha256:f198fca9a1ee1a03922cb06b6c290e7fc6ab97078261276b9373bd650bb8ec96

Observation bd29b6f5-3711-4d88-99ad-f6bbd9fa9895 · outbound

This paper cites Masked-attention mask transformer for universal image segmentation.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Masked-attention mask transformer for universal image segmentation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:02.957777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:02.957777Z digest=sha256:2bb1da3146494a9131a40c969eac93b60b860f94516c196ab65bea19324a97eb

Observation ab19eb58-68c8-41ff-8140-995f6a9f6a23 · outbound

This paper cites The cityscapes dataset for semantic urban scene understanding.

Advancing Visual Large Language Model for Multi-granular Versatile Perception The cityscapes dataset for semantic urban scene understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:02.960867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:02.960867Z digest=sha256:d196b48f203e3063db60d09ff4d573f1f355855d4897a9a512bae55380b77e03

Observation 898316a0-dea6-4eea-9633-15ea60828160 · outbound

This paper cites Phraseclick: toward achieving flexible interactive segmenta- tion by phrase and click.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Phraseclick: toward achieving flexible interactive segmenta- tion by phrase and click

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:02.963638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:02.963638Z digest=sha256:95135458fa0dfad7c81d88363f084f587987671d17734447b16e5fec00a700e9

Observation b1f8d708-1c0e-4e9b-9ad4-979bbb24ecc8 · outbound

This paper cites Vision-language transformer and query generation for refer- ring segmentation.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Vision-language transformer and query generation for refer- ring segmentation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:02.966434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:02.966434Z digest=sha256:a7d79d321ef74fb1444bcc20ae49e7a93bfbfac629bcc1637f040c5d34a9397f

Observation 83ff79e8-4352-442e-be00-2935499567f7 · outbound

This paper cites Open- vocabulary universal image segmentation with maskclip.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Open- vocabulary universal image segmentation with maskclip

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:02.969366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:02.969366Z digest=sha256:dc18916d7afb3bc5c08ae58738bbad974359046620beb901d9fa65e74758e1de

Observation 1c7f3ab4-afa8-4929-a9b5-0f808f00106e · outbound

This paper cites The pascal visual object classes (voc) challenge.

Advancing Visual Large Language Model for Multi-granular Versatile Perception The pascal visual object classes (voc) challenge

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:02.972553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:02.972553Z digest=sha256:8237de444141fd707f76d83b7562a07ab024e3aa38c0f9bbbb7c085549f86737

Observation 7bcd79f2-9a5a-4a80-97f0-ca091d410fc5 · outbound

This paper cites Scott, and Weilin Huang.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Scott, and Weilin Huang

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.879861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:02.975450Z digest=sha256:a97353af0429f60cc7801ee417c88ab83b41d1b6f5b98030a17c407519c6487e

Observation 663cab1d-8cc8-4886-9358-6fe0e6212210 · outbound

This paper cites Prompt- det: Towards open-vocabulary detection using uncurated im- ages.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Prompt- det: Towards open-vocabulary detection using uncurated im- ages

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.872634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:02.978290Z digest=sha256:23e4ac70eb63c2246f7753ba7ae380dfeac13e9bd627cff34bf20e3597e09000

Observation 44550596-7f95-4528-9499-cfb82c36b281 · outbound

This paper cites Instagen: Enhancing object detection by training on syn- thetic dataset.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Instagen: Enhancing object detection by training on syn- thetic dataset

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.865238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:02.981188Z digest=sha256:e2bcfea4d989befaeadd66c7a660f7d51bd3db458fbb5836603e763afca86f90

Observation d86fcecc-3ec4-41d0-83dd-71f50e25360e · outbound

This paper cites Video-r1: Reinforcing video reasoning in mllms, 2025.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Video-r1: Reinforcing video reasoning in mllms, 2025

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.857596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:02.983811Z digest=sha256:56926ed2a9e014eadc47be5b76fad9d91cdef446ad3390f51571aaa63a4710ba

Observation 76ea7b3a-3f72-4747-b11b-86cc895129d4 · outbound

This paper cites Frozen-detr: Enhancing detr with image understanding from frozen foundation models.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Frozen-detr: Enhancing detr with image understanding from frozen foundation models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.849441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:02.987069Z digest=sha256:523d0eff29ed831ac45e033412716452d1c3fb2f8f99a544fd5002ad3a1a7a4e

Observation 4a16c1d6-ff58-4e98-9d1d-9379d7b4e173 · outbound

This paper cites Open-vocabulary object detection via vision and language knowledge distillation.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Open-vocabulary object detection via vision and language knowledge distillation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.841674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:02.989754Z digest=sha256:c30ce903d3f52baeae15a8736acd068165bbe9fa9c0033fd12e3909e310008e6

Observation 99d82295-8421-4552-897d-37e54cda6251 · outbound

This paper cites Dataseg: Taming a universal multi-dataset multi-task segmentation model.Ad- vances in Neural Information Processing Systems, 36, 2024.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Dataseg: Taming a universal multi-dataset multi-task segmentation model.Ad- vances in Neural Information Processing Systems, 36, 2024

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.833711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:02.992355Z digest=sha256:c38e5085e51df3edee0cf2d9a8eb53b6d43d761eb0ec590a4afa6129f88382d2

Observation 6d5f152c-c330-4b77-a032-e1b528425c83 · outbound

This paper cites Textbooks Are All You Need.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Textbooks Are All You Need

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:02.996158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:02.996158Z digest=sha256:d629049fcf85ecfb2e4aea11f83755d55dbea4b42703fef08ea0312e54b87912

Observation 7a0f9581-8649-42bc-a3c2-150d50c2c733 · outbound

This paper cites Llava-uhd: An lmm perceiving any aspect ratio and high- resolution images.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Llava-uhd: An lmm perceiving any aspect ratio and high- resolution images

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.825303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:02.999268Z digest=sha256:e9c826ef895566f2db5b62699cf06fd0b18215abad96d6678daf5c0519d689da

Observation d7750ed9-fcd1-4a89-b2fe-202c81e2919c · outbound

This paper cites Open-vocabulary semantic segmentation with decou- pled one-pass network.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Open-vocabulary semantic segmentation with decou- pled one-pass network

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.817355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.002244Z digest=sha256:b0208d9387fd844211ba49723664d67536387d8d6718d45d368d597fbf2e3d82

Observation 4b126936-388e-4bcb-9fb3-ce6583a596a5 · outbound

This paper cites Mask r-cnn.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Mask r-cnn

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:03.005059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:03.005059Z digest=sha256:38b3e1965d46b932bebedf53b0e39ff8ae569690516c877eb02f71c2ac4c0f0a

Observation 532f1273-a3a6-4c33-8345-f58a67eb9341 · outbound

This paper cites Bi-directional relationship inferring net- work for referring image segmentation.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Bi-directional relationship inferring net- work for referring image segmentation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.804540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.007733Z digest=sha256:ae6cb85df8cc02a1c17175586ea1a9231a38cbecb9423419c85e60d2cd0b0287

Observation f170f6c4-87ad-442b-9138-18c255b56c3c · outbound

This paper cites Densely connected parameter- efficient tuning for referring image segmentation, 2025.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Densely connected parameter- efficient tuning for referring image segmentation, 2025

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.797166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.010900Z digest=sha256:ce157e699bb7cb414a07ee115c32fc7d3e14f434e00b63bd9048bcf8712c5722

Observation 8937b392-bcdb-4806-998f-86ff53dc2c35 · outbound

This paper cites Referring im- age segmentation via cross-modal progressive comprehen- sion.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Referring im- age segmentation via cross-modal progressive comprehen- sion

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.789460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.013622Z digest=sha256:440a4ba712d525b938957f5460f5eb172e73b7f63c959253b9969934e326858f

Observation 4af64b3f-61c2-4889-81eb-fa0ee31fc2fe · outbound

This paper cites Linguistic structure guided context modeling for referring image segmentation.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Linguistic structure guided context modeling for referring image segmentation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.781953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.016386Z digest=sha256:bccd0e1a38104fda6b87d548ad6f1836d154db39bad8ab18c941582357069a5f

Observation 742758de-c6f1-475c-9ffa-76cb73bafa49 · outbound

This paper cites Oneformer: One transformer to rule universal image segmentation.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Oneformer: One transformer to rule universal image segmentation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.774039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.020091Z digest=sha256:e45eb7b6c5a053ed02bdbb34b354a7d11470f64a1fe7115562d76fb67e1fa88e

Observation 31ef1670-998f-420a-ba23-576d05d84415 · outbound

This paper cites Mdetr- modulated detection for end-to-end multi-modal understand- ing.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Mdetr- modulated detection for end-to-end multi-modal understand- ing

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.765376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.022727Z digest=sha256:def995ad4f214612bc3075a38bf8ac187baccadba530e2268775019eb7c45974

Observation d53df0f2-aa30-40d1-9b56-9037c47276d8 · outbound

This paper cites A spoken language dataset of descrip- tions for speech-based grounded language learning.

Advancing Visual Large Language Model for Multi-granular Versatile Perception A spoken language dataset of descrip- tions for speech-based grounded language learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.756673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.025578Z digest=sha256:dfd48ae53829609f4b0fd5ba1a9d3b0e8dbdb5dc8993d2b0097b48fc3795daab

Observation 0b4fbcd1-5e82-4d56-8527-7aebafa74693 · outbound

This paper cites F-vlm: Open-vocabulary object detection upon frozen vision and language models.

Advancing Visual Large Language Model for Multi-granular Versatile Perception F-vlm: Open-vocabulary object detection upon frozen vision and language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.748921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.028433Z digest=sha256:68e4e75c59c653324f6659fd7f4d127697c4968f25eb425e0c01314dad802a51

Observation 8230abef-3c1f-45fa-a6dd-28c842cf370b · outbound

This paper cites LISA: Reasoning Segmentation via Large Language Model.

Advancing Visual Large Language Model for Multi-granular Versatile Perception LISA: Reasoning Segmentation via Large Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:03.031233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:03.031233Z digest=sha256:ae3715c4a5e926cea0f1a9fe05d10395b4c956d4501388f3b7d3e71b0d3f2036

Observation b3c4cae8-f981-4ecc-96ce-7fb7dd9ae564 · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Lisa: Reasoning segmentation via large language model

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.740982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.034435Z digest=sha256:c6036124fd8e8b1ffc2cf68a2e4990567a7f9e53faaaba66bb553975b4ad0ca9

Observation 0c07afee-1d4f-4160-ae0d-529ef1a26bd5 · outbound

This paper cites Discobox: Weakly supervised instance segmentation and semantic correspondence from box super- vision.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Discobox: Weakly supervised instance segmentation and semantic correspondence from box super- vision

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.733055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.037235Z digest=sha256:f3b209138235f6272e5875dc8d72e91543d228e4aa53c8458f2795a51f8b70ad

Observation 63fea573-a68e-4e5b-8d5d-f9b7e661d10f · outbound

This paper cites Vision Transformers Are Good Mask Auto-Labelers.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Vision Transformers Are Good Mask Auto-Labelers

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:25:03.315861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.040147Z digest=sha256:155ddaf0de9fcb650398acfb184dfc119c69a5d811c4f2021bcc7ebf2cbcbad6

Observation b97777ef-b346-48bf-9c22-83ad347d63df · outbound

This paper cites Mask dino: Towards a unified transformer-based framework for object detection and segmentation.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Mask dino: Towards a unified transformer-based framework for object detection and segmentation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.725166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.044358Z digest=sha256:5dc1eff847f691c7b39b1c2b568e4302e7594c780fd49a55691d9869996f697e

Observation c765920b-5ec7-4141-bf16-ba1e6b80c835 · outbound

This paper cites Distilling detr with visual-linguistic knowledge for open-vocabulary object detection.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Distilling detr with visual-linguistic knowledge for open-vocabulary object detection

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:03.047346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:03.047346Z digest=sha256:5f37e45fae4c9cfa5258b8b192e93b33fa1f293916368ffe46d49e3b54c9f34a

Observation 5f13d64f-ef6b-4553-9b6b-3ee39ac0a1b7 · outbound

This paper cites Grounded language-image pre-training.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Grounded language-image pre-training

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.711995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.051037Z digest=sha256:810016c41131682205b94c48b27f7e51c5393dce5a16cf29ff8301cfdc6bc00b

Observation ee04b3ad-bf10-4c3d-9327-b6faafbf605d · outbound

This paper cites Box-supervised instance seg- mentation with level set evolution.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Box-supervised instance seg- mentation with level set evolution

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.703807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.053794Z digest=sha256:6c6ae8a0db0de8bcd3e72e823b23faa1d11dda980b45fbe4023b6ca5c0011599

Observation 8aed0da2-f026-4012-b0c5-292b7ad09795 · outbound

This paper cites Fully convolutional networks for panoptic segmentation with point-based supervision.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Fully convolutional networks for panoptic segmentation with point-based supervision

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.695829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.056437Z digest=sha256:abe076e872e11d05e50636bbd8f504a210eb7ee8041066f93e10aa8ef3fb3028

Observation 00a7a282-fc22-4e44-a7d1-25e6ddf1e20e · outbound

This paper cites A real-time cross-modality correlation fil- tering method for referring expression comprehension.

Advancing Visual Large Language Model for Multi-granular Versatile Perception A real-time cross-modality correlation fil- tering method for referring expression comprehension

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.687610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.059281Z digest=sha256:bdcdeeb82997793675c11eb5ba0c9db6cb20c216b2724ab12986301efb46de39

Observation 4a1657f9-3b3c-4191-a321-b25b7226597a · outbound

This paper cites Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.680245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.061929Z digest=sha256:548127fe4fcd76a9ab2899505e4e47046a09b7857647c893a1f3600db80b4312

Observation 2651bf09-4a8b-45ec-9acb-82ad86de3712 · outbound

This paper cites Gres: Gener- alized referring expression segmentation.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Gres: Gener- alized referring expression segmentation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.672328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.064565Z digest=sha256:861185800d53addbb01308f3128028db815fac7f363a9d92fdcabc1b5b422493

Observation 516d1c06-fccd-472e-9311-c0d9b60eae89 · outbound

This paper cites Visual instruction tuning.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Visual instruction tuning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.663833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.067583Z digest=sha256:0f3adc3c9cd5b05fa8f8c63e49e498971fa6c4694f22bb0b46cda67442367a97

Observation 396ea93a-94bd-43c9-bcfe-5e7eadd99cda · outbound

This paper cites Poly- former: Referring image segmentation as sequential poly- gon generation.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Poly- former: Referring image segmentation as sequential poly- gon generation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.655401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.070394Z digest=sha256:3b3bbb9c6dde01548e09159ae8719e85dc0c9b393a0e250afe6ea951e02a35c4

Observation 0b59ecd3-f4d4-4402-b6b1-8d487fe0098b · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.647545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.073090Z digest=sha256:7b6ede18a7edf7aff0ff0450a132cd84e2d745f11f30c16832facba4dd458e60

Observation 4bc258ef-85e1-45bc-9f4e-ec3f03888c67 · outbound

This paper cites Knowledge-guided pairwise recon- struction network for weakly supervised referring expression grounding.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Knowledge-guided pairwise recon- struction network for weakly supervised referring expression grounding

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.638476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.075993Z digest=sha256:8373f49e0d228653b9d548c8a30847051bcc8665372f8a3804dd6bdc3d502771

Observation b5080384-84b0-44e6-b417-7c5cd23d5ae8 · outbound

This paper cites Learn- ing cross-modal context graph for visual grounding.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Learn- ing cross-modal context graph for visual grounding

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.630597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.078779Z digest=sha256:b3650e528a2b8b1a5d3f21de4955ee37b93bcfd5adf5a20f6e45ccc7d814226b

Observation 11f99a97-e9db-4615-9b27-6c0bb48cc3bd · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Swin transformer: Hierarchical vision transformer using shifted windows

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:03.081686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:03.081686Z digest=sha256:d6a09a18502d3c067d481d46ae139af9e50dbd18bc4d5b69108fd5e1052b7e8a

Observation 90d19aa5-8690-4e28-99a6-18f1f3dec4e4 · outbound

This paper cites Visual- rft: Visual reinforcement fine-tuning, 2025.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Visual- rft: Visual reinforcement fine-tuning, 2025

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.617590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.084391Z digest=sha256:79adf34ff720405aae9759460caeb801bfe152fd9465736f0b7b08689d224178

Observation e9b21f22-e97f-48cb-9b98-250a81250791 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Advancing Visual Large Language Model for Multi-granular Versatile Perception DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:03.086999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:03.086999Z digest=sha256:b0c52e8dbb6731a6dd698a2075517be6d52605df2c04946e4a807e0bb407991e

Observation adf0279d-f848-4aa0-bb46-befbf3392261 · outbound

This paper cites Scaling open-vocabulary object detection.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Scaling open-vocabulary object detection

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.609171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.090622Z digest=sha256:aa229e71d2dd7fe8a1d5b7ffbb1fd97cf72c08a3966dbfbd4377200d1a267bea

Observation 01d20d78-7508-4559-a76f-22cb2c4321b9 · outbound

This paper cites The role of context for object detection and semantic segmentation in the wild.

Advancing Visual Large Language Model for Multi-granular Versatile Perception The role of context for object detection and semantic segmentation in the wild

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.600915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.093349Z digest=sha256:a30226a1f56946680a3be83b2ebdec8f722d1f27e4fe8abb5a64a0214920680d

Observation 50d46028-436f-4fc3-a959-83ce2362efd2 · outbound

This paper cites Mod- eling context between objects for referring expression under- standing.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Mod- eling context between objects for referring expression under- standing

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.592309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.099243Z digest=sha256:e356f57c2cb2d1c55b15ba4547bfa2d514d1206e1bf04b053def923288351789

Observation c58cdb6b-764e-4a4f-bce1-383a37d9f6e5 · outbound

This paper cites Vision-aware text features in referring image segmentation: From object understanding to context understanding, 2024.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Vision-aware text features in referring image segmentation: From object understanding to context understanding, 2024

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.584054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.101957Z digest=sha256:78661846970112bb4d19fbf36e678062b74d27ae823a5eaadfc26204ce9b64aa

Observation 15adb626-08b6-4fb3-8aef-df41b9d5c851 · outbound

This paper cites GLaMM: Pixel Grounding Large Multimodal Model.

Advancing Visual Large Language Model for Multi-granular Versatile Perception GLaMM: Pixel Grounding Large Multimodal Model

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:03.104795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:03.104795Z digest=sha256:420237d16df05a83ca30321f5ec5b72b9e78ada5f3df50a62591815b0967d306

Observation 5cf9db23-a1b3-4486-9bd3-6d6313aaeca9 · outbound

This paper cites PixelLM: Pixel Reasoning with Large Multimodal Model.

Advancing Visual Large Language Model for Multi-granular Versatile Perception PixelLM: Pixel Reasoning with Large Multimodal Model

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:03.107930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:03.107930Z digest=sha256:f2320c27363e3b7a816f2e572ceedb57477ca39975bece088dc39a64d30af234

Observation cc0390e1-8c0b-464c-999f-62c436354f77 · outbound

This paper cites Pixellm: Pixel reasoning with large multimodal model.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Pixellm: Pixel reasoning with large multimodal model

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.574542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.110890Z digest=sha256:19e4611ea86b98970c24c34ea9c7dbe183927ecf7f12a3688848db1fec6577d1

Observation 13f2513a-d917-4376-bb03-c9e5c923fe19 · outbound

This paper cites Grounding of textual phrases in images by reconstruction.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Grounding of textual phrases in images by reconstruction

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.565223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.114638Z digest=sha256:3230044d76b81f3fedb8c744ecf4ec74d8c46111c1634b1f52f62158c4c92ac1

Observation d5a2c00e-8dde-4ada-bc17-2b65970cbd35 · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Objects365: A large-scale, high-quality dataset for object detection

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.556025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.117431Z digest=sha256:e09af4f2092b26819c9beaf21162de0bf0c0eb790b7c303d51fd04908d05f1b1

Observation b2d08439-a2d8-4034-97ca-e4c024525b74 · outbound

This paper cites an unresolved cited work.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:03.120304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:03.120304Z digest=sha256:150309e04ff24f355416675a00447e3794ea141700e5c7992ecc0731408c2f35

Observation ef191971-6719-4529-8669-d85a212a1bb2 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.541868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.123158Z digest=sha256:275f9e1650c07e1254c078e45b88c8752e8cb51aea1e8a73b17fde5eb010319b

Observation cff9cb04-e1f7-441d-ad2e-147517293b21 · outbound

This paper cites Boxinst: High-performance instance segmentation with box annotations.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Boxinst: High-performance instance segmentation with box annotations

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.533787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.125882Z digest=sha256:cb536007e4acb2161ccc3ea3d5ac127783ce6a4454268ea233282b272adb624b

Observation 34336494-a6a6-4418-ab58-be450b0b032f · outbound

This paper cites Max-deeplab: End-to-end panoptic segmentation with mask transformers.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Max-deeplab: End-to-end panoptic segmentation with mask transformers

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:03.128732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:03.128732Z digest=sha256:b243fbc2cbe64ebd2aa5232fb6b4f05ee83cce6efa17c5b313ec353e48f2e0b5

Observation 98f4faab-7e40-497d-b19a-127eb7ff4cb4 · outbound

This paper cites OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion.

Advancing Visual Large Language Model for Multi-granular Versatile Perception OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:03.131454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:03.131454Z digest=sha256:5957c703980d4ac8f7dda5c954fe15c17d5d4b0f0b10e0fe86dea368ea8ce257

Observation bb6d06a3-b170-4103-bd28-90dcb74f6005 · outbound

This paper cites Neighbourhood watch: Refer- ring expression comprehension via language-guided graph attention networks.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Neighbourhood watch: Refer- ring expression comprehension via language-guided graph attention networks

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.518088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.134522Z digest=sha256:bb45322c36619a4998e24f27939d3523ef8839c9de5153432a192178655d36dc

Observation cb440322-a61d-4a06-8915-060afcbbc754 · outbound

This paper cites Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:03.137410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:03.137410Z digest=sha256:cebdb8cdc9e0ef1db4bad79d4ef7d87f10148aae17497a451081a84003fcee5d

Observation afadd51a-cd52-44c8-b374-d943f9aeb7d6 · outbound

This paper cites Cris: Clip- driven referring image segmentation.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Cris: Clip- driven referring image segmentation

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.509852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.140488Z digest=sha256:522b799286c43eb3cc264aa8cdbf6e59a2a4bb278be076d174e0cc0f6838ff31

Observation 25011881-7661-4cc4-ad98-2e6e512407ab · outbound

This paper cites LaSagnA: Language-based Segmentation Assistant for Complex Queries.

Advancing Visual Large Language Model for Multi-granular Versatile Perception LaSagnA: Language-based Segmentation Assistant for Complex Queries

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:03.143395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:03.143395Z digest=sha256:e586ef64c9be86a6ab1408450049d00bb32926b5f0770dd199bd031d35cdd705

Observation 6c1f6087-bed1-425b-84d5-aa0f2c9c261e · outbound

This paper cites Hyperseg: Hybrid segmentation assistant with fine-grained visual perceiver.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Hyperseg: Hybrid segmentation assistant with fine-grained visual perceiver

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.501288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.146653Z digest=sha256:5bd07674ee579552241e066b8f1d91df4e99cdcf69b09b9219ef1169e36414aa

Observation 02f804a7-f5e7-4b3a-ace8-42250d64c4e5 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Chain-of-thought prompting elicits reasoning in large lan- guage models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:03.149351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:03.149351Z digest=sha256:6e6c7484ba5b0f40b4cb099a026365a3df5e215d47f49a1b71b1dd033e3d35ed

Observation 6db5f32d-b051-46b5-9ea5-90634e417c42 · outbound

This paper cites Aligning bag of regions for open- vocabulary object detection.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Aligning bag of regions for open- vocabulary object detection

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.487596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.152081Z digest=sha256:794ca12590717731edd3875945235acd7ec4128dd66bbf7235e8834a74084329

Observation 2c52f9d3-d099-4c71-a893-df38831c04ce · outbound

This paper cites GSVA: Generalized Segmentation via Multimodal Large Language Models.

Advancing Visual Large Language Model for Multi-granular Versatile Perception GSVA: Generalized Segmentation via Multimodal Large Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:03.155595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:03.155595Z digest=sha256:a994da3c2b645426827b78d51c502ad76b6a830a309262f03f73dc04fdb3bce1

Observation 724b5fbe-581d-4ebf-9e0c-5a5490c7dcda · outbound

This paper cites Upsnet: A unified panoptic segmentation network.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Upsnet: A unified panoptic segmentation network

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.480035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.158835Z digest=sha256:e4ebd1d1c5df1a91164f42c65f135f8f67147f23750d84531271732f9ab86c1a

Observation ba103859-1933-47a2-b932-916dd704c33b · outbound

This paper cites A simple baseline for open- vocabulary semantic segmentation with pre-trained vision- language model.

Advancing Visual Large Language Model for Multi-granular Versatile Perception A simple baseline for open- vocabulary semantic segmentation with pre-trained vision- language model

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.472575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.162083Z digest=sha256:92d477126aff4f6b4d6b11f3433a605d6cfece182d14c9b052a856a1af5c650d

Observation ef29d085-269c-400e-9da7-81aeaa72966c · outbound

This paper cites Dynamic graph at- tention for referring expression comprehension.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Dynamic graph at- tention for referring expression comprehension

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.464252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.165628Z digest=sha256:739bb8d6e62d567c32473148b160599587cbf273574a6630369a0e9c1abbba0f

Observation 0fa72d14-b0f8-446d-8546-d554b1f3a4a6 · outbound

This paper cites R1-onevision: Advancing generalized multimodal reasoning through cross- modal formalization, 2025.

Advancing Visual Large Language Model for Multi-granular Versatile Perception R1-onevision: Advancing generalized multimodal reasoning through cross- modal formalization, 2025

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.456045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.170854Z digest=sha256:f886b3777898b2cc0fdadd383d8bea7eb318c9d8f79d21aa4b8d2ef48ba24ba8

Observation c0a577bb-61fe-4fe8-ae83-69622b4af0f6 · outbound

This paper cites Lavt: Language-aware vision transformer for referring image segmentation.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Lavt: Language-aware vision transformer for referring image segmentation

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.447077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.175849Z digest=sha256:0d97294d0030f37bb7f7eb267474d34fe7838536ea3a2768ea0b599570952d72

Observation 4c3979b2-5e45-45c8-9d4d-5aeb617b094f · outbound

This paper cites Modeling context in referring expres- sions.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Modeling context in referring expres- sions

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.438827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.181626Z digest=sha256:e147741c5e3c0e834e5978048942cbcb2b374867a3ea784d502ec355f855ec72

Observation b3d0bdee-5d38-4dca-beba-9523bd29be32 · outbound

This paper cites Modeling context in referring expres- sions.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Modeling context in referring expres- sions

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.430991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.186291Z digest=sha256:78df60c9aa8a012d383e34487d522b654bd2e22c5ee8b7c50552fff635b8eb6d

Observation 55de4384-c3c4-4386-a750-18a9859a40bb · outbound

This paper cites v-clr: View-consistent learning for open-world instance seg- mentation.

Advancing Visual Large Language Model for Multi-granular Versatile Perception v-clr: View-consistent learning for open-world instance seg- mentation

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.422698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.192454Z digest=sha256:d4c71b30ac9b92f9e35cd8d734d74f58f596c5462864f8cd1429546411e4e045

Observation 025244fb-29af-46ee-accf-28ea3e981ce6 · outbound

This paper cites an unresolved cited work.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:25:03.414643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.200344Z digest=sha256:5cf276aea6e265db054e64906777ef3efa62979fd5796c1311bf7eea00a1c43c

Observation 18bcd259-4f19-4ff8-b318-985f7371594f · outbound

This paper cites Grounding referring expressions in images by variational context.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Grounding referring expressions in images by variational context

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:03.203135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:03.203135Z digest=sha256:be70f0429cfda61b15adbd93f21c45c611211e2e6f718589cbc2dcb0f8d4afdb

Observation 765e2d08-ab87-4cbf-bb9c-cbe49607ab6e · outbound

This paper cites A simple framework for open-vocabulary segmentation and detection.

Advancing Visual Large Language Model for Multi-granular Versatile Perception A simple framework for open-vocabulary segmentation and detection

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.401452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.205535Z digest=sha256:7e542d595438299816a979617ae7d9b874452da71a373ea66b49ab7eef01c6df

Observation a3953e7e-616e-4d20-b302-4430a72a2996 · outbound

This paper cites R1-vl: Learning to reason with multimodal large language models via step- wise group relative policy optimization, 2025.

Advancing Visual Large Language Model for Multi-granular Versatile Perception R1-vl: Learning to reason with multimodal large language models via step- wise group relative policy optimization, 2025

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.392598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.207897Z digest=sha256:5c9d70fc7fae8061762b4cf580816e2b4f16e1a9b0afebddf8faa03217fbf7f2

Observation b42c9963-2ac2-42be-af19-2bc9d7ac31f2 · outbound

This paper cites OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding.

Advancing Visual Large Language Model for Multi-granular Versatile Perception OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:03.210189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:03.210189Z digest=sha256:21bc2fc3750b01815dc50ab8f417174087b0c0062c0927bf7567c578efbf6188

Observation 1d326492-bdc2-4deb-9447-6726ca52a6e7 · outbound

This paper cites Psalm: Pixelwise segmentation with large multi-modal model.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Psalm: Pixelwise segmentation with large multi-modal model

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.384291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.212699Z digest=sha256:0060822fdd658e8c681319df9526a66f3e51d3ba0443cf25391343b4127f2afb

Observation c5aaccc8-7ae1-40fb-a4bf-45c33e0a6902 · outbound

This paper cites Semantic under- standing of scenes through the ade20k dataset.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Semantic under- standing of scenes through the ade20k dataset

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.376001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.215135Z digest=sha256:a43b3e822c2ea42f101e41bb9f2a6a2edff1a9e06416c63730ad003369e53bae

Observation 27c1b8fe-6856-40f3-812b-6a0db2ab204d · outbound

This paper cites Generalized decoding for pixel, image, and lan- guage.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Generalized decoding for pixel, image, and lan- guage

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.367658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.217526Z digest=sha256:766b52a56554a7d3522bc625976ed6dfab79b3d758f14a4c0087b3a78be8caed

Observation ee94fd56-822c-4f04-8f19-475ccc1c2842 · outbound

This paper cites Segment everything everywhere all at once.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Segment everything everywhere all at once

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:25:03.359490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:25:03.219873Z digest=sha256:baa4c6bcf48799bd1c85fa632208a94f4ac85e8f0c5d2814724774f67a33715e

Pith citing papers

No inbound Pith citation observations are available.