Pith. sign in

Paper Citation Record · LEDGER

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels

As of 20 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 1 inbound Pith citation observation for arXiv:2505.13788.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.13788 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:16:58.956248Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T15:35:30.549656Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T15:36:33.974443Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact0
  • verified fuzzy39
  • unresolved25
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 09dd84bc-e99a-4a68-a73f-d66b0237e582 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:58.696252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:58.696252Z digest=sha256:14512a9ac44d808550e8d06a7b9c94639b67acfd3a1e568d4757718efc169c79

Observation 1d63c4fd-c8bb-4d54-9981-0daeaea3f22b · outbound

This paper cites Language models are few-shot learners.Advances in Neural Information Process- ing Systems, 2020.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Language models are few-shot learners.Advances in Neural Information Process- ing Systems, 2020

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.874434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.701203Z digest=sha256:40c6993f3fa0ac397f9b9a27336d497cad441c6c6f044495fe4d1350716d00d5

Observation 416fac4e-da57-43a9-a434-2ea8d99542e1 · outbound

This paper cites Coco- stuff: Thing and stuff classes in context.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Coco- stuff: Thing and stuff classes in context

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.862042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.704746Z digest=sha256:d3b0113b224ee9f6325ad8957a59def69921c5d0d7a4c33fcb7f443d5ccdc362

Observation 7f1fb598-e667-4965-96d5-90d1ea1ba6b6 · outbound

This paper cites Masked-attention mask transformer for universal image segmentation.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Masked-attention mask transformer for universal image segmentation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.850081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.708662Z digest=sha256:3a6d226158d825883d50e62d016d0b6e6f9e22aafbfee59e1679fa0749400d74

Observation d15ef60e-8dbe-4eac-a076-3ddc0df88d7b · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Gonzalez, Ion Stoica, and Eric P

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:58.712543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:58.712543Z digest=sha256:7cea09ea43c729d72eec36da5fd25d2b5141169d5d0b31e3e9e5c91af959b5ee

Observation e51618ff-1578-4644-8119-2a7b89ab50d2 · outbound

This paper cites Scaling instruction- finetuned language models.Journal of Machine Learning Research, 25(70):1–53, 2024.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Scaling instruction- finetuned language models.Journal of Machine Learning Research, 25(70):1–53, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:58.716142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:58.716142Z digest=sha256:63be3b6db0b8eb1f1c6ab9c82c38e70f3994c8cad2b0b1d26646c85c48b77efe

Observation d92fd4bc-b7b4-4111-b98f-3405b708bcf1 · outbound

This paper cites InstructBLIP: Towards general-purpose vision-language models with instruction tuning.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels InstructBLIP: Towards general-purpose vision-language models with instruction tuning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.821046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.720068Z digest=sha256:adf833b216447c4368a4a9886d3dcf3e6166353fada9b695b8a30747e1697111

Observation eae0eda7-602a-4faa-b29d-c3e25a17db16 · outbound

This paper cites Vision-language transformer and query generation for refer- ring segmentation.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Vision-language transformer and query generation for refer- ring segmentation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:58.723426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:58.723426Z digest=sha256:82ed27b7d5cba646a4792b788f96b6ab477faa3a9a4833fda8a96e2001799dde

Observation 1a74056c-a561-4845-b66e-2be5685b0bd3 · outbound

This paper cites Segnext: Rethink- ing convolutional attention design for semantic segmenta- tion.Advances in Neural Information Processing Systems, 35:1140–1156, 2022.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Segnext: Rethink- ing convolutional attention design for semantic segmenta- tion.Advances in Neural Information Processing Systems, 35:1140–1156, 2022

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.797079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.727483Z digest=sha256:b1392495c6cc552f42eef20a1496f7a9d632685633ce8e76369a7c8a88a8f069

Observation 19ebaaa3-2ff0-4877-9359-db185a155d36 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Vizwiz grand challenge: Answering visual questions from blind people

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:58.731677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:58.731677Z digest=sha256:93cba0967e64bd81ee379f0e11edeeda3dbc16c9b73239ef1f2cc4564684c865

Observation bcf17a96-50ac-4f76-864f-71763da138fc · outbound

This paper cites Partimagenet: A large, high- quality dataset of parts.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Partimagenet: A large, high- quality dataset of parts

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.770700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.735359Z digest=sha256:e4a11b7afce5aae93b5c74bc4037d569220fcd22d6e578428d34651028626baf

Observation 7df6b6c8-d999-4f66-a59a-7899d8d1774f · outbound

This paper cites Segment any- thing.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Segment any- thing

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.756055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.738830Z digest=sha256:be8c231c7204970ed797ee8e5bf7831621eb0b594fdf53550b060b768684a25e

Observation 863d9aef-d4b2-41dc-8ace-95a32f96da2c · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Lisa: Reasoning segmentation via large language model

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.742348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.742999Z digest=sha256:40e93a9679f5b3f2a3351d1da2f35c8d4e141a705fd436b59ca997fcb1357687

Observation 4e75f98f-0b70-4022-bfb4-3e0cf84c6352 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:58.746568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:58.746568Z digest=sha256:e058f186ca945fe4b15469e4ba0b645d3dcf74d151a683eba100e57f89a6babd

Observation 73940d01-5787-4bff-85ea-6ef3212882e1 · outbound

This paper cites Textbooks Are All You Need II: phi-1.5 technical report.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Textbooks Are All You Need II: phi-1.5 technical report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:58.749931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:58.749931Z digest=sha256:ce350c8b17a1e1c28122a507bc79774f20815c63d0d16471add811b7db6a3f4b

Observation a1633b25-78a9-4ad0-80da-509e450d59a6 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Evaluating Object Hallucination in Large Vision-Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:58.753702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:58.753702Z digest=sha256:d24324c516b592dafecca030ec0ad281a7599c1199dacc18a388a0f0bced774c

Observation 67cc99a7-b38e-4e55-a6be-e8572b9dcf5c · outbound

This paper cites A real-time cross-modality correlation fil- tering method for referring expression comprehension.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels A real-time cross-modality correlation fil- tering method for referring expression comprehension

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.720683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.758472Z digest=sha256:66ffb033ae82e79d0c30bead7172d6d8587e4a5ca87deb9261f0d29c0533ec0e

Observation 5ddddb08-c1a1-4a0f-bef9-93f07c7bac04 · outbound

This paper cites Microsoft coco: Common objects in context.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Microsoft coco: Common objects in context

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.709218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.762633Z digest=sha256:6fded808f1fc66c82cca2ab569b16cc06b6b8a670459cb8b13f151e24bc60c52

Observation 9a237f9f-98aa-47a7-badc-941bc952a127 · outbound

This paper cites Gres: Gen- eralized referring expression segmentation.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Gres: Gen- eralized referring expression segmentation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.697485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.766544Z digest=sha256:20359d96e08a30b8f347ba0ce2d02edb9da9e45c2809db9fe262f35784b70951

Observation 1d092265-6276-44b7-b4fb-6e48a3f90470 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2023.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Visual instruction tuning.Advances in neural information processing systems, 36, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.684592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.770043Z digest=sha256:91a178efd430f71b5755d81cdcda94ffbbfb85f94e73eb4e11feec5d2928fe0d

Observation eff0206e-393d-4f22-9c06-8c4144e9c8c2 · outbound

This paper cites Improved baselines with visual instruction tuning.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Improved baselines with visual instruction tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:58.774321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:58.774321Z digest=sha256:ec26ca11c59e3118cd29f0d3e43ca305c58d4b379460f08cd1044a08652475f6

Observation 74ac52eb-a070-47db-b1c5-8beb11a367c3 · outbound

This paper cites Poly- former: Referring image segmentation as sequential poly- gon generation.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Poly- former: Referring image segmentation as sequential poly- gon generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.663132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.778285Z digest=sha256:8aa69837f29f9d443d93e6d353584144798ad79b2e27c39f5edfb7463abc0041

Observation 68b6d641-7668-4151-ad63-7a44f01197ef · outbound

This paper cites Grounding dino: Marry- ing dino with grounded pre-training for open-set object de- tection.ECCV, 2024.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Grounding dino: Marry- ing dino with grounded pre-training for open-set object de- tection.ECCV, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.651654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.782236Z digest=sha256:2056b10dc24514afad99f1d5567f3d4b49e6a1079cc5f419f37057ca1f88f2e2

Observation aa706191-db2f-45c3-9bbd-aa5100824b10 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vi- sion, pages 216–233.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vi- sion, pages 216–233

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:58.786078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:58.786078Z digest=sha256:51cafa78961bbcea4b53eb5181cf7902e88acd9b4e0fb7b8c37fa03cc36e97ed

Observation 638043a9-af16-4de3-bc14-354e29fa3d0a · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Swin transformer: Hierarchical vision transformer using shifted windows

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:58.789570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:58.789570Z digest=sha256:e701b5371d3ebb82dd98e28581117ff848d9c4005b0a5940028ea66ce3e61453

Observation 0591ee7d-26c4-4864-a719-bf7a4e449334 · outbound

This paper cites Decoupled weight de- cay regularization.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Decoupled weight de- cay regularization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:58.792872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:58.792872Z digest=sha256:3535054cfd2702bd2f0c970be80d3db57aa87b512af353e477b953b0a8f7dc55

Observation 0f545b85-dd49-4eb7-baa9-9a7fef71fbe3 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521,.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:58.796858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:58.796858Z digest=sha256:543d873656ca8a859ef4e11275ea02916963a9678da5eb9ac5ddeab99fedddca

Observation 495f148b-aaee-47c4-975a-2dc6b11ac49c · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Generation and comprehension of unambiguous object descriptions

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.604971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.801387Z digest=sha256:c574f24189292f9b6e28ea619bed26d6c0a1b201c3a555b2e653fb5a63fcc8e9

Observation c378a0cd-837e-4a34-8ecf-503387121e27 · outbound

This paper cites Mod- eling context between objects for referring expression under- standing.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Mod- eling context between objects for referring expression under- standing

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.589914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.805957Z digest=sha256:00735c9920d488a1cb1e8eae0e15731512a7b2d1787d310e780ab17947b6f8ca

Observation d104872e-ec82-4836-94c7-a272f73ef646 · outbound

This paper cites Ground- ing multimodal large language models to the world.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Ground- ing multimodal large language models to the world

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.574831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.810397Z digest=sha256:c1043492a089e53a5a0035094f57671ca75c87e9568bfcc5216505c3a08a8999

Observation 92b54a11-894e-4f1e-8fb4-b30f3ffc4db1 · outbound

This paper cites Perceptiongpt: Effectively fusing visual perception into llm.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Perceptiongpt: Effectively fusing visual perception into llm

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.562430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.815076Z digest=sha256:94bda8949711ea8b71bb3b43283c1d9b0f3862f5ca106b7672b61e078bb358cc

Observation 96ad3284-1722-4692-a652-f040e8780eb8 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Learning transferable visual models from natural language supervi- sion

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:58.818976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:58.818976Z digest=sha256:0d2631ef9e5042c2949d77374df9fafd17ec3cfe0d017b4b432c9869bd0ab9bc

Observation 78df9661-6317-4dbe-a284-58cccc982189 · outbound

This paper cites Vision language models are blind: Failing to translate detailed visual features into words.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Vision language models are blind: Failing to translate detailed visual features into words

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:58.822948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:58.822948Z digest=sha256:8b014877d5f2a0844087e225de6b7670560736821e3c8b0939bdf27dd37122fc

Observation 17f95633-155f-4e49-ada9-1bd882e23a5d · outbound

This paper cites Paco: Parts and attributes of common objects.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Paco: Parts and attributes of common objects

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.543849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.828087Z digest=sha256:ba2c34f1587f25abb65a8db2f293ae23a1f4acaa5c9d0f0beef886b6461eb7dd

Observation 85722242-5d7a-4c35-80c8-adf3ab1399b0 · outbound

This paper cites Learning to lo- calize objects improves spatial reasoning in visual-llms.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Learning to lo- calize objects improves spatial reasoning in visual-llms

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.531877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.831770Z digest=sha256:c0e7ef9f3cc61037496468a4e60405e5d2f6257fc6d03206839c1037a74fc232

Observation 40476545-e1db-4d61-892a-86128a5eefaa · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Glamm: Pixel grounding large multimodal model

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.517584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.835127Z digest=sha256:1081867a14e9a786c8048f6c30acc5b881c04d7dd05374e4d4fbe48b53ef19ac

Observation 632b7d50-0e95-40e9-a834-df1fcc34cd57 · outbound

This paper cites Sam 2: Segment anything in images and videos,.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Sam 2: Segment anything in images and videos,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:58.838824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:58.838824Z digest=sha256:50b6b96b5adf47b10dfe4db915d144d9fd6cc5a359dadabb9351de018df4bd46

Observation b14ef63e-4cde-4dda-8c67-a8806a6a8a8f · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:58.842906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:58.842906Z digest=sha256:548f061bba1ea875839d151d8573d17b401bbcf8f4cd5e713d2466983c96e3af

Observation 6239ae00-5d23-438b-9e9f-3772eb0f898a · outbound

This paper cites Pixellm: Pixel reasoning with large multimodal model.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Pixellm: Pixel reasoning with large multimodal model

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.488989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.847087Z digest=sha256:c0712009f12ff44f443168f1694056ff1e8f694bbf20a3362da0cdbf177af55e

Observation 9ef68c60-9c4e-4480-86bc-f34013e8b54c · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.475176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.850767Z digest=sha256:dcab05aa11fc06e971cf64221fcfac42cb89d1d0573dfcdb7033a227b6514b00

Observation 9a2f73ce-5b07-436c-8fa0-a10b80332c85 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:58.854060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:58.854060Z digest=sha256:efa2f0dff739915d40a0ab7e7d2a84705f89df881ee8242fb5d12907841bdddb

Observation 40b701da-635c-447a-8df8-cb0280a26aa2 · outbound

This paper cites Smith, Daniel Khashabi, and Hannaneh Hajishirzi.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Smith, Daniel Khashabi, and Hannaneh Hajishirzi

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.463018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.857631Z digest=sha256:797ce5543dd78da8f8915d951f09c687fe62ed395b9ff2115f9f19fd66bdeb14

Observation ca5d86b8-6261-4551-b0ac-0113d0e813f6 · outbound

This paper cites Cris: Clip- driven referring image segmentation.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Cris: Clip- driven referring image segmentation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:58.861352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:58.861352Z digest=sha256:9bba6ab105594b7bbf304eda825e01f63597cd2b2da1cb953d316783521b4bc5

Observation f65f6d2a-5017-4d8f-8ad6-a08bb3345e3c · outbound

This paper cites Gsva: Generalized segmentation via multimodal large language models.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Gsva: Generalized segmentation via multimodal large language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.444246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.865422Z digest=sha256:ade327647d6b7ed98dac03c1071489e61c00d3ba5886fd68c5caa4ff6c364234

Observation 49981b7a-c443-49b5-8d62-85ba87fb1b9f · outbound

This paper cites Described object detection: Liberating ob- ject detection with flexible expressions.Advances in Neural Information Processing Systems, 36, 2023.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Described object detection: Liberating ob- ject detection with flexible expressions.Advances in Neural Information Processing Systems, 36, 2023

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.430194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.869287Z digest=sha256:cf44b58ef76e08ddf60f7b07a08cc2b83f1ec65eb2d6226d256afbd691e4f7b0

Observation bf6f09e9-e25a-4ad4-a635-392d2d4b58c6 · outbound

This paper cites u-LLaVA: Unifying Multi-Modal Tasks via Large Language Model.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels u-LLaVA: Unifying Multi-Modal Tasks via Large Language Model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:58.873253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:58.873253Z digest=sha256:6e8346de62b391f2ecd702281b68cf1a134a3565ce66da5d16ed8ef01d5be0a2

Observation 4a11fe01-d78f-439a-a348-5c2bcc1770d3 · outbound

This paper cites Universal instance percep- tion as object discovery and retrieval.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Universal instance percep- tion as object discovery and retrieval

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.416560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.877275Z digest=sha256:16471c8bcb31101c82a32ce6832a86c2c6936f9478ec654743fda136797b1be9

Observation c0f872fa-63cf-4518-87fe-a4e545761500 · outbound

This paper cites Lavt: Language-aware vision transformer for referring image segmentation.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Lavt: Language-aware vision transformer for referring image segmentation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.398952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.881501Z digest=sha256:696891e9138c07437e2b3cddfcdf87a5d7ef9d22b62ef1e103483c3fd9ab454b

Observation 4f0b568c-c5c1-44a7-b430-f2f978e0c9a7 · outbound

This paper cites Cross-modal self-attention network for referring image seg- mentation.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Cross-modal self-attention network for referring image seg- mentation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.383640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.885553Z digest=sha256:d97b47994e93899bb47efdce9d57d1fa263257b3440ddc82137981cc463cf7a5

Observation eb512e66-fed9-4fae-a4e3-ec12d96fd7a8 · outbound

This paper cites Modeling context in referring ex- pressions.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Modeling context in referring ex- pressions

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.367796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.891395Z digest=sha256:80d11e8940cd9e32d2b4b97c56d3af8efae92c12769ecb2b7e6e86fac4fdc976

Observation 1c129d7c-e875-416b-853c-170fdfce50a0 · outbound

This paper cites Mattnet: Modular at- tention network for referring expression comprehension.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Mattnet: Modular at- tention network for referring expression comprehension

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.349089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.897950Z digest=sha256:9929ab2655d35fd897c7cb7b4cc53237df89e9cd843101738d194ea9a5f47036

Observation b05a6990-4bb2-41d3-8a60-0c3e709790b7 · outbound

This paper cites InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:58.901751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:58.901751Z digest=sha256:91bbb4e45754fd4c922f05f4de0b36512a0e72ce2e576dcb16e0b4185c2d097d

Observation 7ba8ef71-1fab-4795-afc3-334ac4aee03c · outbound

This paper cites Psalm: Pixelwise segmentation with large multi-modal model.ECCV, 2024.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Psalm: Pixelwise segmentation with large multi-modal model.ECCV, 2024

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.336296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.906875Z digest=sha256:879dfad079fbbd251cd33ec728456148f203dfb119078d73544a6f96abbcbbfd

Observation 938ef0b2-7e23-4d15-ad74-a1309214d717 · outbound

This paper cites Semantic under- standing of scenes through the ade20k dataset.International Journal of Computer Vision, 127:302–321, 2019.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Semantic under- standing of scenes through the ade20k dataset.International Journal of Computer Vision, 127:302–321, 2019

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.314946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.911111Z digest=sha256:8bf62d4ec04094b88c7f14b9fd4319d936904bb7c303fe3856648e9331b498b0

Observation 66b180ee-7c13-4264-8123-5b927ead9590 · outbound

This paper cites Seqtr: A simple yet universal network for visual grounding.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Seqtr: A simple yet universal network for visual grounding

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.297994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.915235Z digest=sha256:e911ec2f65a39db5360e603c0a7802ca9e179749b05e5e8318ee2bc42f089a18

Observation 2d64df02-9f33-4e03-9224-700845df341e · outbound

This paper cites MiniGPT-4: Enhancing vision-language understanding with advanced large language models.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels MiniGPT-4: Enhancing vision-language understanding with advanced large language models

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.280239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.920314Z digest=sha256:0b1a1b4de8ece4b15e3784a59b4ad83f8c91f0ae93c3e7197e73d32ac07233c4

Observation a9afa6c4-e6a8-4cd6-8197-f5fbcc5d9e04 · outbound

This paper cites Self-supervised multimodal learning: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Self-supervised multimodal learning: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.267109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.925298Z digest=sha256:5dfed33e09a3f1dea829197480bf9d158bed4ed79698432113b98307485d3431

Observation cc76aa36-84ac-4c55-8617-f4d0e434c4f5 · outbound

This paper cites African Bush Elephant.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels African Bush Elephant

Reference 58

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T20:16:59.253827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.929785Z digest=sha256:de7ed7b91bf7cc1dcc59972bcc7f7783127202cd88554e1834b96f3c104fd0c8

Observation 0a367a30-0dd8-450b-9960-592763aefe07 · outbound

This paper cites Example: ”Corgi,” ”Macbook,” ”SUV”.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Example: ”Corgi,” ”Macbook,” ”SUV”

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.240731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.934481Z digest=sha256:8a2fc1617f487577eaf853357ea1dfaaae7a24ce2890a91d9c453b7110da987e

Observation f4753a50-ffbd-4c01-ab14-c275b8ae4886 · outbound

This paper cites Example: ”Dog,” ”Computer,” ”Car”.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Example: ”Dog,” ”Computer,” ”Car”

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.222331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.939015Z digest=sha256:2ddba3a606712f5a2c8ae4344837f82bd9e2dbbc55716360eaf54d571255718f

Observation fb9810c4-6277-45f4-b69a-2ea6d1c1eec5 · outbound

This paper cites an unresolved cited work.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:16:59.200238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.942505Z digest=sha256:d5f8d45ce2f5741e219d3ca1699097d031c04be0bc4522913cdd467acee6ee00

Observation 16aa896e-2974-4248-845d-4ba7c0797225 · outbound

This paper cites an unresolved cited work.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:16:59.183700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.945943Z digest=sha256:6a635f43b7380b597aaf638a4c00229b1787f9df17f721988c9c66988838978a

Observation 725cc3dc-de9d-4460-9e94-aae57275f17c · outbound

This paper cites an unresolved cited work.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:16:59.170441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.949234Z digest=sha256:4c6b104b412c695083a7da594c0dc514ea9fabc7e52bf1f3a5212000bae55de4

Observation e5066cdb-9dd0-498d-950e-f4f2794c6898 · outbound

This paper cites an unresolved cited work.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:16:59.155350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.952886Z digest=sha256:845ca4be6717f0a522d58801b8126703cc1ad0838b32a5fc79eafe087dc5256a

Observation 58c074a2-72e6-4b46-9b35-4cceb24ce2c5 · outbound

This paper cites (2) text answer: it should be one or a few coherent sentences connecting the objects in (1) that answer the question.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels (2) text answer: it should be one or a few coherent sentences connecting the objects in (1) that answer the question

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:59.141753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:16:58.956248Z digest=sha256:7e776c96fa16203f8b19249a89a063a23946c744f0c4d64cf2b47f0cb56982a1

Pith citing papers

Observation dd59e59a-2026-4416-91d0-e3cdf7aa25c2 · inbound

MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes cites this paper.

MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:36:33.976759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T15:35:30.549656Z digest=sha256:0faa30a24d391114a8d8e40d9ab6233649661aa4938f8a5a372310d7373df570