Pith. sign in

Paper Citation Record · LEDGER

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion

As of 8 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 7 inbound Pith citation observations for arXiv:2505.18986.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18986 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:25:43.959038Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T07:07:59.925795Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T10:17:58.075590Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8031d5a1-ae22-493d-b571-74ad3918e70f · outbound

This paper cites Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.647029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:41.306134Z digest=sha256:b900a6303e565d25ac2ce17a297c0b08c7ff7dcb7638cfd3fa1a3cee7c8cad10

Observation 80bd6b94-22d0-4e78-98cc-3191d61c86c3 · outbound

This paper cites MMDetection: Open MMLab Detection Toolbox and Benchmark.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion MMDetection: Open MMLab Detection Toolbox and Benchmark

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:41.407593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:41.407593Z digest=sha256:282c26cb214a27fe001e4f8ac706e6032daba98684e2887a1a3e7758e9ae0e24

Observation 6815a255-5f92-403c-a406-0daa57900b10 · outbound

This paper cites A simple framework for contrastive learning of visual representations.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion A simple framework for contrastive learning of visual representations

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:41.524550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:41.524550Z digest=sha256:470d1a18fc1ca29f36d089e0a8c249020ee6c0f686f6d751198e81a786c01ea3

Observation 7695381b-dbd6-445b-acb7-586bd352201e · outbound

This paper cites Pix2seq: A Language Modeling Framework for Object Detection.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Pix2seq: A Language Modeling Framework for Object Detection

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:41.630304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:41.630304Z digest=sha256:097564978e1927dc6e3cedfb0597ffa662fe8dca0f5012044b27894b50a44ea3

Observation 33c63cf9-3e2e-4215-ada8-85d55d388162 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:41.747798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:41.747798Z digest=sha256:6510fa3f65fd26e37f4eac9454633dbb52ece78327dc10dcc3a826995afb3a2b

Observation 48a071c6-3d4b-4e2f-bffb-25d64aea5eb1 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:41.836297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:41.836297Z digest=sha256:8a550a9e6bb09151b2e19cc7e85cca5641a7d3785ced2ac77d337493ae454572

Observation 288d0ccd-d231-4159-be59-242400cd12f7 · outbound

This paper cites Yolo-world: Real- time open-vocabulary object detection.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Yolo-world: Real- time open-vocabulary object detection

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.612379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:41.984356Z digest=sha256:c435f5635e722e61295800ceba66eba330d798c3c92178b2e68e2121c2d908c1

Observation 83de4954-8d7b-44f0-9aa2-636c8924faf8 · outbound

This paper cites Evaluating Large-Vocabulary Object Detectors: The Devil is in the Details.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Evaluating Large-Vocabulary Object Detectors: The Devil is in the Details

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:42.104267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:42.104267Z digest=sha256:dc83a049a1aecf9205e0697f32dd4169c306a2840a19bb04fb1bf40070f3db9a

Observation b01f4f06-47e3-4fb3-ae57-d385e3b4b817 · outbound

This paper cites Reducing network agnostophobia.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Reducing network agnostophobia

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.596905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:42.210795Z digest=sha256:a4d71b852b83316113d74ec0b2a1e5ea70b73938121e962e3fd16625814437e9

Observation 9db2f7de-1b34-4bd2-9a0c-f139f406c28a · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion An image is worth 16x16 words: Transformers for image recognition at scale

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.582137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:42.302873Z digest=sha256:77c025c079b4fdb6eb198c08cbc709cb29ad15e9d45bb412319f54c5f9d30272

Observation f63d6b4c-fd85-4bc0-ad99-27228d92063a · outbound

This paper cites LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:42.436018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:42.436018Z digest=sha256:6651ce68e40d2ea075cffe7703c6b7d4b67b4017a952c4286a4cee8bbf575ac0

Observation 19dd24c6-c3e8-4896-9945-e7cf4b572175 · outbound

This paper cites Recent advances in open set recognition: A survey.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Recent advances in open set recognition: A survey

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.566449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:42.562637Z digest=sha256:d10e4bf08bec413a3e8dc2b327a465ef129c3f7662fab37ae9d677f2ea5f9a63

Observation d1c5f081-6694-4ff2-8139-7c0bcc87d935 · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Lvis: A dataset for large vocabulary instance segmentation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.550654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:42.643665Z digest=sha256:339c6ab4e42f9de52d32c2f87158cf5e27ed7b61450cac0a8ed4b18222710dda

Observation 129904e7-f0fc-417d-a952-164a0f83f016 · outbound

This paper cites Ow- detr: Open-world detection transformer.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Ow- detr: Open-world detection transformer

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.535337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:42.746308Z digest=sha256:de0e21d1c321ead892a19b1df871a5e6a615ca7cd74c92d2e4d8735c882da7da

Observation b5ca2b6e-0f50-4553-a3d8-05fb99cb394f · outbound

This paper cites Momentum contrast for unsupervised visual representation learning.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Momentum contrast for unsupervised visual representation learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.520086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:42.853037Z digest=sha256:dbf127de88a2fbd8e963412afd2d611c97b684baf90848568755536c3cc7e6f9

Observation 1c8338aa-67bb-4a63-9bc2-2445730ab42b · outbound

This paper cites Mask r-cnn.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Mask r-cnn

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.505146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:42.964370Z digest=sha256:e7f6f556f1ec510eef31d5be7d7198ac694bfdc0c568d9d0bb4e8b8256511fd7

Observation e2ea59b2-12b6-42dc-9e04-128fea24ba2f · outbound

This paper cites T-rex2: Towards generic object detection via text-visual prompt synergy.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion T-rex2: Towards generic object detection via text-visual prompt synergy

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.490204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:43.103990Z digest=sha256:d2da741675b5951d1fcac45a5e2a4a0c8c2a51501fbc6980d83963d23c90841b

Observation d3a8e7ce-0b56-4ead-b1f4-6a1d7596a753 · outbound

This paper cites Segment anything.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Segment anything

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.474379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:43.228584Z digest=sha256:c70276749f38c56984d6b85ee94dc8028459a7e833d76bdc04c91e38912703a8

Observation c0219f15-f889-4c94-a70b-64a2b3560a8c · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Lisa: Reasoning segmentation via large language model

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.458987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:43.350210Z digest=sha256:e5a4f1b79d92dd780172fa974e933b9fefafd0037b7a5eb15865186f4c2afa6c

Observation 95ca5c2d-c1c2-4d79-807c-bc76acb402a1 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion LLaVA-OneVision: Easy Visual Task Transfer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:43.465321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:43.465321Z digest=sha256:a1bce291e983045f1b1609574de8f5c06a0799e420c295f5e0b40a2d2f133c8b

Observation 0a9ca5e0-6b27-41a1-bd80-8c35644b54c7 · outbound

This paper cites Dn-detr: Accelerate detr training by introducing query denoising.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Dn-detr: Accelerate detr training by introducing query denoising

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.443299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:43.514620Z digest=sha256:7d0c6f4c59dda6086eaea27570cd1ee44c6a02053436d6a3bc266b648c8850a0

Observation a9ccb8e4-1767-4134-806f-97bc6c3f423b · outbound

This paper cites Coda: A real-world road corner case dataset for object detection in autonomous driving.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Coda: A real-world road corner case dataset for object detection in autonomous driving

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.426559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:43.613826Z digest=sha256:eea6e2239acbc2e019c0f6b8c070718c8c569faab960509d170c08396d8359ae

Observation 1767c512-b56c-4a29-8389-15275d85d357 · outbound

This paper cites Desco: Learning object recognition with rich language descriptions.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Desco: Learning object recognition with rich language descriptions

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.409036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:43.758188Z digest=sha256:548f27e38923668b197c49d6e501ece708349158f0fc040f00f126f07e2391e8

Observation 6188faa5-3e31-4a7e-a0f4-045ff99fa574 · outbound

This paper cites Grounded language-image pre-training.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Grounded language-image pre-training

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.393433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:43.847663Z digest=sha256:e1865d646a96e5bd037a61c62671c8b7c533cf811bae9b1097477db6bbde7e1e

Observation 1ea8ff82-22a0-4319-a530-fb6425cc897f · outbound

This paper cites Generative region-language pretraining for open-ended object detection.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Generative region-language pretraining for open-ended object detection

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.377953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:43.855977Z digest=sha256:e508761ffe36d29b127f798a50b02d2ea066f27e8598a7597975638d8b1a7df2

Observation b1cd2f1d-0e0a-4292-b2cf-c793afca37fb · outbound

This paper cites Training-free open-ended object detection and segmentation via attention as prompts.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Training-free open-ended object detection and segmentation via attention as prompts

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.361788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:43.862354Z digest=sha256:e300e2b05974b1a867bc9eb29845cf8c04c071f75acf793c57e86e8a64740623

Observation e459fbe4-2f7e-433e-943f-ce29b7572115 · outbound

This paper cites Visual instruction tuning.Neural Information Processing Systems (NeurIPS), 2023.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Visual instruction tuning.Neural Information Processing Systems (NeurIPS), 2023

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.344472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:43.867202Z digest=sha256:4faf3e6a902fa58306771174cc7d26e31940515edecc76d703ade27f2e3428aa

Observation 65e03801-1900-45f7-a0fb-91aff4a1bc80 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.328450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:43.873335Z digest=sha256:bdc093fd5a26b3d0766172d3f96fb34323cf1d80c911854bebfe23cdc40e35ce

Observation 66878a2e-4dc8-44c5-a84a-3a94ab7d4a43 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Swin transformer: Hierarchical vision transformer using shifted windows

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.312770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:43.879076Z digest=sha256:e2c8f61bbd03f45dc4cd1b3726f44841063d73c1a5c2b3ce9ec01ea6c7fbe5fc

Observation 53e798e5-d546-4307-b90d-0fb5d8dc97ad · outbound

This paper cites Capdet: Unifying dense captioning and open-world detection pretraining.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Capdet: Unifying dense captioning and open-world detection pretraining

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.296731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:43.884458Z digest=sha256:3ad1850eb8217125bd8726a3d537521cba846c46806b955fb531d01bf768f79b

Observation 86bf87f0-ecd7-4e52-9836-8ac65fd7ce75 · outbound

This paper cites Scaling open-vocabulary object detection.Neural Information Processing Systems (NeurIPS), 2023.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Scaling open-vocabulary object detection.Neural Information Processing Systems (NeurIPS), 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.281439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:43.889852Z digest=sha256:7e3f9c9435629023a053f1389fde26a5cd58f5b239cfcac83f470c0d80e286e9

Observation d0e3482c-010f-4af8-8e77-014c190f77be · outbound

This paper cites Learning transferable visual models from natural language supervision.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Learning transferable visual models from natural language supervision

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:43.895759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:43.895759Z digest=sha256:e23bfef7f763a83dde2d3f9f358caf23a677f4fd2fd89db4ca7ae2e961282bd4

Observation 1cf861bd-aa92-4732-b410-788eb9ba0f3d · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Faster r-cnn: Towards real-time object detection with region proposal networks

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.256381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:43.901395Z digest=sha256:87dee82a1420f82a95fc4b8d79c13ab8622c03d4698f15b7713c41ad8af16040

Observation 952f7a13-222b-4092-b791-fa62baf447d8 · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:43.907816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:43.907816Z digest=sha256:caa8e5cc8262b0d22ebc939a6cf78369c5514c8475b1f44475f0b46ae557efee

Observation bc5ded50-7edf-402d-a22f-6193f73cc3b1 · outbound

This paper cites Scalability in perception for autonomous driving: Waymo open dataset.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Scalability in perception for autonomous driving: Waymo open dataset

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.239758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:43.914161Z digest=sha256:b21cb87ea314fb161504f7b053984f27c0c9ebef13d5361e88e5839c2bd97353

Observation 46e99aa6-694f-4db8-ab1f-aac1b4d565c3 · outbound

This paper cites OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:43.919389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:43.919389Z digest=sha256:48362cfb928270c5a6b176ab68fd30e2eebdeea8ece5d3da2408d890ac48b5cf

Observation c90909b2-a67e-4f66-8a7d-cde1ca2d1edb · outbound

This paper cites Cogvlm: Visual expert for pretrained language models.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Cogvlm: Visual expert for pretrained language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.223244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:43.926211Z digest=sha256:54cc12ec78ae74ab207d4d5429c3b0b34a69b8eb18de27ac4c14c95133955b63

Observation f802fc6a-8f2c-467f-998a-c86d33c79d40 · outbound

This paper cites Detclipv2: Scalable open-vocabulary object detection pre-training via word-region alignment.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Detclipv2: Scalable open-vocabulary object detection pre-training via word-region alignment

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.207117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:43.931973Z digest=sha256:6c69f87145f94c2d34d9a3c493b2ae7d75fb475412f26f49101688c3666e7b14

Observation a048e2b5-a136-4b19-ae26-7e841210f4b8 · outbound

This paper cites Detclip: Dictionary-enriched visual-concept paralleled pre-training for open-world detection.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Detclip: Dictionary-enriched visual-concept paralleled pre-training for open-world detection

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.188398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:43.937160Z digest=sha256:64dd7db583ebe3bdd8f6fe5de88eb350332bb87be5f3e0a27c3edd08b40ebac4

Observation 0cfd0a67-5e40-40cb-ad5a-0454b65051c3 · outbound

This paper cites Detclipv3: Towards versatile generative open-vocabulary object detection.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Detclipv3: Towards versatile generative open-vocabulary object detection

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.171851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:43.942085Z digest=sha256:0b38586355152de975826101d03018d216cb130b8f29e24eec9d1b16c56b168b

Observation e33a2214-f781-4e90-92e1-5313eccf1022 · outbound

This paper cites Ni, and Heung-Yeung Shum.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Ni, and Heung-Yeung Shum

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.155459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:43.947653Z digest=sha256:ed00ba71638adf66b26aa40747655cceb86b84ca87bab144e4f1e15111f93ac1

Observation bf19aa99-7d41-4b33-830f-9a967b0586c9 · outbound

This paper cites Llava-grounding: Grounded visual chat with large multimodal models.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Llava-grounding: Grounded visual chat with large multimodal models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.140375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:43.953447Z digest=sha256:085c82743852c86014bf81731c89263d8a5ec5c8768d3ed8ab08fc2eaaa453dd

Observation a910885d-f952-4053-a270-93fbdc12f077 · outbound

This paper cites Glipv2: Unifying localization and vision-language understanding.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Glipv2: Unifying localization and vision-language understanding

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.123477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:25:43.959038Z digest=sha256:a17c01e406e47f60aeef1a10b956e59be0be80302c52a74344b80baad61114e6

Pith citing papers

Observation f945f4ea-7141-4be0-a3e6-3f2553c5d7f9 · inbound

Visual Prompt Based Reasoning for Offroad Mapping using Multimodal LLMs cites this paper.

Visual Prompt Based Reasoning for Offroad Mapping using Multimodal LLMs VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:25:51.983415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T19:51:59.102471Z digest=sha256:018a2435c3cc9efbd031c5c10700a3db1fbebe7a3e1e805cc2dcb4b314b1c917

Observation fdd5e044-6917-4c84-b422-71aaa53799e8 · inbound

VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection cites this paper.

VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:30.268838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T01:24:56.236650Z digest=sha256:31113c3a13d63fbe1b906afc95a4fbb03bc640f935f4a2bf1ce6b2655eba2980

Observation d0f78be8-c593-4474-959b-3a09d910a640 · inbound

VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection cites this paper.

VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:00:35.844779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T19:14:12.092478Z digest=sha256:da364879838cf5e0c5508c2e10a8aead8864149b6c89f9f166f504e39665e143

Observation 36545036-448a-4b16-a2fa-2703d6456869 · inbound

VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection cites this paper.

VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:51:48.393215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T01:33:55.317107Z digest=sha256:579cae98b883603d04d47f5d7a5df37a5d3f8f87bae06f3964ef2368602f4bbf

Observation 0e043359-7b35-4c85-b811-1538a58c95d3 · inbound

FoodCHA: Multi-Modal LLM Agent for Fine-Grained Food Analysis cites this paper.

FoodCHA: Multi-Modal LLM Agent for Fine-Grained Food Analysis VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:21:09.469192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T16:11:26.104498Z digest=sha256:22b97d3c12b5b3a01d9320adcf8811cdfaa16db59c9732be18121c66f7788907

Observation 3bae7f73-e0ef-4b89-8e37-602e64d41309 · inbound

3DVLA: Enhancing Vision-Language-Action Models via 3D Spatial and Instance Understanding cites this paper.

3DVLA: Enhancing Vision-Language-Action Models via 3D Spatial and Instance Understanding VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:13:16.671658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T07:07:59.925795Z digest=sha256:c632e4db2af6d30655492f205d440263e16f99d0e087b936b5bf5403bc7aaa83

Observation 0aeb88e6-1235-4112-9af7-b80b89eb9aba · inbound

DrivingAgent: Design and Scheduling Agents for Autonomous Driving Systems cites this paper.

DrivingAgent: Design and Scheduling Agents for Autonomous Driving Systems VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:17:58.077080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T10:05:38.910257Z digest=sha256:f843f6b34f8a03c7db2f072b5ef7f29fe142428c0e0a961f1f2274fbfb67d2e5