Pith. sign in

Paper Citation Record · LEDGER

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion

As of 19 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 7 inbound Pith citation observations for arXiv:2505.18986.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18986 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:25:43.959038Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T07:07:59.925795Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T10:17:58.075590Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8031d5a1-ae22-493d-b571-74ad3918e70f · outbound

This paper cites Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.647029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:25:41.306134Z digest=sha256:d08f27b4720a9d1eac5c148ea26b636417715b2f49b8228110a46b27ed2733b5

Observation 80bd6b94-22d0-4e78-98cc-3191d61c86c3 · outbound

This paper cites MMDetection: Open MMLab Detection Toolbox and Benchmark.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion MMDetection: Open MMLab Detection Toolbox and Benchmark

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:41.407593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:41.407593Z digest=sha256:dc5c1a648e8a0eae9b1c348bcca52f26b881dc86d205af659f446bd4af85e23c

Observation 6815a255-5f92-403c-a406-0daa57900b10 · outbound

This paper cites A simple framework for contrastive learning of visual representations.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion A simple framework for contrastive learning of visual representations

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:41.524550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:41.524550Z digest=sha256:eafe6f0800c9d6a8b04a6e31308ccdf5b0bd67134fd13212c61b5ed0206f2b9b

Observation 7695381b-dbd6-445b-acb7-586bd352201e · outbound

This paper cites Pix2seq: A Language Modeling Framework for Object Detection.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Pix2seq: A Language Modeling Framework for Object Detection

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:41.630304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:41.630304Z digest=sha256:2f8c8e74a01486d0a8da1fce141375bc8f5d63eb86a0094bb7e654b24ad8b097

Observation 33c63cf9-3e2e-4215-ada8-85d55d388162 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:41.747798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:41.747798Z digest=sha256:3b178625452b0492809f871bd4d743b4079abd67aa84a72d2a83785721a29bee

Observation 48a071c6-3d4b-4e2f-bffb-25d64aea5eb1 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:41.836297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:41.836297Z digest=sha256:9da627b05a6b903e655eeab4fa400eff22d0589e44fe67cf5472b68898790666

Observation 288d0ccd-d231-4159-be59-242400cd12f7 · outbound

This paper cites Yolo-world: Real- time open-vocabulary object detection.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Yolo-world: Real- time open-vocabulary object detection

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.612379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:25:41.984356Z digest=sha256:e006c2868b70a5e3b258e4418e083389672b0be792aaa8c00c50027eee60a537

Observation 83de4954-8d7b-44f0-9aa2-636c8924faf8 · outbound

This paper cites Evaluating Large-Vocabulary Object Detectors: The Devil is in the Details.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Evaluating Large-Vocabulary Object Detectors: The Devil is in the Details

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:42.104267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:42.104267Z digest=sha256:e922df46c3e81d3bf88308416d3a6e16f8bf029b14b067f34d8c7135a90bce1a

Observation b01f4f06-47e3-4fb3-ae57-d385e3b4b817 · outbound

This paper cites Reducing network agnostophobia.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Reducing network agnostophobia

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.596905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:25:42.210795Z digest=sha256:c5389bbc1f123ae11b45ae06105f05b211c4c4f200d08e7a7a1fc7e08bd20320

Observation 9db2f7de-1b34-4bd2-9a0c-f139f406c28a · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion An image is worth 16x16 words: Transformers for image recognition at scale

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.582137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:25:42.302873Z digest=sha256:f8993dfd839a06fd92f2ac6cac0c92004f03a7ca799fe74f77dc4165578d7bf1

Observation f63d6b4c-fd85-4bc0-ad99-27228d92063a · outbound

This paper cites LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:42.436018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:42.436018Z digest=sha256:ea067b598404805b353318400f2ba6931f18591c187b0522a5b2061fddbd2eaf

Observation 19dd24c6-c3e8-4896-9945-e7cf4b572175 · outbound

This paper cites Recent advances in open set recognition: A survey.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Recent advances in open set recognition: A survey

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.566449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:25:42.562637Z digest=sha256:1607080384a90ca4fb7d6b0b49ad12dcc9b58279360b111a4f609a2a5c0608b1

Observation d1c5f081-6694-4ff2-8139-7c0bcc87d935 · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Lvis: A dataset for large vocabulary instance segmentation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.550654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:25:42.643665Z digest=sha256:c6b9087fd352113066cbac95d0de66ad01213a3f417a541b0c6d48d6eb7c4722

Observation 129904e7-f0fc-417d-a952-164a0f83f016 · outbound

This paper cites Ow- detr: Open-world detection transformer.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Ow- detr: Open-world detection transformer

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.535337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:25:42.746308Z digest=sha256:f150a516ecbdc3af01a6f6a8795f31af8f2a15360afecdef94f1a7815cfbe2a4

Observation b5ca2b6e-0f50-4553-a3d8-05fb99cb394f · outbound

This paper cites Momentum contrast for unsupervised visual representation learning.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Momentum contrast for unsupervised visual representation learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.520086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:25:42.853037Z digest=sha256:47aa892c9117ce44c7110278130043a90d33d578e8adab0c029f75883fac34da

Observation 1c8338aa-67bb-4a63-9bc2-2445730ab42b · outbound

This paper cites Mask r-cnn.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Mask r-cnn

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.505146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:25:42.964370Z digest=sha256:b224e5f1a9f34f2c0f3f4c0b425aba79e24f641953fb6a0fe12ee4dbed81cbeb

Observation e2ea59b2-12b6-42dc-9e04-128fea24ba2f · outbound

This paper cites T-rex2: Towards generic object detection via text-visual prompt synergy.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion T-rex2: Towards generic object detection via text-visual prompt synergy

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.490204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:25:43.103990Z digest=sha256:24fe39ab843baa05ba9432eadba5a9ac08504d085183acb30ead109659ce6b03

Observation d3a8e7ce-0b56-4ead-b1f4-6a1d7596a753 · outbound

This paper cites Segment anything.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Segment anything

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.474379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:25:43.228584Z digest=sha256:cb6c55be475f4758a5d96caacb2856b5edbf19543f4580c6bfe6242ba453b49b

Observation c0219f15-f889-4c94-a70b-64a2b3560a8c · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Lisa: Reasoning segmentation via large language model

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.458987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:25:43.350210Z digest=sha256:dd25835ea1686a5ba340a724bec920ed264a224c08c2511c39894f9aed257880

Observation 95ca5c2d-c1c2-4d79-807c-bc76acb402a1 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion LLaVA-OneVision: Easy Visual Task Transfer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:43.465321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:43.465321Z digest=sha256:4ff59caba8681cf1e6458f076714d5d940909f5de650d44b1deb17cf4a1907d1

Observation 0a9ca5e0-6b27-41a1-bd80-8c35644b54c7 · outbound

This paper cites Dn-detr: Accelerate detr training by introducing query denoising.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Dn-detr: Accelerate detr training by introducing query denoising

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.443299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:25:43.514620Z digest=sha256:71fbe27b55ae3b1e0bc52e0fa4bef7bf05e99fd2a72cb509c38c5cf98a72b76b

Observation a9ccb8e4-1767-4134-806f-97bc6c3f423b · outbound

This paper cites Coda: A real-world road corner case dataset for object detection in autonomous driving.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Coda: A real-world road corner case dataset for object detection in autonomous driving

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.426559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:25:43.613826Z digest=sha256:080451799460c79842630638924d519c08e3bed2faf67bcd96f05280ed233f16

Observation 1767c512-b56c-4a29-8389-15275d85d357 · outbound

This paper cites Desco: Learning object recognition with rich language descriptions.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Desco: Learning object recognition with rich language descriptions

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.409036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:25:43.758188Z digest=sha256:92565411d1eb093617b40c3ebf92a52c66ffc4f92b0619c4fb2c3dc9ec03a497

Observation 6188faa5-3e31-4a7e-a0f4-045ff99fa574 · outbound

This paper cites Grounded language-image pre-training.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Grounded language-image pre-training

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.393433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:25:43.847663Z digest=sha256:428b2c588b9d261f40c822f48d2bcaf0240474f9784454d420fd19c11dddf648

Observation 1ea8ff82-22a0-4319-a530-fb6425cc897f · outbound

This paper cites Generative region-language pretraining for open-ended object detection.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Generative region-language pretraining for open-ended object detection

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.377953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:25:43.855977Z digest=sha256:79750e9f367377fe6892d8f978c1194040564417d7dcf908e4801a008c445e6d

Observation b1cd2f1d-0e0a-4292-b2cf-c793afca37fb · outbound

This paper cites Training-free open-ended object detection and segmentation via attention as prompts.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Training-free open-ended object detection and segmentation via attention as prompts

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.361788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:25:43.862354Z digest=sha256:38bc7ec65a951261383d6b290743c0724300218eb881a122f1af840c2fa28046

Observation e459fbe4-2f7e-433e-943f-ce29b7572115 · outbound

This paper cites Visual instruction tuning.Neural Information Processing Systems (NeurIPS), 2023.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Visual instruction tuning.Neural Information Processing Systems (NeurIPS), 2023

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.344472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:25:43.867202Z digest=sha256:cfb23f7dab764507d8f1503f9dc012151aeb5d83cba662f19aa1e5b73ec8464b

Observation 65e03801-1900-45f7-a0fb-91aff4a1bc80 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.328450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:25:43.873335Z digest=sha256:da30e41183d67db7540c2c907ec9f3ed6979dfb62f74477d7911ad3073594204

Observation 66878a2e-4dc8-44c5-a84a-3a94ab7d4a43 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Swin transformer: Hierarchical vision transformer using shifted windows

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.312770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:25:43.879076Z digest=sha256:ae3d9114b6b514e0c8abc82ac8258ecf7d90d2f9057133e354908300a36d630a

Observation 53e798e5-d546-4307-b90d-0fb5d8dc97ad · outbound

This paper cites Capdet: Unifying dense captioning and open-world detection pretraining.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Capdet: Unifying dense captioning and open-world detection pretraining

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.296731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:25:43.884458Z digest=sha256:6862424ee8804d22346b7aad14d0683b45400b4bb09610774428a7e3d254d7c4

Observation 86bf87f0-ecd7-4e52-9836-8ac65fd7ce75 · outbound

This paper cites Scaling open-vocabulary object detection.Neural Information Processing Systems (NeurIPS), 2023.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Scaling open-vocabulary object detection.Neural Information Processing Systems (NeurIPS), 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.281439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:25:43.889852Z digest=sha256:adc83ef4ecf8661a8c4168b472c0398cdbb13c60faf58df92de44e7fa170ced8

Observation d0e3482c-010f-4af8-8e77-014c190f77be · outbound

This paper cites Learning transferable visual models from natural language supervision.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Learning transferable visual models from natural language supervision

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:43.895759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:43.895759Z digest=sha256:85c0d82d41658a6b83043582c6fb7c0288ec8d35216c36a272ccb0240954ac8e

Observation 1cf861bd-aa92-4732-b410-788eb9ba0f3d · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Faster r-cnn: Towards real-time object detection with region proposal networks

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.256381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:25:43.901395Z digest=sha256:fbcf9a5591759451816e20ff73ecaccbef9d8540f6c5d14b668cbfb04aee9820

Observation 952f7a13-222b-4092-b791-fa62baf447d8 · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:43.907816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:43.907816Z digest=sha256:dea04389452ccafe23a4621729ad8e8acfe698388e3060cd326c2044da2c5d08

Observation bc5ded50-7edf-402d-a22f-6193f73cc3b1 · outbound

This paper cites Scalability in perception for autonomous driving: Waymo open dataset.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Scalability in perception for autonomous driving: Waymo open dataset

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.239758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:25:43.914161Z digest=sha256:85223553ce8e90d37c713ba2342f06224ce7322690f7f475651f76acdce8f463

Observation 46e99aa6-694f-4db8-ab1f-aac1b4d565c3 · outbound

This paper cites OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:43.919389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:43.919389Z digest=sha256:ab2ddb6a8caf89fb0bfd8b9f1c0bef1ac2458c29d0db72af8250cb1ad2973e92

Observation c90909b2-a67e-4f66-8a7d-cde1ca2d1edb · outbound

This paper cites Cogvlm: Visual expert for pretrained language models.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Cogvlm: Visual expert for pretrained language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.223244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:25:43.926211Z digest=sha256:51f2bbdf192e45950846d617154b15a8fb1e84ed27c70006bc78c3edd081aab1

Observation f802fc6a-8f2c-467f-998a-c86d33c79d40 · outbound

This paper cites Detclipv2: Scalable open-vocabulary object detection pre-training via word-region alignment.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Detclipv2: Scalable open-vocabulary object detection pre-training via word-region alignment

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.207117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:25:43.931973Z digest=sha256:ff2f85fab33f60c49642b86c41fbd4adeab32cbdce8253bccceac41c87525bae

Observation a048e2b5-a136-4b19-ae26-7e841210f4b8 · outbound

This paper cites Detclip: Dictionary-enriched visual-concept paralleled pre-training for open-world detection.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Detclip: Dictionary-enriched visual-concept paralleled pre-training for open-world detection

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.188398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:25:43.937160Z digest=sha256:4935ffeeaaca66fc20869fc0d2f7c74cedd079b250db6cfb2b1670b32ad8ca6d

Observation 0cfd0a67-5e40-40cb-ad5a-0454b65051c3 · outbound

This paper cites Detclipv3: Towards versatile generative open-vocabulary object detection.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Detclipv3: Towards versatile generative open-vocabulary object detection

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.171851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:25:43.942085Z digest=sha256:df8e2e13bce864cc64f10154b8b72254f58cb8a7d430c68fffbcb7fffe6d5e58

Observation e33a2214-f781-4e90-92e1-5313eccf1022 · outbound

This paper cites Ni, and Heung-Yeung Shum.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Ni, and Heung-Yeung Shum

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.155459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:25:43.947653Z digest=sha256:ef3e5994146eb6f15d65ca890192c492e20d0a1d7901445bfef1f6d1db509cd6

Observation bf19aa99-7d41-4b33-830f-9a967b0586c9 · outbound

This paper cites Llava-grounding: Grounded visual chat with large multimodal models.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Llava-grounding: Grounded visual chat with large multimodal models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.140375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:25:43.953447Z digest=sha256:d35d8f7851a2dc580351c50f489270a645110bd484e8035a17245e5680ea5f34

Observation a910885d-f952-4053-a270-93fbdc12f077 · outbound

This paper cites Glipv2: Unifying localization and vision-language understanding.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Glipv2: Unifying localization and vision-language understanding

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.123477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:25:43.959038Z digest=sha256:e5c121b92c7cb7ad9947d99c980ca69e9449f19edbad81444f197249e2010530

Pith citing papers

Observation f945f4ea-7141-4be0-a3e6-3f2553c5d7f9 · inbound

Visual Prompt Based Reasoning for Offroad Mapping using Multimodal LLMs cites this paper.

Visual Prompt Based Reasoning for Offroad Mapping using Multimodal LLMs VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:25:51.983415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T19:51:59.102471Z digest=sha256:5980185d8a5fc82c5578b372331938bd3c7f5f46bdc077512a5fd5846594d1fe

Observation fdd5e044-6917-4c84-b422-71aaa53799e8 · inbound

VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection cites this paper.

VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:30.268838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T01:24:56.236650Z digest=sha256:8aa5f92b24e19023f48ad409375f8c3784bbf04fbf5a577d2e0af27ca39b982d

Observation d0f78be8-c593-4474-959b-3a09d910a640 · inbound

VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection cites this paper.

VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:00:35.844779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T19:14:12.092478Z digest=sha256:71fc7d06cd4f17230a38a36971fe89ab1ce7eef997ced7551b721ae94c5925b3

Observation 36545036-448a-4b16-a2fa-2703d6456869 · inbound

VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection cites this paper.

VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:51:48.393215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T01:33:55.317107Z digest=sha256:5c4734007a546390ae14966e75dead7968ece49d9f7079994b139133cb30b493

Observation 0e043359-7b35-4c85-b811-1538a58c95d3 · inbound

FoodCHA: Multi-Modal LLM Agent for Fine-Grained Food Analysis cites this paper.

FoodCHA: Multi-Modal LLM Agent for Fine-Grained Food Analysis VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:21:09.469192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T16:11:26.104498Z digest=sha256:c6c29e2e85ca0f4800474912613093638d24210af8a9758aeae88241f94a4fb0

Observation 3bae7f73-e0ef-4b89-8e37-602e64d41309 · inbound

3DVLA: Enhancing Vision-Language-Action Models via 3D Spatial and Instance Understanding cites this paper.

3DVLA: Enhancing Vision-Language-Action Models via 3D Spatial and Instance Understanding VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:13:16.671658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T07:07:59.925795Z digest=sha256:6356b19775f2350dad92a29569a0cac2a75c88dc37dba603016cc4a6b8c41766

Observation 0aeb88e6-1235-4112-9af7-b80b89eb9aba · inbound

DrivingAgent: Design and Scheduling Agents for Autonomous Driving Systems cites this paper.

DrivingAgent: Design and Scheduling Agents for Autonomous Driving Systems VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:17:58.077080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T10:05:38.910257Z digest=sha256:cb97c172a18face941e39759d8e2d2801ad3450de1e8fa4ed18993c6e2df3bec