Pith. sign in

Paper Citation Record · LEDGER

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding

As of 20 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2508.20758.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.20758 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:52:56.660665Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy50
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 90c3f9ba-03b2-496a-8cee-cf91dfab34da · outbound

This paper cites Referit3d: Neural listeners for fine-grained 3d object iden- tification in real-world scenes.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Referit3d: Neural listeners for fine-grained 3d object iden- tification in real-world scenes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.935911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.176174Z digest=sha256:c102800ea4a6c1f1d351879212bda2dd170e1105fcc743c114d6d741238aa9f1

Observation 773e90b1-a59f-479b-be30-16b83348fe39 · outbound

This paper cites Tongyi Qianwen.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Tongyi Qianwen

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.919786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.183038Z digest=sha256:5feb721db3b5c5c5e837c79443c5caeea72136739dc618fa5e56c2e82a6d0086

Observation 3f339379-bf52-4e75-9d59-6582480d14ba · outbound

This paper cites VolcEngine.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding VolcEngine

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.903152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.192047Z digest=sha256:3585432fde1129a329082fea523e6c7bfa791d2cad37d708dd4e294b100633f1

Observation 8cc8161b-a39b-469d-befc-5ac3e461c955 · outbound

This paper cites 3djcg: A uni- fied framework for joint dense captioning and visual grounding on 3d point clouds.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding 3djcg: A uni- fied framework for joint dense captioning and visual grounding on 3d point clouds

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.887017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.201622Z digest=sha256:16057a6bbe5f0bb4aa65d3753c8807691783b92ae830f79478d2e2b159a2e1b4

Observation 86c4682c-3d64-415e-b135-20face92b370 · outbound

This paper cites Tparn: A network for enhancing synthetic video quality after 3d-hevc encoding.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Tparn: A network for enhancing synthetic video quality after 3d-hevc encoding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.871506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.210128Z digest=sha256:8903a338d639c1371c206c3834126604d9a368857a01e9b4311ef098b32cfd21

Observation 513346dc-3a1e-4b86-bc93-f09c332b1ce6 · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Scanrefer: 3d object localization in rgb-d scans using natural language

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.855923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.216472Z digest=sha256:5a21737580eec105c2be13fc358dd3f32b285a6199e7029620f35f55fdb4d2bc

Observation d6b8fac5-0394-460e-b024-504529bd8f8c · outbound

This paper cites D 3 net: A unified speaker-listener architecture for 3d dense captioning and visual grounding.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding D 3 net: A unified speaker-listener architecture for 3d dense captioning and visual grounding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.839539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.226555Z digest=sha256:83225da75e8c4066e2f165e69ffc555b7b709ea1492ea7cfcb70a78e35ec27ac

Observation 1ce807c2-d9da-4452-939b-a032805c84fb · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T14:52:56.238082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:52:56.238082Z digest=sha256:ecedd621c4ca81174f3ec9ee41c89d93bdc95a1c9309541b39f7158fb7328262

Observation cf4edee0-1018-48c8-b435-e07501aabaef · outbound

This paper cites Reasoning3D -- Grounding and Reasoning in 3D: Fine-Grained Zero-Shot Open-Vocabulary 3D Reasoning Part Segmentation via Large Vision-Language Models.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Reasoning3D -- Grounding and Reasoning in 3D: Fine-Grained Zero-Shot Open-Vocabulary 3D Reasoning Part Segmentation via Large Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T14:52:56.245767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:52:56.245767Z digest=sha256:4e96878bf1665cb606b1a34076b4b3e5cca1a5d7ba6d27c07a87bb1b51f40006

Observation 8fac060a-2fe1-4f9d-9273-d93698461498 · outbound

This paper cites Objaverse: A universe of annotated 3d objects.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Objaverse: A universe of annotated 3d objects

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T14:52:56.252592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:52:56.252592Z digest=sha256:94f5058dcf073f59d6ca404e5ef6b317e27e56e8d97e1b6738853191505b35fc

Observation e0a5a4f8-7eea-45f5-856a-76641dbd3b81 · outbound

This paper cites Cg-slam: Efficient dense rgb-d slam in a consistent uncertainty-aware 3d gaussian field.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Cg-slam: Efficient dense rgb-d slam in a consistent uncertainty-aware 3d gaussian field

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.812375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.262404Z digest=sha256:eb42e4f8afa41fc4a63a89d87d654c80b7845cc2ff012a1195045d9fb37582c0

Observation 01d8a546-520b-47d5-8182-d29766833499 · outbound

This paper cites Photo-slam: Real- time simultaneous localization and photorealistic mapping for monocular stereo and rgb-d cameras.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Photo-slam: Real- time simultaneous localization and photorealistic mapping for monocular stereo and rgb-d cameras

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.796549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.267935Z digest=sha256:32080207268a812e1a1eabe9c4377d9451ef86337b97e1f56a8bab3c0459b3d5

Observation b164ceec-9a8b-452a-94bf-a3b789245cd6 · outbound

This paper cites Reason3d: Searching and reasoning 3d segmentation via large language model.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Reason3d: Searching and reasoning 3d segmentation via large language model

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.780828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.276702Z digest=sha256:51443696128d41b6a514f45d5ed208baa61c79b85dc108e3c582a9b40fc05ae6

Observation 388e6c81-4929-4e25-b219-5d4a05c7d513 · outbound

This paper cites Text- guided graph neural networks for referring 3d instance segmentation.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Text- guided graph neural networks for referring 3d instance segmentation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.765227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.283603Z digest=sha256:55814333a8fce3dce555cc25dc29d634012afe4328d3156f8016d277bd19697f

Observation 9a9d2c2b-6e7b-405c-a528-531f08b8aaae · outbound

This paper cites Training an open-vocabulary monocular 3d detection model without 3d data.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Training an open-vocabulary monocular 3d detection model without 3d data

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.749130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.289237Z digest=sha256:112f19d42316c0f26027a1323406ff5a01cbb6d7bb76fb8aac7a3a8548aa0c3b

Observation 93b08520-f1ef-42c7-bf64-9bc301848b52 · outbound

This paper cites Multi-view transformer for 3d visual grounding.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Multi-view transformer for 3d visual grounding

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.733468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.294597Z digest=sha256:2574b177e3ce435b32adc663fc539f3a1d94a9315ebb4c0607754c0ce9fa00ff

Observation 5a956406-367b-4218-922a-06e85ec48f78 · outbound

This paper cites Joint semi- supervised and active learning via 3d consistency for 3d object detection.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Joint semi- supervised and active learning via 3d consistency for 3d object detection

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.718020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.301732Z digest=sha256:dd8bc264e2e9f190522105ec4b000318240a19f9888e19a155c4eb9ece2cf286

Observation 66ee71c3-0d79-44d5-9655-744352a45a60 · outbound

This paper cites Bottom up top down detection transformers for language grounding in images and point clouds.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Bottom up top down detection transformers for language grounding in images and point clouds

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.702405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.310233Z digest=sha256:ac1cc106510c4b98ba16cb3794f65447642217b977015cfdae1838d535c50828

Observation 1d0621b7-373a-45ab-8f6c-910a28a78aab · outbound

This paper cites Cplip: zero-shot learning for histopathology with comprehensive vision-language alignment.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Cplip: zero-shot learning for histopathology with comprehensive vision-language alignment

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.686013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.317105Z digest=sha256:6a5cfcba557581e45c3a6d0f5d13f4a580ae93376a9a6943bb89c2c29eb7207a

Observation a0e8b316-c33a-4b98-9152-b00269cf095c · outbound

This paper cites Pointgroup: Dual-set point grouping for 3d instance segmentation.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Pointgroup: Dual-set point grouping for 3d instance segmentation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.669400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.325097Z digest=sha256:7ddb7e440b3d2a45b5194b684ec664a155bb321e3dad9ea12003299e8c16b5a1

Observation 80e8716c-c656-4190-9017-a5075da58bf1 · outbound

This paper cites Lerf: Language embedded radiance fields.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Lerf: Language embedded radiance fields

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.652889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.332470Z digest=sha256:57c1e7d7e411be9e55821497ddaf9ff752b92e83e5b40b84ee3faff4102a7c9a

Observation 12cc7b18-6bdd-4cfc-b193-6bec82c468e3 · outbound

This paper cites Rgb-d based visual slam algorithm for indoor crowd environment.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Rgb-d based visual slam algorithm for indoor crowd environment

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.633220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.340395Z digest=sha256:8c6db9bae75ed3233e7a816419aea82f90bcef385f89c237dad764768b61b21c

Observation 08215bc9-9815-481f-b791-de8ca74e229a · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.614068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.346705Z digest=sha256:10cfb51ba36d68f7ffd43264c01709a2cc0a1a4cb35220f021b96617998d9579

Observation 0b7dc8af-52a2-479e-9600-cd3c2817047a · outbound

This paper cites SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T14:52:56.353243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:52:56.353243Z digest=sha256:8b006ef9c6bb74c2e13b8a6eab42070948fb4e2c4b2eb07b15256d489e0f0645

Observation 3cb6d9ba-c4cf-4537-b51f-7763ddfb548f · outbound

This paper cites Exploring diversity-based active learning for 3d object detection in autonomous driving.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Exploring diversity-based active learning for 3d object detection in autonomous driving

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.596404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.360307Z digest=sha256:2b8dec767e11fb3d54d16aa4ed86d6edbe608edc66cb6f518fb0fb9851395070

Observation 5b63f1d8-e305-40f4-b93b-80656a69654b · outbound

This paper cites Unimel: A unified framework for multimodal entity linking with large language models.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Unimel: A unified framework for multimodal entity linking with large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.575938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.366605Z digest=sha256:ef0690c1cdf774ebaa19cb5f1c82b163f520b0b697a7463e9793b9537b7837a3

Observation 59fa7543-c01d-489a-99ce-ea1df05d9300 · outbound

This paper cites Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T14:52:56.374686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:52:56.374686Z digest=sha256:ced1a759768c80861f8c6849be3541e18005c485df3fb23eb890d85dec2f7eb9

Observation a474cd0f-8075-418d-a653-d005a19d3c64 · outbound

This paper cites Group-free 3d object detection via transformers.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Group-free 3d object detection via transformers

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.551556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.382798Z digest=sha256:f89b4bc7f16842c0d423148f55c896e848d5e57453ccac31b3ad56800703126d

Observation 2894a933-80c8-4244-9a30-d2ddea43f024 · outbound

This paper cites 3d-sps: Single-stage 3d visual grounding via referred point progressive selection.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding 3d-sps: Single-stage 3d visual grounding via referred point progressive selection

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.534460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.389909Z digest=sha256:ba6ba79b77ce185dc1a4e72ebbc2a0a429530a75c71486386be9bb45ac5b6120

Observation d1396919-487d-4234-9a7f-f414182eb64c · outbound

This paper cites Motion detection methods applied on rgb-d images for vehicle classification on the edge computing.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Motion detection methods applied on rgb-d images for vehicle classification on the edge computing

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.515170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.397368Z digest=sha256:c22d3fcea3838a6680a3452b1c5e53543b71d9fd245bc156f7582ac46042bb59

Observation a28df879-ae4f-4b19-aa2d-11a8b414b41b · outbound

This paper cites an unresolved cited work.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:52:57.496701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.403558Z digest=sha256:783600181891b7cbc800abd07ffb0e114a0bb9dd4a6e612f2d1deb104deed729

Observation da167163-0f6b-405c-9577-a43856724288 · outbound

This paper cites Openscene: 3d scene understanding with open vo- cabularies.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Openscene: 3d scene understanding with open vo- cabularies

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.478019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.411890Z digest=sha256:4507cd1d05a70421124bd9926b1eb953709f4e85cacf689170313005649ac8a2

Observation 2dc682c5-8b88-4f47-8743-98ec5722b1ad · outbound

This paper cites Zeetad: Adapting pretrained vision-language model for zero-shot end-to-end temporal action detection.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Zeetad: Adapting pretrained vision-language model for zero-shot end-to-end temporal action detection

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.458714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.424846Z digest=sha256:73a4388ba48c924cd040229d498ac9e2e85876649a9d76a7cd71a39a25dc28df

Observation ab0d0912-7183-4a0a-a9b1-8131e83bb483 · outbound

This paper cites Multi-branch collaborative learning network for 3d visual grounding.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Multi-branch collaborative learning network for 3d visual grounding

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.438240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.435892Z digest=sha256:165128ce1913dd9ed57160f085500df65ab22a06926b43815c6ab4fc0e64282b

Observation fcfbc55b-b25a-424c-ac3b-737ad5591455 · outbound

This paper cites Rgb guided tof imaging system: a survey of deep learning-based methods.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Rgb guided tof imaging system: a survey of deep learning-based methods

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.421056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.443723Z digest=sha256:8bac9c4060d3b6bdad8779d842da42c49400e2abdb59c52bfbfebb7da4fe05fa

Observation 17375a2c-5ed0-4412-a595-9e8541c2f4f0 · outbound

This paper cites Learning transferable visual models from natural language supervision.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Learning transferable visual models from natural language supervision

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.401319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.450536Z digest=sha256:bf562c04e9529ab94db54c48f40a1ad8aa1fac66fb3bb7d7b3c3dd555edf7327

Observation cfc997f9-2134-46ad-a9b9-583463092bc0 · outbound

This paper cites Mask3d: Mask transformer for 3d semantic instance segmentation.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Mask3d: Mask transformer for 3d semantic instance segmentation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.382848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.459017Z digest=sha256:9d630aabb20c8737167ea0c93f6280e1e7b24327c8f3cde82334fc0a85e7e980

Observation 51a7250b-54af-4979-8199-01bf8f27d8f6 · outbound

This paper cites Layer depth denoising and completion for structured-light rgb-d cameras.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Layer depth denoising and completion for structured-light rgb-d cameras

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.366219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.469178Z digest=sha256:1ee702ad768d7683b5b401e1ffd02bbb6e4bbba9b51e42d6130fa4f7cc86c4f5

Observation 75c53b68-3a8c-4847-a567-629750cd4bc3 · outbound

This paper cites FlowCam: Training Generalizable 3D Radiance Fields without Camera Poses via Pixel-Aligned Scene Flow.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding FlowCam: Training Generalizable 3D Radiance Fields without Camera Poses via Pixel-Aligned Scene Flow

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T14:52:56.477553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:52:56.477553Z digest=sha256:47be602b32a4bdb4418b5e5ab96db8d0178a730edf0576f1162d8e18eda0bebb

Observation d2284615-e8ac-48e2-ba2b-a20a4d0bcd39 · outbound

This paper cites Four ways to improve verbo-visual fusion for dense 3d visual grounding.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Four ways to improve verbo-visual fusion for dense 3d visual grounding

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.345561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.486758Z digest=sha256:31b83fa587be10198169fb30305720b4fdfd50ac9a3c13931a5d55d0414e161f

Observation bb55a291-2426-4ee4-b044-f93c6bddcf0a · outbound

This paper cites Soft- group for 3d instance segmentation on point clouds.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Soft- group for 3d instance segmentation on point clouds

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.325286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.501498Z digest=sha256:0a4e2e1e89ce34673fe176978a3688a7dc5ffdb39614039f1f726c3a9bc69a77

Observation c38bbdc2-f353-45f8-9313-bf00ce49c775 · outbound

This paper cites A survey on large language model based autonomous agents.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding A survey on large language model based autonomous agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T14:52:56.508265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:52:56.508265Z digest=sha256:a5f63e4461e137562259c01bb692b62ebbb43c1eca7149515164e03ae15af52a

Observation 2c8136ec-b0cc-408b-98ee-2c570c9c52f3 · outbound

This paper cites Gˆ 3-lq: Marrying hyperbolic alignment with explicit semantic-geometric modeling for 3d visual grounding.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Gˆ 3-lq: Marrying hyperbolic alignment with explicit semantic-geometric modeling for 3d visual grounding

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.286237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.514240Z digest=sha256:ee79b16ca7e41fb30adbac6d2ffcd8684a0b22d1a4f3b1dbdb595e52035f7eaf

Observation 017e607c-924a-4e8a-aa8f-be3fa463b323 · outbound

This paper cites Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T14:52:56.520945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:52:56.520945Z digest=sha256:2c9976ae7efa0c26e97c5f0735befe54c14e12affd695e2608e5cdf0a2582460

Observation 82ece06d-3727-4804-8444-c8647dbc7ae1 · outbound

This paper cites Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.262054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.530843Z digest=sha256:0d8cf75b560108a05f5990163080c9d29981c0f097eb3797925db8cad76da460

Observation 30dee432-b79b-4c08-89ac-51d3e402f5f1 · outbound

This paper cites Eda: Explicit text-decoupling and dense alignment for 3d visual grounding.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Eda: Explicit text-decoupling and dense alignment for 3d visual grounding

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.236688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.540001Z digest=sha256:36facad9a7180df31e0c7579779e17d528f8878a15618a1cbe2febe4ee6c6582

Observation 22945b12-b07a-4909-bd26-67a07473dffc · outbound

This paper cites M-divo: Multiple tof rgb-d cameras enhanced depth-inertial-visual odometry.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding M-divo: Multiple tof rgb-d cameras enhanced depth-inertial-visual odometry

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.209024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.555535Z digest=sha256:85c789961889721c8187c078733ad0b82e9f1b275a5b615d13a6f401121554ff

Observation e338d33b-6e96-42b4-afde-065288427dff · outbound

This paper cites VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T14:52:56.561970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:52:56.561970Z digest=sha256:094193c3516a5f1926acbf3bc939b7a716e91a39963201f3c82a6014e66d31e4

Observation 503a2737-6be6-4e3d-9999-2bc2b96f1b6c · outbound

This paper cites Multi-scale 3d gaussian splatting for anti-aliased rendering.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Multi-scale 3d gaussian splatting for anti-aliased rendering

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.187657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.569700Z digest=sha256:5dc4a9aba026d4b5f11ea15cb4f48eb8d67643be090b39b495b6bde79b01bd93

Observation 41fc1ef3-9624-4cc4-a24c-dae122c638b4 · outbound

This paper cites Llm-grounder: Open-vocabulary 3d visual grounding with large language model as an agent.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Llm-grounder: Open-vocabulary 3d visual grounding with large language model as an agent

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.159395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.576596Z digest=sha256:1bd354856baf8dab6eb1a60bd2d331139b8968f7d7a37b0bb95c695446fcc781

Observation 38e77606-d363-486c-b4f1-dc34bab4e660 · outbound

This paper cites Sat: 2d seman- tics assisted training for 3d visual grounding.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Sat: 2d seman- tics assisted training for 3d visual grounding

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.137228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.586265Z digest=sha256:277f3e7c73a9065f8e6c5153d61f90051132aa4beaa5f99db7760b0a615cf85b

Observation 0b1a98e0-f57e-4432-89bd-2e6ea9d60370 · outbound

This paper cites Visual programming for zero-shot open-vocabulary 3d visual grounding.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Visual programming for zero-shot open-vocabulary 3d visual grounding

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.112222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.592666Z digest=sha256:b07e5ee871ea0e8dd04036d62594385dbcc76e7a61e1bf0908e06eb8cbb7f98a

Observation 8c33f02d-03c2-4d0d-9ab1-80b6b59feb57 · outbound

This paper cites Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual referring.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual referring

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.087980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.599200Z digest=sha256:f7a4015f0fce095579f8acc9689157dac58e1fe5805586426df55436f2e5d723

Observation aa535d2a-e145-4f14-b93e-0b2ed1a7e55e · outbound

This paper cites Vision- language pre-training with object contrastive learning for 3d scene understanding.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Vision- language pre-training with object contrastive learning for 3d scene understanding

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.061599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.604221Z digest=sha256:42e47307d404749f2bfcd04e73f8c2ace9e16abb402a6b9acb9ceee547246096

Observation 2fe3438f-cdb9-40e8-aca6-498ea0278de2 · outbound

This paper cites Earthgpt: A universal multi-modal large language model for multi-sensor image compre- hension in remote sensing domain.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Earthgpt: A universal multi-modal large language model for multi-sensor image compre- hension in remote sensing domain

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.031356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.610673Z digest=sha256:2d0fa0143cf453a099767abcdc59c1295efc3ca33284b0ebf76201e1aa6edeab

Observation 0b728035-ee44-4732-9c1a-821ec498204b · outbound

This paper cites Prototype correlation matching and class-relation reasoning for few-shot medical image segmentation.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Prototype correlation matching and class-relation reasoning for few-shot medical image segmentation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:57.008060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.618028Z digest=sha256:3bdd825bfdbc765ad9d482a72374d7ad6411a3902c802bf49b72d6e9fb7be27d

Observation b9b082a4-124e-49cc-9ee4-fb7012c521dd · outbound

This paper cites Towards clip-driven language-free 3d visual grounding via 2d-3d relational enhancement and consistency.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Towards clip-driven language-free 3d visual grounding via 2d-3d relational enhancement and consistency

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:56.988486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.624278Z digest=sha256:baf2ff39ba0aa40f9b43b3acb2f3049363418a2dc5cfeec94eca416e7a2b1b64

Observation b13345a0-b608-441d-9db4-79cd25ef9a13 · outbound

This paper cites 3dvg-transformer: Relation modeling for visual grounding on point clouds.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding 3dvg-transformer: Relation modeling for visual grounding on point clouds

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:56.967903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.632131Z digest=sha256:8a5369bd3603fb11c0564e4a5675a05f634b5d8b5867012c5e6622fe903b7a14

Observation 40c8c8cd-cadc-47ae-af50-4b537eae3337 · outbound

This paper cites A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T14:52:56.639839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:52:56.639839Z digest=sha256:7b0f47df1c98d8f0eab38e7ad0c1146b9434813e73fb59b8be9559e0b60cee02

Observation 792e4048-406c-456f-9dcf-b83522ea03be · outbound

This paper cites Preventing zero-shot transfer degradation in continual learning of vision- language models.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Preventing zero-shot transfer degradation in continual learning of vision- language models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:56.946876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.651430Z digest=sha256:ffe50bc0374a2b21e3ab33570ef5e01becdf9c5a88243a1c85f9ee531aa5dd47

Observation 10c7c7cc-b94a-4f83-bf7c-0cd4cffb383e · outbound

This paper cites Context-aware 3d object detection from a single image in autonomous driving.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding Context-aware 3d object detection from a single image in autonomous driving

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:52:56.923793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:52:56.660665Z digest=sha256:f0448f5a011650c9671f54162eb59ab66ed5a2842d67e1e5a114af8028868db2

Pith citing papers

No inbound Pith citation observations are available.