Pith. sign in

Paper Citation Record · LEDGER

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding

As of 13 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 7 inbound Pith citation observations for arXiv:2412.04383.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04383 v2

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:29:49.966156Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:14:41.667536Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T08:51:18.441484Z

Reference resolution

72 of 72 outbound references displayed

  • verified exact0
  • verified fuzzy51
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f6cb22ce-43ef-47e8-9ae6-50e6f8b49193 · outbound

This paper cites Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:51.586364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.531960Z digest=sha256:754e4e6ce58cb149d1f1948417ce4f4e93f5468052d4d82b215048105496ddfb

Observation e39fb830-7871-41d2-a4de-a07e3e9af0bf · outbound

This paper cites Look around and refer: 2d synthetic semantics knowledge distillation for 3d visual grounding.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Look around and refer: 2d synthetic semantics knowledge distillation for 3d visual grounding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:51.568113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.538085Z digest=sha256:23695b4f6bb449748eafc33eee448b966d98f4efa31b00056fc3d9832e6b1eae

Observation 112e1648-006b-4229-a063-8415503d7cd7 · outbound

This paper cites Se- mantickitti: A dataset for semantic scene understanding of lidar sequences.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Se- mantickitti: A dataset for semantic scene understanding of lidar sequences

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:51.548363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.545048Z digest=sha256:de6463f2970da5893f7b1261695b5af95b93e667df7aee6b2655b574136f531b

Observation 09d00982-2c1a-4c22-bb0d-b62e8136da9d · outbound

This paper cites Mikasa: Multi-key-anchor & scene-aware transformer for 3d visual grounding.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Mikasa: Multi-key-anchor & scene-aware transformer for 3d visual grounding

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:51.529904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.550779Z digest=sha256:bd28b493213ff7a158974348fa29e81c5508c42dab6f9467aea62c5cd441b3eb

Observation b119eb19-c988-4ca9-aad3-4c57632b50fd · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Scanrefer: 3d object localization in rgb-d scans using natural language

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:49.556206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:49.556206Z digest=sha256:6e3bceaa74d8940fd5dbe27ee4375f01ba21b40e43f94580eeedcb7821062a11

Observation 10a4e4e5-ac24-4d80-9baa-86e6afdcb3a2 · outbound

This paper cites Clip2scene: Towards label-efficient 3d scene understanding by clip.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Clip2scene: Towards label-efficient 3d scene understanding by clip

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:51.497703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.561624Z digest=sha256:6b140f20396364a546c611f96256b662bd8db269d6d7649748bd3e21f4b22085

Observation bda70d72-1882-4358-900e-ff6c7d1a080a · outbound

This paper cites Language conditioned spatial relation reasoning for 3d object grounding.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Language conditioned spatial relation reasoning for 3d object grounding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:51.481767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.567571Z digest=sha256:15848c289c745dabaa7a2246dc0cf65aabcd64bf7103b7b38d1fd4b8b878ba06

Observation 34286f96-c6fc-49c0-9ebf-e438d56a29c9 · outbound

This paper cites Think global, act lo- cal: Dual-scale graph transformer for vision-and-language navigation.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Think global, act lo- cal: Dual-scale graph transformer for vision-and-language navigation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:51.461911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.573242Z digest=sha256:9c0904698a804a3f06409c0239f91e1433e3247f1a59db66c08a22d0203d77e1

Observation e420ad3c-339c-4bf9-8869-c75dcf53a576 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:49.578296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:49.578296Z digest=sha256:5decf392ed0b2f8e13400b8c9955fa151073f9c8f1ebe885fda7f57e55c3e8ce

Observation 8e2ba133-1be6-4810-ab37-5d40571d6053 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:49.584000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:49.584000Z digest=sha256:65af38847c06d34486ee8e2bd6d8401a889c5227101852f366a8952787369d68

Observation 88bcb889-53ae-4c07-a0b9-3b0cbd8eecb7 · outbound

This paper cites Panoptic nuscenes: A large-scale benchmark for lidar panoptic segmentation and tracking.IEEE Robotics and Au- tomation Letters, 7:3795–3802, 2022.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Panoptic nuscenes: A large-scale benchmark for lidar panoptic segmentation and tracking.IEEE Robotics and Au- tomation Letters, 7:3795–3802, 2022

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:51.443224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.590019Z digest=sha256:a2a1e746e4ff0105e9090f8c94622cd9092aaa6bf171d327570fb71439baddfd

Observation b64cd2db-9507-462a-8c4f-02fccf68a985 · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:49.594671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:49.594671Z digest=sha256:05d07fae75ba6407497728458c55a7a06b34d1411c4e880f4d3dc6340038594f

Observation e2e1fbbb-040f-4e9e-addb-3047ac4687d1 · outbound

This paper cites From Cognition to Precognition: A Future-Aware Framework for Social Navigation.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding From Cognition to Precognition: A Future-Aware Framework for Social Navigation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:49.599738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:49.599738Z digest=sha256:4482b1faf4282d38dccf4e08149d41abc9d1306af1c94593b8b515cb427e9a2a

Observation 6dec190a-2741-4e82-98d0-eb3512d7922b · outbound

This paper cites Viewrefer: Grasp the multi-view knowledge for 3d visual grounding.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Viewrefer: Grasp the multi-view knowledge for 3d visual grounding

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:51.420579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.604964Z digest=sha256:f41ab7286e116961d74ddd7744c485e6a00995bcac2f3376435df01e634c7556

Observation 8bcb7c49-8e65-4f33-9d51-52465e746d80 · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding CogVLM2: Visual Language Models for Image and Video Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:49.610925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:49.610925Z digest=sha256:86b63fef511757c233a5c35088c5eab7517a8ba6e3b8cf261ef6a2f29d272765

Observation 6485a535-f2d4-443f-a26a-f8a4259e03f3 · outbound

This paper cites 3d-llm: Inject- ing the 3d world into large language models.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding 3d-llm: Inject- ing the 3d world into large language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:51.399495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.616497Z digest=sha256:6f016a15805d3d5b0d71b1ddaafc2524154f0c5dc2e71283b90c2308e9851271

Observation 2c8dd27c-8fd8-4572-bf9d-59c32111747e · outbound

This paper cites Dhp-mapping: A dense panoptic map- ping system with hierarchical world representation and label optimization techniques.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Dhp-mapping: A dense panoptic map- ping system with hierarchical world representation and label optimization techniques

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:51.380649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.621882Z digest=sha256:b49481b7b72c8cd83b0f6b6ec40178e9478e7cff55cd8a0bdd8909d6c4e9236b

Observation 399943dd-3528-4f6e-9deb-1c421107b5f6 · outbound

This paper cites Text-guided graph neural networks for re- ferring 3d instance segmentation.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Text-guided graph neural networks for re- ferring 3d instance segmentation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:51.354320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.628820Z digest=sha256:cdc8d31ab592eb81df1814e09957fd9ffabedeb4459de65e29fc4db4332bad80

Observation e1054693-672d-46dc-9416-e67e38b96040 · outbound

This paper cites Multi- view transformer for 3d visual grounding.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Multi- view transformer for 3d visual grounding

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:51.327934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.634610Z digest=sha256:5ac1fc25605110434a9e410e7e920306d3a3fdc82f9e4ee54b82b47920a9c03b

Observation 4ceb43ba-c2d3-4c6a-bf3f-06c80b12a1c2 · outbound

This paper cites Assister: As- sistive navigation via conditional instruction generation.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Assister: As- sistive navigation via conditional instruction generation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:51.291471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.640569Z digest=sha256:d32ac4fedb0836c05d14d2909ff3fd946ff259e7059c4dad009524e9fee553b9

Observation b15275ad-dfef-4465-a6e5-95f8daf82222 · outbound

This paper cites Openins3d: Snap and lookup for 3d open-vocabulary instance segmentation.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Openins3d: Snap and lookup for 3d open-vocabulary instance segmentation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:51.263743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.646124Z digest=sha256:7553cb13572fb74466df337670e9ed078fb3fae8486337110149c493c652eeb9

Observation 1fcf8d33-1cb7-46f8-82df-5138d8bec106 · outbound

This paper cites Bottom up top down detection transform- ers for language grounding in images and point clouds.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Bottom up top down detection transform- ers for language grounding in images and point clouds

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:51.239745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.651861Z digest=sha256:16b0fa297e5a60bf0022b9293954b6844d2c22c059dc7cb3d42697ed20237e3b

Observation 1db238a8-95a1-47ca-89bd-fd7281c598f4 · outbound

This paper cites Tenenbaum, Celso Miguel de Melo, Madhava Krishna, Liam Paull, Florian Shkurti, and Antonio Torralba.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Tenenbaum, Celso Miguel de Melo, Madhava Krishna, Liam Paull, Florian Shkurti, and Antonio Torralba

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:51.213955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.657884Z digest=sha256:80957e1e6d86da548a6eb0643e72a12c5ca4592308758ba12f93030194a74687

Observation bc673fdd-578e-43d6-bb3e-26bffbbe837b · outbound

This paper cites Sceneverse: Scaling 3d vision-language learning for grounded scene understanding.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Sceneverse: Scaling 3d vision-language learning for grounded scene understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:51.192061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.663975Z digest=sha256:fbf502163f8d877ca642436805d6d7a7af072a4afe169a2d40cea10c1e2338c2

Observation 8c6db71e-9c4e-4638-996b-df99fea13b19 · outbound

This paper cites Lerf: Language embedded radiance fields.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Lerf: Language embedded radiance fields

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:49.669627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:49.669627Z digest=sha256:f59c04f50042bba044a0c3af7ed6ef0eb487f49402352c43521e36afa26c98ce

Observation 89179189-6ce2-4572-907a-374fcf32d247 · outbound

This paper cites Rethinking range view representation for lidar segmentation.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Rethinking range view representation for lidar segmentation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:51.151696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.676270Z digest=sha256:2d958de053ec6a107cbb84d3e0d3d445c8a074de963d50c754215dbb8dbe4906

Observation 073a566a-f6d6-4ec5-b422-58d7c9c316f8 · outbound

This paper cites Robo3d: Towards robust and reliable 3d perception against corruptions.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Robo3d: Towards robust and reliable 3d perception against corruptions

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:51.130118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.683712Z digest=sha256:cd7c0a50cc923749f55a95e5c0248833a070fdfd7f83b99d8fff5b41c06a7ce5

Observation ac180bbb-69d1-4bdd-9e04-c9f174d84519 · outbound

This paper cites Xvo: Generalized visual odometry via cross-modal self-training.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Xvo: Generalized visual odometry via cross-modal self-training

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:51.107982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.690506Z digest=sha256:3f2bec8a67fe315b660edb0fb25e83aafd7d859087ad9e25cdc5cf94e2fc1a4d

Observation 029e69aa-d9a3-4452-9a27-815cbc6202cf · outbound

This paper cites COARSE3D: Class-Prototypes for Contrastive Learning in Weakly-Supervised 3D Point Cloud Segmentation.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding COARSE3D: Class-Prototypes for Contrastive Learning in Weakly-Supervised 3D Point Cloud Segmentation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:49.696425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:49.696425Z digest=sha256:87f300466753b0ddbc28c4f3bd1a1e0e9bf1b366779084ec3f710b940112f6a8

Observation 577f477f-1bb4-4236-a0fd-d8a96db0c541 · outbound

This paper cites Tfnet: Exploiting temporal cues for fast and accurate lidar semantic segmentation.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Tfnet: Exploiting temporal cues for fast and accurate lidar semantic segmentation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:51.087210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.703674Z digest=sha256:a1cfef7a1eb0c19638e2603f0779751bd766c1a2380390bd677d48094ea8a866

Observation 24df5088-9206-481c-80e8-b93a05f43531 · outbound

This paper cites Uni3DL: Unified Model for 3D and Language Understanding.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Uni3DL: Unified Model for 3D and Language Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:49.711318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:49.711318Z digest=sha256:7e85122007c90541b4b3fad69d8bf8e8083e5f636987775592a981879376cfb7

Observation ef4b1034-0e4f-43cc-9662-f2cf26bfa098 · outbound

This paper cites V oxformer: Sparse voxel transformer for camera- based 3d semantic scene completion.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding V oxformer: Sparse voxel transformer for camera- based 3d semantic scene completion

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:51.067940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.718528Z digest=sha256:b1261fbac72848cc879a3cdf50f9b417dacbfedd7c6fd442c89cd68018abd3f2

Observation b86cca78-eeb9-4c42-97d8-f0c70d79b459 · outbound

This paper cites Is your lidar placement optimized for 3d scene understanding? InAdvances in Neural Information Process- ing Systems, pages 34980–35017, 2024.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Is your lidar placement optimized for 3d scene understanding? InAdvances in Neural Information Process- ing Systems, pages 34980–35017, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:51.047542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.724390Z digest=sha256:a374d643bd57e3cf910266a46fb7545a56db60870a9ff34303cc1dd13b6d9e8f

Observation 7e8d4bb3-e85b-45ad-af75-c4b6c2083e88 · outbound

This paper cites Segment any point cloud sequences by distilling vision foundation models.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Segment any point cloud sequences by distilling vision foundation models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:51.027084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.729975Z digest=sha256:8ffe72abe5f9ae213772692f1ded7236b783722e54eb480162d7811b6679ae85

Observation dca5e6f6-7000-43e5-8a41-7a48fcae46b4 · outbound

This paper cites Deep view synthesis via self-consistent generative network.IEEE Transactions on Multimedia, 24: 451–465, 2021.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Deep view synthesis via self-consistent generative network.IEEE Transactions on Multimedia, 24: 451–465, 2021

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:51.004620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.735681Z digest=sha256:4719a1cb882a1f73e515a6146ab7aab9994414bc3254c9c912c3b948b8a4dcfc

Observation 995845fd-4c67-4fa6-b874-49c8e7a08424 · outbound

This paper cites RayDF: Neural Ray-surface Distance Fields with Multi-view Consistency.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding RayDF: Neural Ray-surface Distance Fields with Multi-view Consistency

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:49.740863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:49.740863Z digest=sha256:ddcdde47e1c120c13b0227f4e345e1673db0169b74d8e5858b8b7715c9f3731c

Observation b69498fd-b26a-48e4-89b7-ef9c5134c419 · outbound

This paper cites PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:49.747976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:49.747976Z digest=sha256:34917920d55d5814954efc464776c17c1f78244cd463fd7ed87422d7115db881

Observation 5a64f816-c135-4b69-91ae-4048d089e496 · outbound

This paper cites Ovir-3d: Open-vocabulary 3d in- stance retrieval without training on 3d data.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Ovir-3d: Open-vocabulary 3d in- stance retrieval without training on 3d data

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:50.984950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.752763Z digest=sha256:e3ed65cb594eae0e0d68875aa422ad1c4cc50c6ff504a833cffbfc998f36a7c1

Observation 3e40ada5-e8f2-47dc-99db-4c7892f0a542 · outbound

This paper cites An Examination of the Compositionality of Large Generative Vision-Language Models.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding An Examination of the Compositionality of Large Generative Vision-Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:49.756784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:49.756784Z digest=sha256:65f8d9082e45d779c4dc7000957f270b7f49d52838bd04dfedb759257a061ccf

Observation 70c56b15-cd52-48d3-88d1-198c23fb25e6 · outbound

This paper cites GLOVER: Generalizable Open-Vocabulary Affordance Reasoning for Task-Oriented Grasping.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding GLOVER: Generalizable Open-Vocabulary Affordance Reasoning for Task-Oriented Grasping

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:49.762251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:49.762251Z digest=sha256:b66996ebe3b9ab23a134f44562363079597bb65146dccf65f94ed53bdd670776

Observation f249905e-ee7b-46fb-b9e9-edb7e074e6d5 · outbound

This paper cites GPT-4 Technical Report.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding GPT-4 Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:49.768946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:49.768946Z digest=sha256:cc751e07dcd2d68d052c37c248eee6c541c32337a058cca51f4c6e5b18684607

Observation 8b0407df-001e-42d8-a7d4-cce5aa3fbb2a · outbound

This paper cites Training lan- guage models to follow instructions with human feedback.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Training lan- guage models to follow instructions with human feedback

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:50.966102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.774004Z digest=sha256:8ba01b98cb9c53855ec5dafd6de0db2267f8a953b946ac1b44c59c1b5bea24de

Observation dc8d5b4c-333c-4d69-8a9d-afced40b54c8 · outbound

This paper cites Openscene: 3d scene understanding with open vocabularies.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Openscene: 3d scene understanding with open vocabularies

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:50.949129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.780484Z digest=sha256:ce393ac7d77da2e305296a78ed874b18db9fcc72a6d14175c0c72f6c4955cc7f

Observation ddc87289-850e-4edb-b6c4-2046e81a1c71 · outbound

This paper cites Multi-branch collaborative learning network for 3d visual grounding.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Multi-branch collaborative learning network for 3d visual grounding

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:50.931562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.785532Z digest=sha256:c5689869a43c98c904599f25d942231f68ec7966cd1c4a8dbca0a7a0d41fa264

Observation d1f1d61a-8e16-429a-9b1b-0a49a261e84b · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Learn- ing transferable visual models from natural language super- vision

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:50.912617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.791050Z digest=sha256:6ffc1ab08c648080a242c3a955300be5a6758db829357a86ad8078545cc26705

Observation a5e725ec-b7d2-41d8-a8d1-95b114346831 · outbound

This paper cites Mask3d: Mask trans- former for 3d semantic instance segmentation.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Mask3d: Mask trans- former for 3d semantic instance segmentation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:50.896064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.797513Z digest=sha256:865973443642e919b5afaa16accd00853dab08d71ba094bb1b39d68d64d1a59c

Observation fdc2c7f6-6fbc-45e9-a74d-44eafa71d734 · outbound

This paper cites Interactive planning using large language models for partially observable robotic tasks.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Interactive planning using large language models for partially observable robotic tasks

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:50.872756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.803278Z digest=sha256:71d337ae4276024f14458242fe896b348371db39a9da93ba327644bbe8bf502d

Observation ef0ee074-2ce8-44f6-9c23-982ab81661a1 · outbound

This paper cites Scalability in perception 19 for autonomous driving: Waymo open dataset.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Scalability in perception 19 for autonomous driving: Waymo open dataset

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:50.850438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.809544Z digest=sha256:55029f7c970fca5279a9b45ec02f3f09d8611502cd60622c870673d694d14f3c

Observation 2f5cc807-e491-4958-8eba-e1721913540a · outbound

This paper cites OpenMask3D: Open-Vocabulary 3D Instance Segmentation.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding OpenMask3D: Open-Vocabulary 3D Instance Segmentation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:49.815015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:49.815015Z digest=sha256:9550cfc7777aae743b4a52e0aba8e8177f04ee2cdb163fb94758233213780ef5

Observation 27a9fa37-85f7-4a3b-821a-bad2a382ee38 · outbound

This paper cites Epmf: Efficient perception-aware multi-sensor fusion for 3d semantic seg- mentation.IEEE Transactions on Pattern Analysis and Ma- chine Intelligence, 2024.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Epmf: Efficient perception-aware multi-sensor fusion for 3d semantic seg- mentation.IEEE Transactions on Pattern Analysis and Ma- chine Intelligence, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:50.823335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.820669Z digest=sha256:ec67c986ce87795a8b392acfe081933092d981369046de46c94ae06e8a5655fc

Observation 18e0165f-6578-4549-9ccf-d210e8576550 · outbound

This paper cites Four ways to improve verbo-visual fusion for dense 3d visual grounding.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Four ways to improve verbo-visual fusion for dense 3d visual grounding

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:50.799762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.827197Z digest=sha256:190b676d0dff1e34beb7efbb2ea82fad7d75422be42737568bd6db7aa413b28b

Observation 164e63c6-291b-4b4f-8736-09cb20cc2964 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:49.833829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:49.833829Z digest=sha256:0b3afb0fcfacaed7274220621627f76bd3347df6f13360cd3dd72c0940381cae

Observation 5fe8600d-c2a5-4150-923b-27215381193f · outbound

This paper cites Embodiedscan: A holistic multi- modal 3d perception suite towards embodied ai.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Embodiedscan: A holistic multi- modal 3d perception suite towards embodied ai

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:50.783374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.841794Z digest=sha256:8e801a9adc28deb944ce74d1e34fdc8d05a142327bed0e5e79b10b2ddec921ca

Observation 70a23182-f22d-453b-86dc-baf0f9660616 · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:49.847278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:49.847278Z digest=sha256:4423fc6b58b46859d23bb480ba46e8d1fa57d5af2fa7d9f74b82ed6ded5c5704

Observation 8ae425a6-1b24-4fb0-a5e2-682d0a79038d · outbound

This paper cites G3-lq: Marry- ing hyperbolic alignment with explicit semantic-geometric modeling for 3d visual grounding.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding G3-lq: Marry- ing hyperbolic alignment with explicit semantic-geometric modeling for 3d visual grounding

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:50.763561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.852642Z digest=sha256:21472970b010e4da2b44b47eda8fea568288611cd9a28cac998ec9a1339107e5

Observation 3816464f-68de-4177-9dc8-b52b3cb27c68 · outbound

This paper cites Distill- ing coarse-to-fine semantic matching knowledge for weakly supervised 3d visual grounding.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Distill- ing coarse-to-fine semantic matching knowledge for weakly supervised 3d visual grounding

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:50.744508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.858032Z digest=sha256:39a4b4a39d8dcdcfd99f7cd6d17dcb380616bf57c7caa27d18b7026475ca1523

Observation af26b7e0-95e1-4112-9cd7-f4aa8947644c · outbound

This paper cites SIR: Multi-view Inverse Rendering with Decomposable Shadow Under Indoor Intense Lighting.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding SIR: Multi-view Inverse Rendering with Decomposable Shadow Under Indoor Intense Lighting

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:49.865801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:49.865801Z digest=sha256:42f36358edbe9c87efec66bb35060472465ee0b5589f1ced7eaf38966dae6dd6

Observation bcd9bb8c-972f-4ecd-bcc7-daf926042dcc · outbound

This paper cites Eda: Explicit text-decoupling and dense alignment for 3d visual grounding.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Eda: Explicit text-decoupling and dense alignment for 3d visual grounding

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:50.723479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.873689Z digest=sha256:efa2f504cb913463da50b135def160e8ee1305e175dd70b793df0ec8ee716116

Observation 0179f253-307d-416e-9eac-3bcda0740004 · outbound

This paper cites VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:49.880892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:49.880892Z digest=sha256:fa7a15c41ec18302e13aec4426d1f5bfff62b96d22cfb4651d083e6548a1cc5c

Observation 653a0a03-c806-4835-958d-e044656a4ebc · outbound

This paper cites 4d contrastive superflows are dense 3d representation learners.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding 4d contrastive superflows are dense 3d representation learners

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:50.705321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.887036Z digest=sha256:bafb8386491e024919597f12dfa6cc966075eccb63f7ef9fc6f332f4fb68a5e5

Observation 09cfb949-faa8-493f-8d28-9f92a25c54c3 · outbound

This paper cites Llm- grounder: Open-vocabulary 3d visual grounding with large language model as an agent.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Llm- grounder: Open-vocabulary 3d visual grounding with large language model as an agent

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:50.684245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.892373Z digest=sha256:54d5288115480994d45395f178b031a79347f436f43f8c524b42a0deadd5b3cd

Observation b48f51be-9a29-4f9c-ab31-2b01212e55ac · outbound

This paper cites Regionplc: Regional point-language contrastive learning for open-world 3d scene understanding.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Regionplc: Regional point-language contrastive learning for open-world 3d scene understanding

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:50.664081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.897530Z digest=sha256:4c6e010abe6fb42a04bec0013a37c4ffd06aba08484bca93be344473e9cad822

Observation e0aac898-37e0-4e83-af92-1859a8ab30c0 · outbound

This paper cites Sat: 2d semantics assisted training for 3d visual grounding.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Sat: 2d semantics assisted training for 3d visual grounding

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:50.644089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.903041Z digest=sha256:f11885657c2f8b35d884894a5f943611604e0c3995628955b1f2a86394f23297

Observation f6643c1a-2542-4fe0-b77b-93fb731b08d8 · outbound

This paper cites Sai3d: Segment any instance in 3d scenes.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Sai3d: Segment any instance in 3d scenes

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:50.625181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.911595Z digest=sha256:c6cf5b5ee7da8cce9e83fe75ae93a1d8f3006a5edac76ce57240c7bab348693f

Observation 9ada1862-c40f-4207-aff6-8449c7cb4d09 · outbound

This paper cites Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual refer- ring.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual refer- ring

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:50.601644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.918472Z digest=sha256:a4453414c14bc6c0799ff148d668b7b7469a7250af4bb4cd4235113107d22401

Observation e287afdb-037c-435b-9b32-7381a7bdcbd6 · outbound

This paper cites Visual programming for zero-shot open-vocabulary 3d visual grounding.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Visual programming for zero-shot open-vocabulary 3d visual grounding

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:50.580719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.924220Z digest=sha256:9caa58bf12374005603dc9d74bf42d9dbecb48d24321db6e86818a148ec73c6f

Observation 17fb0e84-af5a-4e2e-8740-9b5b0858280f · outbound

This paper cites Agent3D-Zero: An Agent for Zero-shot 3D Understanding.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Agent3D-Zero: An Agent for Zero-shot 3D Understanding

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:49.930129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:49.930129Z digest=sha256:27566c530198eb6062b0aa641f8de9fb6f24061449abfbab305443993edcc11f

Observation 226d217a-3e07-4c3b-b688-b2ee95a6c96b · outbound

This paper cites 3dvg- transformer: Relation modeling for visual grounding on point clouds.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding 3dvg- transformer: Relation modeling for visual grounding on point clouds

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:50.559830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.935994Z digest=sha256:342c54b622e1684586323e7d02b9cfc662e9cf5713c56d585e671c37558d39a1

Observation 2a9a1a61-7362-4fcb-9c9b-63f68db3dcad · outbound

This paper cites 3d-vista: Pre-trained transformer for 3d vision and text alignment.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding 3d-vista: Pre-trained transformer for 3d vision and text alignment

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:50.532258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.948385Z digest=sha256:dc3c55813f3dd1f6ff5a598830eaf9189f40506c234c6dc4ef2df0b52511e451

Observation 4c4a031e-7246-46d6-9bd2-547c96541a18 · outbound

This paper cites Unifying 3d vision-language understanding via prompt- 20 able queries.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Unifying 3d vision-language understanding via prompt- 20 able queries

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:50.510643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.953645Z digest=sha256:ecb6201f86e56c64029190f308b267de80478e6425de8d1ad9e8d4dcb223b0ec

Observation 6222d3d7-628d-416b-8490-b8a2311777f3 · outbound

This paper cites Perception-aware multi-sensor fusion for 3d lidar semantic segmentation.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Perception-aware multi-sensor fusion for 3d lidar semantic segmentation

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:29:50.488349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:29:49.959679Z digest=sha256:9dfdb782ce615933b24a4e494a1053164fb1ea41674f198f746ab9b4702919ca

Observation dd747f9d-a8cf-4f27-8846-829062594495 · outbound

This paper cites Robust 3D Semantic Occupancy Prediction with Calibration-free Spatial Transformation.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Robust 3D Semantic Occupancy Prediction with Calibration-free Spatial Transformation

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:49.966156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:49.966156Z digest=sha256:20e998929fb15a83119af2a2f82688537806d5ecc7144a22920a542551f01934

Pith citing papers

Observation 8ae17a44-7ba1-4f05-ab12-cb5d67befba8 · inbound

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation cites this paper.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T15:14:41.667536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:14:41.667536Z digest=sha256:7e66b0d265869718f07b90f8ea2927c737347772601ae43f3ac28ed7cab4e49f

Observation 37219787-af06-4101-8af5-57e158d3730c · inbound

Zero-Shot 3D Visual Grounding from Vision-Language Models cites this paper.

Zero-Shot 3D Visual Grounding from Vision-Language Models SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:16.243728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:16.243728Z digest=sha256:1b76aedbb649c1f2724ba9437b698fab65ae65fe9898a3e9a464000c46301644

Observation d0f27499-30c3-4b67-846b-d64faf7298f9 · inbound

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making cites this paper.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding

Reference 131

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:30.014841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:30.014841Z digest=sha256:51ec04f6f95030627add0d7fe4d991359c6e058b59b2e13d6ecd7cb9bb6a0fa2

Observation e234c58c-8045-4310-aedb-429afb0f45c3 · inbound

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations cites this paper.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:52.340466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:52.340466Z digest=sha256:4a3aa2010d98d947609cec45429dd34afa5a92dab618f288cae55f8580a42b8d

Observation f33928b8-72ce-43a3-a0f8-3f30a67c1475 · inbound

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models cites this paper.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:08.358863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:08.358863Z digest=sha256:e9d0a56540714cc4fad6f0ae988d24f195fc22a804eb121df48fd7b51d2c4a8b

Observation 0b7dc8af-52a2-479e-9600-cd3c2817047a · inbound

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding cites this paper.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T14:52:56.353243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:52:56.353243Z digest=sha256:6dae7d839f5ef55310a4c092d1454455d9c7b665e5c0c8cb54332f1d149faf3f

Observation ed5c3627-0bae-493f-8604-3b51e51a99d6 · inbound

SceneGraphGrounder: Zero-Shot 3D Visual Grounding via Structured Scene Graph Matching cites this paper.

SceneGraphGrounder: Zero-Shot 3D Visual Grounding via Structured Scene Graph Matching SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:51:18.444293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T08:47:25.712575Z digest=sha256:abdeae20d613c14a13a112537c00b9c4d4932aa34591cab34fdd5c75addf3386