Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T21:29:49.966156Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 7 inbound Pith citation observations for arXiv:2412.04383.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T21:29:49.966156Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T15:14:41.667536Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-22T08:51:18.441484Z
72 of 72 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f6cb22ce-43ef-47e8-9ae6-50e6f8b49193 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e39fb830-7871-41d2-a4de-a07e3e9af0bf · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Look around and refer: 2d synthetic semantics knowledge distillation for 3d visual grounding
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 112e1648-006b-4229-a063-8415503d7cd7 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Se- mantickitti: A dataset for semantic scene understanding of lidar sequences
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 09d00982-2c1a-4c22-bb0d-b62e8136da9d · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Mikasa: Multi-key-anchor & scene-aware transformer for 3d visual grounding
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b119eb19-c988-4ca9-aad3-4c57632b50fd · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Scanrefer: 3d object localization in rgb-d scans using natural language
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10a4e4e5-ac24-4d80-9baa-86e6afdcb3a2 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Clip2scene: Towards label-efficient 3d scene understanding by clip
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bda70d72-1882-4358-900e-ff6c7d1a080a · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Language conditioned spatial relation reasoning for 3d object grounding
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 34286f96-c6fc-49c0-9ebf-e438d56a29c9 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Think global, act lo- cal: Dual-scale graph transformer for vision-and-language navigation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e420ad3c-339c-4bf9-8869-c75dcf53a576 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e2ba133-1be6-4810-ab37-5d40571d6053 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88bcb889-53ae-4c07-a0b9-3b0cbd8eecb7 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Panoptic nuscenes: A large-scale benchmark for lidar panoptic segmentation and tracking.IEEE Robotics and Au- tomation Letters, 7:3795–3802, 2022
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b64cd2db-9507-462a-8c4f-02fccf68a985 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2e1fbbb-040f-4e9e-addb-3047ac4687d1 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding From Cognition to Precognition: A Future-Aware Framework for Social Navigation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dec190a-2741-4e82-98d0-eb3512d7922b · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Viewrefer: Grasp the multi-view knowledge for 3d visual grounding
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8bcb7c49-8e65-4f33-9d51-52465e746d80 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding CogVLM2: Visual Language Models for Image and Video Understanding
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6485a535-f2d4-443f-a26a-f8a4259e03f3 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding 3d-llm: Inject- ing the 3d world into large language models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2c8dd27c-8fd8-4572-bf9d-59c32111747e · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Dhp-mapping: A dense panoptic map- ping system with hierarchical world representation and label optimization techniques
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 399943dd-3528-4f6e-9deb-1c421107b5f6 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Text-guided graph neural networks for re- ferring 3d instance segmentation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e1054693-672d-46dc-9416-e67e38b96040 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Multi- view transformer for 3d visual grounding
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4ceb43ba-c2d3-4c6a-bf3f-06c80b12a1c2 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Assister: As- sistive navigation via conditional instruction generation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b15275ad-dfef-4465-a6e5-95f8daf82222 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Openins3d: Snap and lookup for 3d open-vocabulary instance segmentation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1fcf8d33-1cb7-46f8-82df-5138d8bec106 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Bottom up top down detection transform- ers for language grounding in images and point clouds
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1db238a8-95a1-47ca-89bd-fd7281c598f4 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Tenenbaum, Celso Miguel de Melo, Madhava Krishna, Liam Paull, Florian Shkurti, and Antonio Torralba
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bc673fdd-578e-43d6-bb3e-26bffbbe837b · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Sceneverse: Scaling 3d vision-language learning for grounded scene understanding
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8c6db71e-9c4e-4638-996b-df99fea13b19 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Lerf: Language embedded radiance fields
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89179189-6ce2-4572-907a-374fcf32d247 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Rethinking range view representation for lidar segmentation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 073a566a-f6d6-4ec5-b422-58d7c9c316f8 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Robo3d: Towards robust and reliable 3d perception against corruptions
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ac180bbb-69d1-4bdd-9e04-c9f174d84519 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Xvo: Generalized visual odometry via cross-modal self-training
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 029e69aa-d9a3-4452-9a27-815cbc6202cf · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding COARSE3D: Class-Prototypes for Contrastive Learning in Weakly-Supervised 3D Point Cloud Segmentation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 577f477f-1bb4-4236-a0fd-d8a96db0c541 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Tfnet: Exploiting temporal cues for fast and accurate lidar semantic segmentation
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 24df5088-9206-481c-80e8-b93a05f43531 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Uni3DL: Unified Model for 3D and Language Understanding
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef4b1034-0e4f-43cc-9662-f2cf26bfa098 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding V oxformer: Sparse voxel transformer for camera- based 3d semantic scene completion
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b86cca78-eeb9-4c42-97d8-f0c70d79b459 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Is your lidar placement optimized for 3d scene understanding? InAdvances in Neural Information Process- ing Systems, pages 34980–35017, 2024
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7e8d4bb3-e85b-45ad-af75-c4b6c2083e88 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Segment any point cloud sequences by distilling vision foundation models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation dca5e6f6-7000-43e5-8a41-7a48fcae46b4 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Deep view synthesis via self-consistent generative network.IEEE Transactions on Multimedia, 24: 451–465, 2021
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 995845fd-4c67-4fa6-b874-49c8e7a08424 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding RayDF: Neural Ray-surface Distance Fields with Multi-view Consistency
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b69498fd-b26a-48e4-89b7-ef9c5134c419 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a64f816-c135-4b69-91ae-4048d089e496 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Ovir-3d: Open-vocabulary 3d in- stance retrieval without training on 3d data
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3e40ada5-e8f2-47dc-99db-4c7892f0a542 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding An Examination of the Compositionality of Large Generative Vision-Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70c56b15-cd52-48d3-88d1-198c23fb25e6 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding GLOVER: Generalizable Open-Vocabulary Affordance Reasoning for Task-Oriented Grasping
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f249905e-ee7b-46fb-b9e9-edb7e074e6d5 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding GPT-4 Technical Report
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b0407df-001e-42d8-a7d4-cce5aa3fbb2a · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Training lan- guage models to follow instructions with human feedback
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation dc8d5b4c-333c-4d69-8a9d-afced40b54c8 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Openscene: 3d scene understanding with open vocabularies
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ddc87289-850e-4edb-b6c4-2046e81a1c71 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Multi-branch collaborative learning network for 3d visual grounding
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d1f1d61a-8e16-429a-9b1b-0a49a261e84b · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Learn- ing transferable visual models from natural language super- vision
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a5e725ec-b7d2-41d8-a8d1-95b114346831 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Mask3d: Mask trans- former for 3d semantic instance segmentation
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fdc2c7f6-6fbc-45e9-a74d-44eafa71d734 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Interactive planning using large language models for partially observable robotic tasks
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ef0ee074-2ce8-44f6-9c23-982ab81661a1 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Scalability in perception 19 for autonomous driving: Waymo open dataset
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2f5cc807-e491-4958-8eba-e1721913540a · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding OpenMask3D: Open-Vocabulary 3D Instance Segmentation
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27a9fa37-85f7-4a3b-821a-bad2a382ee38 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Epmf: Efficient perception-aware multi-sensor fusion for 3d semantic seg- mentation.IEEE Transactions on Pattern Analysis and Ma- chine Intelligence, 2024
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 18e0165f-6578-4549-9ccf-d210e8576550 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Four ways to improve verbo-visual fusion for dense 3d visual grounding
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 164e63c6-291b-4b4f-8736-09cb20cc2964 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fe8600d-c2a5-4150-923b-27215381193f · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Embodiedscan: A holistic multi- modal 3d perception suite towards embodied ai
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 70a23182-f22d-453b-86dc-baf0f9660616 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ae425a6-1b24-4fb0-a5e2-682d0a79038d · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding G3-lq: Marry- ing hyperbolic alignment with explicit semantic-geometric modeling for 3d visual grounding
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3816464f-68de-4177-9dc8-b52b3cb27c68 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Distill- ing coarse-to-fine semantic matching knowledge for weakly supervised 3d visual grounding
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation af26b7e0-95e1-4112-9cd7-f4aa8947644c · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding SIR: Multi-view Inverse Rendering with Decomposable Shadow Under Indoor Intense Lighting
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcd9bb8c-972f-4ecd-bcc7-daf926042dcc · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Eda: Explicit text-decoupling and dense alignment for 3d visual grounding
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0179f253-307d-416e-9eac-3bcda0740004 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 653a0a03-c806-4835-958d-e044656a4ebc · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding 4d contrastive superflows are dense 3d representation learners
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 09cfb949-faa8-493f-8d28-9f92a25c54c3 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Llm- grounder: Open-vocabulary 3d visual grounding with large language model as an agent
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b48f51be-9a29-4f9c-ab31-2b01212e55ac · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Regionplc: Regional point-language contrastive learning for open-world 3d scene understanding
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e0aac898-37e0-4e83-af92-1859a8ab30c0 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Sat: 2d semantics assisted training for 3d visual grounding
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f6643c1a-2542-4fe0-b77b-93fb731b08d8 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Sai3d: Segment any instance in 3d scenes
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9ada1862-c40f-4207-aff6-8449c7cb4d09 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual refer- ring
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e287afdb-037c-435b-9b32-7381a7bdcbd6 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Visual programming for zero-shot open-vocabulary 3d visual grounding
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 17fb0e84-af5a-4e2e-8740-9b5b0858280f · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Agent3D-Zero: An Agent for Zero-shot 3D Understanding
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 226d217a-3e07-4c3b-b688-b2ee95a6c96b · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding 3dvg- transformer: Relation modeling for visual grounding on point clouds
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2a9a1a61-7362-4fcb-9c9b-63f68db3dcad · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding 3d-vista: Pre-trained transformer for 3d vision and text alignment
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4c4a031e-7246-46d6-9bd2-547c96541a18 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Unifying 3d vision-language understanding via prompt- 20 able queries
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6222d3d7-628d-416b-8490-b8a2311777f3 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Perception-aware multi-sensor fusion for 3d lidar semantic segmentation
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation dd747f9d-a8cf-4f27-8846-829062594495 · outbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding Robust 3D Semantic Occupancy Prediction with Calibration-free Spatial Transformation
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ae17a44-7ba1-4f05-ab12-cb5d67befba8 · inbound
PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37219787-af06-4101-8af5-57e158d3730c · inbound
Zero-Shot 3D Visual Grounding from Vision-Language Models SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0f27499-30c3-4b67-846b-d64faf7298f9 · inbound
AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding
Reference 131
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e234c58c-8045-4310-aedb-429afb0f45c3 · inbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f33928b8-72ce-43a3-a0f8-3f30a67c1475 · inbound
T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b7dc8af-52a2-479e-9600-cd3c2817047a · inbound
SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed5c3627-0bae-493f-8604-3b51e51a99d6 · inbound
SceneGraphGrounder: Zero-Shot 3D Visual Grounding via Structured Scene Graph Matching SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.