Pith. sign in

Paper Citation Record · LEDGER

Zero-Shot 3D Visual Grounding from Vision-Language Models

As of 10 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 1 inbound Pith citation observation for arXiv:2505.22429.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22429 v1

Coverage vector

measured 78 of 78 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:12:19.747542Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T06:41:38.649424Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T06:45:29.442037Z

Reference resolution

78 of 78 outbound references displayed

  • verified exact5
  • verified fuzzy57
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 960747c0-afc7-4659-b22b-4e9984151d05 · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Scanrefer: 3d object localization in rgb-d scans using natural language,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:33.112862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:13.394174Z digest=sha256:a78264963976642a7baf3192e53bee3791606c7b9a0361e38907fc1880e0cc18

Observation 8b9a2380-ea31-4728-a03e-027b0f870747 · outbound

This paper cites RayDF: Neural Ray-surface Distance Fields with Multi-view Consistency.

Zero-Shot 3D Visual Grounding from Vision-Language Models RayDF: Neural Ray-surface Distance Fields with Multi-view Consistency

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:12:21.282429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:13.502388Z digest=sha256:566449fb91edc07d12b5745f9f58f3bb2309675998805bdca1b3da4d3d62c38f

Observation dbfea0d5-1991-4263-ac19-3d58190694cb · outbound

This paper cites Deep view synthesis via self-consistent gen- erative network,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Deep view synthesis via self-consistent gen- erative network,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:32.931597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:13.596606Z digest=sha256:7e4dc8d1b5742f9d2831ddf91141c43b26041430df4af899e796f09fe6414648

Observation e3b704f0-aab2-4e93-86ab-32d1e3515bac · outbound

This paper cites PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation.

Zero-Shot 3D Visual Grounding from Vision-Language Models PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:13.660944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:13.660944Z digest=sha256:d3e49557d4fa2876b849f3dd38a62d35aeff27996cfa86bf3cbf63f3689ddb7c

Observation 57a6e328-f12f-4ab9-af53-7caa8174205f · outbound

This paper cites SIR: Multi-view Inverse Rendering with Decomposable Shadow Under Indoor Intense Lighting.

Zero-Shot 3D Visual Grounding from Vision-Language Models SIR: Multi-view Inverse Rendering with Decomposable Shadow Under Indoor Intense Lighting

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:12:21.087050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:13.744633Z digest=sha256:3c6229588fdd0ca04d93b04a028d4e20a9ef962d4e0d6be36b3b34974c1120e6

Observation 4f6dd310-9524-424d-8221-e830ebea8546 · outbound

This paper cites An Examination of the Compositionality of Large Generative Vision-Language Models.

Zero-Shot 3D Visual Grounding from Vision-Language Models An Examination of the Compositionality of Large Generative Vision-Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:12:20.851836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:13.802829Z digest=sha256:044767d20365136d27c876724553f1170330305182b28b8a975939e9560195ae

Observation aa17f378-4420-4b33-b61d-de83946f313d · outbound

This paper cites Think global, act local: Dual-scale graph transformer for vision-and-language navigation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Think global, act local: Dual-scale graph transformer for vision-and-language navigation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:32.731576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:13.940140Z digest=sha256:9f86d89a0a3d2ee4fd1edfa8b0e9533afaecb51b04450b086ac8b6130d621b3d

Observation 330b875b-8571-4c73-ae2c-62fa7c4aaeb1 · outbound

This paper cites Assister: Assistive navigation via condi- tional instruction generation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Assister: Assistive navigation via condi- tional instruction generation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:32.568017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:14.029025Z digest=sha256:db210898fe9a014b85ccb7f94ef4fefbda4d29fef55e696d5db0b6fe7f3b3baa

Observation 9f53beb9-7803-4995-a8ac-6306a5f28797 · outbound

This paper cites From Cognition to Precognition: A Future-Aware Framework for Social Navigation.

Zero-Shot 3D Visual Grounding from Vision-Language Models From Cognition to Precognition: A Future-Aware Framework for Social Navigation

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:12:20.654751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:14.101940Z digest=sha256:ac3d2bc372a1fa2a55312ebb539c2f7f904f0d9a573f2c32fd9ebea0a8543f83

Observation dd10ce10-0d22-4d89-9acd-dcdaeb9deac4 · outbound

This paper cites Clip2scene: Towards label-efficient 3d scene understanding by clip,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Clip2scene: Towards label-efficient 3d scene understanding by clip,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:32.413691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:14.175634Z digest=sha256:735766694af183c36c0664751aa0a02d2eca301ee7759a005927ef6096bbaa1d

Observation 31d13b7b-2b84-4766-85e6-5592b983dad8 · outbound

This paper cites Robo3d: Towards robust and reliable 3d perception against corruptions,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Robo3d: Towards robust and reliable 3d perception against corruptions,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:32.257194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:14.211674Z digest=sha256:90e822494086896d8c3872138500c936c5383ee5304d76b498526a14ffee6019

Observation c1d22c9f-0892-4398-8716-e2d9ccbc6a59 · outbound

This paper cites Rethinking range view representation for lidar segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Rethinking range view representation for lidar segmentation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:32.093229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:14.269354Z digest=sha256:b094b7a6b2af948768febc93676f62969038987a8b1e598f40554e2b2f0695f8

Observation 2bedc6dc-5cfb-46f0-91a9-202261a576b1 · outbound

This paper cites Xvo: Generalized visual odometry via cross- modal self-training,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Xvo: Generalized visual odometry via cross- modal self-training,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:31.912696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:14.340641Z digest=sha256:9add7a00a4b9347006d24b628ac11fafd82548e03d7888ffd9e09757bb43bd1a

Observation 03bfdf04-2234-4fe5-a716-86bba0b57d4b · outbound

This paper cites COARSE3D: Class-Prototypes for Contrastive Learning in Weakly-Supervised 3D Point Cloud Segmentation.

Zero-Shot 3D Visual Grounding from Vision-Language Models COARSE3D: Class-Prototypes for Contrastive Learning in Weakly-Supervised 3D Point Cloud Segmentation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:14.416591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:14.416591Z digest=sha256:5221e2fdcb83ef9f92846ea662b06e1044ffb60492debc1799c69669fd8f2048

Observation fa261fc4-6dda-44f3-a79d-88fba013e8b5 · outbound

This paper cites Perception-aware multi-sensor fusion for 3d lidar semantic segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Perception-aware multi-sensor fusion for 3d lidar semantic segmentation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:31.722848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:14.467600Z digest=sha256:ae9c03929ea608667fb09402b87663a8c191a4669841539bc3c7a72a8bb207c1

Observation 6b23063d-3182-489e-ae81-83d3d6d434eb · outbound

This paper cites Epmf: Efficient perception-aware multi- sensor fusion for 3d semantic segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Epmf: Efficient perception-aware multi- sensor fusion for 3d semantic segmentation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:31.545958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:14.521529Z digest=sha256:5194a11eb909ce8fa6277e3b0a8da9f968defcae389c76cc5c310dcc563ccb3c

Observation 1040d8e3-f317-4bd5-8225-322ce41fcea5 · outbound

This paper cites Tfnet: Exploiting temporal cues for fast and accurate lidar semantic segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Tfnet: Exploiting temporal cues for fast and accurate lidar semantic segmentation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:31.385278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:14.602509Z digest=sha256:794e484ba53b9873375d64ad722f20cc0fefcc4cbf1c7aa22b8e7e700569cdde

Observation dd2fa089-3b28-4518-a393-274e9128ad3c · outbound

This paper cites Robust 3D Semantic Occupancy Prediction with Calibration-free Spatial Transformation.

Zero-Shot 3D Visual Grounding from Vision-Language Models Robust 3D Semantic Occupancy Prediction with Calibration-free Spatial Transformation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:14.678694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:14.678694Z digest=sha256:9423296d417ddc49f14d2314f2e9455ba234ae4a5319581bfbd5a608964270bf

Observation 6adcb6dd-1a21-4311-bcf5-680846d8444c · outbound

This paper cites Dhp-mapping: A dense panoptic mapping sys- tem with hierarchical world representation and label opti- mization techniques,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Dhp-mapping: A dense panoptic mapping sys- tem with hierarchical world representation and label opti- mization techniques,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:31.170166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:14.771248Z digest=sha256:58dcf66f5c2d5b575c3087a0c861dc7658782397fbf8264fc9aaf4ed96177ba2

Observation dd1c1904-35de-4441-b3eb-dd532efb8937 · outbound

This paper cites Multi-modal data-efficient 3d scene un- derstanding for autonomous drivin,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Multi-modal data-efficient 3d scene un- derstanding for autonomous drivin,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:30.987688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:14.854042Z digest=sha256:36a873383ae5f72a418fbbe42eec102ebae304bcd400b806b60a9b8437625b9e

Observation 77efeec2-affe-46a1-8e6c-5590ca719030 · outbound

This paper cites Dynamiccity: Large-scale 4d occu- pancy generation from dynamic scenes,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Dynamiccity: Large-scale 4d occu- pancy generation from dynamic scenes,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:14.916380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:14.916380Z digest=sha256:1fd1d20e79689e85578b24e9eb3f56e2e3cf05836222f7a017435da1b4918853

Observation 21947037-35b4-4109-83a4-c704eadff963 · outbound

This paper cites Calib3d: Calibrating model preferences for reliable 3d scene understanding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Calib3d: Calibrating model preferences for reliable 3d scene understanding,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:30.705023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:15.013655Z digest=sha256:c76d0bfa7ad41a3eb10fd28b33cf70970e3b2b8f123bf9b0f4a83afe9c047714

Observation 5c6df824-f55a-43df-b72b-d33033b5e4f1 · outbound

This paper cites Bottom up top down detection transform- ers for language grounding in images and point clouds,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Bottom up top down detection transform- ers for language grounding in images and point clouds,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:30.534444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:15.130922Z digest=sha256:7e2b12f9d70aca6b43bdadf5e46d6ee47ead1950c46d7cfcc3d880fd8e505ca0

Observation bc2adfb0-b5bc-4cec-8761-c20b7c790c46 · outbound

This paper cites 3d-vista: Pre-trained transformer for 3d vision and text alignment,.

Zero-Shot 3D Visual Grounding from Vision-Language Models 3d-vista: Pre-trained transformer for 3d vision and text alignment,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:30.315718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:15.204457Z digest=sha256:63d6897c4630f6fff62f9e96eaf877584d0386dbe295ca1f7fd7b342ef5dac83

Observation 7c0380bf-1754-4b5b-91dd-f9bc7301825e · outbound

This paper cites Eda: Explicit text-decoupling and dense align- ment for 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Eda: Explicit text-decoupling and dense align- ment for 3d visual grounding,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:30.141405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:15.273293Z digest=sha256:0483b032839162eae4366f375bb94af2978607131dbbcc13c0e6d4085b5e72f5

Observation d50e43f0-cf9a-41c7-b7b3-64936739d156 · outbound

This paper cites 3dvg-transformer: Relation modeling for vi- sual grounding on point clouds,.

Zero-Shot 3D Visual Grounding from Vision-Language Models 3dvg-transformer: Relation modeling for vi- sual grounding on point clouds,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:30.002253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:15.322438Z digest=sha256:f6a21bda4fe87eabf84dc12d9b2bb0323cce35a634b3179882c8a5e2b4c61b9b

Observation bd712acb-4ec8-4fd0-bc3f-3c2a28ce6d8f · outbound

This paper cites Instancerefer: Cooperative holistic under- standing for visual grounding on point clouds through in- stance multi-level contextual referring,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Instancerefer: Cooperative holistic under- standing for visual grounding on point clouds through in- stance multi-level contextual referring,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:29.817183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:15.400260Z digest=sha256:971ba55f43e0a34f14f2b2ed386e883b37b06c570a6880f28d159b2f8e4adde2

Observation f3b330c4-f015-42f9-86e6-d02a4f2ff0f9 · outbound

This paper cites Multi-branch collaborative learning network for 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Multi-branch collaborative learning network for 3d visual grounding,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:29.648448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:15.451650Z digest=sha256:67499a30fd3d036274da7d27c84033ae983f77bc9986e5e64df86f637c372b7a

Observation 5f07f699-f8bc-441a-8085-dcc8ec68e53d · outbound

This paper cites Semantickitti: A dataset for semantic scene understanding of lidar sequences,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Semantickitti: A dataset for semantic scene understanding of lidar sequences,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:29.457192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:15.491354Z digest=sha256:138d93be05565212b00780ea5423fc8af23e04dcc9b20983369aa09e01fd5fe0

Observation 2f3a2f58-fa72-4e45-9ff5-3861b4926bab · outbound

This paper cites Scalability in perception for autonomous driv- ing: Waymo open dataset,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Scalability in perception for autonomous driv- ing: Waymo open dataset,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:29.255587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:15.564839Z digest=sha256:40729b164099a23d490f209ab7f52163394a57229291fa449c7ca1f49eca500f

Observation 7354791c-6c9e-4f60-956a-af68e73e773c · outbound

This paper cites Panoptic nuscenes: A large-scale bench- mark for lidar panoptic segmentation and tracking,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Panoptic nuscenes: A large-scale bench- mark for lidar panoptic segmentation and tracking,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:29.002623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:15.623601Z digest=sha256:e614d101526a0add60a36e77a15a2f2a22954e7fda031d236bce837b9a7c37b4

Observation 11375c3b-8862-4593-bad2-7761dfc1ac0e · outbound

This paper cites Visual programming for zero-shot open- vocabulary 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Visual programming for zero-shot open- vocabulary 3d visual grounding,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:28.773037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:15.701053Z digest=sha256:e5ffadba8acd89c994299b33c30fc1553b884eb70328b88c1ccc17c156a7d835

Observation a012d746-72fe-44e3-bdeb-0d1aee487572 · outbound

This paper cites Llm-grounder: Open-vocabulary 3d visual grounding with large language model as an agent,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Llm-grounder: Open-vocabulary 3d visual grounding with large language model as an agent,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:28.545584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:15.776796Z digest=sha256:0e4f5b4e0ef7ea769f434f30362417e789348a7ccac6574e085bb721b5842341

Observation c155db64-fdc3-4454-8b88-6e86c1baba4a · outbound

This paper cites Training language models to fol- low instructions with human feedback,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Training language models to fol- low instructions with human feedback,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:28.325098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:15.878975Z digest=sha256:216f0b0f78a84432592534b75220d298c9e9e00a6ce7f43a99d6030585c08aae

Observation 1e1b7fb4-f508-42b6-8fab-a9bf6ea0c806 · outbound

This paper cites GPT-4 Technical Report.

Zero-Shot 3D Visual Grounding from Vision-Language Models GPT-4 Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:15.966014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:15.966014Z digest=sha256:b48f42186dac8bf281d55c17d9bbd56c353dd21afcbaa706a25862d56abd821e

Observation 3579b74e-7593-4f22-8ebb-e28f7c17fb6a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Zero-Shot 3D Visual Grounding from Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:16.033226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:16.033226Z digest=sha256:3d6c9cee082236da7c3e2dbbad69e55c6b46ef7e129cc058c1ba4ed8abe128bd

Observation 5426ae40-e7e6-4ba6-bb60-6b115f0d2f4a · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

Zero-Shot 3D Visual Grounding from Vision-Language Models CogVLM2: Visual Language Models for Image and Video Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:16.099070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:16.099070Z digest=sha256:87084a60dda2717d36b4fcc2256a369faeb53c6c5f6a8b80bb2627cb633298c6

Observation 63c38e4a-105a-4267-8616-ce139f488a22 · outbound

This paper cites Sceneverse: Scaling 3d vision-language learn- ing for grounded scene understanding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Sceneverse: Scaling 3d vision-language learn- ing for grounded scene understanding,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:28.062525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:16.172492Z digest=sha256:c2fbe6a79bb8b1a82f79554b7e5487c55163a5c586958f19634c463c6a921096

Observation 37219787-af06-4101-8af5-57e158d3730c · outbound

This paper cites SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding.

Zero-Shot 3D Visual Grounding from Vision-Language Models SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:16.243728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:16.243728Z digest=sha256:23613577e5f9a8368eca3bd66261982577506a6731db58299234c70c42a28c63

Observation dd2be870-d9b6-4bae-b676-1b4ff513377a · outbound

This paper cites Referit3d: Neural listeners for fine- grained 3d object identification in real-world scenes,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Referit3d: Neural listeners for fine- grained 3d object identification in real-world scenes,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:27.868411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:16.327105Z digest=sha256:0011864c926d0c7896fa38c909410380159a86decb906b1c27d95389e7d5ba76

Observation 1c3d62e2-a284-4093-9485-c07bfd12986b · outbound

This paper cites Viewrefer: Grasp the multi-view knowledge for 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Viewrefer: Grasp the multi-view knowledge for 3d visual grounding,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:27.642877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:16.383282Z digest=sha256:546fb54d51715fd2a829dee99e41afacec8d6c418e9a026d0953edc0b8ef7879

Observation a70094c3-e7ce-4b6c-8fc8-cb4d1670b286 · outbound

This paper cites Multi-view transformer for 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Multi-view transformer for 3d visual grounding,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:27.299693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:16.457874Z digest=sha256:6ccde4d1d4cd0a130c724e5cfde4585920ccf104f4a933519a708c0c070a61f0

Observation 9229b3cc-9dcb-4273-9701-4f45767296f1 · outbound

This paper cites Look around and refer: 2d synthetic se- mantics knowledge distillation for 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Look around and refer: 2d synthetic se- mantics knowledge distillation for 3d visual grounding,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:27.103416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:16.537325Z digest=sha256:71b001034ad0c3ffde3f2a80225509334811be183f2e91c8e18ad5f95ab6b261

Observation d6ce677b-b1b6-4488-b020-89df092fe040 · outbound

This paper cites Sat: 2d semantics assisted training for 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Sat: 2d semantics assisted training for 3d visual grounding,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:26.794171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:16.586571Z digest=sha256:bc54610bb5ea1a921b29c06989c48d55e8752a65471005d738312db76ff1b1e9

Observation 1ca3c590-6dea-4a27-874d-747b1befea47 · outbound

This paper cites Four ways to improve verbo-visual fusion for dense 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Four ways to improve verbo-visual fusion for dense 3d visual grounding,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:26.612752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:16.657203Z digest=sha256:08055aefbb4a3e5acdf98d44a1f14cf1602c2832dda198a063068ef0913e1910

Observation 8b9e7394-52c4-4635-90bc-357d17e6ace1 · outbound

This paper cites Distilling coarse-to-fine semantic matching knowledge for weakly supervised 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Distilling coarse-to-fine semantic matching knowledge for weakly supervised 3d visual grounding,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:26.393051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:16.723540Z digest=sha256:6a93b52db5d8401a179f8ee3ea8026816d233bc0872fe7cead2a026918666429

Observation 8ad91c79-521c-4f81-8bfc-9e2577bc4909 · outbound

This paper cites Unifying 3d vision-language understanding via promptable queries,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Unifying 3d vision-language understanding via promptable queries,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:26.169590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:16.772423Z digest=sha256:9199d6e1622eb44e6ce6ee4c1000f975a9309f412ee046a7591e27e12f66c3f8

Observation f858176d-4543-448f-94e2-5417467643c2 · outbound

This paper cites OVGaussian: Generalizable 3D Gaussian Segmentation with Open Vocabularies.

Zero-Shot 3D Visual Grounding from Vision-Language Models OVGaussian: Generalizable 3D Gaussian Segmentation with Open Vocabularies

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:16.828825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:16.828825Z digest=sha256:c00d0362626c7fe961cbea48cea64c6d05cbb403a05a434aa9ccfe9ed84c6a62

Observation b52141f5-02d0-4ff5-8e29-8998abfa1c46 · outbound

This paper cites Multi-space alignments towards universal lidar segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Multi-space alignments towards universal lidar segmentation,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:25.960756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:16.895168Z digest=sha256:1c9a1d9a26cd3d8c297fda662931cb650d5497906e34a346dc45ff1f023ccdbb

Observation 02ec2c3f-ff86-42ff-8ffb-34b28baecacf · outbound

This paper cites Towards label-free scene understanding by vision foundation models,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Towards label-free scene understanding by vision foundation models,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:25.756569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:16.967579Z digest=sha256:28a33353464b77393763b9bb3ddd5e153d92f027c6df5800666a8207e9ba17e0

Observation 72357686-fe9d-482c-a614-4c8e5f09b4b6 · outbound

This paper cites LiMoE: Mixture of LiDAR Representation Learners from Automotive Scenes.

Zero-Shot 3D Visual Grounding from Vision-Language Models LiMoE: Mixture of LiDAR Representation Learners from Automotive Scenes

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:17.070727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:17.070727Z digest=sha256:d9dbca7eea6eceaea832d3fd446323dbbbf53b36dde639396dd2078991893e31

Observation 197f289d-d06e-4527-8ea0-b7333f8cc82a · outbound

This paper cites GEAL: Generalizable 3D Affordance Learning with Cross-Modal Consistency.

Zero-Shot 3D Visual Grounding from Vision-Language Models GEAL: Generalizable 3D Affordance Learning with Cross-Modal Consistency

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:17.150193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:17.150193Z digest=sha256:0fa40ad12d365ce209b32cc4c0fd369d951ed8380948cf2a32ad3daf819e8810

Observation 030e2204-a25d-4b24-a0b4-3397826df452 · outbound

This paper cites Openscene: 3d scene understanding with open vocabularies,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Openscene: 3d scene understanding with open vocabularies,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:25.487986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:17.213896Z digest=sha256:e3f662e30ac7defb702afc2adededaefd285440ace239d8c448b82e6151142fb

Observation 423967c6-2e0e-4c80-bcf4-565da2f9c4d3 · outbound

This paper cites Lerf: Language embedded radiance fields,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Lerf: Language embedded radiance fields,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:25.225323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:17.283351Z digest=sha256:ae1ea43e8d9c0eadf5e56a5379348a5542666aff43896fc6b00b527e02e17ad3

Observation 94103a13-4e1b-43a6-bbc2-ed6e9e18a290 · outbound

This paper cites Ovir-3d: Open-vocabulary 3d instance retrieval without training on 3d data,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Ovir-3d: Open-vocabulary 3d instance retrieval without training on 3d data,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:25.019581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:17.343736Z digest=sha256:57a8d0a7c68f40d87d7109dc478a66424c4717a07cf6ee5930c447677b737e63

Observation 1f8f6e63-38ef-47e1-98da-a0ea5409aeeb · outbound

This paper cites Agent3D-Zero: An Agent for Zero-shot 3D Understanding.

Zero-Shot 3D Visual Grounding from Vision-Language Models Agent3D-Zero: An Agent for Zero-shot 3D Understanding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:17.411309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:17.411309Z digest=sha256:a7333cb34f89cdad0c9eb51a3d4ccf31b22a4eb48bdbfe5c434d762fb92d3a8f

Observation 1d2c5493-8c26-42f6-90f4-97d8c35ede01 · outbound

This paper cites Regionplc: Regional point-language con- trastive learning for open-world 3d scene understanding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Regionplc: Regional point-language con- trastive learning for open-world 3d scene understanding,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:24.776061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:17.470353Z digest=sha256:855b3c4e0b00307013b59386dd37d45bf0276fef932232e375fe499b04c976cc

Observation 739b960d-0787-420d-8179-33da110aa4cf · outbound

This paper cites OpenMask3D: Open-Vocabulary 3D Instance Segmentation.

Zero-Shot 3D Visual Grounding from Vision-Language Models OpenMask3D: Open-Vocabulary 3D Instance Segmentation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:17.523576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:17.523576Z digest=sha256:dae0d7db7648a493de824a9a2cbffc5934cf0668f5e1b3c63fe30ea62445fe2b

Observation 9fff50e1-dbfb-428b-bd79-ed7ded1e7bfd · outbound

This paper cites Openins3d: Snap and lookup for 3d open- vocabulary instance segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Openins3d: Snap and lookup for 3d open- vocabulary instance segmentation,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:24.563327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:17.579259Z digest=sha256:e90b0d49878ac8c8e4baf20e505750a8e3bc184e2fc4d93858f221255e84acbf

Observation 6b063554-7f08-45ac-8735-ec16104d3204 · outbound

This paper cites Sai3d: Segment any instance in 3d scenes,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Sai3d: Segment any instance in 3d scenes,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:24.321836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:17.637874Z digest=sha256:193aca69fea467557971a6b8e2461e84671b55110fc793fa5f6ee97c00b2e9b1

Observation b8e8a98d-4d75-4d7e-859e-cbf1efa418be · outbound

This paper cites Lasermix for semi-supervised lidar semantic segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Lasermix for semi-supervised lidar semantic segmentation,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:24.131536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:17.710770Z digest=sha256:a641c8b667c57cc9e65686c26eef547e30fae9f12e06f49464d55bec3c52929f

Observation 88a4132e-37ac-4dc3-aec5-a64dc8004e8f · outbound

This paper cites Segment any point cloud sequences by distilling vision foundation models,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Segment any point cloud sequences by distilling vision foundation models,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:23.906258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:17.807087Z digest=sha256:ed81d31cdda6e4ab5d45a6de6039e50d6e69aa7e546c23b8c5c62e3851f65939

Observation c44806ef-9311-4844-9525-3cd2395c8625 · outbound

This paper cites 4d contrastive superflows are dense 3d repre- sentation learners,.

Zero-Shot 3D Visual Grounding from Vision-Language Models 4d contrastive superflows are dense 3d repre- sentation learners,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:23.683419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:17.935512Z digest=sha256:9861994fc5ebdd0e48051f5fee9b3200a208d00e58255c1adb96e5d12ea8d593

Observation 1c3f78e0-9855-4610-a55d-3cdbde6b628c · outbound

This paper cites Frnet: Frustum-range networks for scalable lidar segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Frnet: Frustum-range networks for scalable lidar segmentation,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:23.484256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:18.034869Z digest=sha256:f361431f7095a39676a5d8de1a7e4031817a570cb157c7c39c32b452e396283d

Observation c4b43a23-36bf-481c-ae1a-d60fab1e0d84 · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

Zero-Shot 3D Visual Grounding from Vision-Language Models Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:18.091406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:18.091406Z digest=sha256:9595f4a35b2b5262a5ab2cd3403a6d410cec49693a00c7d63e19ff3996cdb29f

Observation b469ac04-b778-4bfe-b0ae-85bc252010b2 · outbound

This paper cites Uni3DL: Unified Model for 3D and Language Understanding.

Zero-Shot 3D Visual Grounding from Vision-Language Models Uni3DL: Unified Model for 3D and Language Understanding

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:12:20.164538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:18.156553Z digest=sha256:0d10f2fa4d1842b34985892ba5553786fe68fb06523b05cff2b536a1c496d53d

Observation 6171bf57-8155-4ef5-bc3d-97e28edb525c · outbound

This paper cites Conceptfusion: Open-set multi- modal 3d mapping,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Conceptfusion: Open-set multi- modal 3d mapping,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:23.283686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:18.254570Z digest=sha256:a1d3025f5ec8ffeaa7828a01fe208cd7869904a7b58b8462800c7bf17936d4dc

Observation b1e0260d-814e-4607-846a-a2b4f28ac11a · outbound

This paper cites GLOVER: Generalizable Open-Vocabulary Affordance Reasoning for Task-Oriented Grasping.

Zero-Shot 3D Visual Grounding from Vision-Language Models GLOVER: Generalizable Open-Vocabulary Affordance Reasoning for Task-Oriented Grasping

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:18.478933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:18.478933Z digest=sha256:ecff285ed778b5c5cfc6a123b0272f13c13bd9892e0b6b619b44043becade6d9

Observation 11fb35db-9415-452a-8a43-1742b926fdb6 · outbound

This paper cites Interactive planning using large language models for partially observable robotic tasks,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Interactive planning using large language models for partially observable robotic tasks,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:23.161734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:18.706597Z digest=sha256:db914407775a4de669e6221e5dcac88b63215eb415beb01474cd96c6d65faf62

Observation fce09665-8e57-4ecf-ac78-7d026fcc53c8 · outbound

This paper cites 3d-llm: Injecting the 3d world into large language models,.

Zero-Shot 3D Visual Grounding from Vision-Language Models 3d-llm: Injecting the 3d world into large language models,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:22.964761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:18.817427Z digest=sha256:89ebb2b1281bc75693b8b06c91efacf380409d909f42414aa6c686a08c02f1e5

Observation 5fcd2776-fd95-49b2-b879-6e63ed38e8fd · outbound

This paper cites Is your lidar placement optimized for 3d scene understanding?,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Is your lidar placement optimized for 3d scene understanding?,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:22.736693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:18.907550Z digest=sha256:9126c1c77fcc248cdd8f031d42bc3d263e031e813e41e5771eb20588b29c95ba

Observation 7343377b-7b76-4d17-b048-d0c15d351101 · outbound

This paper cites G3-lq: Marrying hyperbolic alignment with explicit semantic-geometric modeling for 3d visual ground- ing,.

Zero-Shot 3D Visual Grounding from Vision-Language Models G3-lq: Marrying hyperbolic alignment with explicit semantic-geometric modeling for 3d visual ground- ing,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:22.557138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:18.980280Z digest=sha256:c8d3b29a762766046557875f397b8bf8128207773fff93f5ab5571045791bf35

Observation 3f863d07-94ef-4de9-ad28-d001bf53d84e · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Learning transferable visual models from natural language supervision,

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:22.328649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:19.127295Z digest=sha256:5912a536d1b259d44eb6e34593d9e75d7b65ed12a6e3cf57d7b242f46047bd5b

Observation edcfabd1-07f9-4cd5-b156-2823a5aa0bc2 · outbound

This paper cites VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding.

Zero-Shot 3D Visual Grounding from Vision-Language Models VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:19.231195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:19.231195Z digest=sha256:2c6b06490a861ea8cf42ac706fb9a60241345dbd45e7284dbe2313b50f8f437c

Observation 6bb7d6ca-a924-4237-8f5e-d5f833206cbb · outbound

This paper cites Text-guided graph neural networks for referring 3d instance segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Text-guided graph neural networks for referring 3d instance segmentation,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:22.080070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:19.353586Z digest=sha256:44365a9b279414f3cf853b7b538daf41940fbd9227f2a61071007537026ab595

Observation 4761a55c-7a2e-4aa9-9ef1-8cf0368ce164 · outbound

This paper cites Mikasa: Multi-key-anchor & scene- aware transformer for 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Mikasa: Multi-key-anchor & scene- aware transformer for 3d visual grounding,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:21.837968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:19.499102Z digest=sha256:4674b5c676997da357c43d3877212355946f8b4e86fc1dc74cbde528f2e9e7b8

Observation 0de609bf-af3c-4a3d-b24c-fc6057d60b72 · outbound

This paper cites Language conditioned spatial relation rea- soning for 3d object grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Language conditioned spatial relation rea- soning for 3d object grounding,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:21.636127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:19.626429Z digest=sha256:c987da705b79326ad8d3eab6ba0a96bb7849b8f282f8c5ff2fcddbf95b7d4354

Observation b9c59bea-69c3-4075-8a2c-08946499f3a6 · outbound

This paper cites Mask3d: Mask transformer for 3d semantic instance segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Mask3d: Mask transformer for 3d semantic instance segmentation,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:21.510030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:12:19.747542Z digest=sha256:2b9a3fde4790a031b8a3835eb96903980855e1d511c7b6f58e7bd77137ae7d11

Pith citing papers

Observation 7e135835-ad52-43f6-9e78-62ef14a3da16 · inbound

PruneGround: Plug-and-play Spatial Pruning for 3D Visual Grounding cites this paper.

PruneGround: Plug-and-play Spatial Pruning for 3D Visual Grounding Zero-Shot 3D Visual Grounding from Vision-Language Models

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-07-01T06:45:29.444056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:41:38.649424Z digest=sha256:b392208c07466ec92471cdf3626c55f2f3e43ff7f32acc59d95938c80bb25388