Pith. sign in

Paper Citation Record · LEDGER

Zero-Shot 3D Visual Grounding from Vision-Language Models

As of 18 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 1 inbound Pith citation observation for arXiv:2505.22429.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22429 v1

Coverage vector

measured 78 of 78 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:12:19.747542Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T06:41:38.649424Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T06:45:29.442037Z

Reference resolution

78 of 78 outbound references displayed

  • verified exact5
  • verified fuzzy57
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 960747c0-afc7-4659-b22b-4e9984151d05 · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Scanrefer: 3d object localization in rgb-d scans using natural language,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:33.112862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:13.394174Z digest=sha256:6b08d72695743e4b364061cf7dba9d8a4754685d1996d64719f13140faa6c434

Observation 8b9a2380-ea31-4728-a03e-027b0f870747 · outbound

This paper cites RayDF: Neural Ray-surface Distance Fields with Multi-view Consistency.

Zero-Shot 3D Visual Grounding from Vision-Language Models RayDF: Neural Ray-surface Distance Fields with Multi-view Consistency

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:12:21.282429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:13.502388Z digest=sha256:81a7223ee91b89f2dcf6fcc1e857b2498732f8006ceac41dbc2c7031ec2c70b0

Observation dbfea0d5-1991-4263-ac19-3d58190694cb · outbound

This paper cites Deep view synthesis via self-consistent gen- erative network,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Deep view synthesis via self-consistent gen- erative network,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:32.931597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:13.596606Z digest=sha256:a00a1b37c19027659c4414ce348fef5baa9a2cc0444842b7de1b937b1ee2ea18

Observation e3b704f0-aab2-4e93-86ab-32d1e3515bac · outbound

This paper cites PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation.

Zero-Shot 3D Visual Grounding from Vision-Language Models PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:13.660944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:13.660944Z digest=sha256:73c911e8d7d608119951fa415f53058a6b9e2cb134aa61a65a1d36048ac908de

Observation 57a6e328-f12f-4ab9-af53-7caa8174205f · outbound

This paper cites SIR: Multi-view Inverse Rendering with Decomposable Shadow Under Indoor Intense Lighting.

Zero-Shot 3D Visual Grounding from Vision-Language Models SIR: Multi-view Inverse Rendering with Decomposable Shadow Under Indoor Intense Lighting

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:12:21.087050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:13.744633Z digest=sha256:aeb8aef4e6c37e3ce3ca6fd8ca2b176cde9b48237b58e5054e4bacd29f44e24d

Observation 4f6dd310-9524-424d-8221-e830ebea8546 · outbound

This paper cites An Examination of the Compositionality of Large Generative Vision-Language Models.

Zero-Shot 3D Visual Grounding from Vision-Language Models An Examination of the Compositionality of Large Generative Vision-Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:12:20.851836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:13.802829Z digest=sha256:04351a6a0cc6e52bfe0b94272ae40d3942dc6db9d6ed26db14793bb86586178d

Observation aa17f378-4420-4b33-b61d-de83946f313d · outbound

This paper cites Think global, act local: Dual-scale graph transformer for vision-and-language navigation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Think global, act local: Dual-scale graph transformer for vision-and-language navigation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:32.731576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:13.940140Z digest=sha256:8a1ef0030da4fa8cd66610fb568f2e2fea334a0b263e85520854c00ebb38747e

Observation 330b875b-8571-4c73-ae2c-62fa7c4aaeb1 · outbound

This paper cites Assister: Assistive navigation via condi- tional instruction generation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Assister: Assistive navigation via condi- tional instruction generation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:32.568017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:14.029025Z digest=sha256:359ed281a768b6ba0f682087f16d58d09d1f1844050427d591e38154bafc5795

Observation 9f53beb9-7803-4995-a8ac-6306a5f28797 · outbound

This paper cites From Cognition to Precognition: A Future-Aware Framework for Social Navigation.

Zero-Shot 3D Visual Grounding from Vision-Language Models From Cognition to Precognition: A Future-Aware Framework for Social Navigation

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:12:20.654751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:14.101940Z digest=sha256:0d244a3328a15c78477ae0f3f8acb2d19199c25ece29e95bec67beccf345a4a9

Observation dd10ce10-0d22-4d89-9acd-dcdaeb9deac4 · outbound

This paper cites Clip2scene: Towards label-efficient 3d scene understanding by clip,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Clip2scene: Towards label-efficient 3d scene understanding by clip,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:32.413691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:14.175634Z digest=sha256:9e9f2e0483e19caf12f67f5f8ddd6bb87a9937b0afdacb5f4808eaf60e8b1c24

Observation 31d13b7b-2b84-4766-85e6-5592b983dad8 · outbound

This paper cites Robo3d: Towards robust and reliable 3d perception against corruptions,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Robo3d: Towards robust and reliable 3d perception against corruptions,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:32.257194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:14.211674Z digest=sha256:5dc849e9f70367d322c0098903503f170a6f433c7212440e3f5bece8950c12b1

Observation c1d22c9f-0892-4398-8716-e2d9ccbc6a59 · outbound

This paper cites Rethinking range view representation for lidar segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Rethinking range view representation for lidar segmentation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:32.093229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:14.269354Z digest=sha256:ad4378b8a560c80e8e90dade251aa9758ef8a27908712093b2183739b1ba1350

Observation 2bedc6dc-5cfb-46f0-91a9-202261a576b1 · outbound

This paper cites Xvo: Generalized visual odometry via cross- modal self-training,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Xvo: Generalized visual odometry via cross- modal self-training,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:31.912696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:14.340641Z digest=sha256:b129afb09e5f655c6b5a5632506bfe2a59ca1e3d97f65fce479e0d693a941f36

Observation 03bfdf04-2234-4fe5-a716-86bba0b57d4b · outbound

This paper cites COARSE3D: Class-Prototypes for Contrastive Learning in Weakly-Supervised 3D Point Cloud Segmentation.

Zero-Shot 3D Visual Grounding from Vision-Language Models COARSE3D: Class-Prototypes for Contrastive Learning in Weakly-Supervised 3D Point Cloud Segmentation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:14.416591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:14.416591Z digest=sha256:f81740c50382739e6f3118576bb5f49462fbfe54b17053f492f1c84635af2c23

Observation fa261fc4-6dda-44f3-a79d-88fba013e8b5 · outbound

This paper cites Perception-aware multi-sensor fusion for 3d lidar semantic segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Perception-aware multi-sensor fusion for 3d lidar semantic segmentation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:31.722848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:14.467600Z digest=sha256:fa03a75920c63c9e1ca633d6ef0aad3ada0b677ad8367783cd7208df31f3edd1

Observation 6b23063d-3182-489e-ae81-83d3d6d434eb · outbound

This paper cites Epmf: Efficient perception-aware multi- sensor fusion for 3d semantic segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Epmf: Efficient perception-aware multi- sensor fusion for 3d semantic segmentation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:31.545958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:14.521529Z digest=sha256:53e7ed4eeaf9c1390be696eb84554614ff80bffb2b5754fac6c9954adfaf7369

Observation 1040d8e3-f317-4bd5-8225-322ce41fcea5 · outbound

This paper cites Tfnet: Exploiting temporal cues for fast and accurate lidar semantic segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Tfnet: Exploiting temporal cues for fast and accurate lidar semantic segmentation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:31.385278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:14.602509Z digest=sha256:51cf942e139cd748d26cbe92cd947b8c125566f494e2b4ea62abf853877c1a0b

Observation dd2fa089-3b28-4518-a393-274e9128ad3c · outbound

This paper cites Robust 3D Semantic Occupancy Prediction with Calibration-free Spatial Transformation.

Zero-Shot 3D Visual Grounding from Vision-Language Models Robust 3D Semantic Occupancy Prediction with Calibration-free Spatial Transformation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:14.678694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:14.678694Z digest=sha256:34366d7c3b6c8850abd5b79137ccc9b9e4ab271c52aafc1718012f28a5c63d17

Observation 6adcb6dd-1a21-4311-bcf5-680846d8444c · outbound

This paper cites Dhp-mapping: A dense panoptic mapping sys- tem with hierarchical world representation and label opti- mization techniques,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Dhp-mapping: A dense panoptic mapping sys- tem with hierarchical world representation and label opti- mization techniques,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:31.170166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:14.771248Z digest=sha256:8a4e97425336cbe5184e127483f0cd153d8d3a94fd1e6b6404ea1460ad6bfebf

Observation dd1c1904-35de-4441-b3eb-dd532efb8937 · outbound

This paper cites Multi-modal data-efficient 3d scene un- derstanding for autonomous drivin,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Multi-modal data-efficient 3d scene un- derstanding for autonomous drivin,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:30.987688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:14.854042Z digest=sha256:695201d14e9cf3656df723b628c6d9787cd74fa77a61441db3c14927cb3b7706

Observation 77efeec2-affe-46a1-8e6c-5590ca719030 · outbound

This paper cites Dynamiccity: Large-scale 4d occu- pancy generation from dynamic scenes,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Dynamiccity: Large-scale 4d occu- pancy generation from dynamic scenes,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:14.916380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:14.916380Z digest=sha256:9049c7748559fc3e22fc67b2dcfb024fd790867fb32ae9fd95048680efea94e6

Observation 21947037-35b4-4109-83a4-c704eadff963 · outbound

This paper cites Calib3d: Calibrating model preferences for reliable 3d scene understanding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Calib3d: Calibrating model preferences for reliable 3d scene understanding,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:30.705023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:15.013655Z digest=sha256:f72f5fa3e064324f488b7e5d7b2676e91f4c055a7e80702ca6a791d174aa98f8

Observation 5c6df824-f55a-43df-b72b-d33033b5e4f1 · outbound

This paper cites Bottom up top down detection transform- ers for language grounding in images and point clouds,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Bottom up top down detection transform- ers for language grounding in images and point clouds,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:30.534444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:15.130922Z digest=sha256:f4233528bf660004d401d690600457e426bb3ffb3fd91ba0224637bdf8be5602

Observation bc2adfb0-b5bc-4cec-8761-c20b7c790c46 · outbound

This paper cites 3d-vista: Pre-trained transformer for 3d vision and text alignment,.

Zero-Shot 3D Visual Grounding from Vision-Language Models 3d-vista: Pre-trained transformer for 3d vision and text alignment,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:30.315718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:15.204457Z digest=sha256:61e525a851c73d26bae7ee4030e4736f27635aca11a0f69bcd3e0d13ec962648

Observation 7c0380bf-1754-4b5b-91dd-f9bc7301825e · outbound

This paper cites Eda: Explicit text-decoupling and dense align- ment for 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Eda: Explicit text-decoupling and dense align- ment for 3d visual grounding,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:30.141405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:15.273293Z digest=sha256:7fbb066f957ed31a9b26216b427f7819b62811de710e04253021cd13867f8a92

Observation d50e43f0-cf9a-41c7-b7b3-64936739d156 · outbound

This paper cites 3dvg-transformer: Relation modeling for vi- sual grounding on point clouds,.

Zero-Shot 3D Visual Grounding from Vision-Language Models 3dvg-transformer: Relation modeling for vi- sual grounding on point clouds,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:30.002253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:15.322438Z digest=sha256:a97738929870eabb616bb7b6c4ca632b432feeab77c815b4693372fb96d53092

Observation bd712acb-4ec8-4fd0-bc3f-3c2a28ce6d8f · outbound

This paper cites Instancerefer: Cooperative holistic under- standing for visual grounding on point clouds through in- stance multi-level contextual referring,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Instancerefer: Cooperative holistic under- standing for visual grounding on point clouds through in- stance multi-level contextual referring,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:29.817183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:15.400260Z digest=sha256:01688fa7fa246ebfc4c234b7cf6b5cc0fcafb7a9d1e428abdf7c487a3bb5d834

Observation f3b330c4-f015-42f9-86e6-d02a4f2ff0f9 · outbound

This paper cites Multi-branch collaborative learning network for 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Multi-branch collaborative learning network for 3d visual grounding,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:29.648448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:15.451650Z digest=sha256:4ece0aa712e47e412776f0a741ac4623b5590377ed7566e5b3ac4559b2ea95da

Observation 5f07f699-f8bc-441a-8085-dcc8ec68e53d · outbound

This paper cites Semantickitti: A dataset for semantic scene understanding of lidar sequences,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Semantickitti: A dataset for semantic scene understanding of lidar sequences,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:29.457192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:15.491354Z digest=sha256:df3c4d418fb9e517eec9c191ad2c389451e7b1d340462d6bf1a4fd6898879eed

Observation 2f3a2f58-fa72-4e45-9ff5-3861b4926bab · outbound

This paper cites Scalability in perception for autonomous driv- ing: Waymo open dataset,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Scalability in perception for autonomous driv- ing: Waymo open dataset,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:29.255587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:15.564839Z digest=sha256:358ff52f6f72dcf5326d572214d93ae2bdeb79e43e0c513c9cb9307c51bb6cdd

Observation 7354791c-6c9e-4f60-956a-af68e73e773c · outbound

This paper cites Panoptic nuscenes: A large-scale bench- mark for lidar panoptic segmentation and tracking,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Panoptic nuscenes: A large-scale bench- mark for lidar panoptic segmentation and tracking,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:29.002623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:15.623601Z digest=sha256:469f7c00f5725fdc6f44707dc6058459b5c8b77f844f0f0d10224c8cd09a3416

Observation 11375c3b-8862-4593-bad2-7761dfc1ac0e · outbound

This paper cites Visual programming for zero-shot open- vocabulary 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Visual programming for zero-shot open- vocabulary 3d visual grounding,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:28.773037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:15.701053Z digest=sha256:24e670c1035857ef0d854f322002301b54a7fe1598ba27557468085680ec2c43

Observation a012d746-72fe-44e3-bdeb-0d1aee487572 · outbound

This paper cites Llm-grounder: Open-vocabulary 3d visual grounding with large language model as an agent,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Llm-grounder: Open-vocabulary 3d visual grounding with large language model as an agent,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:28.545584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:15.776796Z digest=sha256:05d021920d095672abb8b888b45566d338cfe5fb5dfc888a9b82d7636b1495d2

Observation c155db64-fdc3-4454-8b88-6e86c1baba4a · outbound

This paper cites Training language models to fol- low instructions with human feedback,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Training language models to fol- low instructions with human feedback,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:28.325098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:15.878975Z digest=sha256:80650e7c1f23d14fe11f958ca13ecda33ea08cc145a105d3ca1097f9c04db786

Observation 1e1b7fb4-f508-42b6-8fab-a9bf6ea0c806 · outbound

This paper cites GPT-4 Technical Report.

Zero-Shot 3D Visual Grounding from Vision-Language Models GPT-4 Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:15.966014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:15.966014Z digest=sha256:4655780cd0c50a0d9896a4fe530ea0816b17acfb2131659dc320c3f602f9c5cb

Observation 3579b74e-7593-4f22-8ebb-e28f7c17fb6a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Zero-Shot 3D Visual Grounding from Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:16.033226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:16.033226Z digest=sha256:6ce0695f247cfc948cd3956540fdf0dd6bc232288467e96652b4efd1ad0b43d0

Observation 5426ae40-e7e6-4ba6-bb60-6b115f0d2f4a · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

Zero-Shot 3D Visual Grounding from Vision-Language Models CogVLM2: Visual Language Models for Image and Video Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:16.099070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:16.099070Z digest=sha256:1964ee4bb041fe120e69eb7fc30377a25751b77103a960de019edfdd508ecb80

Observation 63c38e4a-105a-4267-8616-ce139f488a22 · outbound

This paper cites Sceneverse: Scaling 3d vision-language learn- ing for grounded scene understanding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Sceneverse: Scaling 3d vision-language learn- ing for grounded scene understanding,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:28.062525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:16.172492Z digest=sha256:d27a74bbdc7a6d72c47568832553c47ea3bc95135c817cfdb65d978b2241f50c

Observation 37219787-af06-4101-8af5-57e158d3730c · outbound

This paper cites SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding.

Zero-Shot 3D Visual Grounding from Vision-Language Models SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:16.243728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:16.243728Z digest=sha256:e53642b58a7a330df8e4cab1a88fa62bc2d3d1c237d11b164ee11cf54ffe2c2a

Observation dd2be870-d9b6-4bae-b676-1b4ff513377a · outbound

This paper cites Referit3d: Neural listeners for fine- grained 3d object identification in real-world scenes,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Referit3d: Neural listeners for fine- grained 3d object identification in real-world scenes,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:27.868411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:16.327105Z digest=sha256:8b9f8070b451448ed561df344b4533210a571229da88f8d102cf7e65260fbfa0

Observation 1c3d62e2-a284-4093-9485-c07bfd12986b · outbound

This paper cites Viewrefer: Grasp the multi-view knowledge for 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Viewrefer: Grasp the multi-view knowledge for 3d visual grounding,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:27.642877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:16.383282Z digest=sha256:47dde7a6aa2718cfe0e7b3099d4ab0ed9bb5c8229a984393e2806fcb1657889f

Observation a70094c3-e7ce-4b6c-8fc8-cb4d1670b286 · outbound

This paper cites Multi-view transformer for 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Multi-view transformer for 3d visual grounding,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:27.299693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:16.457874Z digest=sha256:518ec0fb3917ce42b696631e0cecc27d8e944f332bb59fc8ca07fbd6c1ef0615

Observation 9229b3cc-9dcb-4273-9701-4f45767296f1 · outbound

This paper cites Look around and refer: 2d synthetic se- mantics knowledge distillation for 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Look around and refer: 2d synthetic se- mantics knowledge distillation for 3d visual grounding,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:27.103416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:16.537325Z digest=sha256:33216df4292f9d1fdf110d7750fd57d16014a24dd3355cc9c3b344995441a963

Observation d6ce677b-b1b6-4488-b020-89df092fe040 · outbound

This paper cites Sat: 2d semantics assisted training for 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Sat: 2d semantics assisted training for 3d visual grounding,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:26.794171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:16.586571Z digest=sha256:586db3ab32cb0d027d269e656976f1a5ce0a0aa7f68da744b728df80dc5a099a

Observation 1ca3c590-6dea-4a27-874d-747b1befea47 · outbound

This paper cites Four ways to improve verbo-visual fusion for dense 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Four ways to improve verbo-visual fusion for dense 3d visual grounding,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:26.612752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:16.657203Z digest=sha256:d576f4212252f34ce587f5948ffdd533e4041443589bc829cf7c74265afa462a

Observation 8b9e7394-52c4-4635-90bc-357d17e6ace1 · outbound

This paper cites Distilling coarse-to-fine semantic matching knowledge for weakly supervised 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Distilling coarse-to-fine semantic matching knowledge for weakly supervised 3d visual grounding,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:26.393051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:16.723540Z digest=sha256:8b61e265ea9f719900034c4bb95451e53f5a0923dff4b83dd102e732914a3f3c

Observation 8ad91c79-521c-4f81-8bfc-9e2577bc4909 · outbound

This paper cites Unifying 3d vision-language understanding via promptable queries,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Unifying 3d vision-language understanding via promptable queries,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:26.169590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:16.772423Z digest=sha256:8166c34112651cffec09377923e3380dd006a1ba071a9c3de500e0d2afc1f26a

Observation f858176d-4543-448f-94e2-5417467643c2 · outbound

This paper cites OVGaussian: Generalizable 3D Gaussian Segmentation with Open Vocabularies.

Zero-Shot 3D Visual Grounding from Vision-Language Models OVGaussian: Generalizable 3D Gaussian Segmentation with Open Vocabularies

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:16.828825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:16.828825Z digest=sha256:b0ed8a7b1786aea0472eacba998ab43bd3dc92d32bea6cfc5a76e3fb03bb8899

Observation b52141f5-02d0-4ff5-8e29-8998abfa1c46 · outbound

This paper cites Multi-space alignments towards universal lidar segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Multi-space alignments towards universal lidar segmentation,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:25.960756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:16.895168Z digest=sha256:7aac605e57023d491b9cd7e57f81ec2ae43276131ee3b255523bf7a76e0cfb7b

Observation 02ec2c3f-ff86-42ff-8ffb-34b28baecacf · outbound

This paper cites Towards label-free scene understanding by vision foundation models,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Towards label-free scene understanding by vision foundation models,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:25.756569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:16.967579Z digest=sha256:c55947486871c048c5a30b83d252637248a1c4c68d2cb9ced11cb942a94b789a

Observation 72357686-fe9d-482c-a614-4c8e5f09b4b6 · outbound

This paper cites LiMoE: Mixture of LiDAR Representation Learners from Automotive Scenes.

Zero-Shot 3D Visual Grounding from Vision-Language Models LiMoE: Mixture of LiDAR Representation Learners from Automotive Scenes

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:17.070727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:17.070727Z digest=sha256:4f6be2d246966312e53c47f9814e181e854015ab8a87773b8fd0e512f61594c8

Observation 197f289d-d06e-4527-8ea0-b7333f8cc82a · outbound

This paper cites GEAL: Generalizable 3D Affordance Learning with Cross-Modal Consistency.

Zero-Shot 3D Visual Grounding from Vision-Language Models GEAL: Generalizable 3D Affordance Learning with Cross-Modal Consistency

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:17.150193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:17.150193Z digest=sha256:8b1e6b7d10fffd8a159681efc3b238b68e5fae5fe11555366cb650d1d7b8e534

Observation 030e2204-a25d-4b24-a0b4-3397826df452 · outbound

This paper cites Openscene: 3d scene understanding with open vocabularies,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Openscene: 3d scene understanding with open vocabularies,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:25.487986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:17.213896Z digest=sha256:bba1f0fb7fe53698d08a0fed3f5b5ded2efca32e24617e81edb662cbf6aa373b

Observation 423967c6-2e0e-4c80-bcf4-565da2f9c4d3 · outbound

This paper cites Lerf: Language embedded radiance fields,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Lerf: Language embedded radiance fields,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:25.225323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:17.283351Z digest=sha256:f56f539498ea3bb76c5be5be01c5f4f6e5bcd93f6587f1d2a9c190bc2f07c242

Observation 94103a13-4e1b-43a6-bbc2-ed6e9e18a290 · outbound

This paper cites Ovir-3d: Open-vocabulary 3d instance retrieval without training on 3d data,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Ovir-3d: Open-vocabulary 3d instance retrieval without training on 3d data,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:25.019581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:17.343736Z digest=sha256:9f0b4f1fc03cab0742ee999154652f0f904803d4f438a138a80a1e3c46fed539

Observation 1f8f6e63-38ef-47e1-98da-a0ea5409aeeb · outbound

This paper cites Agent3D-Zero: An Agent for Zero-shot 3D Understanding.

Zero-Shot 3D Visual Grounding from Vision-Language Models Agent3D-Zero: An Agent for Zero-shot 3D Understanding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:17.411309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:17.411309Z digest=sha256:0e98e90e3fddfa824a37ec53a3443bdbe37aa7da967ff0255a7f2b00b6e024b9

Observation 1d2c5493-8c26-42f6-90f4-97d8c35ede01 · outbound

This paper cites Regionplc: Regional point-language con- trastive learning for open-world 3d scene understanding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Regionplc: Regional point-language con- trastive learning for open-world 3d scene understanding,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:24.776061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:17.470353Z digest=sha256:2ebfadb5ff92b23b2258e00a630ebfdfa0838ddf8a0e6144ff5ad3b9c41b40b5

Observation 739b960d-0787-420d-8179-33da110aa4cf · outbound

This paper cites OpenMask3D: Open-Vocabulary 3D Instance Segmentation.

Zero-Shot 3D Visual Grounding from Vision-Language Models OpenMask3D: Open-Vocabulary 3D Instance Segmentation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:17.523576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:17.523576Z digest=sha256:c46470546e72cda6b38f8f555f331cc3e9e8c32a2ac31e89972a9614b3e14c1b

Observation 9fff50e1-dbfb-428b-bd79-ed7ded1e7bfd · outbound

This paper cites Openins3d: Snap and lookup for 3d open- vocabulary instance segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Openins3d: Snap and lookup for 3d open- vocabulary instance segmentation,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:24.563327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:17.579259Z digest=sha256:0e44df619ee10ca4e35a9021c296353950568485ff982d579350202d61ec07a2

Observation 6b063554-7f08-45ac-8735-ec16104d3204 · outbound

This paper cites Sai3d: Segment any instance in 3d scenes,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Sai3d: Segment any instance in 3d scenes,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:24.321836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:17.637874Z digest=sha256:068753c3f92a6428a8fd3854055d99b614be44936240e3303c6df4587b33a498

Observation b8e8a98d-4d75-4d7e-859e-cbf1efa418be · outbound

This paper cites Lasermix for semi-supervised lidar semantic segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Lasermix for semi-supervised lidar semantic segmentation,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:24.131536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:17.710770Z digest=sha256:9bc0cc25a44b3cecbf189a3b98738956809bf0561af2b29416752122c1130343

Observation 88a4132e-37ac-4dc3-aec5-a64dc8004e8f · outbound

This paper cites Segment any point cloud sequences by distilling vision foundation models,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Segment any point cloud sequences by distilling vision foundation models,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:23.906258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:17.807087Z digest=sha256:3d99ad3cce673c6411d8b18ff179ec84bbc9a9202bdc701df345cdf934a90d8d

Observation c44806ef-9311-4844-9525-3cd2395c8625 · outbound

This paper cites 4d contrastive superflows are dense 3d repre- sentation learners,.

Zero-Shot 3D Visual Grounding from Vision-Language Models 4d contrastive superflows are dense 3d repre- sentation learners,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:23.683419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:17.935512Z digest=sha256:bb5bbd8ffdcf58f043de60d03fb74883a546425e2a8e35f26b6a28ab8e2906ff

Observation 1c3f78e0-9855-4610-a55d-3cdbde6b628c · outbound

This paper cites Frnet: Frustum-range networks for scalable lidar segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Frnet: Frustum-range networks for scalable lidar segmentation,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:23.484256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:18.034869Z digest=sha256:d21ac2cc24e74f008fd5da03490528f73f52951342aedfc724a85ef73cd655fd

Observation c4b43a23-36bf-481c-ae1a-d60fab1e0d84 · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

Zero-Shot 3D Visual Grounding from Vision-Language Models Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:18.091406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:18.091406Z digest=sha256:3b049de3ad32f8a88663fee9d8442d683f11819b5f11bca2cd1c522ce31850e6

Observation b469ac04-b778-4bfe-b0ae-85bc252010b2 · outbound

This paper cites Uni3DL: Unified Model for 3D and Language Understanding.

Zero-Shot 3D Visual Grounding from Vision-Language Models Uni3DL: Unified Model for 3D and Language Understanding

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:12:20.164538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:18.156553Z digest=sha256:5fb76cd816e3efd94bf5b9fb6f392f45a72369dbdba55eff48f87930243552b4

Observation 6171bf57-8155-4ef5-bc3d-97e28edb525c · outbound

This paper cites Conceptfusion: Open-set multi- modal 3d mapping,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Conceptfusion: Open-set multi- modal 3d mapping,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:23.283686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:18.254570Z digest=sha256:54b928e2bc739149260a93eb475213bd6eae5e87f7e8e6235d3e095e0fe6a195

Observation b1e0260d-814e-4607-846a-a2b4f28ac11a · outbound

This paper cites GLOVER: Generalizable Open-Vocabulary Affordance Reasoning for Task-Oriented Grasping.

Zero-Shot 3D Visual Grounding from Vision-Language Models GLOVER: Generalizable Open-Vocabulary Affordance Reasoning for Task-Oriented Grasping

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:18.478933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:18.478933Z digest=sha256:2c84e70f8a9ef68c49f393baca31bb0201bf101d3b70eeade50997b19e5d174b

Observation 11fb35db-9415-452a-8a43-1742b926fdb6 · outbound

This paper cites Interactive planning using large language models for partially observable robotic tasks,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Interactive planning using large language models for partially observable robotic tasks,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:23.161734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:18.706597Z digest=sha256:1aa1859c8272c9a899e42bf36d2935ea3ed10fb9ff4bbd0f18cc7b50ac5996ba

Observation fce09665-8e57-4ecf-ac78-7d026fcc53c8 · outbound

This paper cites 3d-llm: Injecting the 3d world into large language models,.

Zero-Shot 3D Visual Grounding from Vision-Language Models 3d-llm: Injecting the 3d world into large language models,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:22.964761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:18.817427Z digest=sha256:2bee1235911a91e19e43876ab73cbcc9155400a9382a423ecd09c7878e926c36

Observation 5fcd2776-fd95-49b2-b879-6e63ed38e8fd · outbound

This paper cites Is your lidar placement optimized for 3d scene understanding?,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Is your lidar placement optimized for 3d scene understanding?,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:22.736693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:18.907550Z digest=sha256:da106a5d320499ae1d9b0b4066ace84a3db3cdddc1e639c88493bfa2d5f93896

Observation 7343377b-7b76-4d17-b048-d0c15d351101 · outbound

This paper cites G3-lq: Marrying hyperbolic alignment with explicit semantic-geometric modeling for 3d visual ground- ing,.

Zero-Shot 3D Visual Grounding from Vision-Language Models G3-lq: Marrying hyperbolic alignment with explicit semantic-geometric modeling for 3d visual ground- ing,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:22.557138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:18.980280Z digest=sha256:be781b1ad9865021fbaadae341fe0178e4fcd1fe8fe9a46f3a8c19a00a8ed49c

Observation 3f863d07-94ef-4de9-ad28-d001bf53d84e · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Learning transferable visual models from natural language supervision,

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:22.328649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:19.127295Z digest=sha256:6fee8e83ebfd18773358f73b82ba72384ac72e48f48c65829755322f1472cbfe

Observation edcfabd1-07f9-4cd5-b156-2823a5aa0bc2 · outbound

This paper cites VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding.

Zero-Shot 3D Visual Grounding from Vision-Language Models VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:19.231195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:19.231195Z digest=sha256:1897634831e57681bbc46512deac5d1786cb0a448bb7cf6dc77f9086818e6dc5

Observation 6bb7d6ca-a924-4237-8f5e-d5f833206cbb · outbound

This paper cites Text-guided graph neural networks for referring 3d instance segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Text-guided graph neural networks for referring 3d instance segmentation,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:22.080070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:19.353586Z digest=sha256:5eda1a6e973c4b3dec4883f5f404032dcb7f22dcc2ec949cc835dfe5e46b88f2

Observation 4761a55c-7a2e-4aa9-9ef1-8cf0368ce164 · outbound

This paper cites Mikasa: Multi-key-anchor & scene- aware transformer for 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Mikasa: Multi-key-anchor & scene- aware transformer for 3d visual grounding,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:21.837968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:19.499102Z digest=sha256:a965878f89fabc383a4fabcd760acd9fb62135b33b422c411411ad658394c4e8

Observation 0de609bf-af3c-4a3d-b24c-fc6057d60b72 · outbound

This paper cites Language conditioned spatial relation rea- soning for 3d object grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Language conditioned spatial relation rea- soning for 3d object grounding,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:21.636127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:19.626429Z digest=sha256:3d876812966f2e81887c5db8ac53236f14a3bab2d4547a0afae33f83a121966c

Observation b9c59bea-69c3-4075-8a2c-08946499f3a6 · outbound

This paper cites Mask3d: Mask transformer for 3d semantic instance segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Mask3d: Mask transformer for 3d semantic instance segmentation,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:21.510030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:12:19.747542Z digest=sha256:8f0ad328f5ff4ec4999f33aad36b0758d8a9ba0206a4ffb18ba8500f016f2ac8

Pith citing papers

Observation 7e135835-ad52-43f6-9e78-62ef14a3da16 · inbound

PruneGround: Plug-and-play Spatial Pruning for 3D Visual Grounding cites this paper.

PruneGround: Plug-and-play Spatial Pruning for 3D Visual Grounding Zero-Shot 3D Visual Grounding from Vision-Language Models

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-07-01T06:45:29.444056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-07-01T06:41:38.649424Z digest=sha256:001e04ee55b57b6b6925aa6a36265c101b456fe452457bca7d5aeb63816f8c67