Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:58:28.684524Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 98 of 98 outbound references and 0 inbound Pith citation observations for arXiv:2504.19500.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:58:28.684524Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
98 of 98 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 38265fd2-bf84-4720-96fa-a5b0ea65c3b5 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7af9f3da-38f5-4351-adda-614c0e9f1eb5 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1129ef08-b7be-4e72-9665-e905db17e6b2 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Arkitscenes: A diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f562cf5e-636a-4c0b-b3ab-0d724ee87149 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding 3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e1a7024-04a5-46b2-a928-2df69eb28251 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Matterport3D: Learning from RGB-D Data in Indoor Environments
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f243a64a-31d8-4373-89c8-3621803e3f75 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Scanrefer: 3d object localization in rgb-d scans using natural language
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3ecf1a9-bb75-4eac-8699-52c42db0fe19 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Clip2scene: Towards label-efficient 3d scene under- standing by clip
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bdb4a93-88d3-41b1-8541-cde6b9893e4d · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Scan2cap: Context-aware dense captioning in rgb-d scans
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bae31890-7edd-4389-bb9b-24c44fe48eaf · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Unit3d: A unified transformer for 3d dense captioning and visual grounding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82d1f0b0-fe58-4cd9-b00a-feb4a2712f9a · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Cat-seg: Cost aggregation for open-vocabulary semantic segmentation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae81c91b-f8dd-4f11-842a-efd0bbdbde31 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding 4d spatio-temporal convnets: Minkowski convolutional neural networks
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a3fe26b-3a03-4195-8ac8-c442519e6c38 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Pointcept: A codebase for point cloud perception research
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a62495bc-7196-446e-bccd-5b26d4b46df1 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Spconv: Spatially sparse convolu- tion library
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5963e24-c8fe-4519-b45d-98e65533c850 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Scannet: Richly- annotated 3d reconstructions of indoor scenes
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f39a7d42-c843-4e03-b11b-1ee301b2097c · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Procthor: Large-scale embodied ai using procedural generation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afccfc09-1231-4ba9-a200-5ef83d94d643 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdb49ad8-7c48-4a96-a0b9-25dccc7d53d6 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Pla: Language-driven open-vocabulary 3d scene understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d370514-93bd-4b7f-8b55-b7057e849b38 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ec085f1-155b-4afa-bb46-02e593b72e16 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Foundation models in robotics: Applications, challenges, and the fu- ture.The International Journal of Robotics Research, page 02783649241281508, 2023
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8fa74c8-7c73-4319-9885-0f9a1f250a74 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 113ea330-bde1-4bf4-82b9-dbdbc6b2b2ad · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Scaling open-vocabulary image segmentation with image-level labels
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ad92dc0-61ae-4bf1-b538-acceed9a0f69 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding 3d semantic segmentation with submanifold sparse convolutional networks
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 372c28a3-b51a-4ea6-9ff5-6384925891ee · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Concept- graphs: Open-vocabulary 3d scene graphs for perception and planning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 582c9234-5cd1-4879-b6e5-f2e013dbc122 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0eecbbc-e8c3-4072-8d78-529428d28ee4 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Masked autoencoders are scalable vision learners
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 594ef030-55c7-4a45-a9fa-a3be8db2e5d7 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding 3d-llm: Injecting the 3d world into large language models.NIPS, 36:20482– 20494, 2023
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5fa8346b-1306-4eb7-84e3-367c51b2094d · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding 3d-llm: Injecting the 3d world into large language models.NIPS, 36:20482– 20494, 2023
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e6acafed-ca28-4cc4-8994-295ac792e9b1 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Multiply: A multisensory object- centric embodied large language model in 3d world
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cb1c4f3d-9dac-49bd-a38e-d5fe81667875 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Exploring data-efficient 3d scene understanding with 9 contrastive scene contexts
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c3347c05-161b-4a7a-a0a2-1faf833703b1 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Chat-scene: Bridging 3d scene and large language models with object identifiers
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c142d170-516e-4582-b7a4-dc9b33389961 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding An Embodied Generalist Agent in 3D World
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa3ecb45-f31e-42e1-80e0-522fbace92c0 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Unveiling the mist over 3d vision-language under- standing: Object-centric evaluation with chain-of-analysis
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 70ea3f8e-dc5b-481c-b524-e9134698b88d · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Spatio-temporal self-supervised representation learning for 3d point clouds
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c2ea155b-0282-4f18-978a-81bc7ffe347f · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Openins3d: Snap and lookup for 3d open-vocabulary instance segmentation
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1e2b678d-1710-4cda-abf8-f2a3d9957f97 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61084017-02c1-4262-82c2-96f5b4016848 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Scaling up visual and vision-language representation learning with noisy text supervision
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aacf438-0730-4f42-b12e-d41953f1f5fe · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Pointgroup: Dual-set point grouping for 3d instance segmentation
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 34bed0b6-56f7-4cba-8cd6-b1201b1dbd7a · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Open-vocabulary 3d semantic segmentation with foundation models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0fd09de1-f7aa-4287-8c72-5ba5e4390ef3 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Robin3D: Improving 3D Large Language Model via Robust Instruction Tuning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35f40d81-70a7-4c1b-a238-fb1638c67e70 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Segment any- thing
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a4d29b7b-f67e-4039-9868-c3deb215616d · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Language-driven Semantic Segmentation
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9de54dd7-6ae3-4056-b34b-b9ea586d3c8d · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Maniptrans: Efficient dexterous bimanual manipula- tion transfer via residual learning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5d9069cf-f25c-4a11-90ce-35b19d493b9d · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Dense multimodal alignment for open-vocabulary 3d scene understanding
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 688f14b2-f57c-4dc0-8e2a-115204457f8c · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Visual instruction tuning, 2023
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e607342-8aa6-4f45-ae19-cda0499f74ee · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Building interactable replicas of complex articulated objects via gaussian splatting
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c6e99081-f65f-4393-a427-f06713ed89dd · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Movis: Enhancing multi-object novel view synthesis for indoor scenes
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 093762cf-f983-484f-a4f2-8698fcd957e8 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Ovir-3d: Open-vocabulary 3d instance retrieval without training on 3d data
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 91930c99-fbf9-4179-abcd-b596c95a3ee6 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding SQA3D: Situated Question Answering in 3D Scenes
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 730252d7-a215-49ee-a4ca-d892a4913615 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding When llms step into the 3d world: A survey and meta-analysis of 3d tasks via multi-modal large language models.arXiv preprint arXiv:2405.10255, 2024
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6912e089-9502-4e95-aa27-ee1993a738d3 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Multiscan: Scalable rgbd scanning for 3d environments with articulated objects
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9194ca88-c5a0-42ab-be8d-cf5e3fdaa6db · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Occupancy-mae: Self-supervised pre-training large- scale lidar point clouds with masked occupancy autoencoders
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9f32ff46-4822-446f-8071-bcfbe9a42857 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Open3dis: Open-vocabulary 3d instance segmentation with 2d mask guidance
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc28536c-1b91-4aae-930c-617f0eaffacc · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Phyrecon: Physically plausible neural scene recon- struction
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4091d38f-f5dc-43f4-852f-4dd29d617813 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Decompositional neural scene reconstruction with generative diffusion prior
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 39eff41d-26a6-4f37-bfbc-33368219cf26 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Open X-Embodiment: Robotic Learning Datasets and RT-X Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33105af5-5026-44b2-be79-73cf95d79b90 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Gpt-4 with vision (gpt-4v) system card, 2023
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ae3d28a7-a157-4660-89f5-16801ab81e64 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Openscene: 3d scene understanding with open vocabularies
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3514242c-189b-4bcd-bc05-da6a5b38b782 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Shapellm: Universal 3d object understanding for embodied interaction
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ace3405f-dfcc-4ab0-a731-891d66dbcd05 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Learning transferable visual models from natural language supervision
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 586de2a0-e055-4eea-82fc-c8877e04dfda · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ae228f7-2896-4632-a111-684dc6cc0104 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Language- grounded indoor 3d semantic segmentation in the wild
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8f9f7f0e-c425-4a0a-8179-3a8ccf316d7f · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Mask3d: Mask trans- former for 3d semantic instance segmentation
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e28baa2c-f938-4f11-bb3c-0c7783fa0351 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Open- mask3d: open-vocabulary 3d instance segmentation
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5ea2272d-ba49-492e-986b-bb6386fa7a91 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Minigpt-3d: Efficiently aligning 3d point clouds with large language models using 2d priors
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 02010a7b-12b0-40bd-969f-927ecd489d55 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Geomae: Masked geometric target prediction for self-supervised point cloud pre-training
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 49319da8-1243-4d6d-bd92-2e87331dc8d9 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Rio: 3d object instance re- localization in changing indoor environments
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 95ed88bf-6851-408c-80c2-e2c66443995a · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Groupcontrast: Semantic-aware self-supervised representation learning for 3d understanding
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c0346087-45f8-4c33-a945-cbd6859c46c9 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Open vocabulary 3d scene understanding via geometry guided self-distillation
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 76d9c71c-1a21-49ef-acf2-7581fad9fca2 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Embodiedscan: A holistic multi- modal 3d perception suite towards embodied ai
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21d9f7da-6e7b-4e93-a61c-0013f03eac27 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding T-MAE: Temporal Masked Autoencoders for Point Cloud Representation Learning
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e646850d-569a-4142-ad03-301e36e20fab · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d82e6fe-eefa-401f-bad3-276b0dcb4e4e · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Masked scene contrast: A scalable framework for unsuper- vised 3d representation learning
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 048627c5-2f22-4c35-ab1e-10b4d916d323 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Sed: A simple encoder-decoder for open-vocabulary semantic segmentation
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 414850d3-e67e-4afa-91be-4c5a5c20aa32 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Pointcontrast: Unsupervised pre- training for 3d point cloud understanding
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 65bcb96e-43bf-4878-a2bb-8704191bd398 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Simmim: A simple framework for masked image modeling
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4b521be6-91bb-449e-b4e4-a1983d23d469 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding SAMPro3D: Locating SAM Prompts in 3D for Zero-Shot Instance Segmentation
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2ef24f6-82f9-4f86-82eb-3d7139be65fd · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Maskclus- tering: View consensus based mask graph clustering for open- vocabulary 3d instance segmentation
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 51be96d3-3c84-41b3-a576-df7cf73126f6 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding 3D Vision and Language Pretraining with Large-Scale Synthetic Data
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1db6b825-c282-48f1-ae6a-5f986b6b4a3a · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Gd-mae: gener- ative decoder for mae pre-training on lidar point clouds
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c867954f-d97d-4a02-9867-858d4c34012e · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding 3D-GRAND: A Million-Scale Dataset for 3D-LLMs with Better Grounding and Less Hallucination
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d03856e-9818-46a7-a458-b49748010c4b · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Regionplc: Regional point-language contrastive learning for open-world 3d scene understanding
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ffeb84ee-446c-4192-84aa-2aba3d2f0704 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding SAM3D: Segment Anything in 3D Scenes
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc681d9b-ced2-466f-be6e-78865afb9a16 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Metascenes: Towards automated replica creation for real-world 3d scans
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bd840b72-6f36-4b46-9177-058e9f49875b · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Point-bert: Pre-training 3d point cloud transformers with masked point modeling
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2f04313e-6117-4f6a-b617-739a653158ae · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Clip2: Contrastive language-image- point pretraining from real-world point cloud data
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b6b0a255-c531-48da-a783-950469e10e30 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Vision-language pre-training with object contrastive learning for 3d scene understanding
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 47a460c8-59af-407b-955b-4b9d46f13f22 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Multi3drefer: Grounding text description to multiple 3d ob- jects
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 67577ae0-edce-4ec2-9339-d407355d4165 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Self-supervised pretraining of 3d features on any point-cloud
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 159917a2-244a-47dc-afd8-8c86931c3136 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding 3D-VLA: A 3D Vision-Language-Action Generative World Model
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8e42cae-a947-4fa0-ae2d-91ff5174f736 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Structured3d: A large photo-realistic dataset for structured 3d modeling
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34df5384-c2b2-43e5-b637-b368bd378e4e · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Closed-Loop Open-Vocabulary Mobile Manipulation with GPT-4V
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a19a8cf6-9e8b-4835-8b33-a52bee2c0729 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Extract free dense labels from clip
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b818cc9b-dc45-458b-a33f-35c5b4cfde2f · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Uni3D: Exploring Unified 3D Representation at Scale
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad00c17a-0951-4836-86e7-8fd63036a62a · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Detecting twenty-thousand classes using image-level supervision
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b56dc2ec-4b53-482b-aa78-8b770496dfb8 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c3cb4cf-79a4-4f90-9499-8c9ddaafe77e · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding 3d-vista: Pre-trained transformer for 3d vision and text alignment
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2c389b2c-4926-4c6a-a8a3-b04307b12597 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Unifying 3D Vision-Language Understanding via Promptable Queries
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de0e1657-967c-477f-8c07-6cb2d9b273b3 · outbound
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Unresolved cited work
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
No inbound Pith citation observations are available.