Pith. sign in

Paper Citation Record · LEDGER

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding

As of 18 August 2026, this Paper Citation Record lists 98 of 98 outbound references and 0 inbound Pith citation observations for arXiv:2504.19500.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.19500 v1

Coverage vector

measured 98 of 98 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:58:28.684524Z

measured 98 of 98 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

98 of 98 outbound references displayed

  • verified exact1
  • verified fuzzy48
  • unresolved47
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 38265fd2-bf84-4720-96fa-a5b0ea65c3b5 · outbound

This paper cites Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:27.064439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:27.064439Z digest=sha256:5317e0229acf647976a4b796bc220efd989daf37456f384a415a2ea44dfd337b

Observation 7af9f3da-38f5-4351-adda-614c0e9f1eb5 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:27.068692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:27.068692Z digest=sha256:181066760a6ac11647f35e4306131c8ca2e56a8eb4431e0b65d0596ab5587429

Observation 1129ef08-b7be-4e72-9665-e905db17e6b2 · outbound

This paper cites Arkitscenes: A diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Arkitscenes: A diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:27.072585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:27.072585Z digest=sha256:d5ded1a8261e816a5e371e781b39d84ed91ffb9f617c76acd28604a36977e5ae

Observation f562cf5e-636a-4c0b-b3ab-0d724ee87149 · outbound

This paper cites 3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding 3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:27.077595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:27.077595Z digest=sha256:f320d2492f00ca7df90898a86149e1bae1d4c0bbd54ab394626b49c968010680

Observation 0e1a7024-04a5-46b2-a928-2df69eb28251 · outbound

This paper cites Matterport3D: Learning from RGB-D Data in Indoor Environments.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Matterport3D: Learning from RGB-D Data in Indoor Environments

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:27.082358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:27.082358Z digest=sha256:01db3d2ffa7e19a0b716bea32967bf091f083ee2b177ae9d0058222d78c3dcc1

Observation f243a64a-31d8-4373-89c8-3621803e3f75 · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Scanrefer: 3d object localization in rgb-d scans using natural language

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:27.086644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:27.086644Z digest=sha256:b79ec95565a284fc5e5bfa216f6889714ba789f297d9057044bcfdf90cc00bb8

Observation b3ecf1a9-bb75-4eac-8699-52c42db0fe19 · outbound

This paper cites Clip2scene: Towards label-efficient 3d scene under- standing by clip.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Clip2scene: Towards label-efficient 3d scene under- standing by clip

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:27.090601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:27.090601Z digest=sha256:bf4edf7c38033043bbc89d7593d361fc81804801d2b43eef3d96fdc832e2b606

Observation 5bdb4a93-88d3-41b1-8541-cde6b9893e4d · outbound

This paper cites Scan2cap: Context-aware dense captioning in rgb-d scans.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Scan2cap: Context-aware dense captioning in rgb-d scans

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:27.095168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:27.095168Z digest=sha256:4646e712b81d6fecbd6b6730d07666a55cc3b0713da741d2e6d1072f7ce14ca1

Observation bae31890-7edd-4389-bb9b-24c44fe48eaf · outbound

This paper cites Unit3d: A unified transformer for 3d dense captioning and visual grounding.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Unit3d: A unified transformer for 3d dense captioning and visual grounding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:27.134760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:27.134760Z digest=sha256:16867dc95d7e329ab0698cb410c1bd6b021a7870aa9c085003162807497595eb

Observation 82d1f0b0-fe58-4cd9-b00a-feb4a2712f9a · outbound

This paper cites Cat-seg: Cost aggregation for open-vocabulary semantic segmentation.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Cat-seg: Cost aggregation for open-vocabulary semantic segmentation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:27.207037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:27.207037Z digest=sha256:ee60ca49d235f3eb33a5f37967d5fa981ad483fa90f70af691d5394ed5f5c062

Observation ae81c91b-f8dd-4f11-842a-efd0bbdbde31 · outbound

This paper cites 4d spatio-temporal convnets: Minkowski convolutional neural networks.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding 4d spatio-temporal convnets: Minkowski convolutional neural networks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:27.357563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:27.357563Z digest=sha256:2277090f5e6632c173cffa6c9686c15f344a4eee9dc1a8669956b3814aa4c8a1

Observation 7a3fe26b-3a03-4195-8ac8-c442519e6c38 · outbound

This paper cites Pointcept: A codebase for point cloud perception research.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Pointcept: A codebase for point cloud perception research

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:27.395473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:27.395473Z digest=sha256:c3bf2688dcab1684315edec9f8070b94ad89bb161f326e5d3a3794dd79cb559a

Observation a62495bc-7196-446e-bccd-5b26d4b46df1 · outbound

This paper cites Spconv: Spatially sparse convolu- tion library.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Spconv: Spatially sparse convolu- tion library

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:27.399178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:27.399178Z digest=sha256:ae4e1b658653b8328045302ba7c0850bbce3fb625a96dceebcee1b0c9f968a31

Observation a5963e24-c8fe-4519-b45d-98e65533c850 · outbound

This paper cites Scannet: Richly- annotated 3d reconstructions of indoor scenes.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Scannet: Richly- annotated 3d reconstructions of indoor scenes

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:27.402574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:27.402574Z digest=sha256:1ec26e130a80137dbe4b9881f79c5b4fa8019c084ffb0bc334a6c5ed9ef06844

Observation f39a7d42-c843-4e03-b11b-1ee301b2097c · outbound

This paper cites Procthor: Large-scale embodied ai using procedural generation.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Procthor: Large-scale embodied ai using procedural generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:27.405998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:27.405998Z digest=sha256:3d4bd2a1f813ef220b5be107d2d79f93f6528cc6007cb871a2028c99fa5f9766

Observation afccfc09-1231-4ba9-a200-5ef83d94d643 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:27.409042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:27.409042Z digest=sha256:dcce52466ff18ac4c7430d15b5d08755b4be98cbad3ced4e08f69306004a0863

Observation fdb49ad8-7c48-4a96-a0b9-25dccc7d53d6 · outbound

This paper cites Pla: Language-driven open-vocabulary 3d scene understanding.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Pla: Language-driven open-vocabulary 3d scene understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:27.412739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:27.412739Z digest=sha256:9f891ed3991b969572c78fc3a0825ec1865de4a3e829897646e093191db89a83

Observation 1d370514-93bd-4b7f-8b55-b7057e849b38 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:27.416304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:27.416304Z digest=sha256:54fd92ea3ae1d6671722ae5d366f7f7d806d1a263f4fe8bafe17ae7ba4ad2a84

Observation 9ec085f1-155b-4afa-bb46-02e593b72e16 · outbound

This paper cites Foundation models in robotics: Applications, challenges, and the fu- ture.The International Journal of Robotics Research, page 02783649241281508, 2023.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Foundation models in robotics: Applications, challenges, and the fu- ture.The International Journal of Robotics Research, page 02783649241281508, 2023

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:27.420081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:27.420081Z digest=sha256:45e9111797b507a73e18664b0129895c1af2570f009e19408e5f6cf01f07b1f9

Observation f8fa74c8-7c73-4319-9885-0f9a1f250a74 · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:27.422768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:27.422768Z digest=sha256:455850856d78e94c85dddecedf92187ba0f3bf01a1692fa5eccc64aa1c9ac9e5

Observation 113ea330-bde1-4bf4-82b9-dbdbc6b2b2ad · outbound

This paper cites Scaling open-vocabulary image segmentation with image-level labels.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Scaling open-vocabulary image segmentation with image-level labels

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:27.426532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:27.426532Z digest=sha256:b59a879ed8f0f089794757a6ddc6d3e335042a0c0e2d680588e58972e2f0aee1

Observation 8ad92dc0-61ae-4bf1-b538-acceed9a0f69 · outbound

This paper cites 3d semantic segmentation with submanifold sparse convolutional networks.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding 3d semantic segmentation with submanifold sparse convolutional networks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:27.429715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:27.429715Z digest=sha256:a0b66ee8ba5418d3968dbd5cfd5a8283112127f36af4e3af9ae1575d9c953dfc

Observation 372c28a3-b51a-4ea6-9ff5-6384925891ee · outbound

This paper cites Concept- graphs: Open-vocabulary 3d scene graphs for perception and planning.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Concept- graphs: Open-vocabulary 3d scene graphs for perception and planning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:27.432839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:27.432839Z digest=sha256:9650ea83e64926d7165395ca2fbe1e9fbe33e32305a3a496e79c85f8de9b7434

Observation 582c9234-5cd1-4879-b6e5-f2e013dbc122 · outbound

This paper cites Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:27.435522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:27.435522Z digest=sha256:46cb6c021c756d6aca36eb58a22a466ae07e53568d303acf96f7ec44ffbd56cd

Observation f0eecbbc-e8c3-4072-8d78-529428d28ee4 · outbound

This paper cites Masked autoencoders are scalable vision learners.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Masked autoencoders are scalable vision learners

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:31.090300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:27.439119Z digest=sha256:7c01ccdf243ca8078693dd05ae9c6c5a5fa821367771c4e1e01630eccff49ddf

Observation 594ef030-55c7-4a45-a9fa-a3be8db2e5d7 · outbound

This paper cites 3d-llm: Injecting the 3d world into large language models.NIPS, 36:20482– 20494, 2023.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding 3d-llm: Injecting the 3d world into large language models.NIPS, 36:20482– 20494, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:31.078766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:27.442536Z digest=sha256:dbe5fef8c033cf25898255079c7a8276a29e30ef0b754959f92894f1e17b4558

Observation 5fa8346b-1306-4eb7-84e3-367c51b2094d · outbound

This paper cites 3d-llm: Injecting the 3d world into large language models.NIPS, 36:20482– 20494, 2023.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding 3d-llm: Injecting the 3d world into large language models.NIPS, 36:20482– 20494, 2023

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:31.067969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:27.490563Z digest=sha256:59e4ff432432db9dfa401582f7ea970cfd355dce2195f960a67b87f8836a9696

Observation e6acafed-ca28-4cc4-8994-295ac792e9b1 · outbound

This paper cites Multiply: A multisensory object- centric embodied large language model in 3d world.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Multiply: A multisensory object- centric embodied large language model in 3d world

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:31.056964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:27.538214Z digest=sha256:e6f745b587a1144515aa145806c97e98718ee0847256936e2e6fab854bfc8aca

Observation cb1c4f3d-9dac-49bd-a38e-d5fe81667875 · outbound

This paper cites Exploring data-efficient 3d scene understanding with 9 contrastive scene contexts.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Exploring data-efficient 3d scene understanding with 9 contrastive scene contexts

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:31.044004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:27.602360Z digest=sha256:e8d153073969bcdc86757c8df931890006b259b55718c1cb213fb8f6f62796af

Observation c3347c05-161b-4a7a-a0a2-1faf833703b1 · outbound

This paper cites Chat-scene: Bridging 3d scene and large language models with object identifiers.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Chat-scene: Bridging 3d scene and large language models with object identifiers

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:30.836791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:27.695904Z digest=sha256:bd6498653cbcb46f6758841d11e081f292e6f1593adee0322dca3692a05605fb

Observation c142d170-516e-4582-b7a4-dc9b33389961 · outbound

This paper cites An Embodied Generalist Agent in 3D World.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding An Embodied Generalist Agent in 3D World

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:27.699704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:27.699704Z digest=sha256:2342b5a6672c39289412fb7eaa413ba21017366b69a3eb5932696d87c61d3466

Observation fa3ecb45-f31e-42e1-80e0-522fbace92c0 · outbound

This paper cites Unveiling the mist over 3d vision-language under- standing: Object-centric evaluation with chain-of-analysis.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Unveiling the mist over 3d vision-language under- standing: Object-centric evaluation with chain-of-analysis

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:30.827381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:27.702747Z digest=sha256:20f7fab4d138385a0c533f9e5627867aad7a846e7e6a71c4d0075e7f69628715

Observation 70ea3f8e-dc5b-481c-b524-e9134698b88d · outbound

This paper cites Spatio-temporal self-supervised representation learning for 3d point clouds.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Spatio-temporal self-supervised representation learning for 3d point clouds

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:30.815516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:27.706285Z digest=sha256:e6893df44cefe1a837ac49c9463f031285ac4517c3b5ee7e56bd89ea6cc56cd1

Observation c2ea155b-0282-4f18-978a-81bc7ffe347f · outbound

This paper cites Openins3d: Snap and lookup for 3d open-vocabulary instance segmentation.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Openins3d: Snap and lookup for 3d open-vocabulary instance segmentation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:30.804575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:27.709530Z digest=sha256:cd7066b6b1a743a7f47b858b91a99c3c7d8875fc981e2f85c64f4258312dbede

Observation 1e2b678d-1710-4cda-abf8-f2a3d9957f97 · outbound

This paper cites SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:27.712861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:27.712861Z digest=sha256:ac852348310c8b7f2b8871222202a4a12b19299c909e3fe17b82f11b129565ac

Observation 61084017-02c1-4262-82c2-96f5b4016848 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Scaling up visual and vision-language representation learning with noisy text supervision

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:27.716730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:27.716730Z digest=sha256:bbbd65baeac75f96f5b0980e484b7fe5a29eee16ef3d7c459e55a5cd5f719695

Observation 6aacf438-0730-4f42-b12e-d41953f1f5fe · outbound

This paper cites Pointgroup: Dual-set point grouping for 3d instance segmentation.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Pointgroup: Dual-set point grouping for 3d instance segmentation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:30.786978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:27.720904Z digest=sha256:521781812be9ac81849287e8bc9e4e402c5737d8794a1d9afa963327536ab86a

Observation 34bed0b6-56f7-4cba-8cd6-b1201b1dbd7a · outbound

This paper cites Open-vocabulary 3d semantic segmentation with foundation models.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Open-vocabulary 3d semantic segmentation with foundation models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:30.776123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:27.724975Z digest=sha256:820181734581dbabd10f051046b546bf83c0a18ed4204d652c4a85055191ebda

Observation 0fd09de1-f7aa-4287-8c72-5ba5e4390ef3 · outbound

This paper cites Robin3D: Improving 3D Large Language Model via Robust Instruction Tuning.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Robin3D: Improving 3D Large Language Model via Robust Instruction Tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:27.728279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:27.728279Z digest=sha256:8e5673882380dcd7752ad64e461a3e54a557da9f9e32b48e2489c3fda510956e

Observation 35f40d81-70a7-4c1b-a238-fb1638c67e70 · outbound

This paper cites Segment any- thing.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Segment any- thing

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:30.707360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:27.762323Z digest=sha256:a469f1747b1f417c5c3665512e458d72831b34b7c94226339a8974f7765bca65

Observation a4d29b7b-f67e-4039-9868-c3deb215616d · outbound

This paper cites Language-driven Semantic Segmentation.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Language-driven Semantic Segmentation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:27.807699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:27.807699Z digest=sha256:7bb966312d4fbea2ec739a8b975a59df759d6f7dce69778d9da08314ab2ada85

Observation 9de54dd7-6ae3-4056-b34b-b9ea586d3c8d · outbound

This paper cites Maniptrans: Efficient dexterous bimanual manipula- tion transfer via residual learning.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Maniptrans: Efficient dexterous bimanual manipula- tion transfer via residual learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:30.635965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:27.855758Z digest=sha256:e02ed98c12726f6d148b7be79e22cdc242fb7eea6d5d824cc64b1a1927f1d74f

Observation 5d9069cf-f25c-4a11-90ce-35b19d493b9d · outbound

This paper cites Dense multimodal alignment for open-vocabulary 3d scene understanding.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Dense multimodal alignment for open-vocabulary 3d scene understanding

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:30.625079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:27.933238Z digest=sha256:7f123391195c94efaa1a56a72db1b0fa3e4f16d6f540c29ec5e865a4a526f27b

Observation 688f14b2-f57c-4dc0-8e2a-115204457f8c · outbound

This paper cites Visual instruction tuning, 2023.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Visual instruction tuning, 2023

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:28.008889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:28.008889Z digest=sha256:c20669a806159de9810649d20c25f5a96dfdf4889a7e9c7507b910c27543716c

Observation 5e607342-8aa6-4f45-ae19-cda0499f74ee · outbound

This paper cites Building interactable replicas of complex articulated objects via gaussian splatting.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Building interactable replicas of complex articulated objects via gaussian splatting

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:30.607648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.012417Z digest=sha256:9a0b02d06c06b252cde64f2f329a53805b98bbe1e473272a59f70fb04aef0b9a

Observation c6e99081-f65f-4393-a427-f06713ed89dd · outbound

This paper cites Movis: Enhancing multi-object novel view synthesis for indoor scenes.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Movis: Enhancing multi-object novel view synthesis for indoor scenes

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:30.596817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.016033Z digest=sha256:e226286014dc2cdde78b016f90c7e255b443c28d68239441c395e390ef8f781c

Observation 093762cf-f983-484f-a4f2-8698fcd957e8 · outbound

This paper cites Ovir-3d: Open-vocabulary 3d instance retrieval without training on 3d data.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Ovir-3d: Open-vocabulary 3d instance retrieval without training on 3d data

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:30.586444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.020079Z digest=sha256:1580cc5c4c1ace4ba9b6c5f1350bcc4b8b255828fc6f2ec6cc9b304ff344751c

Observation 91930c99-fbf9-4179-abcd-b596c95a3ee6 · outbound

This paper cites SQA3D: Situated Question Answering in 3D Scenes.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding SQA3D: Situated Question Answering in 3D Scenes

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:28.023924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:28.023924Z digest=sha256:448ec36111aa35136228900739251b87a6b130abc4101b19a76f1f5ba741afbc

Observation 730252d7-a215-49ee-a4ca-d892a4913615 · outbound

This paper cites When llms step into the 3d world: A survey and meta-analysis of 3d tasks via multi-modal large language models.arXiv preprint arXiv:2405.10255, 2024.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding When llms step into the 3d world: A survey and meta-analysis of 3d tasks via multi-modal large language models.arXiv preprint arXiv:2405.10255, 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:28.027799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:28.027799Z digest=sha256:4eaab1b0066950237b5b91779e7a21ccab696931a004be73e3eb0e9e61de973f

Observation 6912e089-9502-4e95-aa27-ee1993a738d3 · outbound

This paper cites Multiscan: Scalable rgbd scanning for 3d environments with articulated objects.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Multiscan: Scalable rgbd scanning for 3d environments with articulated objects

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:30.502185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.031218Z digest=sha256:65106a9bb4ae60f4d78491068141263cd19e59335051c3c2581191d1a5f5331e

Observation 9194ca88-c5a0-42ab-be8d-cf5e3fdaa6db · outbound

This paper cites Occupancy-mae: Self-supervised pre-training large- scale lidar point clouds with masked occupancy autoencoders.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Occupancy-mae: Self-supervised pre-training large- scale lidar point clouds with masked occupancy autoencoders

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:30.373872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.034598Z digest=sha256:a8d96d3332f0a70d0cb9b282029fca5e802a44ed461db226441969ffc9bffc75

Observation 9f32ff46-4822-446f-8071-bcfbe9a42857 · outbound

This paper cites Open3dis: Open-vocabulary 3d instance segmentation with 2d mask guidance.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Open3dis: Open-vocabulary 3d instance segmentation with 2d mask guidance

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:28.037963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:28.037963Z digest=sha256:6e5eb0d2bd0291b06d6e3708c5da85ccd06569a8f087aad20d12c4a1636f03e9

Observation cc28536c-1b91-4aae-930c-617f0eaffacc · outbound

This paper cites Phyrecon: Physically plausible neural scene recon- struction.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Phyrecon: Physically plausible neural scene recon- struction

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:30.315409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.041035Z digest=sha256:d9eb161f41b62ac7b2fa4a7cffa91fa2efd18583f03ec18e42c188cf0e2d81b0

Observation 4091d38f-f5dc-43f4-852f-4dd29d617813 · outbound

This paper cites Decompositional neural scene reconstruction with generative diffusion prior.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Decompositional neural scene reconstruction with generative diffusion prior

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:30.305959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.044502Z digest=sha256:0d2aa4d10b570c850138ebc3919218a302d1c744d6b2e685ff0acb6a772a71bc

Observation 39eff41d-26a6-4f37-bfbc-33368219cf26 · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:28.047947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:28.047947Z digest=sha256:f03b6450480a85691a30761f356b2bb99d16edf176d98aae48495776e001977c

Observation 33105af5-5026-44b2-be79-73cf95d79b90 · outbound

This paper cites Gpt-4 with vision (gpt-4v) system card, 2023.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Gpt-4 with vision (gpt-4v) system card, 2023

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:30.294773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.051524Z digest=sha256:db88a157994126c25a1eb3f4cdc62d490947065854a71551909ebf6262eef394

Observation ae3d28a7-a157-4660-89f5-16801ab81e64 · outbound

This paper cites Openscene: 3d scene understanding with open vocabularies.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Openscene: 3d scene understanding with open vocabularies

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:30.281720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.055025Z digest=sha256:75af93db8fdee4b63c4fc12c4e3c89eeb10da2e7eee17cfd1c4edd6c4ac0bf42

Observation 3514242c-189b-4bcd-bc05-da6a5b38b782 · outbound

This paper cites Shapellm: Universal 3d object understanding for embodied interaction.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Shapellm: Universal 3d object understanding for embodied interaction

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:30.120521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.057533Z digest=sha256:16a6a69939736e6895479d889b27505e412d73b0faa3d6ea1f4961ab18ca54a2

Observation ace3405f-dfcc-4ab0-a731-891d66dbcd05 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Learning transferable visual models from natural language supervision

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:30.079804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.060634Z digest=sha256:b148de8c8f5fc4063569570ffd35f666e1b2eafa84cbaae4f404adf67722bb51

Observation 586de2a0-e055-4eea-82fc-c8877e04dfda · outbound

This paper cites Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:28.063996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:28.063996Z digest=sha256:a6cd1b582bf7c281f213299c69d1cdeb43007d586c70d8d795b34679b7053c3e

Observation 5ae228f7-2896-4632-a111-684dc6cc0104 · outbound

This paper cites Language- grounded indoor 3d semantic segmentation in the wild.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Language- grounded indoor 3d semantic segmentation in the wild

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:30.069393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.067023Z digest=sha256:7ca7624e64b82e07b46c2b83fec7c4992fe82f259512bdb3664b71cf25bfc60c

Observation 8f9f7f0e-c425-4a0a-8179-3a8ccf316d7f · outbound

This paper cites Mask3d: Mask trans- former for 3d semantic instance segmentation.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Mask3d: Mask trans- former for 3d semantic instance segmentation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:30.058426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.069852Z digest=sha256:b818539c8caf126386bbc5fb5f1ccf1080ba6ebb0189310e2a432f143a595582

Observation e28baa2c-f938-4f11-bb3c-0c7783fa0351 · outbound

This paper cites Open- mask3d: open-vocabulary 3d instance segmentation.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Open- mask3d: open-vocabulary 3d instance segmentation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:30.047630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.073061Z digest=sha256:10e754fbf67656365130e2b2f16e097812b9a9b8748fdebe9315640e7ab89ffe

Observation 5ea2272d-ba49-492e-986b-bb6386fa7a91 · outbound

This paper cites Minigpt-3d: Efficiently aligning 3d point clouds with large language models using 2d priors.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Minigpt-3d: Efficiently aligning 3d point clouds with large language models using 2d priors

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:30.035585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.086230Z digest=sha256:97d7aae4a5e70d518269de7d7bbfd9441bbc0c2131a5c2217967a0e80ccf74de

Observation 02010a7b-12b0-40bd-969f-927ecd489d55 · outbound

This paper cites Geomae: Masked geometric target prediction for self-supervised point cloud pre-training.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Geomae: Masked geometric target prediction for self-supervised point cloud pre-training

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:29.942005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.136299Z digest=sha256:04019245e0e8317b3e2a328b37567da69858cec6587dceb5199a2f953bc2f857

Observation 49319da8-1243-4d6d-bd92-2e87331dc8d9 · outbound

This paper cites Rio: 3d object instance re- localization in changing indoor environments.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Rio: 3d object instance re- localization in changing indoor environments

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:29.761638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.160645Z digest=sha256:1646f4039b55336afe80d57478129e4f19fc58891e152e5a987c5fc0b60d70fa

Observation 95ed88bf-6851-408c-80c2-e2c66443995a · outbound

This paper cites Groupcontrast: Semantic-aware self-supervised representation learning for 3d understanding.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Groupcontrast: Semantic-aware self-supervised representation learning for 3d understanding

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:29.742221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.202180Z digest=sha256:741b30c20b5b53715cce99edcc952a2cb5906b9b908b7afaac6bcc88146f650d

Observation c0346087-45f8-4c33-a945-cbd6859c46c9 · outbound

This paper cites Open vocabulary 3d scene understanding via geometry guided self-distillation.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Open vocabulary 3d scene understanding via geometry guided self-distillation

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:29.731826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.246949Z digest=sha256:d9201ca3f98ce3055ee2764f95e25f1ff48c0f44427a7362ea2bfbc259ecc2cf

Observation 76d9c71c-1a21-49ef-acf2-7581fad9fca2 · outbound

This paper cites Embodiedscan: A holistic multi- modal 3d perception suite towards embodied ai.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Embodiedscan: A holistic multi- modal 3d perception suite towards embodied ai

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:28.297466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:28.297466Z digest=sha256:d3608634a82b9e0a6aecd30799247d87d6be4b6e74f1eb0b5da74498dfa87caf

Observation 21d9f7da-6e7b-4e93-a61c-0013f03eac27 · outbound

This paper cites T-MAE: Temporal Masked Autoencoders for Point Cloud Representation Learning.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding T-MAE: Temporal Masked Autoencoders for Point Cloud Representation Learning

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-08-16T05:58:28.857342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.326334Z digest=sha256:0d765c383bde039268a65448d758a876311a255866cebb96ace7b86a0fcf4ffc

Observation e646850d-569a-4142-ad03-301e36e20fab · outbound

This paper cites CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:28.375228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:28.375228Z digest=sha256:ba1fb12b38cb5435c13275648b72d749ca974e858228f584b187f24c3147d8a7

Observation 4d82e6fe-eefa-401f-bad3-276b0dcb4e4e · outbound

This paper cites Masked scene contrast: A scalable framework for unsuper- vised 3d representation learning.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Masked scene contrast: A scalable framework for unsuper- vised 3d representation learning

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:29.715108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.407106Z digest=sha256:20b497112e986d8c2fa0e167951296b5878285d32b2ba9ee8693966b5b2ff0f8

Observation 048627c5-2f22-4c35-ab1e-10b4d916d323 · outbound

This paper cites Sed: A simple encoder-decoder for open-vocabulary semantic segmentation.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Sed: A simple encoder-decoder for open-vocabulary semantic segmentation

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:29.704158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.410563Z digest=sha256:291696615a65d82cafe41e8b757025605c9b1e4d9659cad925f8b0e8fabe65be

Observation 414850d3-e67e-4afa-91be-4c5a5c20aa32 · outbound

This paper cites Pointcontrast: Unsupervised pre- training for 3d point cloud understanding.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Pointcontrast: Unsupervised pre- training for 3d point cloud understanding

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:29.693562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.413712Z digest=sha256:29b8d433f08f0f1bc79021a9fa671a41b7b9e4b858a13e80f8a31c787fc67db0

Observation 65bcb96e-43bf-4878-a2bb-8704191bd398 · outbound

This paper cites Simmim: A simple framework for masked image modeling.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Simmim: A simple framework for masked image modeling

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:29.612618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.416866Z digest=sha256:4fd3565b2059423585f337d61d12a3407332324ee9be2f92d37c57ce9db45dac

Observation 4b521be6-91bb-449e-b4e4-a1983d23d469 · outbound

This paper cites SAMPro3D: Locating SAM Prompts in 3D for Zero-Shot Instance Segmentation.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding SAMPro3D: Locating SAM Prompts in 3D for Zero-Shot Instance Segmentation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:28.420256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:28.420256Z digest=sha256:c91c2fce0995446c63468ad6bee2301aab7963bb9ce1181d46af7fe67803f536

Observation b2ef24f6-82f9-4f86-82eb-3d7139be65fd · outbound

This paper cites Maskclus- tering: View consensus based mask graph clustering for open- vocabulary 3d instance segmentation.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Maskclus- tering: View consensus based mask graph clustering for open- vocabulary 3d instance segmentation

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:29.483251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.423244Z digest=sha256:2d5930baae6e003605a21e48a0708d517b238060cc0406214706133f9cb42161

Observation 51be96d3-3c84-41b3-a576-df7cf73126f6 · outbound

This paper cites 3D Vision and Language Pretraining with Large-Scale Synthetic Data.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding 3D Vision and Language Pretraining with Large-Scale Synthetic Data

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:28.426729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:28.426729Z digest=sha256:4acd70a7e14d6ff5dd1a2c0852a551f57bdd4a97f9b056b8deae041f9c8945f8

Observation 1db6b825-c282-48f1-ae6a-5f986b6b4a3a · outbound

This paper cites Gd-mae: gener- ative decoder for mae pre-training on lidar point clouds.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Gd-mae: gener- ative decoder for mae pre-training on lidar point clouds

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:28.430295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:28.430295Z digest=sha256:11c1e570c735ef0d7649ca401a3d7c2ad3b27f3001b6f59485bb02ed4e88815b

Observation c867954f-d97d-4a02-9867-858d4c34012e · outbound

This paper cites 3D-GRAND: A Million-Scale Dataset for 3D-LLMs with Better Grounding and Less Hallucination.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding 3D-GRAND: A Million-Scale Dataset for 3D-LLMs with Better Grounding and Less Hallucination

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:28.433338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:28.433338Z digest=sha256:827d0223c669ae9b2ae4ea9b0c30814e57e6e0aee4cb2fb9bd6185eac0356672

Observation 9d03856e-9818-46a7-a458-b49748010c4b · outbound

This paper cites Regionplc: Regional point-language contrastive learning for open-world 3d scene understanding.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Regionplc: Regional point-language contrastive learning for open-world 3d scene understanding

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:29.467599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.436902Z digest=sha256:c9f87343d8dfea06261ae73318ef5f713905007477e2e3993120570fc53c4963

Observation ffeb84ee-446c-4192-84aa-2aba3d2f0704 · outbound

This paper cites SAM3D: Segment Anything in 3D Scenes.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding SAM3D: Segment Anything in 3D Scenes

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:28.440274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:28.440274Z digest=sha256:d75fab05650dac35529fb5764cbc9c0b5ced15af55599aae68948a70275ccada

Observation fc681d9b-ced2-466f-be6e-78865afb9a16 · outbound

This paper cites Metascenes: Towards automated replica creation for real-world 3d scans.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Metascenes: Towards automated replica creation for real-world 3d scans

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:29.457227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.443762Z digest=sha256:cc9bcc17ceb963e291849dab0ab985e06b039bcf46467fc9717a9cdcaa7c9eec

Observation bd840b72-6f36-4b46-9177-058e9f49875b · outbound

This paper cites Point-bert: Pre-training 3d point cloud transformers with masked point modeling.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Point-bert: Pre-training 3d point cloud transformers with masked point modeling

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:29.446462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.447024Z digest=sha256:f0055f99552f456c8f3ee944c9a80646ae18b32ec6218135433c22753ada61b1

Observation 2f04313e-6117-4f6a-b617-739a653158ae · outbound

This paper cites Clip2: Contrastive language-image- point pretraining from real-world point cloud data.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Clip2: Contrastive language-image- point pretraining from real-world point cloud data

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:29.436137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.450240Z digest=sha256:571e85fd30b0f05dfc75f1a62104360c01cf22e7a423be043e7ea1f915e9cae9

Observation b6b0a255-c531-48da-a783-950469e10e30 · outbound

This paper cites Vision-language pre-training with object contrastive learning for 3d scene understanding.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Vision-language pre-training with object contrastive learning for 3d scene understanding

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:29.425860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.453381Z digest=sha256:2cab4a9c233e030a90d418dac54df579fca1461b70d977b549a913fa22357ac5

Observation 47a460c8-59af-407b-955b-4b9d46f13f22 · outbound

This paper cites Multi3drefer: Grounding text description to multiple 3d ob- jects.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Multi3drefer: Grounding text description to multiple 3d ob- jects

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:29.415316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.456692Z digest=sha256:a53e0239fb71a842419d4aa4e12484b83669ec47addb62727eb1d1e977218c25

Observation 67577ae0-edce-4ec2-9339-d407355d4165 · outbound

This paper cites Self-supervised pretraining of 3d features on any point-cloud.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Self-supervised pretraining of 3d features on any point-cloud

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:29.384716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.459958Z digest=sha256:fb475e55d7faa6d81246518263280b3d4b923dc429da54d5313ddffda63155d1

Observation 159917a2-244a-47dc-afd8-8c86931c3136 · outbound

This paper cites 3D-VLA: A 3D Vision-Language-Action Generative World Model.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:28.463384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:28.463384Z digest=sha256:db1663fced975cc0fe6394d965fa386ba719148573158b8a7734147952b68e92

Observation f8e42cae-a947-4fa0-ae2d-91ff5174f736 · outbound

This paper cites Structured3d: A large photo-realistic dataset for structured 3d modeling.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Structured3d: A large photo-realistic dataset for structured 3d modeling

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:28.466561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:28.466561Z digest=sha256:6c13778b964f1047a825305b9c95b9471f277621dec78555f42bbc6b1f8a1e88

Observation 34df5384-c2b2-43e5-b637-b368bd378e4e · outbound

This paper cites Closed-Loop Open-Vocabulary Mobile Manipulation with GPT-4V.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Closed-Loop Open-Vocabulary Mobile Manipulation with GPT-4V

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:28.469942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:28.469942Z digest=sha256:e03da80d5c1c957e3b301bfccd1e52cee327d19a4ee205dd9a65bce67dde6fff

Observation a19a8cf6-9e8b-4835-8b33-a52bee2c0729 · outbound

This paper cites Extract free dense labels from clip.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Extract free dense labels from clip

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:29.225113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.472858Z digest=sha256:d76552470959a15ffbe657db4453122ffc071250b428d7b6ced9ff6ce26292f1

Observation b818cc9b-dc45-458b-a33f-35c5b4cfde2f · outbound

This paper cites Uni3D: Exploring Unified 3D Representation at Scale.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Uni3D: Exploring Unified 3D Representation at Scale

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:28.475958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:28.475958Z digest=sha256:39db86acbd4e057609bf1aafc5eb97433a658adcef2f92ebc0c35c57792f3317

Observation ad00c17a-0951-4836-86e7-8fd63036a62a · outbound

This paper cites Detecting twenty-thousand classes using image-level supervision.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Detecting twenty-thousand classes using image-level supervision

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:29.124539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.502078Z digest=sha256:8436e2ad186ce0c9eb41ed6725509f4a254ad56dda80d4bd759e9c61f4d091c6

Observation b56dc2ec-4b53-482b-aa78-8b770496dfb8 · outbound

This paper cites LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:28.555565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:28.555565Z digest=sha256:45789eb1c2e94e3c4d7f66bd285e01116c0201139de1e9d65928239580fb3ada

Observation 9c3cb4cf-79a4-4f90-9499-8c9ddaafe77e · outbound

This paper cites 3d-vista: Pre-trained transformer for 3d vision and text alignment.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding 3d-vista: Pre-trained transformer for 3d vision and text alignment

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:58:29.062110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:28.627676Z digest=sha256:ba8a536a3cdb77e904af8f016b6dde2ee3335bcfdf38f9d417e5490555305b09

Observation 2c389b2c-4926-4c6a-a8a3-b04307b12597 · outbound

This paper cites Unifying 3D Vision-Language Understanding via Promptable Queries.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Unifying 3D Vision-Language Understanding via Promptable Queries

Reference 97

Resolution
malformed identifier
no resolver link, observed 2026-08-16T05:58:28.684524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:28.684524Z digest=sha256:162504bab4f3d010a036cb1cdda94da77375c1a0a30e440cf48414d10f19aa6f

Observation de0e1657-967c-477f-8c07-6cb2d9b273b3 · outbound

This paper cites an unresolved cited work.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Unresolved cited work

Reference 2021

Resolution
parse uncertain
raw_fallback, observed 2026-08-16T05:58:30.953066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:58:27.667351Z digest=sha256:80d004b7ff55bb704bf94a67e4ef5b32e988e6d09fe00491be226550cf6644a2

Pith citing papers

No inbound Pith citation observations are available.