Pith. sign in

Paper Citation Record · LEDGER

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation

As of 22 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2411.13243.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13243 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:46:10.072459Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact2
  • verified fuzzy22
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cf1a2954-5bf4-4875-b748-477ce9e0d64c · outbound

This paper cites 3d semantic parsing of large-scale indoor spaces.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation 3d semantic parsing of large-scale indoor spaces

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.835898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.835898Z digest=sha256:d70530923d8496575414d796512effdf245ad06e59518bbb93649040d99587d5

Observation 78a22aa5-8226-45e1-8b61-3f5f734538fb · outbound

This paper cites Emerging properties in self-supervised vision transformers.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Emerging properties in self-supervised vision transformers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.841512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.841512Z digest=sha256:f3cb88c72d006cbe2fbcd2a32f71d4e6a69f4198b75cc4207287b71f213d4bc3

Observation 7f36bd1a-307a-4595-8e54-35d1243464f8 · outbound

This paper cites Masked-attention mask transformer for universal image segmentation.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Masked-attention mask transformer for universal image segmentation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.786369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:46:09.846666Z digest=sha256:1758ee9498e1de677072854e95d508a0bc6513d435b1804503303ba8b6b9d8fc

Observation 558d4a84-abc6-47f0-b362-5249f61db796 · outbound

This paper cites Per-pixel classification is not all you need for semantic segmentation.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Per-pixel classification is not all you need for semantic segmentation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.769872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:46:09.851979Z digest=sha256:222f974a36ae12b68be67173e8e2962c7042283be3d897a4dc9673aeb2958a09

Observation 857ed3f8-de2b-469a-b7f1-25dbbf8bdc2b · outbound

This paper cites Transductive zero-shot learning for 3d point cloud classification.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Transductive zero-shot learning for 3d point cloud classification

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.754268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:46:09.857085Z digest=sha256:f10e8417c98fbfe71e6669f37a97d900e61b5e0a09bddd0176d37213a4206e57

Observation 11d8c0be-a93f-452b-acb4-bfe4445bf919 · outbound

This paper cites CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.861897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.861897Z digest=sha256:92d02f5a0e85451a3ba43dea35f9c0032246c7c6b030e1f8a3d075a439177521

Observation 721146ed-8144-491d-9086-aaa2a1988be3 · outbound

This paper cites 4d spatio-temporal convnets: Minkowski convolutional neural networks.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation 4d spatio-temporal convnets: Minkowski convolutional neural networks

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.738908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:46:09.867561Z digest=sha256:cd0cb2b6908468bb0f2f3a77e41867b6353c91ba80a85058defbf2bef51bb40c

Observation 83eb9f6b-6e09-4c81-a1ba-d9d7c3f3e2b1 · outbound

This paper cites Spconv: Spatially sparse convolution library.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Spconv: Spatially sparse convolution library

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.871855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.871855Z digest=sha256:30b47bdf7bb7222c77f6f03c87f0b9fccacb2b0f55c73db1ca51913ef3aee98c

Observation b822299a-bca6-46ac-ba91-9d1bbdf56dd3 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.876484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.876484Z digest=sha256:2fa124420b1db5fdd4bbe72bba4a16233c4467c73e4cfc183205bd7ea68813ea

Observation efd47bee-b905-467c-9dc8-1e812a919cf2 · outbound

This paper cites Decoupling zero-shot semantic segmen- tation.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Decoupling zero-shot semantic segmen- tation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.704441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:46:09.881197Z digest=sha256:0d18576f4c6093ac165e2a8691229ae979e8b77f9d5b5902ed3aa58faf3a5618

Observation dcf081a5-4de0-44a2-b913-ffb064558750 · outbound

This paper cites Pla: Language-driven open-vocabulary 3d scene understanding.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Pla: Language-driven open-vocabulary 3d scene understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.885601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.885601Z digest=sha256:b5dad77584728cf9ff0a3f8eda48f1438203cdf3c9cc7748355af4a99507e734

Observation ea3da3bf-6a19-47d2-b36d-189d3cc877cb · outbound

This paper cites Open-vocabulary universal image segmentation with maskclip.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Open-vocabulary universal image segmentation with maskclip

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.678562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:46:09.890230Z digest=sha256:dcc6923a13151f2c8d3318f66d744b32671b00aa0425ecd1c2e214f8a07b7027

Observation d060d7ef-3829-4316-9136-70ade5a9db14 · outbound

This paper cites Scaling open-vocabulary image segmen- tation with image-level labels.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Scaling open-vocabulary image segmen- tation with image-level labels

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.662409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:46:09.894688Z digest=sha256:91e1b4fa4801deb4e209c1a2a4e4fff1723a0d3ea61cdd52f1b09f172830aebd

Observation e5f36373-22e4-4d0a-8927-2f1f0ca5b3ce · outbound

This paper cites Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.898818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.898818Z digest=sha256:8f0788eda671b938cc88df96ba9258d53eeb31a1f9ec7f68cbe58c2a951a171e

Observation 9a355519-75d1-4bb9-9e98-f979e2d83829 · outbound

This paper cites Semantic Abstraction: Open-World 3D Scene Understanding from 2D Vision-Language Models.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Semantic Abstraction: Open-World 3D Scene Understanding from 2D Vision-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.903856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.903856Z digest=sha256:0b877ce44834a383f8105a1c339c042527e33738065787c418fe95d8437b9260

Observation 1ad8db9a-f275-43a1-af8e-67e9ed00dcbc · outbound

This paper cites UniM-OV3D: Uni-Modality Open-Vocabulary 3D Scene Understanding with Fine-Grained Feature Representation.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation UniM-OV3D: Uni-Modality Open-Vocabulary 3D Scene Understanding with Fine-Grained Feature Representation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.908857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.908857Z digest=sha256:6896d75cac2717d5fb0b04db73bb406e994a179e21d5046aeb8edacead47796a

Observation 0f93fb09-3cd1-4cad-958b-cfe34be54ea7 · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.913908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.913908Z digest=sha256:297a29ee766375a1e018677714002323b4c366efc8a41317506d6f2325672f60

Observation 80fb5555-0805-41b3-b501-bfb86c64254c · outbound

This paper cites Denoising diffusion probabilistic models.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Denoising diffusion probabilistic models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.647307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:46:09.919063Z digest=sha256:b90438d4e949e5844cc3ae53c8febbe084e692111092d2919d664cf8c86744ae

Observation dab2d84f-8365-4fc7-a82a-dbc632566b50 · outbound

This paper cites Clip2point: Transfer clip to point cloud classification with image-depth pre-training.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Clip2point: Transfer clip to point cloud classification with image-depth pre-training

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.923653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.923653Z digest=sha256:3f7907ce5a08bd8f3eb59fb93a6d1bb016c3d46224f659e09310ee8dc0eb8ce2

Observation 62cfd89d-462a-499a-b6f1-6203bb873b54 · outbound

This paper cites Open-vocabulary 3d semantic segmentation with foundation models.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Open-vocabulary 3d semantic segmentation with foundation models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.622689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:46:09.928234Z digest=sha256:993b5dfc9ceb94594760a7691b092e7bf7459b2abd316ac7418eccbc6d6428ce

Observation 609eb4ee-0363-47ab-915a-ce8ad7489197 · outbound

This paper cites Diffusion Models for Open-Vocabulary Segmentation.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Diffusion Models for Open-Vocabulary Segmentation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.932975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.932975Z digest=sha256:ceeed23f6eb66ae5d6bde6edf88a0b65615970c0ee89fdaa6c73b352ee7cc93b

Observation ce77b63d-a1f3-4478-8f2a-52061a56d790 · outbound

This paper cites Segment anything.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Segment anything

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.937672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.937672Z digest=sha256:281aa076045e72519972b5c50090954631227804bf9b2dda10e01c4629c14a7d

Observation 39e53374-2591-465a-91eb-1659c74ed2d4 · outbound

This paper cites Language-driven Semantic Segmentation.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Language-driven Semantic Segmentation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.941912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.941912Z digest=sha256:9c33411c1640eb06cc6efa6bf77c68a5805963ce75a7464ffa429ebc5c5e6f0d

Observation 52d6224d-8473-404d-976d-a8d9d99b8466 · outbound

This paper cites TagCLIP: Improving Discrimination Ability of Open-Vocabulary Semantic Segmentation.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation TagCLIP: Improving Discrimination Ability of Open-Vocabulary Semantic Segmentation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.946613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.946613Z digest=sha256:64a34900dc326258793cf55f3d750ba7b991e08ea7515fed45934d411bebaf3b

Observation fc44eab2-f161-4e25-aa9f-037048dc1e9d · outbound

This paper cites Guiding text-to-image diffusion model towards grounded generation.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Guiding text-to-image diffusion model towards grounded generation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.598419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:46:09.951597Z digest=sha256:582fbd4c603e16f166a681d7f2b78dbf407361827071f34f5db152533e468e2b

Observation 64846e98-c1f8-4e76-b350-034dbd85ba83 · outbound

This paper cites Magic3d: High-resolution text-to-3d content creation.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Magic3d: High-resolution text-to-3d content creation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.582927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:46:09.955911Z digest=sha256:9b0da1fb567f3294cac9cf365fbc71bef6f03ad0ad71c37a14825051371c5562

Observation 19c529cf-e808-47ab-b667-9314e3f8ea13 · outbound

This paper cites Weakly Supervised 3D Open-vocabulary Segmentation.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Weakly Supervised 3D Open-vocabulary Segmentation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.960220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.960220Z digest=sha256:08ecfed887d8834b9ace3c79ed080064da39f425b7560917c3608a752283e1d1

Observation 7077cfea-7048-4b40-b00e-047fd27ca418 · outbound

This paper cites Segment any point cloud sequences by distilling vision foundation models.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Segment any point cloud sequences by distilling vision foundation models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.568087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:46:09.964909Z digest=sha256:1f835bb1a2be36fd5809955cb1df8af239b592ec28731a8e90f20e00871f2c61

Observation 27448eb0-aeb7-43a9-93c4-9fbc36512e56 · outbound

This paper cites Decoupled Weight Decay Regularization.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Decoupled Weight Decay Regularization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.969515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.969515Z digest=sha256:edd977a26dc583d73692818a21e8feebda7ead9a046429859e830a40ab18e99d

Observation 395aba57-d766-42d7-873b-cb4d2912e21b · outbound

This paper cites Open-vocabulary Semantic Segmentation with Frozen Vision-Language Models.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Open-vocabulary Semantic Segmentation with Frozen Vision-Language Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-12T16:46:10.182714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:46:09.974385Z digest=sha256:37efd462ef47d322282a6f1ebd5a03bd49365c1a3c411319ceff967d42f873e9

Observation eac43e42-c907-41be-928d-0f8f12efa93e · outbound

This paper cites Generative zero-shot learning for semantic segmentation of 3d point clouds.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Generative zero-shot learning for semantic segmentation of 3d point clouds

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.553002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:46:09.978968Z digest=sha256:827ea85925c00da55bfcd87d9162ca57a5be628a6d5a8454fc5bb8c1f41eb787

Observation 5dd0dbd7-89bc-4da2-bd26-15f2e8b8e4bb · outbound

This paper cites Nerf: Representing scenes as neural radiance fields for view synthesis.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Nerf: Representing scenes as neural radiance fields for view synthesis

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.983560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.983560Z digest=sha256:1b0ebc08568c4189416440f95fd1a15108f74bf5bb03abfbb475764632129077

Observation 7ba131b5-fb58-4b05-8bdc-28e47a8432b9 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation DINOv2: Learning Robust Visual Features without Supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.987758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.987758Z digest=sha256:f0aacf5c95e5045ec6a91dcdc61728f06af11d44dee2622a6169456eb9353772

Observation 0b2204d1-82d8-403f-9f6d-1d93b0da5614 · outbound

This paper cites Openscene: 3d scene understanding with open vocabularies.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Openscene: 3d scene understanding with open vocabularies

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.993708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.993708Z digest=sha256:f161d8a787aebb1475041423ee6b7ec6eb1c0f12253908d0b581151524df42bc

Observation eff91478-3e75-408c-bdb3-659c1841506e · outbound

This paper cites Dreamfusion: Text-to-3d using 2d diffusion.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Dreamfusion: Text-to-3d using 2d diffusion

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.519250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:46:09.998127Z digest=sha256:85e62e63efa0768977f5b7eeb9b4e09d9af2ba95f684569d321fa82eef5d0f49

Observation a5f85f7f-74aa-48fa-8f51-5fd6ca11687b · outbound

This paper cites Learning transferable visual models from natural language supervision.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Learning transferable visual models from natural language supervision

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:10.003200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:10.003200Z digest=sha256:116a80841b7b313ad3b7c73473bf376b87d83b6923b733732912000ab968da17

Observation 93103c53-289e-4781-9f83-f5c2f706ebd1 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation High-resolution image synthesis with latent diffusion models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:10.007716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:10.007716Z digest=sha256:617e01b89dc93f12fdd4cb242b6e298097171e5cfc8894575e701cf46a0e8b63

Observation 5f773dd6-48f4-4386-964b-6e205659d548 · outbound

This paper cites Language-grounded indoor 3d semantic segmentation in the wild.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Language-grounded indoor 3d semantic segmentation in the wild

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:10.012269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:10.012269Z digest=sha256:c9ecbfaa2a5f75fd6941a7e8fae1d151cd3e96533cd4c9114f3d9ffa67b44840

Observation 5828134d-e411-40d1-97c9-7e6bb1c6729e · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Photorealistic text-to-image diffusion models with deep language understanding

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.474877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:46:10.017134Z digest=sha256:4282211372f21b176c4b6eecc546f3e0886fe3833e8efab5673c1d8721c45f9c

Observation c446b89e-62ed-4580-aed1-8e2b4e6153ed · outbound

This paper cites Dreamgaussian: Generative gaussian splatting for efficient 3d content creation.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Dreamgaussian: Generative gaussian splatting for efficient 3d content creation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.459203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:46:10.022245Z digest=sha256:296a55b6b576858ee0f50886219a8b476089518317fef2d01fe047251cb84999

Observation 0a64c565-b35d-4415-8c8b-44523533772f · outbound

This paper cites Diffumask: Synthesizing images with pixel-level annotations for semantic segmentation using diffusion models.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Diffumask: Synthesizing images with pixel-level annotations for semantic segmentation using diffusion models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.443737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:46:10.026793Z digest=sha256:b82d8dbc0335ae568e7c7ef6136c158a68213bde52ea275ec284ba8a525cec68

Observation 1a8875c1-c52a-43e6-ab82-b6701ee2bbdf · outbound

This paper cites Point Transformer V3: Simpler, Faster, Stronger.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Point Transformer V3: Simpler, Faster, Stronger

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:10.031420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:10.031420Z digest=sha256:18514ffbd5f3c2f3c5a7f25e0ddbaf944332cdb305b0d9a798f7c0a35ef3d35f

Observation 3d8920c5-7953-46fd-9a1e-21730a93cf52 · outbound

This paper cites Point transformer v2: Grouped vector attention and partition-based pooling.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Point transformer v2: Grouped vector attention and partition-based pooling

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.428777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:46:10.036429Z digest=sha256:41f9bbad03c4082f5d3a428f726ae23eefd21da2357281a97ce2556740451386

Observation ef8baa7b-3fbc-47c3-9165-7d2b2eee86a2 · outbound

This paper cites 3D Open-Vocabulary Panoptic Segmentation with 2D-3D Vision-Language Distillation.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation 3D Open-Vocabulary Panoptic Segmentation with 2D-3D Vision-Language Distillation

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-12T16:46:10.131526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:46:10.041152Z digest=sha256:a5fe44dc0352293eb6474042e544a68d0765b82445b00e71a9ecb27ac57419e4

Observation dac75b9e-69b9-4853-92eb-f52db788d026 · outbound

This paper cites Open-vocabulary panoptic segmentation with text-to-image diffusion models.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Open-vocabulary panoptic segmentation with text-to-image diffusion models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:10.045841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:10.045841Z digest=sha256:646252b7a0c53898e939e2a0ae949319ea5b8340ed4774586d10e4c38c7ed446

Observation 03402bea-bfa2-402e-ae9a-2c662e1c0345 · outbound

This paper cites Side adapter network for open-vocabulary semantic segmentation.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Side adapter network for open-vocabulary semantic segmentation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.403769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:46:10.050407Z digest=sha256:c6a1723c3f75f29724aae12df5d1559b59e9e4f29e40087ed594f8d92bd64946

Observation 689b8e2e-4542-48bd-8dce-93e5a4082c8c · outbound

This paper cites RegionPLC: Regional Point-Language Contrastive Learning for Open-World 3D Scene Understanding.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation RegionPLC: Regional Point-Language Contrastive Learning for Open-World 3D Scene Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:10.054878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:10.054878Z digest=sha256:6cfcf041d731d22b6342fd88641f13cb7918da0ef910cba6b9d7ac0b6b43aac3

Observation b7a64b4e-a6dd-4b62-ad84-357088974853 · outbound

This paper cites Vit-gpt2 image captioning.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Vit-gpt2 image captioning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.388705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:46:10.059407Z digest=sha256:051c1718e896df3698435ab596c4c7d597301261b0368530c80c583b2332e91a

Observation ab28d98a-0014-4cbd-9097-9372a76f5ca7 · outbound

This paper cites Clip-fo3d: Learning free open-world 3d scene representations from 2d dense clip.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Clip-fo3d: Learning free open-world 3d scene representations from 2d dense clip

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.373333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:46:10.063707Z digest=sha256:4f82ad2b472e766629c81394982f9b8f0c30b7d2b6fb2f6a638b141d1119d14e

Observation c7d84e53-af3b-4033-a054-292f1cff11ed · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Adding conditional control to text-to-image diffusion models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:10.067942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:10.067942Z digest=sha256:7690a190e03556ad734c568f31e50b9864672d64685ba08bb7d4630ed794604c

Observation 247efaf2-45f9-4e37-913c-a97bce88aaf8 · outbound

This paper cites Point transformer.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Point transformer

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.348859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T16:46:10.072459Z digest=sha256:c49aadbf6b91f387ced8245b12a8a35e4c0d951fc15fc44c5323f01ebc124990

Pith citing papers

No inbound Pith citation observations are available.