Pith. sign in

Paper Citation Record · LEDGER

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation

As of 10 August 2026, this Paper Citation Record lists 100 of 113 outbound references and 0 inbound Pith citation observations for arXiv:2502.02548.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02548 v2

Coverage vector

measured 100 of 113 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:52:14.978847Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 113 outbound references displayed

  • verified exact0
  • verified fuzzy37
  • unresolved63
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 00e19b1d-8dcb-4811-945b-7b9a620ece04 · outbound

This paper cites https://huggingface.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation https://huggingface

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.562340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.562340Z digest=sha256:0050d2511e796e373fb66e2c52f2175249a9157cae3aaefa67b6fab44bbd29fd

Observation 0de5cbea-bd04-4fda-8e9d-fbbf712568fe · outbound

This paper cites GPT-4 Technical Report.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.567886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.567886Z digest=sha256:81b8f3df4d89f48c6aff3a1e7aae027d71876ad7a8ececc245803e646f1162a0

Observation 0e29b170-02c0-4df9-bfb0-671e370e82f9 · outbound

This paper cites an unresolved cited work.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.572923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.572923Z digest=sha256:c5b8bc09a5d4c026a089c23249ed28d0b81ac9cb8a68242781f5bcdbededca46

Observation 3e1d639c-f057-45ae-9757-e97df963a3fd · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Scanqa: 3d question answering for spatial scene understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.577648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.577648Z digest=sha256:6df5bdfaf16a4b9ee06c8926a38abd049de7a2df762eebb6e38771e8985c7de8

Observation 4c4b582d-deda-451e-a389-2e911d10581c · outbound

This paper cites Qwen Technical Report.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Qwen Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.581601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.581601Z digest=sha256:7c1ba54cdf7445088576ee9d37b307c1f8b071dd06ca5ccfd916d9c47599f953

Observation 0ba80f78-5e9c-4d59-b1a4-d70e6a00b12e · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.586476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.586476Z digest=sha256:3052a0aedfb0b504ba3bae81cb9cfa9297e37f42920d7931296526710181e7c4

Observation 3639b949-16f6-4b95-9744-eef372a2b615 · outbound

This paper cites Arkitscenes: A diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Arkitscenes: A diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.590946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.590946Z digest=sha256:e31a4cae943631a76c96cd6c73240c5dacb1ad7c43e3940aac253e15e57806e7

Observation 29b0a2c2-4039-427f-8956-128928764e1c · outbound

This paper cites Audiolm: A language modeling approach to audio generation.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Audiolm: A language modeling approach to audio generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.595991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.595991Z digest=sha256:b0d5b9733fe7b3af5bfe03041d9d0ec8ff98fb4fefdeca56d0600260c06d7b45

Observation d8d940d8-b557-43f7-b49d-8964fbdf988e · outbound

This paper cites Large-scale machine learning with stochas- tic gradient descent.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Large-scale machine learning with stochas- tic gradient descent

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.600157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.600157Z digest=sha256:ca55262548be2a46ee31b3ddbac603fbd7e146059b35b81a628fc9109713e871

Observation ad369f0b-c399-48fa-95d1-ffc270cccd52 · outbound

This paper cites Language Models are Few-Shot Learners.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Language Models are Few-Shot Learners

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.604539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.604539Z digest=sha256:fc99baf6864ccdfe102c33e0d2a2e32fe31fd883b2020aa471732384b82ba4f4

Observation 14e98da3-0684-481e-b882-6cfd28da27f6 · outbound

This paper cites Coyo-700m: Image-text pair dataset.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Coyo-700m: Image-text pair dataset

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.609153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.609153Z digest=sha256:35fcd890180db4cce26f4da010083c01f62a12fe98280c47fe04a3745e01e293

Observation 221ae02b-dd6b-4cb6-81c1-a3bb34e3218c · outbound

This paper cites End- to-end object detection with transformers.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation End- to-end object detection with transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.613366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.613366Z digest=sha256:15bfc1aa45384de1c38e32e6de0740c528d2c74984376eccd3281e37d5d2e631

Observation 52057327-0ffd-42f7-ae4d-bb2d7947d518 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Emerging properties in self-supervised vision transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.617847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.617847Z digest=sha256:eaa98bd2db4607e442ef402d6f416b15145847e4f9bf6312d8c7b288b16d2dd6

Observation a12d3e7a-41c0-4542-990d-1f922873e342 · outbound

This paper cites Matterport3d: Learning from rgb-d data in indoor environments.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Matterport3d: Learning from rgb-d data in indoor environments

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.622107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.622107Z digest=sha256:f41aced79791bd7b06de40b7ccb6bf1bc656cd556b473a8ed2d8469d80ddd419

Observation 0b88250b-1c5f-434a-9913-e08b0c1daf13 · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Scanrefer: 3d object localization in rgb-d scans using natural language

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.626298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.626298Z digest=sha256:bb76bef4d2c740317e929114cd08bd799f5eb1a87dea9dea23f70797545a56fd

Observation c003712d-e347-4713-aeb7-707cc7e791a3 · outbound

This paper cites Pali: A jointly-scaled multilingual language-image model.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Pali: A jointly-scaled multilingual language-image model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.630797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.630797Z digest=sha256:bcd51561f1999c0652f7de3cd53e07c42736323183aa25302873e96e47cd9fae

Observation bfe2f05f-8f82-4f3c-94e2-f244efeae5c5 · outbound

This paper cites Per-pixel classification is not all you need for semantic segmentation.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Per-pixel classification is not all you need for semantic segmentation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.634940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.634940Z digest=sha256:6674661114f2ff3fb4cc9ae688c08c5d870283440b98b38a121173c4450cb842

Observation f96db358-1b11-4f9d-b955-f74a13d91fa5 · outbound

This paper cites Masked-attention mask transformer for universal image segmentation.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Masked-attention mask transformer for universal image segmentation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.639318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.639318Z digest=sha256:ca29815225d3c3e58a860051d07d2876d6574ce62f2b1244f7ac80eac79340df

Observation 2f5ad14a-f21c-4b58-9ee0-b4ff22b0adde · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Gonzalez, Ion Stoica, and Eric P

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.643885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.643885Z digest=sha256:3fc1b2491f89376c6b8c97b84587ef278ca363f1bcac6bf02a1a28e1b2d69030

Observation 6dd0943b-d02f-406f-a495-24fdc474956f · outbound

This paper cites Cat- seg: Cost aggregation for open-vocabulary semantic seg- mentation.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Cat- seg: Cost aggregation for open-vocabulary semantic seg- mentation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.648054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.648054Z digest=sha256:d3e4c1cde4c5d0e83b7f17ed9ea944fda641fa532cd20ae484568ee6c0d9565c

Observation 0cd8e49c-106c-4e32-8721-6a31597b6b33 · outbound

This paper cites 4d spatio-temporal convnets: Minkowski convolutional neural networks.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation 4d spatio-temporal convnets: Minkowski convolutional neural networks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.652521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.652521Z digest=sha256:853739d0a02bf675cd60476b1bb4da62fc5e9e354766fcb14f4e1bcba18b9f1f

Observation 4fef109b-e5a2-4649-af72-6c41e31eaafb · outbound

This paper cites Scaling instruction- finetuned language models.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Scaling instruction- finetuned language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.656602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.656602Z digest=sha256:94cbe023cf96faf5d035b03f83eda9ec355c3aa720a834a87f6f2efa6ff67298

Observation cf04d106-fd56-40a3-a7d0-198382c3c0c7 · outbound

This paper cites Pointcept: A codebase for point cloud perception research.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Pointcept: A codebase for point cloud perception research

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.660694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.660694Z digest=sha256:197053a76def40b1d1e51e523cbaa81a44ece42a371122bd55721e867e1cf10a

Observation dcbaa040-18d9-4e97-9961-aa7248af709c · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.664581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.664581Z digest=sha256:81d82612dd5b578fb34251f1bdc6773a294a372f352d3c256d2f2ef5aab6a15b

Observation 853ff7e2-95c8-42b9-907e-7b55d17ea351 · outbound

This paper cites Procthor: Large-scale embodied ai using procedural generation.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Procthor: Large-scale embodied ai using procedural generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.668572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.668572Z digest=sha256:3758688eadb2feeee20196bbefd7d145eee86284aaf804a3c2044a507855a954

Observation a9b89b0c-746c-44c2-9241-9152b147eca2 · outbound

This paper cites Pengi: An audio language model for audio tasks.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Pengi: An audio language model for audio tasks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.672837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.672837Z digest=sha256:db76ecc8def72f6e9fa7184aeccf0ad38ddc49a5694b7a9f778cec1ed44257ba

Observation 874ef941-752f-4003-b50c-849466114b63 · outbound

This paper cites Pla: Language-driven open- vocabulary 3d scene understanding.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Pla: Language-driven open- vocabulary 3d scene understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.677012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.677012Z digest=sha256:051c135650985daa9d2c742f911fffc5b1a2c6bfc595b02de0f6f7951f8fc655

Observation 90edc446-c1b2-4171-a472-a8a6645caadf · outbound

This paper cites The Llama 3 Herd of Models.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation The Llama 3 Herd of Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.681025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.681025Z digest=sha256:cc68f7c7ceaaca116ad104d08dd4b9404600ddb01230f7c70764d932b602c5ef

Observation d2a0e2a6-2a47-4af7-a4d8-2eaa6b1782f0 · outbound

This paper cites Efficient graph-based image segmentation.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Efficient graph-based image segmentation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.685595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.685595Z digest=sha256:fdeedda93e33d931d237727bb15c4a3fd1220f579bf156936c5087a450a79038

Observation e7bb280a-4088-4615-a105-42f71715cb38 · outbound

This paper cites Dat- acomp: In search of the next generation of multimodal datasets.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Dat- acomp: In search of the next generation of multimodal datasets

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.689680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.689680Z digest=sha256:f9966d1339833ff1685c3fa31b209ec1066579411fff67d4debd0c568aa2b378

Observation affb818f-864c-437a-8742-acb3439b38eb · outbound

This paper cites Scal- ing open-vocabulary image segmentation with image-level labels.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Scal- ing open-vocabulary image segmentation with image-level labels

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.693601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.693601Z digest=sha256:86f51e4e3b7cb7efa22ee7a22557a3329009b9ca8d2095b0bf6042a71c3cf2e1

Observation 82c286eb-35b4-447f-9fd1-96c11635699c · outbound

This paper cites Imagebind one embedding space to bind them all.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Imagebind one embedding space to bind them all

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.697384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.697384Z digest=sha256:134de3cbec60632cdca7c1e368777173dc4f7e7013241f1dd16c0b25c4c3de13

Observation 80b722c7-3dac-4852-8e14-757d560df0e8 · outbound

This paper cites 3d semantic segmentation with submanifold sparse convolutional networks.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation 3d semantic segmentation with submanifold sparse convolutional networks

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.701027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.701027Z digest=sha256:c25ec48ac21486bf6c25e10a123f708a681ca3958c1a599e8207a2f28b5fa110

Observation 275aa17b-4b30-422c-8f52-a1ebc99847df · outbound

This paper cites Open- vocabulary object detection via vision and language knowl- edge distillation.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Open- vocabulary object detection via vision and language knowl- edge distillation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.705118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.705118Z digest=sha256:ae37191d11991cac2e925d810161b8880e33b70b29df2f1b8a631a9613507b23

Observation 29f12a23-69f1-4df5-9e6f-7afb377fbebb · outbound

This paper cites RegionGPT: Towards Region Understanding Vision Language Model.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation RegionGPT: Towards Region Understanding Vision Language Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.708786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.708786Z digest=sha256:2d91b55dd83e0499dfc6ef506ea46edc90fdcaa6817c73a486bd63206a077da5

Observation 06895320-52ee-4f98-90a0-bc1db054f058 · outbound

This paper cites Deep residual learning for image recognition.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Deep residual learning for image recognition

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.712783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.712783Z digest=sha256:9410267b87b0409ab207021bfbd34ca1b9591ffa5fedad8b5a152645ef18706e

Observation a3949f46-5039-46ba-9463-e09527f2f633 · outbound

This paper cites Denoising diffu- sion probabilistic models.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Denoising diffu- sion probabilistic models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.716714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.716714Z digest=sha256:ad064e4ad3d479c8de43a0327ac0c54d5469ed8980edbdae5df3b30292b6b674

Observation cecbf8e7-2c25-4899-93dc-25671c5a93bb · outbound

This paper cites Scaling up vision-language pre-training for image captioning.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Scaling up vision-language pre-training for image captioning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.720613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.720613Z digest=sha256:363a10a661e3187edbc062d9ba986ab4190df6ab1006f72e1e277bfddc5dd59f

Observation a74ad816-a918-4526-bef7-94ef9b703b70 · outbound

This paper cites An Embodied Generalist Agent in 3D World, 2023.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation An Embodied Generalist Agent in 3D World, 2023

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.724428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.724428Z digest=sha256:2464cb983731c6cf578ce3e0aac814972b0e5c0076186a71cde7e9b21e59a9d9

Observation cfc6e54a-bb57-437c-9fab-32911b613dc9 · outbound

This paper cites Segment3d: Learning fine-grained class-agnostic 3d segmentation without manual labels.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Segment3d: Learning fine-grained class-agnostic 3d segmentation without manual labels

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:16.268078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.728313Z digest=sha256:7105bae7454bbc22dd47f1536226485aaad06fe8000a611818a3815bfb053b41

Observation bfa33195-a310-44db-ab51-a169911857d5 · outbound

This paper cites Open-set image tagging with multi-grained text supervision.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Open-set image tagging with multi-grained text supervision

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:16.252677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.732287Z digest=sha256:3f1faa6568135c060327028e028dc28cbdee19ef8974033ca24654c750992044

Observation daf2bf4e-38a0-4e7a-8daa-9dfb26f4a3d0 · outbound

This paper cites Openins3d: Snap and lookup for 3d open-vocabulary instance segmentation.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Openins3d: Snap and lookup for 3d open-vocabulary instance segmentation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:16.237410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.737048Z digest=sha256:dece29a909b46356d1f548f7dcf6ed7bc9b1044daefc67b33c6f1fd60ea8ec25

Observation b9c9593d-abe8-4039-91b1-fee9aa4083a0 · outbound

This paper cites Oneformer: One transformer to rule universal image segmentation.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Oneformer: One transformer to rule universal image segmentation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.741058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.741058Z digest=sha256:190cfe6dcdb62a4acd99263897d098d7377daf28219edb001be3285caea4773e

Observation 4a45b972-ec9e-4c94-aa03-1ec35f0ccc6b · outbound

This paper cites Scen- eVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding, 2024.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Scen- eVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding, 2024

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:16.210733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.745098Z digest=sha256:71a0f1f126184e9186fc46e76011c5af4cee7c7881e059b6c6dae95137fc8fe7

Observation e412da27-9b92-4f97-a127-2174cc369ab5 · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.749054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.749054Z digest=sha256:7867558dc35851cf04ec2e5845288068dc86007d0bebb7b9e7d24f4affdbcd41

Observation 97cd7bfa-79b8-40f6-83a8-69edb7645e72 · outbound

This paper cites Mistral 7B.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Mistral 7B

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.753930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.753930Z digest=sha256:fad2132e328a2d5bdb4a4c89a811315591b07a690bcf77d5799ecbb597d180ae

Observation a96f6ac5-9a3b-4aa2-b53e-1f4f34c88b0c · outbound

This paper cites Mixtral of Experts.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Mixtral of Experts

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.758083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.758083Z digest=sha256:907327ba3eacae1e64afe402bf866d4efb8a199efa13ccf66e5f0919fc08b8ff

Observation 42d0e6dc-74c8-426e-a0cf-35e1a991544c · outbound

This paper cites Open-vocabulary 3d semantic segmentation with foundation models.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Open-vocabulary 3d semantic segmentation with foundation models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:16.183318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.762576Z digest=sha256:983d3e1c0c47f7d1cb23567b80c4f1127342bbab120211b8f548fccbc35ba1e2

Observation 917ac727-ee91-446c-8584-4546bcfa6273 · outbound

This paper cites In defense of lazy visual grounding for open-vocabulary semantic segmentation.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation In defense of lazy visual grounding for open-vocabulary semantic segmentation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:16.167880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.766649Z digest=sha256:724f8ca43019e63ff8d84c9fe2c9c53a9cd97790d18935660bcec0aed7069ab9

Observation 8d77640e-e6b6-4ee5-a438-891fe4fa787c · outbound

This paper cites Segment any- thing.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Segment any- thing

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:16.153388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.770955Z digest=sha256:469b7954ac6ecc3487b34d25c15273203f4a2346aad2522a853534d6cdad3bb4

Observation 4c119a91-cee5-4b8d-ac9f-154d3a7c83aa · outbound

This paper cites Language-driven semantic seg- mentation.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Language-driven semantic seg- mentation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:16.138563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.775337Z digest=sha256:6b70c95216a0c2b308661ed424158646406d73bed06bd898bf581cc6670277c2

Observation 66a26ccf-570b-4106-bad7-c82817705a69 · outbound

This paper cites Semantic-sam: Segment and recognize anything at any gran- ularity.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Semantic-sam: Segment and recognize anything at any gran- ularity

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:16.122887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.779553Z digest=sha256:665428ee3cef930a9ff2ade3ab449e17a4b10c03bf182a7d72774810740d6064

Observation bd214a08-26b4-432c-a489-3fb65d964819 · outbound

This paper cites an unresolved cited work.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:52:16.107193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.783713Z digest=sha256:9171225d7b549c897f710d5d801784e4beb13b7da635a313356af27a20ebe79e

Observation 4d3bdca8-f72a-4e4d-8435-bdcf36ef16e2 · outbound

This paper cites an unresolved cited work.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:52:16.090482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.788044Z digest=sha256:20b5900c69080149de6158d5e59c67c119c39e05d120c6820d92a45a867b6411

Observation 7d1262d9-2858-43e1-ad4d-6566f6f7459e · outbound

This paper cites What If We Recaption Billions of Web Images with LLaMA-3?.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation What If We Recaption Billions of Web Images with LLaMA-3?

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.792125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.792125Z digest=sha256:7a06513a8880af59b00ed29fb5ac1cc46b35e357e1601d58e9a32fa5426433d2

Observation 8d8d16bd-341f-4c4e-84f8-b04f7a8fb787 · outbound

This paper cites Open-vocabulary semantic segmentation with mask-adapted clip.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Open-vocabulary semantic segmentation with mask-adapted clip

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:16.073384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.796758Z digest=sha256:47878a53e3ec512316da377f441a98344691f08cd5dd17409472ba3dc7f7fa7e

Observation cafe5ede-5d40-403f-86f5-78a1cc20a36f · outbound

This paper cites Improved baselines with visual instruction tuning.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Improved baselines with visual instruction tuning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:16.057811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.801329Z digest=sha256:da7bc999560547136283d9360b784ef0d64f81993f47d558da5c9a2a6e88e8a3

Observation 62176497-bbee-443c-b9d7-63956413a7b2 · outbound

This paper cites Visual instruction tuning.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Visual instruction tuning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:16.042655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.805634Z digest=sha256:14336a7f983867fae5cc5dbec34be91d2c5381f723779b81e8fcaf05138c99d6

Observation 944f78de-385d-4098-9ecf-8240a8a9d096 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.809674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.809674Z digest=sha256:6a5be5ec830bbb815efcff0f2e2a82adac8811521b53ac228a8da43e6b786de3

Observation a038eb75-8d16-46f8-a685-26915a8ddbf2 · outbound

This paper cites Mmscan: A multi-modal 3d scene dataset with hierarchical grounded language annotations.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Mmscan: A multi-modal 3d scene dataset with hierarchical grounded language annotations

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:16.027246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.813845Z digest=sha256:6196e616c41da0e56545a3380266f65d2dbe0dec8bd8466f661d4dca928435f5

Observation 8a65f695-9d47-4f26-9e95-0e56413c5246 · outbound

This paper cites Multiscan: Scalable rgbd scanning for 3d environments with articulated objects.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Multiscan: Scalable rgbd scanning for 3d environments with articulated objects

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:16.011839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.817951Z digest=sha256:8477ceefc79fefcc162f11346063e210fe3bb52fe959a5aa0f8dcd776e49f3fa

Observation 0ef770a6-86a3-494f-bfaa-2c53ab59a942 · outbound

This paper cites V-net: Fully convolutional neural networks for volumetric medical image segmentation.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation V-net: Fully convolutional neural networks for volumetric medical image segmentation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.821919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.821919Z digest=sha256:707d758f1307277b6092f013e20a3ca3e3354c05637b14bf36a6cbf2c3a89bc8

Observation 4f8d6c61-81f5-4fba-bf9e-c89c1417b0c3 · outbound

This paper cites Silc: Improving vision language pretraining with self-distillation.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Silc: Improving vision language pretraining with self-distillation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.983110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.826052Z digest=sha256:40057cdff0a5ad6ccdd927b153030d769511b85d65f40cb998a411573552fea0

Observation d331ff8d-a1fd-4598-8940-d1f297e3f0e6 · outbound

This paper cites Isbnet: a 3d point cloud instance segmentation network with instance- aware sampling and box-aware dynamic convolution.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Isbnet: a 3d point cloud instance segmentation network with instance- aware sampling and box-aware dynamic convolution

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.967578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.830000Z digest=sha256:9dae5d6f580942749d56f83c4bd95bea2e7e625b3f32940d7ef823c53d96aeac

Observation 52f7004c-a463-42a1-bf5f-11bcb7e40aa8 · outbound

This paper cites Open3dis: Open-vocabulary 3d instance segmentation with 2d mask guidance.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Open3dis: Open-vocabulary 3d instance segmentation with 2d mask guidance

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.951605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.833920Z digest=sha256:f6c2f2fb0a6b80fe5e2c75e7195642575e544ef7747cbed1273a70bcb335c9a7

Observation db49205e-1102-46ab-ae86-1f6b00356dcc · outbound

This paper cites Improved denoising diffusion probabilistic models.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Improved denoising diffusion probabilistic models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.838076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.838076Z digest=sha256:2dddbebf5cf1554bd3431e64192d4b674a0b709a9a32eaef537b59de83ab2d82

Observation 437c7dd3-6051-45ba-9b30-a6969133c011 · outbound

This paper cites an unresolved cited work.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:52:15.926916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.842102Z digest=sha256:1f518f4aa4cc296f7e357faf4ce7551f77da558b20644a1a9d364b44c2394aff

Observation 227a1564-54f1-4295-aa9e-21da05878fa6 · outbound

This paper cites Openscene: 3d scene understanding with open vocabularies.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Openscene: 3d scene understanding with open vocabularies

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.911926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.846000Z digest=sha256:eea6a552cf283f6282433507a9b30dc56b7bb5e98167db48e42422b540e0e000

Observation e3969332-b124-442c-9ea3-239664801938 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.850056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.850056Z digest=sha256:f8c8e88be4ed227165ba2d0c1ed142b6de8512bbaae341839912df5654ff004c

Observation 22966ade-2886-49b3-abe2-24aae0b57603 · outbound

This paper cites Language models are un- supervised multitask learners.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Language models are un- supervised multitask learners

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.895454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.854046Z digest=sha256:cf1881bfb4439b75c490b8996cf402476959bbfd26ccad3a0ee0c629f5f78769

Observation de4e563d-3480-4b5d-8558-0a78c389876e · outbound

This paper cites Learning transferable visual models from natural language supervision.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Learning transferable visual models from natural language supervision

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.858027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.858027Z digest=sha256:a0a1ef72e9791a67dbfa34ae153d872629a811458fd24c79ec3aef4de1ce6189

Observation dbc9ddaf-8ff0-46df-b640-d0ec3ed21df0 · outbound

This paper cites Habitat-matterport 3d dataset (hm3d): 1000 large-scale 3d environments for embodied ai.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Habitat-matterport 3d dataset (hm3d): 1000 large-scale 3d environments for embodied ai

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.868396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.862319Z digest=sha256:61daec741618995d278236eae575c0b96970371fb3cbf1646c42cbb2e80eef16

Observation 7e015112-44c9-4cd7-b4a4-ab2e5fb1ec0d · outbound

This paper cites Denseclip: Language-guided dense prediction with context- aware prompting.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Denseclip: Language-guided dense prediction with context- aware prompting

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.852879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.866163Z digest=sha256:d846d927a40bd6b40f6ebb4ea65e36af5f444f3bedacc39d6bd9e6a655429c6e

Observation 680b86ee-8ea0-4327-b090-2a045849b70e · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Glamm: Pixel grounding large multimodal model

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.838090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.870527Z digest=sha256:0582587de1cc76985ac9a98f16e53cf924fa58a98b07ea7b68657f7796a2c6fe

Observation da088572-6e7b-41b3-bd14-4a7eebcd8b9d · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation SAM 2: Segment Anything in Images and Videos

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.874411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.874411Z digest=sha256:b81aab72538e67ef20282652c1bf1d03e2cafce77ec508c3a9784ae573582074

Observation bea92ce4-0e2f-48d5-bfc5-03a1add15019 · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.878648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.878648Z digest=sha256:f5dda85d586930966188c3f7f23b6ddcfe0c21f11cbc6f86e7848b341a9ded25

Observation 7479ea9b-c09b-45e0-9c16-0c470f089aaf · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation High-resolution image synthesis with latent diffusion models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.882773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.882773Z digest=sha256:dee3ef6e6a8cf4bfbdffe34b71da442f692bd9e14c7bac02d8aaff0ecb9096bd

Observation 0ac63485-9b44-4471-86aa-71f807630c69 · outbound

This paper cites Language- grounded indoor 3d semantic segmentation in the wild.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Language- grounded indoor 3d semantic segmentation in the wild

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.812509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.887128Z digest=sha256:1ed225c6f509788c8765767ee10ec7dda5a334de9257469d29b727876690c0dd

Observation b75cad6e-adb1-4754-95cd-024c8708cfb9 · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.891226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.891226Z digest=sha256:e08302289cc528ff4a091135c239b3a45515cf0b5133c93bf315d2f786bbeb7c

Observation eba8f216-4af8-4c68-a3f2-efbf85eb62a0 · outbound

This paper cites A multi-view stereo benchmark with high- resolution images and multi-camera videos.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation A multi-view stereo benchmark with high- resolution images and multi-camera videos

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.796445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.895671Z digest=sha256:369331782011b77e885d1e95cec6223e8e06c65d4491e1c5e5065a0c07c20f56

Observation bd752179-a75a-4e49-ba6f-49596dfcaea6 · outbound

This paper cites Laion-5b: An open large-scale dataset for train- ing next generation image-text models.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Laion-5b: An open large-scale dataset for train- ing next generation image-text models

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.781293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.899890Z digest=sha256:b49a88f6e8cfe4f23fab067d90d1976846551e772ae5954895d7f1347da92d06

Observation d72dade2-eaf3-41b5-99fb-0a94abe84dc8 · outbound

This paper cites Mask3d: Mask trans- former for 3d semantic instance segmentation.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Mask3d: Mask trans- former for 3d semantic instance segmentation

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.766568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.904387Z digest=sha256:8deedca10c35697b9f140686fd85ec67eaa783ef17c61ee8bf65ccc4c6a398dd

Observation 23570bcc-5127-4d60-a5d3-c548bbd66bc3 · outbound

This paper cites Clip-fields: Weakly supervised semantic fields for robotic memory.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Clip-fields: Weakly supervised semantic fields for robotic memory

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.751438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.908510Z digest=sha256:d328c909d278fc36b0043e2d99b971600ad91db704ba5962d7625b03ac741eec

Observation 2ba6bcdc-5af8-442d-adf5-0b97d7c0c9fc · outbound

This paper cites Super-convergence: Very fast training of neural networks using large learning rates.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Super-convergence: Very fast training of neural networks using large learning rates

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.736995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.912701Z digest=sha256:e84fdf8416e8367cc4f41a2e7d6e3445907c801a4fab20af6d988d2da37df504

Observation 9053103e-006e-4379-8a24-b4af73fafbb2 · outbound

This paper cites Denois- ing diffusion implicit models.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Denois- ing diffusion implicit models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.916736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.916736Z digest=sha256:07ccedf5cbc37ab30b7c592c02b363aa6ab3143561daa5ff464670ed22a85051

Observation 6e7b02f5-b163-4c87-a68f-c672a457b773 · outbound

This paper cites Score-based generative modeling through stochastic differential equa- tions.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Score-based generative modeling through stochastic differential equa- tions

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.712154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.920752Z digest=sha256:fd2ae42a4909553316d65d4869e1201204fbe602cfb1fd7abf370d14b8f5931c

Observation dc84ceaa-8ca9-4668-ae2d-b27557d85d5b · outbound

This paper cites Open- mask3d: open-vocabulary 3d instance segmentation.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Open- mask3d: open-vocabulary 3d instance segmentation

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.697981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.924759Z digest=sha256:920121cdd7aa8bc4347cfd5d9d41ef3983b70a8a7da42b1f3b45a5d4f8b36820

Observation 90173931-af52-49d2-824e-73679f5b0cc6 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Gemini: A Family of Highly Capable Multimodal Models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.928740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.928740Z digest=sha256:14c4cd5484e7b7cdd59e182a4e6a7aed0b4e0168ffc2b4f04309ecf642c6c7fe

Observation af2fa7ac-9477-476a-a3d5-343a25669593 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.933150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.933150Z digest=sha256:ed17db05b3dcf9da1a1c4644c269de21b439d9fd1eb1e8d540fefb3394161fa4

Observation 3713a9ab-650f-44a5-9ac9-244e55a82ef9 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation LLaMA: Open and Efficient Foundation Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.937675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.937675Z digest=sha256:fe8d0d9e27f9102ac465a20d1bd8e492a9e4261cc3253508c92c783efac41b96

Observation f4414da6-319f-434a-9e94-603cb6c7a9c9 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.942007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.942007Z digest=sha256:0108ef2d365a9947c3aa194eb60fa472c2dab6d3a91b774d88a33ec1dd681317

Observation 188efec1-78d5-45b2-b1e5-ce0fcf114127 · outbound

This paper cites Rio: 3d object instance re- localization in changing indoor environments.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Rio: 3d object instance re- localization in changing indoor environments

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.684118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.946399Z digest=sha256:d8b85df8fc7c551933f3dc3dd637fa139e5446f56d4db76d2f06548fcb576acf

Observation 6a5fa897-59a9-4aa5-9e79-71927bfd4a50 · outbound

This paper cites Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learn- ing framework.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learn- ing framework

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.670541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.950455Z digest=sha256:75508ba52366bb03dd9fb15417c56cee3a2e95186acda8df5b2cd251caa388e8

Observation 36b75abf-d93b-40da-a930-dae7958998e8 · outbound

This paper cites EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI, 2023.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI, 2023

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.655917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.954462Z digest=sha256:0af28b07d8ffb04d512142248fd892b80335d7a622c26a7c3b4bfcb926e29347

Observation 1e765161-49db-4cfa-b061-16ce0c0d2cd3 · outbound

This paper cites Towards Large-scale 3D Representation Learning with Multi-dataset Point Prompt Training.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Towards Large-scale 3D Representation Learning with Multi-dataset Point Prompt Training

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.958546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.958546Z digest=sha256:3e8641e0f14493a2e9b93abb464358f8f2a0ad8b883dade899e336d335af4a0c

Observation 6d02b364-431b-4fcb-9a34-24d44812b2f1 · outbound

This paper cites Groupvit: Semantic segmentation emerges from text supervision.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Groupvit: Semantic segmentation emerges from text supervision

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.641360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.962797Z digest=sha256:9b1932655497186c8e99d2114284c2f5ca8dafdd5cb3fb7bf90a06c2d7ccfd3f

Observation 8efbfbf8-8e95-46fb-9a5c-67288c2bc6d8 · outbound

This paper cites Qwen2 Technical Report.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Qwen2 Technical Report

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.966702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.966702Z digest=sha256:05da954b4f156db8bdd8b2a5c16d0ee12ab9e8c2241965793935e0e5717fb626

Observation 9b00930e-bdc8-4255-a703-030886c09744 · outbound

This paper cites Regionplc: Regional point-language contrastive learning for open-world 3d scene understanding.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Regionplc: Regional point-language contrastive learning for open-world 3d scene understanding

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.627270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.971103Z digest=sha256:3949236d3f8bbae30d95658c177868ba92a5b909779293c6025be216fe8c85f8

Observation 3b50bde3-04dd-4cc0-8a10-3c8987d93792 · outbound

This paper cites Scannet++: A high-fidelity dataset of 3d in- door scenes.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Scannet++: A high-fidelity dataset of 3d in- door scenes

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.612383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.975066Z digest=sha256:0b2b909f847b93ffbcdbb1d6679704c1cda6e00e73a69b441599cd063c9be30c

Observation 5879e755-4333-4f1e-968f-6e636d3a7a22 · outbound

This paper cites Sai3d: Segment any instance in 3d scenes.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Sai3d: Segment any instance in 3d scenes

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.598253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:52:14.978847Z digest=sha256:ba4f0ff36fe3974f5e484569ad3dfa76b6319cc7b076f2ae60d4142d5e61f9eb

Pith citing papers

No inbound Pith citation observations are available.