Pith. sign in

Paper Citation Record · LEDGER

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding

As of 20 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2506.22817.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22817 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:04:26.911809Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact1
  • verified fuzzy50
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d862c8fb-0622-4778-bee2-d0b2969728d3 · outbound

This paper cites Open-YOLO 3D: Towards Fast and Accurate Open-Vocabulary 3D Instance Segmentation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open-YOLO 3D: Towards Fast and Accurate Open-Vocabulary 3D Instance Segmentation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:22.010410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:22.010410Z digest=sha256:75556edb83d7b1f4d9b3ce68a47cd4cfc54433bb73a54c7075009d38cf7a52d8

Observation 0a5c8f2f-b8cc-417e-b8bc-0075048536cb · outbound

This paper cites Matterport3d: Learning from rgb-d data in indoor environments.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Matterport3d: Learning from rgb-d data in indoor environments

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:36.605437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:22.071577Z digest=sha256:e2e29ed71ae19959aae6afcf62f21ba7ce8bfdada1bd6f7b310808821e044774

Observation 817ea282-ef21-49ab-90b6-975c251dccff · outbound

This paper cites Clip2scene: Towards label-efficient 3d scene understanding by clip.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Clip2scene: Towards label-efficient 3d scene understanding by clip

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:22.157362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:22.157362Z digest=sha256:bf8eb4a1057d8727b32a13ffeec1998b84862fc499a19729bec8779baf03436d

Observation 221a3ce0-75cf-48e6-acc7-849d3156425a · outbound

This paper cites Yolo-world: Real-time open-vocabulary object detection.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Yolo-world: Real-time open-vocabulary object detection

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:36.427717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:22.249266Z digest=sha256:2c7750ef92f11b404d95ef667bffcd82f3e456bf650e9d45ddec3997f4793481

Observation aade28f8-34cc-4f69-8720-4b9de7222f8d · outbound

This paper cites 4d spatio-temporal convnets: Minkowski convolutional neu- ral networks.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding 4d spatio-temporal convnets: Minkowski convolutional neu- ral networks

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:36.252491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:22.355844Z digest=sha256:5ac77a78ced3a7085ba51e7601fcb767e269ebbd86f45096ed4e727f2ce4bb0b

Observation e325d129-4696-43da-a644-0f31ed7de852 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:36.057704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:22.437904Z digest=sha256:0d0b08553caa0633a13851c84a0a49e9977ef2be30c23aa348069913b309b26f

Observation 2c98312e-0c7c-4ea7-ab69-f040c085cf36 · outbound

This paper cites Pla: Language-driven open- vocabulary 3d scene understanding.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Pla: Language-driven open- vocabulary 3d scene understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:35.820884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:22.513194Z digest=sha256:455d411bfb5629475230ac946d540205c4dd773d7a54c001455daedc07e13a97

Observation d1c40434-2502-4382-a8c3-10198e2ea9ce · outbound

This paper cites A density-based algorithm for discovering clusters in large spatial databases with noise.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding A density-based algorithm for discovering clusters in large spatial databases with noise

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:35.597980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:22.618688Z digest=sha256:d6b9d02f63e34263740652a75bb6402fd78f308cf3368153b82b20ec7df9222e

Observation da4eb8b9-77c3-4f66-a972-72f2d89e7aac · outbound

This paper cites Efficient graph-based image segmentation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Efficient graph-based image segmentation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:35.336113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:22.699936Z digest=sha256:2976ac21ad21dbe792a93f529e8a5bdaafbd6107005c24fa96a74152cb24bd75

Observation 1f8aed7b-1828-4be3-91f1-360252aefda2 · outbound

This paper cites Scal- ing open-vocabulary image segmentation with image-level labels.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Scal- ing open-vocabulary image segmentation with image-level labels

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:35.102671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:22.777482Z digest=sha256:5a58e8bd8b261187c28879cb4be3d1933eefc10dd094b7803689277b2cefe960

Observation 85ed6f15-1795-40cc-a14d-7ed5193835f3 · outbound

This paper cites Sam-guided graph cut for 3d instance segmentation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Sam-guided graph cut for 3d instance segmentation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:34.842643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:22.821605Z digest=sha256:52a264ac435971512d879909e390488b161b98c858815eac18a54fec77a57803

Observation c56b82a9-481e-4cc5-b0d2-88f168114bda · outbound

This paper cites spacy 2: Natural lan- guage understanding with bloom embeddings, convolutional neural networks and incremental parsing.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding spacy 2: Natural lan- guage understanding with bloom embeddings, convolutional neural networks and incremental parsing

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:34.620721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:22.909272Z digest=sha256:4cd25e0e5af35b885b5176ae151770bbbf535d4b833b6d46717cb898e89f6cf9

Observation e2ae8e89-bc58-444b-9583-e40b705df8d1 · outbound

This paper cites Open-Set Image Tagging with Multi-Grained Text Supervision.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open-Set Image Tagging with Multi-Grained Text Supervision

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:04:27.162531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:22.991770Z digest=sha256:d3b3cb74cb599b14cbb1f65b2dc225af61d23185dd4dd94293f45f144c426959

Observation c22408fb-30cc-454f-ba23-2975b517077a · outbound

This paper cites Odin: A single model for 2d and 3d segmentation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Odin: A single model for 2d and 3d segmentation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:34.404545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:23.080658Z digest=sha256:75304df89def88c2a3762ee40dd4d22e77997b73f46cff8c44434836f86e0fd9

Observation e9ab477e-a527-40b0-9369-b7991f928d63 · outbound

This paper cites Con- ceptfusion: Open-set multimodal 3d mapping.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Con- ceptfusion: Open-set multimodal 3d mapping

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:34.160072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:23.155746Z digest=sha256:efe4f603cb71ef8065cc1dc16ff000e6ccda20a59e03018a05adff38126b26fd

Observation 02e28432-4ef9-4e2a-9420-90ca782c20ab · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:33.833358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:23.203242Z digest=sha256:a5366dc5b105f5680f5aceea5f8457901f586042192aa7281cd836bd7bb9ef4f

Observation 0aedb603-85e6-4dfb-8e34-285348d27db4 · outbound

This paper cites Pointgroup: Dual-set point group- ing for 3d instance segmentation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Pointgroup: Dual-set point group- ing for 3d instance segmentation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:33.536514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:23.293731Z digest=sha256:80416f4fb2feedc9c5444a8bd9fccbf18f2686cc84d45a84f3cc8f81bf9109b6

Observation c1f5d4af-1562-41c8-9058-30f798b0e6a8 · outbound

This paper cites Open-vocabulary 3d semantic segmentation with foundation models.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open-vocabulary 3d semantic segmentation with foundation models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:33.256176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:23.352475Z digest=sha256:097f22020b4c7bcece3f523d2a38441b3bdd03d77abf0ce730cd9dcf12908f65

Observation 481c95b7-e3f8-4e1a-8427-60d192cda968 · outbound

This paper cites Segment any- thing.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Segment any- thing

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:32.977129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:23.440915Z digest=sha256:34105452bf801cc7f556b237d95607015447813b8aec497350439b8bbd30d2e0

Observation d55d3f15-4868-472e-90ca-007d5cf55a39 · outbound

This paper cites Segment Any 3D Object with Language.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Segment Any 3D Object with Language

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:23.521003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:23.521003Z digest=sha256:c4f944642bdf4899717762317e80ed5aecd9fdd25b2eeeba226324cd2bbd8aa1

Observation 87f496db-a38c-4b33-be66-2d9be5002250 · outbound

This paper cites Language-driven semantic seg- mentation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Language-driven semantic seg- mentation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:32.711837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:23.650481Z digest=sha256:2c6a9d28fdcb94167bf7f6a0a7f569dd8b177a73f71293b74290d29e894ec337

Observation c277b8fd-9aff-4d86-a887-ac667d61faeb · outbound

This paper cites Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:32.445243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:23.704502Z digest=sha256:ca5890418b70cafb17a2336a9a85741b44c65e128aa9a51ba366949e7ac3c4bc

Observation ed9a6c5e-9f19-464d-81d0-5ecd83f37dae · outbound

This paper cites Grounded language-image pre-training.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Grounded language-image pre-training

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:32.237398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:23.809703Z digest=sha256:a0b341ad26f1c5ee0ac8d454cd1d4d2d7e201a3e9a3cdd25f0496823b6dffb97

Observation 2d752310-2e3d-42d9-9dcc-5ae62c6f966b · outbound

This paper cites Decap: Decoding clip latents for zero-shot captioning via text-only 2 training.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Decap: Decoding clip latents for zero-shot captioning via text-only 2 training

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:32.035731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:23.876577Z digest=sha256:e39f4cae4c42dc885ab76680c676b926d86939286210955916a086ce14d8b4b8

Observation ef56cedc-8d59-4cab-be92-20c9776cffaf · outbound

This paper cites Open-vocabulary semantic segmentation with mask-adapted clip.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open-vocabulary semantic segmentation with mask-adapted clip

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:23.977911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:23.977911Z digest=sha256:07d5ad5adabbc57e6d190afa88081bb9c8d8b5f3a9aa98bfe94d6a0a7d0b5995

Observation 49a463a6-d0dd-44dc-a926-ee60422fd869 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:24.039336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:24.039336Z digest=sha256:8dfa67c00bb01b70614874615d3ee34d915aa964c799ecded9daf0891de6c157

Observation 2bef844e-217b-4ee7-89b5-d1e83e7972a0 · outbound

This paper cites Ovir-3d: Open-vocabulary 3d in- stance retrieval without training on 3d data.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Ovir-3d: Open-vocabulary 3d in- stance retrieval without training on 3d data

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:31.758417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:24.101366Z digest=sha256:8bcd0c6677cb388d88bd55007756c5e7224af328e057029aa60982012b813986

Observation 78e18653-9ae1-4f41-ba5d-378c4fac1c9e · outbound

This paper cites An end-to- end transformer model for 3d object detection.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding An end-to- end transformer model for 3d object detection

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:24.177736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:24.177736Z digest=sha256:d091c73139000d23c5909eaf68277c7bb4ed65ef59e86700b29bc4ad90e5c2a3

Observation 1ded09a3-37f7-4e31-a942-01732b4775f8 · outbound

This paper cites Isbnet: a 3d point cloud instance segmentation network with instance- aware sampling and box-aware dynamic convolution.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Isbnet: a 3d point cloud instance segmentation network with instance- aware sampling and box-aware dynamic convolution

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:31.504554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:24.257808Z digest=sha256:3db07676f9e092979f6a7645aedd3fe9f5080e698a8d8a89b220bfbef922f77d

Observation ff8a22de-058e-4002-811f-b1881c0ff054 · outbound

This paper cites Open3dis: Open-vocabulary 3d instance segmentation with 2d mask guidance.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open3dis: Open-vocabulary 3d instance segmentation with 2d mask guidance

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:31.288532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:24.338008Z digest=sha256:615be825f58b4cec01d92ea6e954522309b9e554ec70c126c5008c919b9057b2

Observation 6ea5db78-d815-4c22-a769-3b417885b848 · outbound

This paper cites V oxel cloud connectivity segmentation- supervoxels for point clouds.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding V oxel cloud connectivity segmentation- supervoxels for point clouds

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:31.020530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:24.406111Z digest=sha256:842e147f595279a69ad67b83aa1241f1d466c8c96e9312a98b8befe9fc502216

Observation 6d8875fd-fabc-4f8b-b423-6a146be19250 · outbound

This paper cites Openscene: 3d scene understanding with open vocabularies.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Openscene: 3d scene understanding with open vocabularies

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:30.783522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:24.455441Z digest=sha256:182a19c4d58d5504932269cced8cd1b92ecf8a9be97799d189ebff3f717492af

Observation 41261cd9-0f92-471a-be01-3c8e93769ee4 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:24.595455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:24.595455Z digest=sha256:e0b802edd7dd4b78e54fd84c398dfd1b4812a3b75d0ecf9f20e8069959b19254

Observation f8a9ab7c-b2b5-4246-b961-4ba0d4178294 · outbound

This paper cites Pointnet++: Deep hierarchical feature learning on point sets in a metric space.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Pointnet++: Deep hierarchical feature learning on point sets in a metric space

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:30.530755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:24.712330Z digest=sha256:ccb50faaa6a4aed1b913bb4a2c3382115f44ec5d8fe58f5749494454dcce668d

Observation 1245972d-d8c0-4a64-99c6-c2ec725bae19 · outbound

This paper cites Pointnext: Revisiting pointnet++ with improved training and scaling strategies.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Pointnext: Revisiting pointnet++ with improved training and scaling strategies

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:30.233921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:24.813661Z digest=sha256:e65150d5a04874b60c8f2b1b73cc9a9ad591782397ad07fa06b2d5e4ad9d4838

Observation e783c82f-3083-4223-ae58-2b2dabd020fd · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Learn- ing transferable visual models from natural language super- vision

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:30.018565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:24.906952Z digest=sha256:db84fa4b1304c3ba40a11c8b964b93b106d7fd9fac49c75c1afb6a58eb9a1e71

Observation ce3af4e4-8fb1-45ef-843c-4c957316bc50 · outbound

This paper cites Language- grounded indoor 3d semantic segmentation in the wild.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Language- grounded indoor 3d semantic segmentation in the wild

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:29.777998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:24.978354Z digest=sha256:182dceeda633897e74a4e3ea348583aa403253b9fa1c0d1c80255eee0bee66a0

Observation 87199b5b-123c-4058-af14-2459ac33dc25 · outbound

This paper cites Dense multimodal alignment for open-vocabulary 3d scene understanding.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Dense multimodal alignment for open-vocabulary 3d scene understanding

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:29.587167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:25.074672Z digest=sha256:d7532e4cf60230e595e4780f40b058583b0da3292cd4a8f1e753f95f7fa1743a

Observation caba26fd-c95e-4145-a951-3415bfd46ee5 · outbound

This paper cites Mask3d: Mask trans- former for 3d semantic instance segmentation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Mask3d: Mask trans- former for 3d semantic instance segmentation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:29.417674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:25.178981Z digest=sha256:c7e76b593cd3e476975d0025e489da1b4d29c355ccaaed909e74a8039e574887

Observation 144baabf-9704-493e-af1d-e4065f70590c · outbound

This paper cites Pointr- cnn: 3d object proposal generation and detection from point cloud.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Pointr- cnn: 3d object proposal generation and detection from point cloud

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:25.277779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:25.277779Z digest=sha256:b622ca0eba729c1ed5becda6e4e80c494638a22a07583a7c962d5f2b90436a03

Observation e1fc3a7d-c0f7-4299-946c-047ecd3e4457 · outbound

This paper cites The Replica Dataset: A Digital Replica of Indoor Spaces.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding The Replica Dataset: A Digital Replica of Indoor Spaces

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:25.352234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:25.352234Z digest=sha256:09a0a9486550541e15d16c386ae98c7df622626365403c812310c3309bb7c443

Observation 29eb27ca-f877-4a1e-afe7-a95ac6682dbb · outbound

This paper cites Open- mask3d: open-vocabulary 3d instance segmentation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open- mask3d: open-vocabulary 3d instance segmentation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:29.244663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:25.418913Z digest=sha256:9d6bfc259df74116b38dadc0c9a0fc7c960846aff3e1a579066c9a6d60697ed9

Observation 06d30a65-8d4b-412b-b94d-c319ffb436ab · outbound

This paper cites Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:29.152314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:25.484739Z digest=sha256:50d64fc2e6c58a49ea41c266fcbfcf67555c1d058bde8f92e74fbe7af1562f74

Observation 6e55810b-e7e8-4d45-a634-fb19c99fbd84 · outbound

This paper cites Open vocabulary 3d scene under- standing via geometry guided self-distillation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open vocabulary 3d scene under- standing via geometry guided self-distillation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:29.046440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:25.564953Z digest=sha256:4b3d8bc18a1a2a1f53e1d4fb4c941fb13d47ba3bef152576136c02111ec368fd

Observation c2c0fcc5-b4fe-4ca7-943f-3b736b1ba0d6 · outbound

This paper cites Detr3d: 3d object detection from multi-view images via 3d-to-2d queries.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Detr3d: 3d object detection from multi-view images via 3d-to-2d queries

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:28.937754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:25.629129Z digest=sha256:86fd02a2bc124d0cc70f77556a8e022e3ffd03f00cbb113066069b4dfed6b4b1

Observation f98a4414-c302-4e2e-bc9c-b6af28403af1 · outbound

This paper cites Uni3detr: Unified 3d detection trans- former.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Uni3detr: Unified 3d detection trans- former

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:28.837411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:25.674003Z digest=sha256:5821f81a17a5f8194e345df0cb978f34a7e6330870f341cef8a422b26f87d1e2

Observation d53afc04-3ee3-4b33-aaa1-191d221b7f27 · outbound

This paper cites Point transformer v3: Simpler faster stronger.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Point transformer v3: Simpler faster stronger

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:28.715963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:25.735877Z digest=sha256:5e9f40fc117e0a0c062d4c64bc9b7ee8c16b91f59ade26f34c53bc2dbb4382b4

Observation b13d381d-c0aa-4ba0-933c-e9943c23fb9b · outbound

This paper cites Open-vocabulary panop- 3 tic segmentation with text-to-image diffusion models.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open-vocabulary panop- 3 tic segmentation with text-to-image diffusion models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:28.633199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:25.810601Z digest=sha256:92fb2206759f26c71ff17010cbf8547e4cd043824b37cd62b0ba3da3b90a4e1c

Observation fc4587d7-fd21-485f-be34-8dd8e54e0d60 · outbound

This paper cites Paconv: Position adaptive convolution with dy- namic kernel assembling on point clouds.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Paconv: Position adaptive convolution with dy- namic kernel assembling on point clouds

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:28.543457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:25.853406Z digest=sha256:99cc3ace37dd6402d99a9f74c30f0fc8c1109d1cbc0b6ba8e848363b6d0b71e9

Observation a1465d24-4171-442b-b92f-ff475a070587 · outbound

This paper cites SAMPro3D: Locating SAM Prompts in 3D for Zero-Shot Instance Segmentation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding SAMPro3D: Locating SAM Prompts in 3D for Zero-Shot Instance Segmentation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:25.946778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:25.946778Z digest=sha256:e6cf13094e407e4f9ee9055a64b1bba1e55bcc7d4cd57ea2e2eec92f90f2e8cb

Observation 25cb3969-22a2-40b6-ad59-feb63a1b2f00 · outbound

This paper cites A unified framework for 3d scene understanding.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding A unified framework for 3d scene understanding

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:28.453467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:26.007884Z digest=sha256:322d837d1204278c6f50a4e7e2588ce615b5548516bb871b2806d1348846e3f5

Observation 1bad85a9-7d29-4045-b462-e393dfbb24f5 · outbound

This paper cites Regionplc: Regional point-language contrastive learning for open-world 3d scene understanding.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Regionplc: Regional point-language contrastive learning for open-world 3d scene understanding

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:28.365669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:26.049587Z digest=sha256:26ca3781c7e2b14d51bdcae7478056a3c04d704d3a3ca664cd9006cd597db975

Observation 86e4464a-03c7-4444-9c8e-6d82710e2d65 · outbound

This paper cites Sa3dip: Segment any 3d instance with potential 3d priors.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Sa3dip: Segment any 3d instance with potential 3d priors

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:28.252504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:26.130338Z digest=sha256:dc1c17968c1b4c57399a837cfaa0c5ab29c8d21d59c3537d0ac1cc05429124cc

Observation 57d72f2d-c14b-4814-b926-729e614e05bd · outbound

This paper cites SAM3D: Segment Anything in 3D Scenes.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding SAM3D: Segment Anything in 3D Scenes

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:26.211071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:26.211071Z digest=sha256:bdf6ffc84c617ea8be4532d222a6f0e5eb59fce60c944e3d72d2db7dd523cc47

Observation 4c503099-ae0c-4611-8aa9-392a2abfb99c · outbound

This paper cites Point deformable network with enhanced nor- mal embedding for point cloud analysis.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Point deformable network with enhanced nor- mal embedding for point cloud analysis

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:28.148885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:26.321851Z digest=sha256:7275b29ee134b65578b8ccd10409117e2b9f9b2638a8bc387f7a56026a190da6

Observation a4edc5ae-8f26-490b-8637-de3c652bf272 · outbound

This paper cites Sai3d: Segment any instance in 3d scenes.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Sai3d: Segment any instance in 3d scenes

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:28.062390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:26.399770Z digest=sha256:8dcdd942b1b9fff9c1830ebce0caab6492dfb3d1bc3d8935631a33017881a73e

Observation 87f8f197-2039-4b13-a046-9793e04d82f7 · outbound

This paper cites Convolutions die hard: Open-vocabulary seg- mentation with single frozen convolutional clip.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Convolutions die hard: Open-vocabulary seg- mentation with single frozen convolutional clip

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:27.968280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:26.485535Z digest=sha256:921c689afd4e25fae6daab3b002b36dcb1c9c31b26eb4b7ced2801710e181d8e

Observation d9e1f414-50e4-40cc-afdc-03647a7ff35b · outbound

This paper cites Recognize anything: A strong image tagging model.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Recognize anything: A strong image tagging model

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:27.875800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:26.585203Z digest=sha256:ff283e33c86632bab6e34d7c10d9e7492e96bbe622e677c55e888ccb04946707

Observation 4399e95b-2d94-4dc2-a7c2-d401a625b62f · outbound

This paper cites Point transformer.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Point transformer

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:27.789413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:26.664878Z digest=sha256:4dfa33cb56e601dc231db438bf2aba20c37a8316ac048cefec54c3a4f8da2fa1

Observation 87774ca9-b572-4089-94c6-198c6dd3e27e · outbound

This paper cites Se- ssd: Self-ensembling single-stage object detector from point cloud.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Se- ssd: Self-ensembling single-stage object detector from point cloud

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:27.655081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:26.765134Z digest=sha256:7547b8ee60dc207a594c3d0b75af5df605d20c967026495832af25d565eae3aa

Observation 9dd8fe42-6988-4e8c-aafe-49a356da40d7 · outbound

This paper cites Detecting twenty-thousand classes using image-level supervision.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Detecting twenty-thousand classes using image-level supervision

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:27.511718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:26.846855Z digest=sha256:8fc7666495dd2de01c6039fa30577d17d1a53929febef30c16784bbc0d434e62

Observation 0eefbddc-203e-45da-a043-65961e46fda9 · outbound

This paper cites Open-vocabulary 3d semantic segmentation with text-to-image diffusion models.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open-vocabulary 3d semantic segmentation with text-to-image diffusion models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:27.346606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:04:26.911809Z digest=sha256:a87614f5c197b2fa868bf9da6e179978ecc483c8f7fd0b06c900fad8ab89d891

Pith citing papers

No inbound Pith citation observations are available.