Pith. sign in

Paper Citation Record · LEDGER

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding

As of 9 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2506.22817.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22817 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:04:26.911809Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact1
  • verified fuzzy50
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d862c8fb-0622-4778-bee2-d0b2969728d3 · outbound

This paper cites Open-YOLO 3D: Towards Fast and Accurate Open-Vocabulary 3D Instance Segmentation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open-YOLO 3D: Towards Fast and Accurate Open-Vocabulary 3D Instance Segmentation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:22.010410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:22.010410Z digest=sha256:fa3c34ec1bb7ffeede5a853cc3cb0c947b0a9ae02d2d0c0ebf9b5570d78bd169

Observation 0a5c8f2f-b8cc-417e-b8bc-0075048536cb · outbound

This paper cites Matterport3d: Learning from rgb-d data in indoor environments.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Matterport3d: Learning from rgb-d data in indoor environments

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:36.605437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:22.071577Z digest=sha256:91eb788827b5ee2148a60bcd277fe06b2017d563248e7bbd0b9b4c57b51b300b

Observation 817ea282-ef21-49ab-90b6-975c251dccff · outbound

This paper cites Clip2scene: Towards label-efficient 3d scene understanding by clip.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Clip2scene: Towards label-efficient 3d scene understanding by clip

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:22.157362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:22.157362Z digest=sha256:9e4ac3f4d519ce7b90237c0263535051c17cb23460fc3288387cc639879efbd2

Observation 221a3ce0-75cf-48e6-acc7-849d3156425a · outbound

This paper cites Yolo-world: Real-time open-vocabulary object detection.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Yolo-world: Real-time open-vocabulary object detection

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:36.427717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:22.249266Z digest=sha256:faf19f3280c7d7289e2ed529b42b4489d3ff8fdc2fdc2924f704a39e42b2ecb1

Observation aade28f8-34cc-4f69-8720-4b9de7222f8d · outbound

This paper cites 4d spatio-temporal convnets: Minkowski convolutional neu- ral networks.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding 4d spatio-temporal convnets: Minkowski convolutional neu- ral networks

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:36.252491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:22.355844Z digest=sha256:df5348848335f05b730073085a34e87d623afac22adb83e64b3f771987bb5e2a

Observation e325d129-4696-43da-a644-0f31ed7de852 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:36.057704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:22.437904Z digest=sha256:cb0dd0b8ff14abd230757a6eda075e7c02dee36f9db15cd8c72899678c599d08

Observation 2c98312e-0c7c-4ea7-ab69-f040c085cf36 · outbound

This paper cites Pla: Language-driven open- vocabulary 3d scene understanding.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Pla: Language-driven open- vocabulary 3d scene understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:35.820884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:22.513194Z digest=sha256:c9a76703412ff4cb519c83a0b76213d0d03001ef694629791c76466bf124bf0c

Observation d1c40434-2502-4382-a8c3-10198e2ea9ce · outbound

This paper cites A density-based algorithm for discovering clusters in large spatial databases with noise.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding A density-based algorithm for discovering clusters in large spatial databases with noise

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:35.597980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:22.618688Z digest=sha256:165b616490513b81d1418f7a9abd30d79b254f11754574e400ef70953017dd49

Observation da4eb8b9-77c3-4f66-a972-72f2d89e7aac · outbound

This paper cites Efficient graph-based image segmentation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Efficient graph-based image segmentation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:35.336113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:22.699936Z digest=sha256:03b763eedb658014533dfaa818d366444b86b2dd89d447fe3ee06ab80a630c49

Observation 1f8aed7b-1828-4be3-91f1-360252aefda2 · outbound

This paper cites Scal- ing open-vocabulary image segmentation with image-level labels.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Scal- ing open-vocabulary image segmentation with image-level labels

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:35.102671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:22.777482Z digest=sha256:ff3e1aea9506804836beeaaba49d9c743d2f249e92fb4ab2b3284b016d7f938d

Observation 85ed6f15-1795-40cc-a14d-7ed5193835f3 · outbound

This paper cites Sam-guided graph cut for 3d instance segmentation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Sam-guided graph cut for 3d instance segmentation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:34.842643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:22.821605Z digest=sha256:d0cb32464c66552c58cfab7e64b6ac012fda03c90081202fd17be95e8a0c6948

Observation c56b82a9-481e-4cc5-b0d2-88f168114bda · outbound

This paper cites spacy 2: Natural lan- guage understanding with bloom embeddings, convolutional neural networks and incremental parsing.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding spacy 2: Natural lan- guage understanding with bloom embeddings, convolutional neural networks and incremental parsing

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:34.620721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:22.909272Z digest=sha256:b50b1a83055cb3a3577224d324831ad497f559704c25d28e1d76347ae80c2c01

Observation e2ae8e89-bc58-444b-9583-e40b705df8d1 · outbound

This paper cites Open-Set Image Tagging with Multi-Grained Text Supervision.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open-Set Image Tagging with Multi-Grained Text Supervision

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:04:27.162531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:22.991770Z digest=sha256:5bac95b4e1366f720f24372abde2518fb58f4d5dd743cd1fda433d3d14227233

Observation c22408fb-30cc-454f-ba23-2975b517077a · outbound

This paper cites Odin: A single model for 2d and 3d segmentation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Odin: A single model for 2d and 3d segmentation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:34.404545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:23.080658Z digest=sha256:b2cfa8aa47050337db39b4ef8808951e7162791fb3513ed8d9de82f3d42f7e84

Observation e9ab477e-a527-40b0-9369-b7991f928d63 · outbound

This paper cites Con- ceptfusion: Open-set multimodal 3d mapping.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Con- ceptfusion: Open-set multimodal 3d mapping

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:34.160072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:23.155746Z digest=sha256:48b46c1e5ee526763c83e5efc44f88751614bbaa1fa030c0f2edba2dfb9d68f8

Observation 02e28432-4ef9-4e2a-9420-90ca782c20ab · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:33.833358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:23.203242Z digest=sha256:fa59976ac195330ee6f7e324bd2b4df073f64bbde25ef1b4ae01217f72857edc

Observation 0aedb603-85e6-4dfb-8e34-285348d27db4 · outbound

This paper cites Pointgroup: Dual-set point group- ing for 3d instance segmentation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Pointgroup: Dual-set point group- ing for 3d instance segmentation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:33.536514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:23.293731Z digest=sha256:95d6b3271c55bef5bce8742c5fa91c260deeb10e3f83803fa658877662e6ce78

Observation c1f5d4af-1562-41c8-9058-30f798b0e6a8 · outbound

This paper cites Open-vocabulary 3d semantic segmentation with foundation models.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open-vocabulary 3d semantic segmentation with foundation models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:33.256176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:23.352475Z digest=sha256:987d5478b0949456c65762935a7340e0607f73933f290be12b0a1f8892d338c0

Observation 481c95b7-e3f8-4e1a-8427-60d192cda968 · outbound

This paper cites Segment any- thing.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Segment any- thing

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:32.977129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:23.440915Z digest=sha256:cdf673b3ca5984d3dd54bad945fb30380c6e0d4c9ac146409e6b09bbab2ac2eb

Observation d55d3f15-4868-472e-90ca-007d5cf55a39 · outbound

This paper cites Segment Any 3D Object with Language.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Segment Any 3D Object with Language

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:23.521003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:23.521003Z digest=sha256:ca85aac6927dcfbda143550abe0afb49f0043699ab9e32caae1d988e58fc6b55

Observation 87f496db-a38c-4b33-be66-2d9be5002250 · outbound

This paper cites Language-driven semantic seg- mentation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Language-driven semantic seg- mentation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:32.711837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:23.650481Z digest=sha256:c835ec5d1bd35c282fb2ade48554fd0c33c16dacdc50ab7270958f72b28ca7fd

Observation c277b8fd-9aff-4d86-a887-ac667d61faeb · outbound

This paper cites Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:32.445243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:23.704502Z digest=sha256:d9bceb7cca77b8d3467df5f8efcd367263d283d22a3424cee8192841212fa340

Observation ed9a6c5e-9f19-464d-81d0-5ecd83f37dae · outbound

This paper cites Grounded language-image pre-training.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Grounded language-image pre-training

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:32.237398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:23.809703Z digest=sha256:e705ea8beed34a01830fd8eda53b8c8fb61e3d759ee84fb3972a67a7e634ddfd

Observation 2d752310-2e3d-42d9-9dcc-5ae62c6f966b · outbound

This paper cites Decap: Decoding clip latents for zero-shot captioning via text-only 2 training.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Decap: Decoding clip latents for zero-shot captioning via text-only 2 training

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:32.035731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:23.876577Z digest=sha256:dfa28bd5ab71bf502dc1ee4dde9ecbac3fa65b7c3f912d1c6d9405596c369dae

Observation ef56cedc-8d59-4cab-be92-20c9776cffaf · outbound

This paper cites Open-vocabulary semantic segmentation with mask-adapted clip.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open-vocabulary semantic segmentation with mask-adapted clip

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:23.977911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:23.977911Z digest=sha256:827ae078f3a96c7e24a18d9a6e8741acbd2422449576e40261ea7b85ce07f0d8

Observation 49a463a6-d0dd-44dc-a926-ee60422fd869 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:24.039336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:24.039336Z digest=sha256:0f7bed8e4158c5a60ba19aab76ea8abf3379d68f184d775263174f2efd919e57

Observation 2bef844e-217b-4ee7-89b5-d1e83e7972a0 · outbound

This paper cites Ovir-3d: Open-vocabulary 3d in- stance retrieval without training on 3d data.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Ovir-3d: Open-vocabulary 3d in- stance retrieval without training on 3d data

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:31.758417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:24.101366Z digest=sha256:e956f1cfdf64a5dad62c41a6dcd1758dc2da291538e759f6204dd9cc902450c7

Observation 78e18653-9ae1-4f41-ba5d-378c4fac1c9e · outbound

This paper cites An end-to- end transformer model for 3d object detection.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding An end-to- end transformer model for 3d object detection

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:24.177736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:24.177736Z digest=sha256:a320c8fe2d825f840514d06dff4326b4cc990dec0bafab954735fc2cf85f8758

Observation 1ded09a3-37f7-4e31-a942-01732b4775f8 · outbound

This paper cites Isbnet: a 3d point cloud instance segmentation network with instance- aware sampling and box-aware dynamic convolution.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Isbnet: a 3d point cloud instance segmentation network with instance- aware sampling and box-aware dynamic convolution

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:31.504554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:24.257808Z digest=sha256:f254cac0a7ae398f60d45d9803925c02941114f0cd8ab660fc9b643d3108457f

Observation ff8a22de-058e-4002-811f-b1881c0ff054 · outbound

This paper cites Open3dis: Open-vocabulary 3d instance segmentation with 2d mask guidance.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open3dis: Open-vocabulary 3d instance segmentation with 2d mask guidance

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:31.288532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:24.338008Z digest=sha256:2fd45bff4cc605fe7c40ff01f1a0474a27c962d2f1616d1a481111529186f129

Observation 6ea5db78-d815-4c22-a769-3b417885b848 · outbound

This paper cites V oxel cloud connectivity segmentation- supervoxels for point clouds.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding V oxel cloud connectivity segmentation- supervoxels for point clouds

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:31.020530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:24.406111Z digest=sha256:3a8815850e413e4003fd23fbb6a0b44d607041bda469271a780db9e090e29759

Observation 6d8875fd-fabc-4f8b-b423-6a146be19250 · outbound

This paper cites Openscene: 3d scene understanding with open vocabularies.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Openscene: 3d scene understanding with open vocabularies

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:30.783522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:24.455441Z digest=sha256:a085a53223b2bb4a6adb916dd0fff84a7c7e144848374d04f3a7412c72ae8e60

Observation 41261cd9-0f92-471a-be01-3c8e93769ee4 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:24.595455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:24.595455Z digest=sha256:427052b7b27db9182084f94f05848add3fe39be0b824fba306def9379e53c7d9

Observation f8a9ab7c-b2b5-4246-b961-4ba0d4178294 · outbound

This paper cites Pointnet++: Deep hierarchical feature learning on point sets in a metric space.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Pointnet++: Deep hierarchical feature learning on point sets in a metric space

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:30.530755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:24.712330Z digest=sha256:af3dcf35d057a949c29541de03f47bca0a1eb1a232ac2a9e125d8f0771a0c8c6

Observation 1245972d-d8c0-4a64-99c6-c2ec725bae19 · outbound

This paper cites Pointnext: Revisiting pointnet++ with improved training and scaling strategies.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Pointnext: Revisiting pointnet++ with improved training and scaling strategies

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:30.233921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:24.813661Z digest=sha256:ca39fab3250db5db82cf05390a18081a0e72d9e8ac195e197914ac1e313e509b

Observation e783c82f-3083-4223-ae58-2b2dabd020fd · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Learn- ing transferable visual models from natural language super- vision

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:30.018565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:24.906952Z digest=sha256:b8181de5873fb5269be0017464579cca0ea3931d5ab41200c9edb8cae7a2a87a

Observation ce3af4e4-8fb1-45ef-843c-4c957316bc50 · outbound

This paper cites Language- grounded indoor 3d semantic segmentation in the wild.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Language- grounded indoor 3d semantic segmentation in the wild

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:29.777998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:24.978354Z digest=sha256:bb2bbed5cf6ae7a88c1f32b721bcfe19c14c2d69fdb253b82e14a37166be3586

Observation 87199b5b-123c-4058-af14-2459ac33dc25 · outbound

This paper cites Dense multimodal alignment for open-vocabulary 3d scene understanding.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Dense multimodal alignment for open-vocabulary 3d scene understanding

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:29.587167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:25.074672Z digest=sha256:75c0eeede29e600957c092c686e3bdce2f4faf43832fa762b93a4e0a39749125

Observation caba26fd-c95e-4145-a951-3415bfd46ee5 · outbound

This paper cites Mask3d: Mask trans- former for 3d semantic instance segmentation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Mask3d: Mask trans- former for 3d semantic instance segmentation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:29.417674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:25.178981Z digest=sha256:6ce985cbc2b04843e9188f093b95fa50d497c66927ff9f4d16c3f383cc6df548

Observation 144baabf-9704-493e-af1d-e4065f70590c · outbound

This paper cites Pointr- cnn: 3d object proposal generation and detection from point cloud.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Pointr- cnn: 3d object proposal generation and detection from point cloud

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:25.277779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:25.277779Z digest=sha256:a802ca0ffb0f3877430641058bb8fc34f72b38a557ccce86e72699cf176de5cd

Observation e1fc3a7d-c0f7-4299-946c-047ecd3e4457 · outbound

This paper cites The Replica Dataset: A Digital Replica of Indoor Spaces.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding The Replica Dataset: A Digital Replica of Indoor Spaces

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:25.352234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:25.352234Z digest=sha256:5a6ae23cf0bc895fe8f6f857a1fb0bd13e61ec3c6f8b993c99991449c65aea4d

Observation 29eb27ca-f877-4a1e-afe7-a95ac6682dbb · outbound

This paper cites Open- mask3d: open-vocabulary 3d instance segmentation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open- mask3d: open-vocabulary 3d instance segmentation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:29.244663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:25.418913Z digest=sha256:57220e70aea7434e88c2b8cbe87693e2e8b43052c0e08e1907cbe92d250086f3

Observation 06d30a65-8d4b-412b-b94d-c319ffb436ab · outbound

This paper cites Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:29.152314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:25.484739Z digest=sha256:7079d199e13f34b055d6b61156449053f34078776d200f39df77a9709c4bc599

Observation 6e55810b-e7e8-4d45-a634-fb19c99fbd84 · outbound

This paper cites Open vocabulary 3d scene under- standing via geometry guided self-distillation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open vocabulary 3d scene under- standing via geometry guided self-distillation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:29.046440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:25.564953Z digest=sha256:dcc35135a78405198a1d7c97427226ae04e3125b45148b08de7cb6313fc65556

Observation c2c0fcc5-b4fe-4ca7-943f-3b736b1ba0d6 · outbound

This paper cites Detr3d: 3d object detection from multi-view images via 3d-to-2d queries.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Detr3d: 3d object detection from multi-view images via 3d-to-2d queries

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:28.937754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:25.629129Z digest=sha256:17e9b139b24fc47a31279d4721295d7214b1e4f3058444c28dae6f8ed6bd7015

Observation f98a4414-c302-4e2e-bc9c-b6af28403af1 · outbound

This paper cites Uni3detr: Unified 3d detection trans- former.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Uni3detr: Unified 3d detection trans- former

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:28.837411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:25.674003Z digest=sha256:d87df2d284b139c7e088fbe4bf08951ffc3e55988b22b313217262519074bd38

Observation d53afc04-3ee3-4b33-aaa1-191d221b7f27 · outbound

This paper cites Point transformer v3: Simpler faster stronger.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Point transformer v3: Simpler faster stronger

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:28.715963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:25.735877Z digest=sha256:ec5b5611d78b12ca291bd7e5033b1d49351ca58b41d1128b092f4d8b6bdaf404

Observation b13d381d-c0aa-4ba0-933c-e9943c23fb9b · outbound

This paper cites Open-vocabulary panop- 3 tic segmentation with text-to-image diffusion models.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open-vocabulary panop- 3 tic segmentation with text-to-image diffusion models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:28.633199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:25.810601Z digest=sha256:e85d26681c782e205b1a6512e064a3bc89179b2f0a48032bb298b606bc7568ee

Observation fc4587d7-fd21-485f-be34-8dd8e54e0d60 · outbound

This paper cites Paconv: Position adaptive convolution with dy- namic kernel assembling on point clouds.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Paconv: Position adaptive convolution with dy- namic kernel assembling on point clouds

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:28.543457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:25.853406Z digest=sha256:a74f6d55cfb6c2db85297499958ad373a79c242ba17d33c8ea399380a5691d15

Observation a1465d24-4171-442b-b92f-ff475a070587 · outbound

This paper cites SAMPro3D: Locating SAM Prompts in 3D for Zero-Shot Instance Segmentation.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding SAMPro3D: Locating SAM Prompts in 3D for Zero-Shot Instance Segmentation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:25.946778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:25.946778Z digest=sha256:504c3230110e31c194a99235636a3eefe817170d050a2b87b0d72d56e770890c

Observation 25cb3969-22a2-40b6-ad59-feb63a1b2f00 · outbound

This paper cites A unified framework for 3d scene understanding.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding A unified framework for 3d scene understanding

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:28.453467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:26.007884Z digest=sha256:a1a03fc7b34b363bc59198d86dc94bf0749264ac60b0aac2fab704c2c37b8bb4

Observation 1bad85a9-7d29-4045-b462-e393dfbb24f5 · outbound

This paper cites Regionplc: Regional point-language contrastive learning for open-world 3d scene understanding.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Regionplc: Regional point-language contrastive learning for open-world 3d scene understanding

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:28.365669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:26.049587Z digest=sha256:cdaebbfdcb473021e8bdea2d5b3919b88832793635e769d59d0aea9b6eb8372a

Observation 86e4464a-03c7-4444-9c8e-6d82710e2d65 · outbound

This paper cites Sa3dip: Segment any 3d instance with potential 3d priors.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Sa3dip: Segment any 3d instance with potential 3d priors

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:28.252504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:26.130338Z digest=sha256:d30c13b48ae921c6e179b4291acfbb4afd7f6bc1ac7cf729651d80fe0dc6497d

Observation 57d72f2d-c14b-4814-b926-729e614e05bd · outbound

This paper cites SAM3D: Segment Anything in 3D Scenes.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding SAM3D: Segment Anything in 3D Scenes

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:26.211071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:26.211071Z digest=sha256:18dae75b015611a71374ee83e8ca754de60e43bff5e9c4663fd77a3260ef612e

Observation 4c503099-ae0c-4611-8aa9-392a2abfb99c · outbound

This paper cites Point deformable network with enhanced nor- mal embedding for point cloud analysis.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Point deformable network with enhanced nor- mal embedding for point cloud analysis

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:28.148885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:26.321851Z digest=sha256:edda74f868c846f2e8dfa850082ddc85dfc04bd3e740cb7a84b89f379c799b27

Observation a4edc5ae-8f26-490b-8637-de3c652bf272 · outbound

This paper cites Sai3d: Segment any instance in 3d scenes.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Sai3d: Segment any instance in 3d scenes

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:28.062390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:26.399770Z digest=sha256:275107498940ce96452a2766f9de33039347331a3e6ac106f0bc63f25081e073

Observation 87f8f197-2039-4b13-a046-9793e04d82f7 · outbound

This paper cites Convolutions die hard: Open-vocabulary seg- mentation with single frozen convolutional clip.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Convolutions die hard: Open-vocabulary seg- mentation with single frozen convolutional clip

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:27.968280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:26.485535Z digest=sha256:2d661c834d008d54809779bdedb3429c9b6b951c9c2e21bdd38e3f8ee911166b

Observation d9e1f414-50e4-40cc-afdc-03647a7ff35b · outbound

This paper cites Recognize anything: A strong image tagging model.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Recognize anything: A strong image tagging model

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:27.875800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:26.585203Z digest=sha256:4f70fa6d77fc0d34e931e657e6e35406abbb3ce1e73d3062636964cab0bcab80

Observation 4399e95b-2d94-4dc2-a7c2-d401a625b62f · outbound

This paper cites Point transformer.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Point transformer

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:27.789413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:26.664878Z digest=sha256:622e4bde722922145504f68a5abbccdd291cca79c2f02fb6708baa4c250cb793

Observation 87774ca9-b572-4089-94c6-198c6dd3e27e · outbound

This paper cites Se- ssd: Self-ensembling single-stage object detector from point cloud.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Se- ssd: Self-ensembling single-stage object detector from point cloud

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:27.655081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:26.765134Z digest=sha256:f6a82c6b945b3d49cc17460d47bc523ed3d11890d17bbd9e1db2e8012b5ccb01

Observation 9dd8fe42-6988-4e8c-aafe-49a356da40d7 · outbound

This paper cites Detecting twenty-thousand classes using image-level supervision.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Detecting twenty-thousand classes using image-level supervision

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:27.511718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:26.846855Z digest=sha256:ff7c9719fd91e13ee003a046d1f17c874c9ad6d4c15960cba4364d5beffba336

Observation 0eefbddc-203e-45da-a043-65961e46fda9 · outbound

This paper cites Open-vocabulary 3d semantic segmentation with text-to-image diffusion models.

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding Open-vocabulary 3d semantic segmentation with text-to-image diffusion models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:27.346606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:04:26.911809Z digest=sha256:7a1d5cde421796dad813986006e86967247d0ab84375f8a9ee21084a4181eb4e

Pith citing papers

No inbound Pith citation observations are available.