Pith. sign in

Paper Citation Record · LEDGER

Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 59 inbound Pith citation observations for arXiv:2309.00615.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.00615 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 59 of 59 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:03:16.395096Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:09:29.321784Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c7239216-2861-4a65-ace9-fbe23b8d32f4 · inbound

LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention cites this paper.

LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 109

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T23:07:42.622601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-14T23:07:42.245641Z digest=sha256:00f1a623273f396f97cb39329a07687dc794700807b716514a7de5dce7d75d9c

Observation 2b2ea280-db75-4de4-bd15-237819b722e0 · inbound

SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models cites this paper.

SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:03:26.883172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T03:03:26.723464Z digest=sha256:0af50110cd6134450118adf51f0269c311205e85b263f8f3bdb6d776f67b97f4

Observation 63e58c97-9553-4247-933f-e64629c117d9 · inbound

MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems? cites this paper.

MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems? Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T01:29:30.121405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T01:29:30.032408Z digest=sha256:d435ede6c937e51a9ba70b01ef09767d7a2f47fd2edcebc256e674106638b367

Observation 080577c8-4f50-4d7f-a6ab-4f3b2ac2b91e · inbound

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models cites this paper.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:01:54.075378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:56b6a6d8f4c905caaaae54261b359ae1ef056c70aac559d4dabf015c64c86192

Observation ee4d18be-d8a8-4d3a-8433-57f5ae56d34d · inbound

LLaVA-OneVision: Easy Visual Task Transfer cites this paper.

LLaVA-OneVision: Easy Visual Task Transfer Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:23:49.552473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:7228eb846016e46fa3d120b73d2e505b3dccaa8277e04d58b667c0892f3188d9

Observation 6b4cb0ed-a327-4aff-827a-e27a4f09a8dc · inbound

Polymath: A Challenging Multi-modal Mathematical Reasoning Benchmark cites this paper.

Polymath: A Challenging Multi-modal Mathematical Reasoning Benchmark Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:05:47.810561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-23T20:03:38.336841Z digest=sha256:5ca554e3c2af14d0675b05a13db227e8014de2152af0a4221727066d6e81f6e5

Observation e5f36373-22e4-4d0a-8927-2f1f0ca5b3ce · inbound

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation cites this paper.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.898818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.898818Z digest=sha256:8f0788eda671b938cc88df96ba9258d53eeb31a1f9ec7f68cbe58c2a951a171e

Observation 79aa4343-2b7f-46b8-9a49-10707bb728af · inbound

Multimodal 3D Reasoning Segmentation with Complex Scenes cites this paper.

Multimodal 3D Reasoning Segmentation with Complex Scenes Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T15:47:37.115653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:47:37.115653Z digest=sha256:b7f4496a5cee2c4c172d5a7c37aba2405e5ea8aebf29b6a54c443e60e0f65e9d

Observation c42e3d4c-4679-49c7-b7ef-9aff57d943c2 · inbound

MICAS: Multi-grained In-Context Adaptive Sampling for 3D Point Cloud Processing cites this paper.

MICAS: Multi-grained In-Context Adaptive Sampling for 3D Point Cloud Processing Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T13:35:40.800734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:35:40.800734Z digest=sha256:2448590f7734a9b6424ad168f564ad44f3f8ba7f25385c35bf762720c2814f3c

Observation d718883a-8478-4927-8f57-37d366037e8d · inbound

SAR3D: Autoregressive 3D Object Generation and Understanding via Multi-scale 3D VQVAE cites this paper.

SAR3D: Autoregressive 3D Object Generation and Understanding via Multi-scale 3D VQVAE Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T12:52:10.167237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:52:10.167237Z digest=sha256:4aae4504ba83d8e9ece2ddd10c1803a6344c5e7b87088f725203453b33f19493

Observation e390967d-dcdd-425c-ac04-c8266c754a76 · inbound

Lift3D Foundation Policy: Lifting 2D Large-Scale Pretrained Models for Robust 3D Robotic Manipulation cites this paper.

Lift3D Foundation Policy: Lifting 2D Large-Scale Pretrained Models for Robust 3D Robotic Manipulation Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:12.864682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:12.864682Z digest=sha256:c02d969bba57de264a271cea906385b95cfdfa583ecac0c27c8442cfc7745178

Observation a4448104-fa27-41da-bf9f-d48299e08c4d · inbound

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences cites this paper.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.836304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.836304Z digest=sha256:d893f54350c236591cb3f60d7643ef6ebcc85f1332654bee9604bc14e1d65b4f

Observation c601d7a0-ab0d-4afa-92ed-e20d9e70784c · inbound

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models cites this paper.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.815043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.815043Z digest=sha256:369c95c85ba93c27b2f02e4da4a0d0ed6356ec9954eb6db6af32f997774cdc8f

Observation 296da122-4568-4268-9d3a-54963c89ba5a · inbound

Towards Modality Generalization: A Benchmark and Prospective Analysis cites this paper.

Towards Modality Generalization: A Benchmark and Prospective Analysis Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T04:55:17.081012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:55:17.081012Z digest=sha256:97108009efc9130c96e0399e34bbb1830bc7125d791dc059434017f097277e87

Observation 088d3bb2-e42d-4c43-84b7-89d1750a40bf · inbound

GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models cites this paper.

GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:55.301412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:55.301412Z digest=sha256:37c4aa397a40acd181dd302fb36f179124244487c7aa01a45401efdb6b2f25e5

Observation 3a9a78d0-17a9-47c7-9e3f-86ad738ffc39 · inbound

Hyperbolic Contrastive Learning for Hierarchical 3D Point Cloud Embedding cites this paper.

Hyperbolic Contrastive Learning for Hierarchical 3D Point Cloud Embedding Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:56.250068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:56.250068Z digest=sha256:3504f8e0fd911fb2df4a88ac61df28989b3f658c67ca8464315bb885d9d8f732

Observation ea930fc1-ced0-4f12-be11-471d0703ff7e · inbound

CL3DOR: Contrastive Learning for 3D Large Multimodal Models via Odds Ratio on High-Resolution Point Clouds cites this paper.

CL3DOR: Contrastive Learning for 3D Large Multimodal Models via Odds Ratio on High-Resolution Point Clouds Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T21:48:26.988637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:48:26.988637Z digest=sha256:dc6d4574a192345fce3c12284c3af9d04ee97ffc355cd6ed5411bbcbe264ad44

Observation 147724f4-f3a0-4a5f-928d-2feeb035c551 · inbound

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models cites this paper.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.517200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.517200Z digest=sha256:5364337a460e3321b7e8e914c284649003331166721f76b102234a24a1e7fa8a

Observation 6d5d7674-a9f3-4190-b90c-9bcdb95d3c78 · inbound

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step cites this paper.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.195022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.195022Z digest=sha256:f861eaf8fb5dc7222eaa4bc789c9e527ab5fad5404663790087426e1dbdc608a

Observation 24ffe4b5-2dae-4888-a3ad-20a8aa220429 · inbound

Scaling LLaNA: Advancing NeRF-Language Understanding Through Large-Scale Training cites this paper.

Scaling LLaNA: Advancing NeRF-Language Understanding Through Large-Scale Training Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T12:03:16.395096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:03:16.395096Z digest=sha256:8efb29eeb781ea4f72645b42a5933bdbdbb25e2cba8cea2005a822627fffd6fa

Observation 57daca8a-8f45-4737-a16f-2604620fa634 · inbound

VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models cites this paper.

VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:47.123124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:33:47.123124Z digest=sha256:4846f3ea944403deb447b2014b37dbdc4f73693aa84de1041e1c3ee6edf47560

Observation 4c30594e-b689-472e-bb37-e0bffc51e018 · inbound

PICO: Reconstructing 3D People In Contact with Objects cites this paper.

PICO: Reconstructing 3D People In Contact with Objects Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T10:39:54.386245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:39:54.386245Z digest=sha256:c6d0d802cf9dc4ba996a3e86a547ef583cd40548d91492ef0822ed959282c063

Observation 582c9234-5cd1-4879-b6e5-f2e013dbc122 · inbound

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding cites this paper.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:27.435522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:27.435522Z digest=sha256:46cb6c021c756d6aca36eb58a22a466ae07e53568d303acf96f7ec44ffbd56cd

Observation fb9951c6-01da-41c6-b870-43e634eac887 · inbound

SVL: Empowering Spiking Neural Networks for Efficient 3D Open-World Understanding cites this paper.

SVL: Empowering Spiking Neural Networks for Efficient 3D Open-World Understanding Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-22T02:14:31.402477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T02:12:29.058875Z digest=sha256:1427aeec2ce0eddf6a8d3c56248b3cd82bd77dacf40717b1677a2ef945ec45d1

Observation 1fb51303-b186-4311-81fd-ed0387cc064f · inbound

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts cites this paper.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:13.736700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:13.736700Z digest=sha256:3ed6d693adbcdc3f3803561ea3527422efaa743905422d15119db649876f18fc

Observation 6dcc3e57-1339-43a5-89b4-aa16ea337daa · inbound

Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs cites this paper.

Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:25:44.471372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:25:44.471372Z digest=sha256:65ae86f348f9dcecd390cccaa61ab62240905622e9b39c5f51739965ca23f9f8

Observation a0b0bc7b-2015-4230-ba85-88eeb169b2dd · inbound

MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning cites this paper.

MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:48.831388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:48.831388Z digest=sha256:8fa63ab9bc2eae7b7164b26456355f9ff6cbaea808fbf4eb659ebfdfb15cee06

Observation 8c99c45f-392a-4f2b-81e0-ccb75022e912 · inbound

Aligning Proteins and Language: A Foundation Model for Protein Retrieval cites this paper.

Aligning Proteins and Language: A Foundation Model for Protein Retrieval Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:49.128007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:50:49.128007Z digest=sha256:20d2e8d851fd3a36277e62bc51717f759e2fdbec46f03e735190aee845c52c70

Observation b9d52e38-0094-444f-aa1f-65c4f97ce632 · inbound

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations cites this paper.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:51.681819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:51.681819Z digest=sha256:adeba55594b7cad346271753a7182011d81c3407be6d5778c3a16df81ff720eb

Observation 44a72786-e440-4241-b3d0-a9436f35572d · inbound

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation cites this paper.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.369243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.369243Z digest=sha256:76f589cbbb22122bf024af56f1da6a2d962539ae8e787afe9cfe32fe5b364c0c

Observation 18c0dbd6-b05f-4c15-a4df-c38a6f616566 · inbound

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding cites this paper.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:59.359288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:59.359288Z digest=sha256:565d9fbf2b6fdc68dd1b96bd445f052ae489863f2520615947fbe0560860382a

Observation 09301636-1c60-43b7-806a-0de568cf3f9b · inbound

Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models cites this paper.

Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T16:42:13.251847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:42:13.251847Z digest=sha256:69a03401af91e20309761a9f3fb869f37501161ef17db46d553bf4068935f530

Observation cb278c4f-bce7-469a-87bf-a312ab79e322 · inbound

NeuroVoxel-LM: Language-Aligned 3D Perception via Dynamic Voxelization and Meta-Embedding cites this paper.

NeuroVoxel-LM: Language-Aligned 3D Perception via Dynamic Voxelization and Meta-Embedding Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T13:48:02.658757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:48:02.658757Z digest=sha256:15a022657de064b3c6357f9210cad170cd033ae0a607f850b3e18b1e2ebeebc9

Observation 024d5114-7247-4ace-a7d7-45e71e900995 · inbound

RelMap: Reliable Spatiotemporal Sensor Data Visualization via Imputative Spatial Interpolation cites this paper.

RelMap: Reliable Spatiotemporal Sensor Data Visualization via Imputative Spatial Interpolation Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T05:45:38.656582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:45:38.656582Z digest=sha256:6f7c9c5401492a7084cf464374230f5ba9da4f3e1234a169106b625ff500c6ba

Observation d63d3400-c36f-4480-b124-c841d9bcdf57 · inbound

MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh cites this paper.

MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T05:46:02.903719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:46:02.903719Z digest=sha256:b37867badd70d59f597857f9f23558ece595b063d0e2726179490ab5628bd679

Observation 5d699655-fa0b-4b15-8b7d-d26d123166e9 · inbound

Correspondence as Video: Test-Time Adaption on SAM2 for Reference Segmentation in the Wild cites this paper.

Correspondence as Video: Test-Time Adaption on SAM2 for Reference Segmentation in the Wild Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T21:53:38.546874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:53:38.546874Z digest=sha256:51b1d1f458314cec37bcba4c8f24ff8ebc9a7f8d253a8124e22f817f874fb097

Observation 5f32692a-f101-4fee-a90b-a1b6bcbbc49b · inbound

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints cites this paper.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:09.736041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:07:09.736041Z digest=sha256:a0d89aa893f8a0a1143fde451645c295ad2cbf77f682232544f6ac3ca0972465

Observation 84f5614f-93ca-4b8c-8136-58b77355c6a4 · inbound

Calibrated Multimodal Representation Learning with Missing Modalities cites this paper.

Calibrated Multimodal Representation Learning with Missing Modalities Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:10:22.660828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T22:08:07.217659Z digest=sha256:387e1383c6cfae981c8c94cc44e19d42e194612898dbaf3d3659e1a26d723537

Observation e61121aa-809c-44e4-a755-e5b43e9b7d83 · inbound

Rethinking Multimodal Few-Shot 3D Point Cloud Segmentation: From Fused Refinement to Decoupled Arbitration cites this paper.

Rethinking Multimodal Few-Shot 3D Point Cloud Segmentation: From Fused Refinement to Decoupled Arbitration Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T12:50:01.034749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:50:01.034749Z digest=sha256:dbfa5cef7c0efac7c78172f35119ea419139d02861c4f357a4621a608acd55c6

Observation 26024ed0-c9be-4d31-be8a-7c3e922ab0e1 · inbound

Pointy - A Lightweight Transformer for Point Cloud Foundation Models cites this paper.

Pointy - A Lightweight Transformer for Point Cloud Foundation Models Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:30:01.609235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T13:28:23.934522Z digest=sha256:fd9631230186a37609a3a7abf0926286b14800e59abe1ad8c8fa929135eb7bb3

Observation 860ced58-9fcd-4dd9-a5c6-08df71a734a3 · inbound

Feeling the Space: Egomotion-Aware Video Representation for Efficient and Accurate 3D Scene Understanding cites this paper.

Feeling the Space: Egomotion-Aware Video Representation for Efficient and Accurate 3D Scene Understanding Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T09:30:22.462415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T09:30:03.668178Z digest=sha256:e27b9264b1749ba94a69e190644c25072cdd274230b2c477df4bf29665dd9096

Observation fc0d30ca-9ace-47ee-aae3-8a25e85ce2a6 · inbound

Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding cites this paper.

Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T17:54:04.904102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:54:04.904102Z digest=sha256:4aeb1bca33098be374ec1cc86d0dbc9b1f7c2610546a8cc04beea63876dd58b8

Observation 2a95df81-557a-4628-91a3-784c68deacc3 · inbound

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM cites this paper.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:32:59.409437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:7d03b93205fb5bfc5c38778237c0c621073fe433dac892b366ca8f82f5d052d3

Observation 5ec2897e-d7c3-4518-a28f-f3a1f3b05512 · inbound

RGB-Pointmap Pretraining for Unified 3D Scene Understanding cites this paper.

RGB-Pointmap Pretraining for Unified 3D Scene Understanding Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T21:23:17.900121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T21:19:49.421653Z digest=sha256:736e81f02233a629245e0eba96dab9f57319cc31f5b511186d40f8d06b7a1a9b

Observation 8ad60e09-435a-4793-aa0d-2884ed868abf · inbound

Efficient3D: A Unified Framework for Adaptive and Debiased Token Reduction in 3D MLLMs cites this paper.

Efficient3D: A Unified Framework for Adaptive and Debiased Token Reduction in 3D MLLMs Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:48:11.404582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T19:45:33.950587Z digest=sha256:e12e7288d1bbddfcbfdf2d0c87c62bb9c45415b14a0a809a57cb65b6cc1fbb6f

Observation bb3797ac-0573-40bb-893f-d23c12f67e7d · inbound

Beyond Surface Artifacts: Capturing Shared Latent Forgery Knowledge Across Modalities cites this paper.

Beyond Surface Artifacts: Capturing Shared Latent Forgery Knowledge Across Modalities Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:30:59.344899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T18:06:06.896112Z digest=sha256:e4f4ab263e940345faadbb55a8260fa34878fc1ea829de3edf8442a54167dd76

Observation 22b41bc1-497c-41a9-90d5-cfed256071a6 · inbound

Reinforcing 3D Understanding in Point-VLMs via Geometric Reward Credit Assignment cites this paper.

Reinforcing 3D Understanding in Point-VLMs via Geometric Reward Credit Assignment Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:04:18.054828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-09T23:00:14.031709Z digest=sha256:45b0bc5c9509b6b6b9189c3b31062c61f82cf60bbc715f65a7b846a9bfbb4750

Observation 41006236-384f-48d0-967e-987897c84767 · inbound

SGSoft: Learning Fused Semantic-Geometric Features for 3D Shape Correspondence via Template-Guided Soft Signals cites this paper.

SGSoft: Learning Fused Semantic-Geometric Features for 3D Shape Correspondence via Template-Guided Soft Signals Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:53:15.141947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T11:48:55.446684Z digest=sha256:34a15707d86f693e65f2e68b1718b268fd38304c29f6f4924a86821694a90898

Observation bb6a6b9a-3818-4e0a-a863-38a7bfce053d · inbound

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning cites this paper.

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 80

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:18:57.763666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T01:23:40.564561Z digest=sha256:71e4fc07556c53466521322e1ffcdff101ee32e9fe0b4bbfcc02ceba7253bc41

Observation e7e808af-abb6-456d-a443-0ebefd3e80f8 · inbound

Occ-VLM: Occupancy Grounded Vision Language Model for Indoor Scene Understanding cites this paper.

Occ-VLM: Occupancy Grounded Vision Language Model for Indoor Scene Understanding Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T02:59:25.799065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T18:40:20.588652Z digest=sha256:ecb3f1ea71d5224e9fe548e5e231479ed4c9f210cb0dc51edda94ee55a189caf

Observation 083afcaf-f843-4667-976e-a20b19ff9d03 · inbound

3D-PLOT-LLM: Part-Level Object Tokens for 3D Large Language Models cites this paper.

3D-PLOT-LLM: Part-Level Object Tokens for 3D Large Language Models Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:09:29.323841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-26T18:25:55.011125Z digest=sha256:17c7edc08f80e0bb2ee6df8a1dfdc3826a5943871e9afcf3ca322926592b8ab1

Observation cc7bc106-3e0b-4cde-94ca-d2f97f99e77a · inbound

MV-GEL: Language-Driven Multi-View Geometric Entity Localization on Meshes cites this paper.

MV-GEL: Language-Driven Multi-View Geometric Entity Localization on Meshes Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:05:40.451224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-01T05:56:47.849311Z digest=sha256:69c92658af2c160093a1656b8670862046f237bca25a2ea69693a5b8aac48482

Observation 85be3b90-7c38-49e4-8bab-740661559d48 · inbound

MV-GEL: Language-Driven Multi-View Geometric Entity Localization on Meshes cites this paper.

MV-GEL: Language-Driven Multi-View Geometric Entity Localization on Meshes Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T09:23:10.452175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:23:10.452175Z digest=sha256:28bff69ac121271878fbd97ae16507035a34afccc2cfbdecaedc884543ecd2fe

Observation d1acf887-f64c-4ed8-a250-ec3305992d20 · inbound

SpaceEra++: A Unified Framework Towards 3D Spatial Reasoning in Video cites this paper.

SpaceEra++: A Unified Framework Towards 3D Spatial Reasoning in Video Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:18:37.260308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T16:16:41.412451Z digest=sha256:00eba586c9c7143b7e4c66646a982800fc67d2bc9c17d4d5696e67139f51cc0b

Observation 131fa64a-d148-48f8-88e0-f269398cb505 · inbound

Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation cites this paper.

Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T06:23:28.101393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T06:23:28.101393Z digest=sha256:41136f7e733474f02f6b4002871fb4e4bcbf0d83def534ca45666ccec10aedaa

Observation 4ba8b9c3-c4ab-43e9-9f7c-f875897ca14e · inbound

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding cites this paper.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.446109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.446109Z digest=sha256:3ba38ff8f26480a56c12966e2aa19fc98d69ded48d8e4170836d3fcdbe7d5986

Observation 4cac77ab-1356-46cc-99b1-f15769faffd5 · inbound

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding cites this paper.

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T04:28:21.948953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:28:21.948953Z digest=sha256:c344405bf916a3273018a773ea7829c178b99d0800cf514d579671fb027b7d19

Observation af998083-3948-4289-a462-9bf20c369917 · inbound

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding cites this paper.

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:40.838444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:30:40.838444Z digest=sha256:414317f3ed3ae98cf2caf499c6d8466745cd1b7cab05dd78cf665942efcd2af3

Observation 9f9aacc4-5617-4525-b65e-333599fa3935 · inbound

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding cites this paper.

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:06.824915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:06.824915Z digest=sha256:74086e397aa7569d86667bc9ded7f87644b5894117081cf573155821fbab03cf