Pith. sign in

Paper Citation Record · LEDGER

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer

As of 19 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 6 inbound Pith citation observations for arXiv:2501.01163.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.01163 v2

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:38:21.624165Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:42:46.762264Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T01:00:51.232818Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact2
  • verified fuzzy36
  • unresolved29
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 888363c1-dff3-4a7b-8693-7180a188f73d · outbound

This paper cites Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:22.962313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.304040Z digest=sha256:2b54dab45324a704e8c19b4129c9b759d02af2e55c47ca78572955fbc562721e

Observation d6842355-1d5a-444b-bedb-5d9dd5945449 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:38:21.312544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:38:21.312544Z digest=sha256:d58120020c1fb4bef3d470de2142b9ec8733292e387f8c2ea1b9afdb05892afc

Observation 92779771-0203-4486-88c0-4f1ba3691bd8 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Scanqa: 3d question answering for spatial scene understanding

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:22.855494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.317517Z digest=sha256:5b90951123d09f1498502154e79617e7c6a6ba0ddf0801420f8a6fcdb38170c5

Observation 03a0081c-632f-4dc4-a330-e3e5497fac99 · outbound

This paper cites Lan- guage models are few-shot learners.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Lan- guage models are few-shot learners

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:38:21.321204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:38:21.321204Z digest=sha256:fb1e8069be8ec1546dab38f316c31ea0e80491f0d962e6f9170b4b5ce075034d

Observation 71bf9048-bb15-4c38-90d4-c84ba2015eba · outbound

This paper cites Lan- guage models are few-shot learners.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Lan- guage models are few-shot learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T22:38:21.325205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:38:21.325205Z digest=sha256:ae4f6c97521687a1b972a5f1cff1b91a7aa936f6926737b2cd6253a974ee2e8c

Observation 32d1d7ea-e996-4c19-9215-90674afccb5c · outbound

This paper cites 3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer 3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:22.778103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.329044Z digest=sha256:1b1971fa527ebbadc6afc5312facb6d2077f1a2c3e1e3c432e29fc773bf9ed4f

Observation 7b180116-770f-45d4-9321-a6de978b9974 · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natu- ral language.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Scanrefer: 3d object localization in rgb-d scans using natu- ral language

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:22.704309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.332860Z digest=sha256:465731a3fa67288504690fae52a23ccb856fccbdab319c87c4dff2a8a8a1484a

Observation ad158610-2e48-438e-8253-f05e8e44240a · outbound

This paper cites D3Net: A Unified Speaker-Listener Architecture for 3D Dense Captioning and Visual Grounding.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer D3Net: A Unified Speaker-Listener Architecture for 3D Dense Captioning and Visual Grounding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T22:38:21.340614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:38:21.340614Z digest=sha256:d587839c17057c88aa92e3e567ab00ac6113b2980008732583ceede1fafb841b

Observation 3efbaee2-fb55-4c10-b2bd-ef18754cbc09 · outbound

This paper cites End-to-end 3d dense captioning with vote2cap-detr.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer End-to-end 3d dense captioning with vote2cap-detr

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:22.643116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.344824Z digest=sha256:efe2d60ffa2ce6982dfa63a5d0c533582beef46cc67a053f72edc01ab4b01f15

Observation 55c2b2e9-af23-4074-b04f-751de9427cae · outbound

This paper cites Vote2Cap-DETR++: Decoupling Localization and Describing for End-to-End 3D Dense Captioning.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Vote2Cap-DETR++: Decoupling Localization and Describing for End-to-End 3D Dense Captioning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:38:21.845911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.354622Z digest=sha256:4f507e8e1c30811b35683de9e679a0daf1bd4ee5e4f62d67527c90448f56d10b

Observation 00233a35-9796-4d66-8433-75abe511b726 · outbound

This paper cites Ll3da: Visual interactive instruction tuning for omni-3d understand- ing reasoning and planning.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Ll3da: Visual interactive instruction tuning for omni-3d understand- ing reasoning and planning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:22.577170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.361785Z digest=sha256:9290e81218d0c18cdce220f333b8550acf80389fc3407a53adb377ee62d4a5ba

Observation 20e8b86a-ef1a-4642-9790-05f549fdf540 · outbound

This paper cites Grounded 3D-LLM with Referent Tokens.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Grounded 3D-LLM with Referent Tokens

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:38:21.365509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:38:21.365509Z digest=sha256:cb13dec3b01ed7f3010e082bf547c82b86046a9ad47b2640355d2a76bd96876a

Observation af20d30c-6eac-48c3-9ff3-7cb34d3ec6fe · outbound

This paper cites Scan2cap: Context-aware dense captioning in rgb-d scans.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Scan2cap: Context-aware dense captioning in rgb-d scans

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:22.486472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.369436Z digest=sha256:5403a9b9c527889e14fcff4cf9f6769f1dd8406995413bf1ed9e07061fecf284

Observation e3f787c6-4610-428c-a1cd-d692e906e049 · outbound

This paper cites Unit3d: A unified trans- former for 3d dense captioning and visual grounding.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Unit3d: A unified trans- former for 3d dense captioning and visual grounding

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:22.437822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.372993Z digest=sha256:3f6a0e0cae8f43231dd105abf53b708740ec9efdf09ef8f00f26ac71ecba8466

Observation 9eacf99b-44d8-4fd2-83cb-613ea48634de · outbound

This paper cites Palm: Scaling language modeling with pathways.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Palm: Scaling language modeling with pathways

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:38:21.376268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:38:21.376268Z digest=sha256:04b9fd2502e7bcd43de7af9378a1e6b716aa121d28ae66d1a6a6534ced13d125

Observation d3025561-fdaf-4a2a-a96e-cc0fe084e127 · outbound

This paper cites 4d spatio-temporal convnets: Minkowski convolutional neural networks.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer 4d spatio-temporal convnets: Minkowski convolutional neural networks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:38:21.379646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:38:21.379646Z digest=sha256:b2209b234f3d5cb504277a51bd871ceb5e205353ae7a76a9fe9d752db7035a80

Observation ca065221-c100-446b-abc7-1e11dd280b46 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:38:21.383295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:38:21.383295Z digest=sha256:5d413dab11d1954286b529191b41b580a09038d5dad03250984ad7c8e999569a

Observation 33c19405-ee4b-474a-9f14-9859cc79fad0 · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning,.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Instructblip: Towards general- purpose vision-language models with instruction tuning,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T22:38:21.387006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:38:21.387006Z digest=sha256:3ed299af3f3cba2aa2bbd11ad5cb1c76363858435e665050ed6f8234e500a615

Observation 3d7cfe24-58bf-4cf6-a63b-513c3b0e429a · outbound

This paper cites A density-based algorithm for discovering clusters in large spatial databases with noise.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer A density-based algorithm for discovering clusters in large spatial databases with noise

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T22:38:21.390503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:38:21.390503Z digest=sha256:c311eb64018b7467c257dae97af0ca2c6b37c5c31d63fccb17a0ec78c093e65b

Observation 49e2eb89-8bb6-4a7e-8531-20716a010ce5 · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:38:21.395235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:38:21.395235Z digest=sha256:ed8509f0dc8fed89fbf028acd5de29b5615b29f464e580ef2ceef0586e3c8845

Observation 13c49fe8-ad71-4d8c-b695-fb10e5cada30 · outbound

This paper cites Segpoint: Segment any point cloud via large language model.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Segpoint: Segment any point cloud via large language model

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:22.359552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.425392Z digest=sha256:aa3151750e6aba6ddec5e5bcc3d672621bc3c154885865a564a7636812c1b2df

Observation 36cc0668-c536-488c-9f56-9ac26d530e83 · outbound

This paper cites Training Compute-Optimal Large Language Models.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Training Compute-Optimal Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:38:21.440041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:38:21.440041Z digest=sha256:b57851e496a33db46ac0f22a0279d14f0e9d04ac7f6dfdf2abb8529e496ed478

Observation 6ffa7714-17bc-441a-89b7-b30ecb338cab · outbound

This paper cites 3d-llm: In- 9 jecting the 3d world into large language models.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer 3d-llm: In- 9 jecting the 3d world into large language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:22.313257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.453086Z digest=sha256:199d44999dfe065b857cde629b102c868ba151728320d0d077c8ccc60ef455f0

Observation e2a88b52-3c15-43dc-ae05-bbc3781585b8 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer LoRA: Low-Rank Adaptation of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:38:21.464197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:38:21.464197Z digest=sha256:c200b55d0ec6f49feb901777fb987f73c7d7bd1099ed4b23a397c3dab6cd2124

Observation aeda7315-d3a4-4bad-a110-25fbff11d612 · outbound

This paper cites Chat-scene: Bridging 3d scene and large language models with object identifiers.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Chat-scene: Bridging 3d scene and large language models with object identifiers

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:22.276860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.468717Z digest=sha256:7648aff4b0104e0ee58e799ab6647efdffc41118083edad5dfdf53fe7bdb752f

Observation f6ce3afe-e49e-4af4-8e7d-189065fe2983 · outbound

This paper cites An Embodied Generalist Agent in 3D World.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer An Embodied Generalist Agent in 3D World

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T22:38:21.473130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:38:21.473130Z digest=sha256:d3250d7c719c05213b7b179ac5b090940bfb4295213dd735d081ed22c7aef740

Observation 71b3d911-595e-4caf-8769-8212050dfea6 · outbound

This paper cites Reason3D: Searching and Reasoning 3D Segmentation via Large Language Model.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Reason3D: Searching and Reasoning 3D Segmentation via Large Language Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T22:38:21.477005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:38:21.477005Z digest=sha256:0b00121ae8b2fccb27d71459a74b7d2159566e256e71011988b5c92b0ac8f0b7

Observation a2197ecb-b74c-468d-8667-43ef23731294 · outbound

This paper cites Text-guided graph neural networks for re- ferring 3d instance segmentation.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Text-guided graph neural networks for re- ferring 3d instance segmentation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:22.239976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.481578Z digest=sha256:2d300cce0dee95a8b253d904b92e389c161a664654b0fc50dbdd2ad049555935

Observation c2d8902b-5f69-495d-9ba4-d24797693deb · outbound

This paper cites Multi- view transformer for 3d visual grounding.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Multi- view transformer for 3d visual grounding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:22.166653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.484693Z digest=sha256:af331b5be05dc20f72d04cd2b56315b9e361542c208d6045d5b8409843b7757b

Observation 57b31f50-339a-444c-95b7-ff41fff1ab5a · outbound

This paper cites Oneformer: One transformer to rule universal image segmentation.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Oneformer: One transformer to rule universal image segmentation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:22.127546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.487999Z digest=sha256:8843ca0e5cf4e8a6be9a27ac79e0d44d03017722332b53707487728fbb0895a5

Observation ee970a3b-4264-43ca-a5e0-a1e0b207c96a · outbound

This paper cites MORE: Multi-Order RElation Mining for Dense Captioning in 3D Scenes.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer MORE: Multi-Order RElation Mining for Dense Captioning in 3D Scenes

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:38:21.770912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.491774Z digest=sha256:664046f2fdcbed393b086f568050888d6bdaaa83138eaab5e4dcc389c781e302

Observation 2208f299-cb9a-48c0-ae37-1373c9c0b1a4 · outbound

This paper cites Tod3cap: Towards 3d dense captioning in outdoor scenes.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Tod3cap: Towards 3d dense captioning in outdoor scenes

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:22.116245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.495053Z digest=sha256:cc05193ed2426a6c6c1ac6802932aff47f4a7f0803a7d180a85de02c1688bd3c

Observation 64d0f10c-3d58-4d1d-a8fa-441bfe7b70a7 · outbound

This paper cites Context-aware alignment and mutual masking for 3d- language pre-training.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Context-aware alignment and mutual masking for 3d- language pre-training

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:22.106042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.499102Z digest=sha256:af696ff4efd41ea32c61add55f87c6a4f6337c4774800ae1d7f6a966bedde8f7

Observation b4d81e63-0fe6-4ca5-a2a6-a410ccbb8ed0 · outbound

This paper cites Bi-directional contextual attention for 3d dense captioning.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Bi-directional contextual attention for 3d dense captioning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:22.096317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.504199Z digest=sha256:5ab78fb22d3f77b6b55ee4eacb9a5717d90f6070d35d38ddedeace11b0375ac4

Observation e28d0ccb-4417-4116-b318-c24fd5fa7d20 · outbound

This paper cites Oneformer3d: One transformer for unified point cloud segmentation.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Oneformer3d: One transformer for unified point cloud segmentation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:22.086848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.508129Z digest=sha256:a7e75adf436b2b822f29da59ce73b96e8fd53252ed8cf8872698a3a1c629ebe6

Observation a7455c13-ae39-4dc1-af48-37254b701cf9 · outbound

This paper cites Mask-attention-free transformer for 3d in- stance segmentation.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Mask-attention-free transformer for 3d in- stance segmentation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:22.076558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.511433Z digest=sha256:b4d577f5d11e6c709be734d79ecea202f78d069ab77b24495f6671fd3c71da3f

Observation 82af0fb9-3ffe-443a-9cf5-2b0b752b3645 · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Lisa: Reasoning segmentation via large language model

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:22.065918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.515719Z digest=sha256:16bcfd59c2b735977d2ca7055af0c17f6d5ebee4d65ecd3b6ef214f603ba8259

Observation 88d3bb10-aa68-42c7-a584-8ca93419ae35 · outbound

This paper cites Large-scale point cloud semantic segmentation with superpoint graphs.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Large-scale point cloud semantic segmentation with superpoint graphs

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:22.054980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.519220Z digest=sha256:ffe939a3e202139498c5a0617b40b8cf868ceea77fb8d28d8117e6ec2b1845c7

Observation 12db2c66-d551-4c76-8f57-443b017ba61a · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T22:38:21.522463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:38:21.522463Z digest=sha256:3f1481e0d7352624e0a9dad08ce0bd542c2b8beb2ea73e0a158c2080158dd967

Observation 9fd1d6e1-c061-46c3-930f-599d1d9cbb0b · outbound

This paper cites Visual Instruction Tuning.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Visual Instruction Tuning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T22:38:21.526491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:38:21.526491Z digest=sha256:2bd0d6a6faec2efd5fcff5193cb9666e92a5c43d6c8f6519d177b0df88a4288a

Observation f7fa5907-153b-4eb8-ae1f-ac33bc4e2416 · outbound

This paper cites Sparsebev: High-performance sparse 3d object de- tection from multi-camera videos.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Sparsebev: High-performance sparse 3d object de- tection from multi-camera videos

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:22.044494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.530014Z digest=sha256:a584939ac41a085105d8dc036bae562540003c2293eaa195af99efab6f1552d3

Observation 4ff97781-8532-42f5-a51e-c7a6bea2ffd3 · outbound

This paper cites Improved baselines with visual instruction tuning.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Improved baselines with visual instruction tuning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:22.034319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.534340Z digest=sha256:017fc7f541288be6a41d5c3716e3196f8018adda6075b31d39e076203470108d

Observation 1c5c1887-09f9-4ddb-af7a-ec474210340f · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90quality, 2023.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Vicuna: An open-source chatbot impressing gpt-4 with 90quality, 2023

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:22.023125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.537855Z digest=sha256:fb1548b5e4fa30a29304fec811330226ec38274ea800099fd752484ee434abb6

Observation 88c484b1-9a0e-4456-9a8e-bce52bb091a6 · outbound

This paper cites SQA3D: Situated Question Answering in 3D Scenes.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer SQA3D: Situated Question Answering in 3D Scenes

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T22:38:21.541249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:38:21.541249Z digest=sha256:7cf3e627cb4f03ad652ca6d486698eba48e294e218903927a3598d02bc6576c7

Observation 64daf037-cda0-4116-95d2-d465018b5676 · outbound

This paper cites Clip-guided vision-language pre-training for question answering in 3d scenes.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Clip-guided vision-language pre-training for question answering in 3d scenes

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:22.013185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.545029Z digest=sha256:af82d381c717a78e79f14ebaad618d93f3cdc77b857b6c5a56828a5a4fbbeab6

Observation 4a2aae5a-0b39-4954-a83a-b6418e4523ed · outbound

This paper cites Openscene: 3d scene understanding with open vocabularies.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Openscene: 3d scene understanding with open vocabularies

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:22.002593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.548433Z digest=sha256:94751e3a7f9843a4e9d35c1ffebf849ffe22805560beff061939f99a6b7e34f9

Observation 0b459fd9-c9a4-4ff5-8553-64531e4efe7b · outbound

This paper cites Pointnet++: Deep hierarchical feature learning on point sets in a metric space.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Pointnet++: Deep hierarchical feature learning on point sets in a metric space

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:21.991393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.552351Z digest=sha256:2efce54b08caa940abb6c5e56fc22f5aa134b0c7e17cf6745d62f88dd95390ba

Observation 56fdd562-5f10-4944-9426-7270d0d4b1f7 · outbound

This paper cites X- refseg3d: Enhancing referring 3d instance segmentation via structured cross-modal graph neural networks.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer X- refseg3d: Enhancing referring 3d instance segmentation via structured cross-modal graph neural networks

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:21.980346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.555749Z digest=sha256:49cc05e7d5385b9ffd0772349452fe3bf3af9abb48c6dc3cebb567c4cbc9124b

Observation 466c5048-85b4-4814-9487-483ccaa36e9f · outbound

This paper cites Language models are unsu- pervised multitask learners.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Language models are unsu- pervised multitask learners

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T22:38:21.560097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:38:21.560097Z digest=sha256:d2cb2db99c48e86df074b30edc288e62513b04132f4a6819cd76ebd788e9f62b

Observation 1f541a13-c5cd-4d78-bd9c-d10e5dec5e4f · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Learning transferable visual models from natural language supervi- sion

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T22:38:21.563660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:38:21.563660Z digest=sha256:af79d2035552742e980b7175bd888cf2975090d2365442d8a21f943f8fb74478

Observation f856169c-932c-428c-9c52-6408d637adfe · outbound

This paper cites Language- grounded indoor 3d semantic segmentation in the wild.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Language- grounded indoor 3d semantic segmentation in the wild

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T22:38:21.567100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:38:21.567100Z digest=sha256:8ff5d24bc0ee4a235000ba9cd96bd75cc18af13d79bfb0006189d5868bd79e89

Observation b213143b-3088-4bfc-8a8e-adf2d5a51114 · outbound

This paper cites Mask3d: Mask trans- former for 3d semantic instance segmentation.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Mask3d: Mask trans- former for 3d semantic instance segmentation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:21.949511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.570410Z digest=sha256:fab1fb6c822fa80cb0af6142baffe0ea8b3b4817258a2cd027d06b8aab143478

Observation d101734e-f423-4331-9cf6-be9cd14347b7 · outbound

This paper cites Superpoint transformer for 3d scene instance segmentation.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Superpoint transformer for 3d scene instance segmentation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:21.938590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.573546Z digest=sha256:b6b125d1911b5b78e071a8591301c892e364a9153a60e6df2e88ada9ce42a9b9

Observation a496371f-7daf-445e-a63e-a0b161bc5da1 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Gemini: A Family of Highly Capable Multimodal Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T22:38:21.576753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:38:21.576753Z digest=sha256:2e509fac237a0135725c490010a6ad1e109cee9bf91999c7c81ac0108ee79686

Observation 0fa5dcfa-41e3-4cbe-985b-b1a04a730b44 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T22:38:21.580213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:38:21.580213Z digest=sha256:a296d70c3daa175c23d08c7a7701502f16bbaeb333a0d0c8868d0f73e489e48c

Observation d9b82dab-2c4e-4a4d-8b13-ae2085e4023f · outbound

This paper cites Four ways to improve verbo-visual fusion for dense 3d visual grounding.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Four ways to improve verbo-visual fusion for dense 3d visual grounding

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:21.927085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.583315Z digest=sha256:96fc1a404ad4fcaf4af231cecac93893c1f60fbb63d938b57f2ac79af925b5d7

Observation deeb111e-a2ff-488b-a6de-806879e07375 · outbound

This paper cites Spatiality-guided Transformer for 3D Dense Captioning on Point Clouds.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Spatiality-guided Transformer for 3D Dense Captioning on Point Clouds

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T22:38:21.586317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:38:21.586317Z digest=sha256:61fb37976b5ba22c843a3b01ea10704feda4fd1ce2ccafc40746ae38bc1b5053

Observation f3c05a8e-d8d7-418d-8d22-2533a7dbf8bf · outbound

This paper cites 3d-stmn: Dependency- driven superpoint-text matching network for end-to-end 3d referring expression segmentation.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer 3d-stmn: Dependency- driven superpoint-text matching network for end-to-end 3d referring expression segmentation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:21.914969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.590789Z digest=sha256:95c16fca919714a3db0eaccfeb5b395b01091d8ddff9bcc71b8f7fd17680ed01

Observation e00940a1-fa89-420c-abb2-07f3d0ab5619 · outbound

This paper cites PointLLM: Empowering Large Language Models to Understand Point Clouds.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer PointLLM: Empowering Large Language Models to Understand Point Clouds

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T22:38:21.594099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:38:21.594099Z digest=sha256:94b5c0d745d1b017c346e91c0d57f57119cf69beaf388dcf8bf0014faa13cc6c

Observation 16d1c125-bdcd-428e-9e68-1a8d68657e75 · outbound

This paper cites Sat: 2d semantics assisted training for 3d visual grounding.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Sat: 2d semantics assisted training for 3d visual grounding

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:21.904223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.597693Z digest=sha256:c30d4772b538f430f3b4c72c35c4a42dd3c9d973887e1c294261141d3445f442

Observation 2a36d1f9-6cbc-4428-ba3f-63cf1e97cb01 · outbound

This paper cites X-trans2cap: Cross- modal knowledge transfer using transformer for 3d dense captioning.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer X-trans2cap: Cross- modal knowledge transfer using transformer for 3d dense captioning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:21.893500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.600989Z digest=sha256:781a0d6266e47aba1153c55687e3dcbeceb8a3494c823ec598c8e3a972951930

Observation 4a4369bd-209d-4062-87b9-7a32805e4a8f · outbound

This paper cites Agent3D-Zero: An Agent for Zero-shot 3D Understanding.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Agent3D-Zero: An Agent for Zero-shot 3D Understanding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T22:38:21.604533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:38:21.604533Z digest=sha256:d963e926bdb07c5aba90fda6d2c29976681fb48b3d50e5a10c21a51054ea06d1

Observation f85724c3-f37a-4faf-a3ca-4ebca065201a · outbound

This paper cites OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T22:38:21.608432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:38:21.608432Z digest=sha256:afc533c4250d9191ed038ca3f2f73906044a9987b17b24d0a911eac66467c080

Observation 4a3f066d-daef-475d-8890-51c36c419dac · outbound

This paper cites Multi3drefer: Grounding text description to multiple 3d ob- jects.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Multi3drefer: Grounding text description to multiple 3d ob- jects

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:21.881949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.611886Z digest=sha256:4b822b9783f6d95131a778e0a9ad7fb781a1801de77ec0fd3483c9c7874e2998

Observation 8e5d5896-b9de-4d90-a357-7632483ddb84 · outbound

This paper cites Contextual Modeling for 3D Dense Captioning on Point Clouds.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Contextual Modeling for 3D Dense Captioning on Point Clouds

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T22:38:21.616056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:38:21.616056Z digest=sha256:475008478a2c04f0e88051f10106d98102fdd8674684cdad1e80f99f71c90d01

Observation 20cf3ff7-b3d8-47a0-8303-cf08584a3259 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T22:38:21.620323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:38:21.620323Z digest=sha256:9e7227380bdfeb3a1ee553094972824b70d0733620472086f77cfee98defee49

Observation 3291e147-aded-4c1a-a830-24c6d28a20f9 · outbound

This paper cites 3d-vista: Pre-trained transformer for 3d vision and text alignment.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer 3d-vista: Pre-trained transformer for 3d vision and text alignment

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:38:21.866923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.624165Z digest=sha256:6be7cd6951ac23780cd63dc54ce9cd30df783c0e2ca21317cf191cc9bdb8ac4e

Observation 4f1d636c-c6c8-451c-bee0-2bc6b5e634ea · outbound

This paper cites an unresolved cited work.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer Unresolved cited work

Reference 440

Resolution
parse uncertain
raw_fallback, observed 2026-08-10T22:38:22.925089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:38:21.308636Z digest=sha256:4dfc03f82ccdb83d7d8103d0296064eff0a0761a0a7b90f92a94124cd00adf57

Pith citing papers

Observation b735ca13-2ae1-4ac0-a97c-862e8cef45de · inbound

LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding cites this paper.

LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding 3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:46.762264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:46.762264Z digest=sha256:739b9e0e9331dbd3b89560a3e6b93af549625889b1b976409c516cf42ac56d22

Observation c441afbe-0c1b-4609-a8a9-d55ab5a1967a · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence 3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:34:36.864611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T08:34:36.824053Z digest=sha256:ac98c5812a9749519cae1f40afa67d220caf8da35f5fc0dc5586399a9a774918

Observation 16dbd855-9a29-4172-8260-00ea79ed4957 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence 3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:00:51.236919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T00:59:13.826054Z digest=sha256:9adf00ab500ff45af4327215197ca4e5ca8d275760c0138c11d26975594b1936

Observation cfa85fcf-2cdf-4e1d-8934-7ba0409f6ef4 · inbound

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing cites this paper.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing 3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.796197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.796197Z digest=sha256:eb4bb44ccaae4d98cffc841a0c2c1dff5283b1d9b1a15b17f64232698cfed7aa

Observation 82d902a7-3186-4720-a85d-6957cb10f36d · inbound

MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding cites this paper.

MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding 3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:51:21.546294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T17:25:31.097385Z digest=sha256:4ef9f631f5637bbb32887651a3e9b512ba0b07b29cbbce7a30e1d25b4dbca3e1

Observation 08779163-fc63-47d5-92c3-8e12b46253ba · inbound

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models cites this paper.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models 3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.320043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.320043Z digest=sha256:d134f85b1eef43c8a33ebff9e88f188899f00b784cb2c127b1f1a947491a42c0