Pith. sign in

Paper Citation Record · LEDGER

LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 54 inbound Pith citation observations for arXiv:2409.18125.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.18125 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 54 of 54 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T04:54:02.129840Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:39:56.857028Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ea7272f7-bad2-4ed9-978f-d9fab1682bfe · inbound

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces cites this paper.

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 107

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T09:27:44.102329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T09:27:43.919941Z digest=sha256:ebd24fb50e30cff9549e3df51b646e12171f67aabc41b1c3b51fb076aea0f8c8

Observation fbe72696-13d7-474c-bec0-7d43145b806e · inbound

SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model cites this paper.

SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:12:19.853033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T06:12:19.643111Z digest=sha256:b7ec4b22de324bf45b1f7206138d94413a6be5ce27a3b20ea5a8a4d5d720170b

Observation 923e75a7-626e-4223-bbb4-39188fb0d440 · inbound

Revisiting 3D LLM Benchmarks: Are We Really Testing 3D Capabilities? cites this paper.

Revisiting 3D LLM Benchmarks: Are We Really Testing 3D Capabilities? LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:02.129840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T04:54:02.129840Z digest=sha256:871e1dd9b2da62e59d76b1d5f8c878713523a9a6edf363f26825e0720845b90b

Observation b26fd254-490e-46e6-b794-35981a26791a · inbound

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts cites this paper.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:14.423259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:14.423259Z digest=sha256:bb5ca8b2136f0054c54d2ac171c035610cde1056db10069233f05e006bc9cbdc

Observation 6ac13126-0ca5-4894-912f-2fb74e07e39c · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:34:36.868541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T08:34:36.824053Z digest=sha256:271d7253ccd60ec94ed3acb576e81be4f898c9673690b04246204cfa70a03c8d

Observation 506780a9-b225-41c3-82f0-71e1a9ab68f3 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:00:51.268241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T00:59:13.826054Z digest=sha256:28b9540378ef9909e161abbca50f598e6c08b192d072b83cf04e2779aa17e154

Observation e9eba971-990b-4037-98c2-afd389d58a14 · inbound

Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames cites this paper.

Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:20.069930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:20.069930Z digest=sha256:ad1292f974bd22a6b42a12e668c5f411a768f0fdcf913374b752cb8af8f68416

Observation 7046edbd-9db3-4a06-8615-8c6489c47695 · inbound

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces cites this paper.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:16.730134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:16.730134Z digest=sha256:d50ce22833901dc0fb794552bfd62e1f9e45ae0dc978dca031b2df3e1a4d7f36

Observation 264df855-651c-4abf-a942-66b432d95dc4 · inbound

Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs cites this paper.

Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T10:25:48.356054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:25:48.356054Z digest=sha256:debe0942921d918ff972b7a2af011fcf702c480e60c00e80202d626d57cc1eee

Observation 7ab239f0-022e-42d6-bc5d-4a37fbd8672b · inbound

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing cites this paper.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.588064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.588064Z digest=sha256:f3821ccee3460380f77a2001bf8ff0a4e112fbb8d4a4c2a2be237b0ee366bef2

Observation 50cb826f-7ff1-4bd9-9c9f-50f230d11047 · inbound

Pts3D-LLM: Studying the Impact of Token Structure for 3D Scene Understanding With Large Language Models cites this paper.

Pts3D-LLM: Studying the Impact of Token Structure for 3D Scene Understanding With Large Language Models LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:20:37.480142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:20:37.480142Z digest=sha256:496e10549e3017e4d5cb5dd9aa9421432d9d0fa4244798a203c00e1fc170d323

Observation c8499945-4b1d-4d4d-89e3-eaac8adcd815 · inbound

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making cites this paper.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:27.599419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:27.599419Z digest=sha256:c44618e013669a02ebc7e8ee94cb9c062e397fe82d8a4e355a5ab6b4565d44c9

Observation baa76daf-d0cc-4046-867a-c8c4e18cb39f · inbound

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering cites this paper.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:12.924934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:12.924934Z digest=sha256:fd23226bc5e944e43fbb23eb491d7aff54f7b598feccbab11af820720a3593ca

Observation b4048b41-fcc1-4dee-a24f-dc6722516e0a · inbound

3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding cites this paper.

3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T10:49:47.146212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:49:47.146212Z digest=sha256:b9c1c880a25105363731379b6980bbbdb77d716e905405a17bdaa74b106d6b8e

Observation 3cb81618-21a3-405b-8dac-28d034d2a138 · inbound

PySeizure: A single machine learning classifier framework to detect seizures in diverse datasets cites this paper.

PySeizure: A single machine learning classifier framework to detect seizures in diverse datasets LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T22:18:05.150726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:18:05.150726Z digest=sha256:a127ff4062413f8e50d948d3a158e4c61531bba0bb4a23817e24406c8df06118

Observation 8fb6f571-f63e-4466-985d-9a52c58cb6d2 · inbound

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models cites this paper.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.649376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.649376Z digest=sha256:ca589589bf2adc933c948b97e6604eeab404ad949b61738d020336bb1b3a47e9

Observation d947ec6d-b7b5-4312-a67d-8132fdde69d8 · inbound

GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert cites this paper.

GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T11:38:13.467839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:38:13.467839Z digest=sha256:b2701cc55f32c4fbe601d08e51bc20c1053e1b2f25768c41d33433b4d333a7d8

Observation 8b515b32-dab1-44d5-ace8-129bf2380628 · inbound

POMA-3D: The Point Map Way to 3D Scene Understanding cites this paper.

POMA-3D: The Point Map Way to 3D Scene Understanding LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:30:11.559924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T20:27:27.347592Z digest=sha256:02392879f77c493790cd8b3949354c7a1d44f5c6249417493aca4c20a31b8b11

Observation 9ece6334-1468-428b-868d-cd6c9d1295ac · inbound

Boosting Reasoning in Large Multimodal Models via Activation Replay cites this paper.

Boosting Reasoning in Large Multimodal Models via Activation Replay LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:09:03.922650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T05:05:48.682057Z digest=sha256:0ff7ed43a64c0df8b912f2139ea95cb14b6528da67643aa95dfeed34ac29b303

Observation 7ea85dfe-8b0f-466b-b79d-b8a9c07fb2f3 · inbound

Vision-Language Memory for Spatial Reasoning cites this paper.

Vision-Language Memory for Spatial Reasoning LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-03T20:15:37.516145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:15:37.516145Z digest=sha256:ba9e730ba4d3899b61562d1eda46de1ae71bbe9f613253c5cacb4ccb362c2725

Observation 6e8a4b7d-4e8b-426a-888d-96b9dceb4bfa · inbound

SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition cites this paper.

SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:59:04.240779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T04:54:59.903644Z digest=sha256:6cb5cb7a211259e7aa7d7d11bfdb1fb515bd88dc5d41ebd596787d35d8f317da

Observation fef0e861-6800-4855-80ee-a03f7a27a63b · inbound

SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving cites this paper.

SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:21:31.269741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T12:17:54.325055Z digest=sha256:7c8a152b797aa2bd5b5a164aa8ab3aeb149fd81a32204b4a8a064b91d063b45c

Observation 0c2d3562-4638-4175-90ac-aafe8eccde47 · inbound

SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving cites this paper.

SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-03T17:07:45.189496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:07:45.189496Z digest=sha256:5213c3e8e199f1662984e428d6950ffbb6ccb8bc1e08a884ac76168df86b6728

Observation 725628d7-5b50-4f3c-9e4c-88091d41cc87 · inbound

OpenGround: Planning-based Online Perception for Open-World 3D Visual Grounding cites this paper.

OpenGround: Planning-based Online Perception for Open-World 3D Visual Grounding LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-03T13:46:30.457026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:46:30.457026Z digest=sha256:6225cb2c9729161b2fc72e49b728febede19b9db5e0ece3358f6799d500a8d01

Observation bf521ba6-1604-4758-b59a-e5446d40d410 · inbound

Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models cites this paper.

Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:26:26.916148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T18:25:21.621268Z digest=sha256:302ffa94c87008c4561b4fd1f5e0c073369a12a42263df0dcf6a0a733be95d3f

Observation 46529768-f4c3-4547-991f-32b41b4863c2 · inbound

Boosting MLLM Spatial Reasoning with Geometrically Referenced 3D Scene Representations cites this paper.

Boosting MLLM Spatial Reasoning with Geometrically Referenced 3D Scene Representations LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:35:56.050224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T14:31:03.909336Z digest=sha256:3ca99a2758da40454fe74c85aacef3e57962aa8b00baccec8f9dc437a4f51cb7

Observation 78606104-cf6a-4120-8cd3-97f27074cf27 · inbound

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models cites this paper.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.995051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.995051Z digest=sha256:48bed314e6cf69574356e0045cc40fdf7dc8acd6dcc1ef2bc53afa37ef02c53b

Observation 57ad9ea2-ad06-4ac9-a736-0307cecf070c · inbound

VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations cites this paper.

VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-13T23:42:44.515158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:42:44.515158Z digest=sha256:832a14ca1397ce0984f6b3e88cc5d7a617ea55d401bbc1a61d4e2e729fa41d59

Observation 829037dc-1aad-4ab7-8b1b-df4face8d24e · inbound

Feeling the Space: Egomotion-Aware Video Representation for Efficient and Accurate 3D Scene Understanding cites this paper.

Feeling the Space: Egomotion-Aware Video Representation for Efficient and Accurate 3D Scene Understanding LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T09:30:22.395942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T09:30:03.668178Z digest=sha256:1ea0f09601acf1d8caf7d95d778b5c5cdc50d3a87e6b015e64b61afabc22185b

Observation c438ce0f-c7a5-40bd-9c12-975277b5176a · inbound

Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning cites this paper.

Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:09:35.116047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T00:08:22.362342Z digest=sha256:68e8e5f75e298dd1d05c6e87060271c1db1c0f4277b16a33b63de2efe49f1e15

Observation 0691fced-5885-4c15-adcb-4003ed9e02a6 · inbound

3D-IDE: 3D Implicit Depth Emergent cites this paper.

3D-IDE: 3D Implicit Depth Emergent LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:38:11.453485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T22:34:04.833557Z digest=sha256:d60e1cd7f18cc08ed950fc2854d7e2c5130e08dd909fbb71d93d7ef10f11c585

Observation 3a769354-f6c4-494b-89d9-0e9f76277df1 · inbound

EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs cites this paper.

EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:43:22.828897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T22:41:09.840792Z digest=sha256:e7e7f2765fc6063002d24fb1250c96db48bd1f7c288eb7ad6acbcc6ecf735c53

Observation e8389b68-bbd3-450c-a0e2-ab92616c3637 · inbound

EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs cites this paper.

EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-13T14:39:55.177552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T14:39:55.177552Z digest=sha256:fe46d6eccb26519b5b644f1b3ec5fdaf6aa800d8098a0f8ac0bd5092baf61f4c

Observation 2b917ffe-6cf8-440d-804d-e651014702dc · inbound

UniMesh: Unifying 3D Mesh Understanding and Generation cites this paper.

UniMesh: Unifying 3D Mesh Understanding and Generation LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:56:47.338138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T06:55:42.679323Z digest=sha256:6fef92661b1f79d35ba04632de9cffdad26c46e57c83f10296502997aef736db

Observation a819ce1a-ddab-48f3-8388-30975d06eab5 · inbound

Flame3D: Zero-shot Compositional Reasoning of 3D Scenes with Agentic Language Models cites this paper.

Flame3D: Zero-shot Compositional Reasoning of 3D Scenes with Agentic Language Models LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T02:11:15.854529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T02:09:09.684786Z digest=sha256:83b3d8930ea6797d885c7c5a51e2d6ef6be9731d52c8d99c97734348a48b2e91

Observation 05ddfdf3-00fb-458a-aa00-d13e6dd877d0 · inbound

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT cites this paper.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:31:25.837388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:c9def1d2c2624b887a472dff00e313dfbf627de1b598d7968f10220662f692a4

Observation b2e9ba17-1ab5-4936-a7cf-ef96753ab33d · inbound

Unlocking Dense Metric Depth Estimation in VLMs cites this paper.

Unlocking Dense Metric Depth Estimation in VLMs LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T19:23:40.971811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T19:20:04.468206Z digest=sha256:99694ecc123025b95562143ef1eb49d903662f6ffb39078c2286b70a8770c3cd

Observation c7c1b832-f5bc-48bb-9d9c-d284f8a40baf · inbound

Unlocking Dense Metric Depth Estimation in VLMs cites this paper.

Unlocking Dense Metric Depth Estimation in VLMs LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 86

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T07:59:51.117080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T07:54:52.926995Z digest=sha256:3d51fd15ec1fff36ae4eef3f2469d3f95ab6fe1f13b8cbd63fd9a0060505e553

Observation c568eb3c-3fee-4e39-bcc1-a6c77c8e4a7e · inbound

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models cites this paper.

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T05:28:04.610018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T05:27:30.938311Z digest=sha256:5bf08c2393f2365c80af08c1db87e258d71693f14c170e53fbaaa79b03e031fe

Observation d8ccdd82-7f95-41e5-8e78-e71a7b07fca7 · inbound

Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models cites this paper.

Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T13:44:40.748826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T13:40:59.330788Z digest=sha256:77744f80fc21c23fac8d9891518e2fb130f7fe45759366e9d82275472bdc7b85

Observation 4350b7bb-d452-4ce5-aed8-9627f153003b · inbound

Perceive-then-Plan: Layout-as-Policy for Monocular 3D Scene Layout Estimation cites this paper.

Perceive-then-Plan: Layout-as-Policy for Monocular 3D Scene Layout Estimation LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:14:01.279614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T23:11:59.282627Z digest=sha256:cc89a387ccb8361733dfe196213ec744a17a1c104ac283e6625ed11f2223522d

Observation 19ace1ec-6239-45a8-9243-d82e6f3fa22b · inbound

Geometry-Aware Representation Denoising for Robust Multi-view 3D Reconstruction cites this paper.

Geometry-Aware Representation Denoising for Robust Multi-view 3D Reconstruction LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:14:01.824595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T23:08:52.333329Z digest=sha256:51b0b896ed0ff189c4aa46542a8a3b478da8682642e77c4469afa26cb85110ab

Observation c55ddd02-4122-41fc-81b9-b0161d46f9b0 · inbound

Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning cites this paper.

Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:13.622650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T07:47:52.739735Z digest=sha256:4d556c6fdc45378d01256ade261ca17f71adf236fc1b5b66090761b0fcfd859b

Observation a27de77b-f796-415c-8585-73b8247538b8 · inbound

Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators cites this paper.

Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.162004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:25:30.998989Z digest=sha256:d8b399aec4eec61812656309aebb7603eb89b06ee51da5bd2ef555baea508e68

Observation 75ee56c0-637b-4033-af2f-4716bda53048 · inbound

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks cites this paper.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 91

Resolution
malformed identifier
arxiv_id, observed 2026-07-03T01:27:30.682419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:4fec0ae1269eec67798fac20d18a01beba5c6571d4f7c030306f861159d381ce

Observation 7eb1e127-ffc6-40f8-ab3c-a070e54e1fa4 · inbound

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models cites this paper.

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 86

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:08:55.586365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T01:42:30.005911Z digest=sha256:24d9a0654f34183cc55b50da4682df7b71bf1ba507e0931b877cf2a39733e58d

Observation 2d49d67e-8080-4b9b-8b8c-83bc07d1cde2 · inbound

Occ-VLM: Occupancy Grounded Vision Language Model for Indoor Scene Understanding cites this paper.

Occ-VLM: Occupancy Grounded Vision Language Model for Indoor Scene Understanding LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T02:59:25.806993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T18:40:20.588652Z digest=sha256:377b75274bb3c379acbe1b5af5cdf029fe7c573f61f80d8d16160ac705d69298

Observation 63d561c5-8a01-4c23-84f7-bc3d0e40f601 · inbound

3D-PLOT-LLM: Part-Level Object Tokens for 3D Large Language Models cites this paper.

3D-PLOT-LLM: Part-Level Object Tokens for 3D Large Language Models LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:09:29.331044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T18:25:55.011125Z digest=sha256:c5c8248a44ebd94ae4212d9ab99476737430082bde5bde241d747a2574353408

Observation 0fcf1d7a-fc6c-4154-9227-0f0858810e3d · inbound

CVSBench: A Comprehensive Benchmark for Cross-view Spatial Reasoning and Dreaming cites this paper.

CVSBench: A Comprehensive Benchmark for Cross-view Spatial Reasoning and Dreaming LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:49:42.526702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T10:54:49.774490Z digest=sha256:f6d3cec1008d79a3af42ad6e664cf5e321bb3af25554038fc726d451174a1956

Observation 5b1597f1-6ccf-4f79-951a-0010d5918c3a · inbound

ObsGraph: Hierarchical Observation Representation for Embodied Reasoning and Exploration cites this paper.

ObsGraph: Hierarchical Observation Representation for Embodied Reasoning and Exploration LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T15:39:56.858556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T01:30:05.294887Z digest=sha256:2b605f0da91ce9b970b9a88514d3d424ec332f37b0451191044edbc1f0499449

Observation c41b482a-6788-499c-a549-4641c73976b2 · inbound

Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence cites this paper.

Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:31.366247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:31.366247Z digest=sha256:8c0f7a96d82287898d05c9802b1db9cb1d97c2b067c844b5824de38d366c6cb7

Observation d4fbc204-770b-4f7d-8a80-f4f86b9f6505 · inbound

Towards Enhancing 3D Spatial Reasoning in Medical Multimodal Large Language Models cites this paper.

Towards Enhancing 3D Spatial Reasoning in Medical Multimodal Large Language Models LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T03:31:46.403528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:31:46.403528Z digest=sha256:7ff06a390f589cdd439daa1824f236514a192e66e3d98d7efaf85eb7def3c9ce

Observation 18e083b6-e336-4298-b764-a396607929ae · inbound

IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer cites this paper.

IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T13:07:03.267487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:07:03.267487Z digest=sha256:369e20cf56c50cde42872c1bbc618a4df4de17afacf3a30ac68289b6f90658e3

Observation 8fd70c46-7875-4f27-8e60-647ad55a4dc6 · inbound

ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? cites this paper.

ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-01T09:09:24.877243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T09:09:24.877243Z digest=sha256:0bfe9acc94f2798c9c433a449f7eda83ccf97576f7e672c99f7b24740bb491b2