Pith. sign in

Paper Citation Record · LEDGER

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2412.00493.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00493 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:46:13.494993Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T01:00:51.358037Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 31acac97-89dd-432e-855f-aef3ea07a7bd · inbound

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts cites this paper.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:13.494993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:13.494993Z digest=sha256:a9dcd5e0f6b8ceaaa3e109bc1561b78113bdf14bdd4f89917b19720a6f3007f3

Observation 0c60e752-4b9b-46f1-8d89-b0217515b7c5 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:34:36.872348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T08:34:36.824053Z digest=sha256:23ea5af2c1ae73482e2bb63eaf2797645feac99b45e3b5334c3bc026b7fe0656

Observation a4c06c78-5058-4645-811b-0778b0d90f73 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:00:51.361125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T00:59:13.826054Z digest=sha256:99423279e0727bebf1ba10989713483260057909bd42f49060cbd4a6c47103ce

Observation 3eabc801-a4d0-4b7b-ae40-b2bc915861bf · inbound

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing cites this paper.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.802988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.802988Z digest=sha256:70609e9edef8de3bc2d53118863dc0d209755f6154f8573f2d3f7afb1ccf69c8

Observation 2d0cd3b5-f8a4-4fb0-95d8-c5b13fa3842e · inbound

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture cites this paper.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:24.704510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:24.704510Z digest=sha256:213b580882fe2ba3f5873bf48a055e30d29e3eae7d07283f0d532e95202cfbda

Observation 9f03306a-48ba-4fb2-813d-0922a73fea03 · inbound

MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding cites this paper.

MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:51:19.088377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:25:31.097385Z digest=sha256:b0689046f4019c08627eea01d8dda91613dbd273e7b6610cb65709bb73551d2e

Observation 6f6f27bb-e7d2-4e8a-8ae7-ca3dc3495693 · inbound

A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning cites this paper.

A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:41:01.809714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:53:52.775348Z digest=sha256:ac6c7e1bbcdff500135fc4d3d1661346cd6b302f0f14797d06ff2d1494b63535

Observation 40688cb6-5a6f-412d-afde-7fe87b94e9ca · inbound

Seeing Once is Enough? Online Geometry-Aware Token Pruning for 3D Question Answering cites this paper.

Seeing Once is Enough? Online Geometry-Aware Token Pruning for 3D Question Answering Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-11T21:50:49.625343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:50:49.625343Z digest=sha256:4fb7533b4d79b3a86385ef30a376bcbed731d2cb37fc6e885209d2c7bd471981