Pith. sign in

Paper Citation Record · LEDGER

Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2312.08168.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.08168 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:59:11.056801Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T02:44:27.661049Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f82e33e3-62e5-412a-8276-a6204f0f24b7 · inbound

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding cites this paper.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T05:26:19.602076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:26:19.602076Z digest=sha256:3d54e4f062059a858a5d3aaacc5fec0d51af1022846e7bc53822172b6f036c1b

Observation 27f7e11d-5cba-4ba7-9b8b-aa69f5a921d3 · inbound

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences cites this paper.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.858454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.858454Z digest=sha256:0d96b6f124722d1d61e384d95ee81303ddb2a02c4e2aba1b18ccc7928d1e6283

Observation 9ea74b37-bbbf-49b7-9edf-769a8dc1b8c2 · inbound

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation cites this paper.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.333277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.333277Z digest=sha256:7d093ea99f3b6624620c1c4625799a94e20341d771004ed6dd36ffeaa67afcd3

Observation f217f645-8772-4c3f-85ee-db2aaee3defc · inbound

ObjVariantEnsemble: Advancing Point Cloud LLM Evaluation in Challenging Scenes with Subtly Distinguished Objects cites this paper.

ObjVariantEnsemble: Advancing Point Cloud LLM Evaluation in Challenging Scenes with Subtly Distinguished Objects Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T11:54:12.155293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:54:12.155293Z digest=sha256:7a262b4f1660645a3f551b17dc56a8bfd7d168355e89195bc87416ca7cc367a5

Observation f4b82864-43c3-4d00-96bb-a77d840c7f12 · inbound

3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene Understanding cites this paper.

3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene Understanding Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T04:45:36.679239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:45:36.679239Z digest=sha256:d99c5c7971749626435fce99c53a1a6f665c4934f96fd892a0b4d9614dbdf613

Observation a3309f5f-4320-43eb-b769-a4b8ef653210 · inbound

GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models cites this paper.

GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:55.326487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:55.326487Z digest=sha256:5b647cb416ece94cb2c93b6afd76c5e1275a6416605b016fdb046f4fb2f42ded

Observation 907e174d-fdc8-404d-b5db-bf8aa789b7cb · inbound

CL3DOR: Contrastive Learning for 3D Large Multimodal Models via Odds Ratio on High-Resolution Point Clouds cites this paper.

CL3DOR: Contrastive Learning for 3D Large Multimodal Models via Odds Ratio on High-Resolution Point Clouds Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T21:48:27.009762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:48:27.009762Z digest=sha256:0c38eccdac63fdb3ec12b62d9818c7cc8925a303392bd7bb959565e6561f5d86

Observation b171e546-d8cd-4076-8a8f-7972e223fc3f · inbound

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation cites this paper.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.056801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.056801Z digest=sha256:f09e393b6f7808fc501b5150312df846a5d5b1f462b01d3ab47eaa43df02fdc0

Observation bd07887c-f2d2-405e-b30e-2d9dce6fc3e9 · inbound

LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding cites this paper.

LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:46.893381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:46.893381Z digest=sha256:d667198c6a5b2039dc13a9b52e9d83d832fbb595b8c5041ac7274aaa95435779

Observation f1d2de7e-81d4-459f-9cda-97874fd4f2d5 · inbound

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts cites this paper.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:17.141359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:17.141359Z digest=sha256:60423a04e200ca5ce99bfda7ff28cf4b88036aa279e13e8517919bd1fbb174f6

Observation 83ed7cf1-98a9-455c-92d6-50f5faa34813 · inbound

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces cites this paper.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.043939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.043939Z digest=sha256:9e7370136bd5454a4875ae2af577c4f9805e2b92946dd1e9e71401415eae2b64

Observation 0f9edb34-9b72-4b97-8b5d-91ae6ec18f30 · inbound

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation cites this paper.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.681855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.681855Z digest=sha256:2a5e4c04f46d70ff9c636f41100ab96071ec59cc87badd0fea741388387efdc5

Observation 88a6554e-19af-47e0-8da7-d9bf491cdb57 · inbound

StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling cites this paper.

StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T19:37:06.509669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:37:06.509669Z digest=sha256:b61e5f9c30dcbede26fc5915e7ac26bc549c8cef7f0711c92a49e5c430789def

Observation 14f17569-c587-46af-bf64-25ab994e7398 · inbound

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding cites this paper.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:00.188535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:00.188535Z digest=sha256:8531dfe81cf52040c18999fef9c04d83844a382192caf930a3de7dae154c9bc1

Observation fc0b60dd-6355-4a7b-af04-d789863647da · inbound

City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete Learning cites this paper.

City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete Learning Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T16:45:18.210032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:45:18.210032Z digest=sha256:bc89833b17e26f19c652af3dd8b2bafcebfbacac4e23504b5bffb27eb164ad2f

Observation d3b55d4d-7093-4d20-9716-1656db672a69 · inbound

Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models cites this paper.

Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T16:42:13.269825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:42:13.269825Z digest=sha256:5cd691e6f24e6816f2be63dd8eadd2a8545c6b20b0fa4dbe730cfe2af0561397

Observation 062f83a3-560f-41e2-8e9d-11384709116e · inbound

Spatial 3D-LLM: Exploring Spatial Awareness in 3D Vision-Language Models cites this paper.

Spatial 3D-LLM: Exploring Spatial Awareness in 3D Vision-Language Models Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:14:16.888231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:14:16.888231Z digest=sha256:072ef667828084e19f6a92f444c4b82149ae78c0e763b0b78cc90ef6db64a8fc

Observation f61fc6d8-17f6-4dfd-9042-44b6db630081 · inbound

B4DL: A Benchmark for 4D LiDAR LLM in Spatio-Temporal Understanding cites this paper.

B4DL: A Benchmark for 4D LiDAR LLM in Spatio-Temporal Understanding Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:11:55.990699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T00:09:57.236162Z digest=sha256:d3c7a856660705763017f92d75d991d4d828d0e27853f70223a0cbca36144d94

Observation 364d4a58-010a-4834-b8f3-01f1140893dd · inbound

Flame3D: Zero-shot Compositional Reasoning of 3D Scenes with Agentic Language Models cites this paper.

Flame3D: Zero-shot Compositional Reasoning of 3D Scenes with Agentic Language Models Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:11:15.851366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T02:09:09.684786Z digest=sha256:9b8da13a6b49fa4437f31d252ab453eab15d183f3ef4acacb6e2a6a8ea0e4254

Observation bc3c1cb1-049d-446a-9367-efe54c960990 · inbound

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models cites this paper.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.662561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T02:44:00.608590Z digest=sha256:84983141302c974040f2e958de05d56fe5514e57900d1949475be191288b153d

Observation 6d645eff-9322-418b-bf88-1cc0453a0ee0 · inbound

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models cites this paper.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:488ca5e1abdae1ab5918ca6ac2ffa9e1c5ca0bd66b9f7efa03683b497769c8e5

Observation 4675c47a-7b5a-4ab3-b24f-83a706477a0e · inbound

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding cites this paper.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.530176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.530176Z digest=sha256:50ae82e294aaf6c9ce37a2f7fd0980dfcca1bead7a7bc05868eb5ebc5fc0bdf6

Observation f77f55a7-9fcd-4930-a28a-9d075308a302 · inbound

SocietyBench: Forecasting Counterfactual Social-World Evolution cites this paper.

SocietyBench: Forecasting Counterfactual Social-World Evolution Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 147

Resolution
unresolved
no resolver link, observed 2026-08-05T04:17:18.840883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:17:18.840883Z digest=sha256:d3c6daddd192daee96bd719ebfa374bc72f31de7a8f82ec8530961fcdc52c204

Observation f4d0eec4-0369-41df-81bd-2fd1c884db49 · inbound

SocietyBench: Forecasting Counterfactual Social-World Evolution cites this paper.

SocietyBench: Forecasting Counterfactual Social-World Evolution Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 147

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:41.225884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:41.225884Z digest=sha256:5440c4e6693055f09cc562b19773a6034ee14ff5743704c1cd4cddc4d347f0b1

Observation 166595a1-2bb6-4f83-8d6a-ed60a2d91047 · inbound

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding cites this paper.

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T04:28:22.542036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:28:22.542036Z digest=sha256:b3fae27eca2ac50c33743334bb37700bc2242db5893fa8207ae016e5dbb234c0

Observation 5c4ff52c-5c96-4b82-8145-5347e70b6aad · inbound

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding cites this paper.

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:40.863904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:30:40.863904Z digest=sha256:0644fbc92705d3c3626ea1b0f78a38a52f2301dc9e430f010e3399ae1c5b888f

Observation b2bc0ca0-a3d6-4508-b248-06b766cd70ca · inbound

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding cites this paper.

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:06.857543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:06.857543Z digest=sha256:30e05f990473391ccac4d5423cb7ee7a638b1d093c76185a9bc11c171672e622

Observation 80814b0b-2ab0-46c3-a529-d5a44720240d · inbound

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models cites this paper.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.215429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.215429Z digest=sha256:fff6b26afe23d27f0c7fc5f31057766c31b88b9db4bd7013972b0c355f79c008