Pith. sign in

Paper Citation Record · LEDGER

Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2312.08168.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.08168 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:59:11.056801Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T02:44:27.661049Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f82e33e3-62e5-412a-8276-a6204f0f24b7 · inbound

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding cites this paper.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T05:26:19.602076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:26:19.602076Z digest=sha256:c40d6d418199b95c56fd49ec73343da72628a1f40a49e114bbd505a07ca8ba82

Observation 27f7e11d-5cba-4ba7-9b8b-aa69f5a921d3 · inbound

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences cites this paper.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.858454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.858454Z digest=sha256:7a50fcf9804e1ddf04e004df9ba7b87faca114ee6cc303dc39eefc3c6104cf3d

Observation 9ea74b37-bbbf-49b7-9edf-769a8dc1b8c2 · inbound

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation cites this paper.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.333277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.333277Z digest=sha256:c8ebe666da436d41a7b94ea88b2fa535353487c5e89108f6981b8621c5596b09

Observation f217f645-8772-4c3f-85ee-db2aaee3defc · inbound

ObjVariantEnsemble: Advancing Point Cloud LLM Evaluation in Challenging Scenes with Subtly Distinguished Objects cites this paper.

ObjVariantEnsemble: Advancing Point Cloud LLM Evaluation in Challenging Scenes with Subtly Distinguished Objects Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T11:54:12.155293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:54:12.155293Z digest=sha256:96f4cb84fb59a593abc70b54ab1a4a7fba01d7c159990a7e93561861acf96b74

Observation f4b82864-43c3-4d00-96bb-a77d840c7f12 · inbound

3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene Understanding cites this paper.

3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene Understanding Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T04:45:36.679239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:45:36.679239Z digest=sha256:287d40a0f51265a5cdbb82a616fad83cc27194c2f1a515a342dec36c8742e691

Observation a3309f5f-4320-43eb-b769-a4b8ef653210 · inbound

GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models cites this paper.

GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:55.326487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:55.326487Z digest=sha256:feac07ac1e7d7cf270a5952eff7b75eae9f7f76e35d761145dc3294118588f5c

Observation 907e174d-fdc8-404d-b5db-bf8aa789b7cb · inbound

CL3DOR: Contrastive Learning for 3D Large Multimodal Models via Odds Ratio on High-Resolution Point Clouds cites this paper.

CL3DOR: Contrastive Learning for 3D Large Multimodal Models via Odds Ratio on High-Resolution Point Clouds Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T21:48:27.009762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:48:27.009762Z digest=sha256:fb3eaa1070fe0b9ff6303580301646e5039d850e593343a8c0cd29f9cc916964

Observation b171e546-d8cd-4076-8a8f-7972e223fc3f · inbound

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation cites this paper.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.056801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.056801Z digest=sha256:07bd83eec29c7067d5f99465ea52caedec4e3cf8632ca858ab191896b3155cb4

Observation bd07887c-f2d2-405e-b30e-2d9dce6fc3e9 · inbound

LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding cites this paper.

LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:46.893381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:46.893381Z digest=sha256:ca3d338657a045003c4b8df38f93ba338a9a76311bc19cdcdd574e732215b3a9

Observation f1d2de7e-81d4-459f-9cda-97874fd4f2d5 · inbound

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts cites this paper.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:17.141359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:17.141359Z digest=sha256:73e5be020f43f26447cde719ca57ae2fdf0a9f791b6774674a1355b22d23491c

Observation 83ed7cf1-98a9-455c-92d6-50f5faa34813 · inbound

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces cites this paper.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.043939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.043939Z digest=sha256:67465ea2c2db164cc91522442729e5f86ab58f80cc534e03e94024f7235b14df

Observation 0f9edb34-9b72-4b97-8b5d-91ae6ec18f30 · inbound

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation cites this paper.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.681855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.681855Z digest=sha256:9df49773abe9f6b17b41426f0ae1e263aecc7bf275123816d67630531ad59619

Observation 88a6554e-19af-47e0-8da7-d9bf491cdb57 · inbound

StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling cites this paper.

StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T19:37:06.509669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:37:06.509669Z digest=sha256:77de0adf6334672e6c3c09a043cc8df371982c221ff6959a5b3a5b0947300ec8

Observation 14f17569-c587-46af-bf64-25ab994e7398 · inbound

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding cites this paper.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:00.188535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:00.188535Z digest=sha256:bfd3aed3265c8225a9b65d2c87d9f1206a553cc6309ff62953f48ac747fbbd7e

Observation fc0b60dd-6355-4a7b-af04-d789863647da · inbound

City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete Learning cites this paper.

City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete Learning Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T16:45:18.210032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:45:18.210032Z digest=sha256:a109251d33044dfc36927bfa2e8d5f6bcce2789807eafc7a2dc77856990a39f7

Observation d3b55d4d-7093-4d20-9716-1656db672a69 · inbound

Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models cites this paper.

Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T16:42:13.269825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:42:13.269825Z digest=sha256:95ab94102642d0dbe62bd34366494fb80a6a406877c51e92272966cade869d74

Observation 062f83a3-560f-41e2-8e9d-11384709116e · inbound

Spatial 3D-LLM: Exploring Spatial Awareness in 3D Vision-Language Models cites this paper.

Spatial 3D-LLM: Exploring Spatial Awareness in 3D Vision-Language Models Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:14:16.888231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:14:16.888231Z digest=sha256:641bf7ccf05f97867fb45fad166ab974195ceaa5c99743f13d641bc30c29d453

Observation f61fc6d8-17f6-4dfd-9042-44b6db630081 · inbound

B4DL: A Benchmark for 4D LiDAR LLM in Spatio-Temporal Understanding cites this paper.

B4DL: A Benchmark for 4D LiDAR LLM in Spatio-Temporal Understanding Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:11:55.990699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T00:09:57.236162Z digest=sha256:8b52598a1372e33b6cdc35bbbf640c912df30ba38fa799377370ffa55ac2bb98

Observation 364d4a58-010a-4834-b8f3-01f1140893dd · inbound

Flame3D: Zero-shot Compositional Reasoning of 3D Scenes with Agentic Language Models cites this paper.

Flame3D: Zero-shot Compositional Reasoning of 3D Scenes with Agentic Language Models Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:11:15.851366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T02:09:09.684786Z digest=sha256:c4bf7fd69b086a55eded70a30faa8ad5f4e560023cf2aeec8cbadf13cb20d39e

Observation bc3c1cb1-049d-446a-9367-efe54c960990 · inbound

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models cites this paper.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.662561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T02:44:00.608590Z digest=sha256:74f84860886a8c1f34047a8ec77c4b8d1e0b4fd4257d5f0b6f6a53226676cc21

Observation 6d645eff-9322-418b-bf88-1cc0453a0ee0 · inbound

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models cites this paper.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:a2c4d467c72295fdb071b92fa5d173516f27c92a92846f0a5824ec4b29b8995d

Observation 4675c47a-7b5a-4ab3-b24f-83a706477a0e · inbound

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding cites this paper.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.530176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.530176Z digest=sha256:a303d675ffbaef6c8db2e3a9de9ef5f43285e4a5faf8ded45967a3d18f4c8794

Observation f77f55a7-9fcd-4930-a28a-9d075308a302 · inbound

SocietyBench: Forecasting Counterfactual Social-World Evolution cites this paper.

SocietyBench: Forecasting Counterfactual Social-World Evolution Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 147

Resolution
unresolved
no resolver link, observed 2026-08-05T04:17:18.840883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:17:18.840883Z digest=sha256:1b70730ec00046e388ac99c9cb551229db6a016a6675429c8becd75eb58a2dc6

Observation f4d0eec4-0369-41df-81bd-2fd1c884db49 · inbound

SocietyBench: Forecasting Counterfactual Social-World Evolution cites this paper.

SocietyBench: Forecasting Counterfactual Social-World Evolution Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 147

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:41.225884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:41.225884Z digest=sha256:2dc1bcdda0ff1d02d6c5965997833fa6e5bac929f9100446b54e256ee616d0f1

Observation 166595a1-2bb6-4f83-8d6a-ed60a2d91047 · inbound

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding cites this paper.

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T04:28:22.542036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:28:22.542036Z digest=sha256:61984dca7ad86af4bb0425d58d9120aa4f516f653453a7525f6b101067b3e5a2

Observation 5c4ff52c-5c96-4b82-8145-5347e70b6aad · inbound

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding cites this paper.

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:40.863904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:30:40.863904Z digest=sha256:381f8df431aef4b155b886f843075beb105c5061dd14cdabacf1562fb78cfad1

Observation b2bc0ca0-a3d6-4508-b248-06b766cd70ca · inbound

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding cites this paper.

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:06.857543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:06.857543Z digest=sha256:6982394b8b03c3d847753411635aa3a6b31116c1c9ddf77b1f9e09f7a4949c7d

Observation 80814b0b-2ab0-46c3-a529-d5a44720240d · inbound

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models cites this paper.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.215429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.215429Z digest=sha256:1252bdc4a40a79a559b7407bf8e6798af472332e1c570f1694a59b46a6cbaee2