Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T04:35:37.029873Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 7 inbound Pith citation observations for arXiv:2412.01292.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T04:35:37.029873Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:22:03.376214Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T10:49:47.240788Z
48 of 48 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3fe63a2a-e12e-4997-9012-8109844ba71e · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cce39ea2-b811-48b5-afb0-58f06112dc11 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Scanqa: 3d question answering for spatial scene understanding
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 89625eb6-3a0a-47af-b56b-0c07ee467f36 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Qwen Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f011e68-62db-40b9-9a65-35802a41da18 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8314fb60-969c-494a-8404-e6af4f4ace65 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 766b527c-8945-45cc-85b4-3158c4abfe4d · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Scanrefer: 3d object localization in rgb-d scans using natural language
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99481ef7-81e6-43aa-8a80-6e604f803d84 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences $A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ed7d7af-6a74-46ef-a9c0-42411919e95f · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Ll3da: Visual interactive instruction tuning for omni-3d understand- ing reasoning and planning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 55b59e00-9332-4f1b-85b8-fea3868d9847 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences End-to-end 3d dense captioning with vote2cap- detr
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0084ae39-f534-471a-b858-6b40fbd4d9d1 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95bbf4d7-0bfb-4bee-9130-8cbbc9b3daf7 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Scannet: Richly-annotated 3d reconstructions of indoor scenes
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7b8680e-eed0-4471-94fc-ec625400368c · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 831ae781-da79-4ceb-86bc-b97f7354b054 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4448104-fa27-41da-bf9f-d48299e08c4d · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9598c877-e40e-4dca-a05f-1c8b005f0efc · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences 3d-llm: Injecting the 3d world into large language models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a5b9862-ebd9-4560-9457-c18768df34dd · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Chat-scene: Bridging 3d scene and large language models with object identifiers, 2024
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 27f7e11d-5cba-4ba7-9b8b-aa69f5a921d3 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8e95ab9-4675-4a26-8643-a5b8ab002919 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences An Embodied Generalist Agent in 3D World
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc713f0b-4911-4c64-bd7e-847baf796a85 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16476d8f-cf5d-460d-8cbd-6c8c3c118703 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac1239c6-4c5f-47d1-a80f-5c66a1dc739c · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Context-aware alignment and mutual masking for 3d-language pre-training
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 91ac9420-3e8a-4900-b510-e51b59bc384a · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Beyond the nav-graph: Vision-and-language navigation in continuous environments
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f90ef089-cb8b-405a-a583-f9d18e022e63 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28576c4a-b2b0-42e9-9c68-fe7aef66df65 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Rouge: A package for automatic evaluation of summaries
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4249674-4ce6-4727-a50f-667fa6c79017 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Visual instruction tuning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b1d5d94-7ef6-46d9-b3b1-145ecc5d3287 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Decoupled Weight Decay Regularization
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03f47964-be3c-47ca-a0ed-e811e786e0ab · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Openscene: 3d scene understanding with open vocabularies
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation aa0ba61a-c73e-45dd-8cca-f1cafe004031 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Pointnet++: Deep hierarchical feature learning on point sets in a metric space
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cf8fe14-0928-4d78-bda2-90b5bf9784aa · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d32fbdfe-714f-49aa-8c0e-604f83201fc7 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b67ca78-061a-4d9c-bb91-f968e9452ec5 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Eye movements in iconic visual search
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4a93a63b-c80b-4b5e-9944-eeee7bd933a5 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Mask3d: Mask transformer for 3d semantic instance segmentation
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49ba5686-d19b-4bfd-9813-167bdebcc65a · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Indoor scene segmen- tation using a structured light sensor
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5eec47bb-6fef-436a-9893-525d27d01bd3 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Sun rgb-d: A rgb-d scene understanding benchmark suite
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3ae5547-944b-4072-b5ab-31fefc9fe72a · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Fgprompt: fine-grained goal prompting for image-goal navigation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2abc01b9-451d-48fa-9b09-c1b2d4935959 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D Priors
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d38280e9-3764-4b09-9095-8f114c10ff3e · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences LLaMA: Open and Efficient Foundation Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fd286fb-d35a-4bb2-86ca-7f27d0c4447f · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences A feature-integration theory of attention
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 44949439-52ab-4c0d-89a1-1e7d694004e8 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Cider: Consensus-based image description evalu- ation
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6070df36-5bfd-4529-b8ac-c1d978427a70 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Rio: 3d object instance re-localization in changing indoor environments
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 470c7e11-cbf2-4bdb-b133-de04097cc0c4 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a82ddc5-e9f0-4d88-8fdc-c8425cae31b4 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences OccLLaMA: An Occupancy-Language-Action Generative World Model for Autonomous Driving
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb40708d-d3b5-4e09-aa2b-cb2dba6761f4 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences What attributes guide the deployment of visual attention and how do they do it? Nature reviews neuroscience, 5(6):495–501, 2004
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6e6ffa06-16de-4985-b0fd-5779beaf94e6 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences PointLLM: Empowering Large Language Models to Understand Point Clouds
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8aca1262-4a18-4809-adfd-c15946878935 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences LiDAR-LLM: Exploring the Potential of Large Language Models for 3D LiDAR Understanding
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ba662bc-0477-4522-9029-b31ae0f3edb1 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11dde869-e4c9-4999-a23d-7b200624bf83 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Uni3d: A unified baseline for multi-dataset 3d object detection
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 83f99ab1-5a8a-4afa-a952-dbb517c0ee36 · outbound
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences 3D-VLA: A 3D Vision-Language-Action Generative World Model
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e8b58bb-e232-48c0-9cc4-1cc51e9b19ae · inbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9faf74b7-9151-42b8-a42e-61024e8edca4 · inbound
LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91e291e6-f4df-43a5-88c0-2cb0ff3dd88f · inbound
A Large-Scale Referring Remote Sensing Image Segmentation Dataset and Benchmark LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 537570de-39ae-44c7-b1c1-77720a093e87 · inbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5587a349-bc45-4921-9800-e501f3c2a3f7 · inbound
GaussianVLM: Scene-centric 3D Vision-Language Models using Language-aligned Gaussian Splats for Embodied Reasoning and Beyond LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 966be097-444d-475d-91d4-6aabacd26c1e · inbound
3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c02b4ead-fde1-4b93-92b0-417ce4d8e435 · inbound
Nav-R1: Reasoning and Navigation in Embodied Scenes LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.