Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:22:03.380823Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 6 inbound Pith citation observations for arXiv:2505.04911.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:22:03.380823Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-02T14:16:39.649823Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T16:39:57.806668Z
57 of 57 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e86eb903-a87a-4e2f-9bf5-fc7b6fbca0c8 · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1d12989b-4334-4396-9791-e2829238a99f · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Scanqa: 3d question answering for spatial scene understanding
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9605b7ab-546f-4a40-87e2-f74e0f8f9b1d · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Language models are few-shot learners
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12bd605f-b8da-4fb9-8a85-6ae7a8c52104 · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Scanrefer: 3d object localization in rgb-d scans using natural language
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c77c582b-ceee-4e72-acfd-2c9c91dd2bb9 · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Language conditioned spatial relation reasoning for 3d object grounding
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3a7b681-8ca4-4306-b2c5-817da5cf11a2 · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models End-to-end 3d dense captioning with vote2cap-detr
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 62a696d9-2606-4d80-8f36-054def1b838e · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Ll3da: Visual interactive instruction tuning for omni-3d understand- ing reasoning and planning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 19477877-9569-4c18-9160-a0cba3708f80 · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models V ote2cap-detr++: Decoupling localization and describing for end-to-end 3d dense captioning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddab9ca1-522b-407d-be4f-3b801bf96576 · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Scan2cap: Context-aware dense captioning in rgb-d scans
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bb8a2e0d-e198-4779-91a4-10508fb01ecb · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Zero-Shot Video Question Answering with Procedural Programs
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e62622ba-bc4b-46ce-8b0e-a6f2555e6e67 · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 97757f96-34c9-402a-b4fa-55bc56daca0e · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Bundlefusion: Real-time globally consistent 3d reconstruction using on-the-fly surface reintegration
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 73d615a7-3f14-4ce8-9fae-f1fe4c60a5d1 · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Lan- guage to map: Topological map generation from natural lan- guage path instructions
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6ab4f79a-5cf5-4133-b3a8-8ed308676a1d · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c85059b-8684-45bd-bf00-4b842fe50835 · outbound
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ea930b38-9c25-4aa5-b206-aa8d476712b7 · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models 3d-llm: Inject- ing the 3d world into large language models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6c524b60-e4dc-4b7c-9069-91dc333d6502 · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Vtimellm: Empower llm to grasp video moments
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc4bb9a6-c3aa-46c9-8ffe-59e94d7d12ea · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Chat-scene: Bridging 3d scene and large language models with object identifiers
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 42dea5ca-3c0f-4f88-bcd5-4a9a421e304e · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models An embodied generalist agent in 3d world
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 08f25081-872c-45dc-8d72-7b56624a3114 · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Multi- view transformer for 3d visual grounding
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation feba9d65-0368-403c-9c7d-88c1af251926 · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models More: Multi-order relation mining for dense captioning in 3d scenes
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5fc57cdc-e54a-4e75-9c31-5352eecc6956 · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Chat-univi: Unified visual representation em- powers large language models with image and video under- standing
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 409bcd8b-76b4-459e-a068-4c1804712b8d · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Context-aware alignment and mutual masking for 3d- language pre-training
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5873b649-ecec-4ba3-b20b-2261d623c629 · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Large language models are zero-shot reasoners
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 22f3ec89-007b-4aff-a790-82c9e75bbebb · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models VideoINSTA: Zero-shot Long Video Understanding via Informative Spatial-Temporal Reasoning with LLMs
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e386186-e9eb-44c0-be87-a92f0d04684e · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6210f3c7-3150-4d1c-8a14-250ccb5f87f9 · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models SQA3D: Situated Question Answering in 3D Scenes
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecff8102-9088-4f9f-9f71-53aed34d39fa · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d79439f8-c82d-4088-9bde-eab93ec71203 · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Egoschema: A diagnostic benchmark for very long- form video language understanding
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a9d9fd3d-cb6b-4465-8f5a-79d7cdf080b5 · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Large Language Models Know Your Contextual Search Intent: A Prompting Framework for Conversational Search
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbb705d2-7184-44fc-ad5d-4e7927e915e2 · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Morevqa: Exploring modular reason- ing models for video question answering
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 977261d6-c621-4116-9ad3-86628a6ccff0 · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 838691c0-94e7-48c6-b3d8-a0537adfce99 · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Clip-guided vision-language pre-training for question answering in 3d scenes
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 83064e19-086e-483f-937f-3acc166af40d · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Learn- ing transferable visual models from natural language super- vision
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 117e1819-8a16-454e-b764-f39d5d31578e · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Traveler: A multi-lmm agent frame- work for video question-answering
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 617918b5-632b-4f33-b8ac-2610ed8fa4ec · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 058ab9a4-c8b9-4231-a66f-c261116658f5 · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Four ways to improve verbo-visual fusion for dense 3d visual grounding
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 11f386b9-62d8-44ef-b37e-3e39a81c415f · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Vamos: Versatile action models for video understanding
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3dec7b14-3bf8-4c59-a6e4-ea69e32e4772 · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Videoagent: Long-form video understanding with large language model as agent
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 73ba2d20-d31d-4ac9-b98f-ffd17c52e0f2 · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Language models with im- age descriptors are strong few-shot video-language learners
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1e267989-021d-45ee-acc4-cd5c5fb3e726 · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models 3DRP-Net: 3D Relative Position-aware Network for 3D Visual Grounding
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e35fece7-918a-49a2-a437-e9dea1860898 · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Distill- ing coarse-to-fine semantic matching knowledge for weakly supervised 3d visual grounding
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9be380f3-06e8-44da-92d1-7c819c976757 · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57f36b66-8254-4401-8b8b-d408b660de53 · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdc57e75-91e2-4629-951f-215073b6f9c0 · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Emergent Abilities of Large Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7c67614-691b-4381-b798-a71f74fdb12e · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Retrieval-based video language model for efficient long video question answering
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0477564-9838-4399-9a69-280a98f48bce · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Depth anything: Unleashing the power of large-scale unlabeled data
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cd0b139d-0cd1-4d05-b8a0-31aff332231e · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Self-chained image-language model for video localization and question answering
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a80e4e70-9be1-4353-ae75-2bd9952b0c13 · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models X-trans2cap: Cross-modal knowledge transfer using transformer for 3d dense caption- ing
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5026e5e9-11f5-4cbb-aea8-2b920f8e913a · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models 3dgraphllm: Combin- ing semantic graphs and large language models for 3d scene understanding, 2024
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b4df0afd-5be2-4674-9677-a4a40005770a · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models A Simple LLM Framework for Long-Range Video Question-Answering
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 411d878c-4c64-4dfa-9de8-1318d557804d · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60b24c58-b124-4ebe-914c-b2bf96d06de6 · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Multi3drefer: Grounding text description to multiple 3d ob- jects
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fc122c2b-4a4e-443a-98a1-d9ac1e2b8279 · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models 3dvg- transformer: Relation modeling for visual grounding on point clouds
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e8b58bb-e232-48c0-9cc4-1cc51e9b19ae · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b11c3423-8cde-4b5f-a60b-8b72dcd0781d · outbound
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b5082d10-da5c-4036-998f-6a28e8f99f7a · outbound
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Unresolved cited work
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8e256f1d-9d5c-47c7-a611-d24fdaf9da67 · inbound
Flame3D: Zero-shot Compositional Reasoning of 3D Scenes with Agentic Language Models SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cbafe699-c661-4036-a0bf-7b6d2b5e497a · inbound
SPATIOROUTE: Dynamic Prompt Routing for Zero-Shot Spatial Reasoning SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cffbfe24-0b3a-4e70-ac24-9e6f7f0ccbae · inbound
Reason, Then Re-reason: Cross-view Revisiting Improves Spatial Reasoning SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e4675a4d-f28b-497f-9265-d25caba2b4a3 · inbound
Agentic Collaborative Cognition for Zero-Shot 3D Understanding SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b2aa3805-169c-484a-9519-0475390ff47f · inbound
Agentic Collaborative Cognition for Zero-Shot 3D Understanding SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ffce045f-5314-4807-ace6-774b901ddb76 · inbound
OmniView-Space: Reinforcing Spatial Reasoning via Multi-Perspective Spatial Mapping SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.