Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T14:02:18.803304Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 88 of 88 outbound references and 3 inbound Pith citation observations for arXiv:2411.15714.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T14:02:18.803304Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T18:53:00.943821Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T16:18:54.098490Z
88 of 88 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d6c7504f-9516-4f5f-b774-72047c101214 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3eaba14-492b-4cdc-ae2d-120fe54ba347 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58a61e6b-a608-41f5-946d-d143b49a35fa · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Towards in-context scene understanding
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60fd053f-fe81-4fb7-89fa-20945be29804 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond MAPLM: A real-world large-scale vision-language benchmark for map and traffic scene under- standing
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b117f13b-f66a-4f9d-927c-15e4d3827d26 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond SpatialVLM: Endow- ing vision-language models with spatial reasoning capabili- ties
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b20ee71-6525-4ef2-b534-c7759cc61d09 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Poly- Diffuse: Polygonal shape reconstruction via guided set diffu- sion models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 020dc9fc-4d84-4baa-bf18-e566e74af60d · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f52cd7bf-92f4-4007-965f-8c41948b3f5b · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond The cityscapes dataset for semantic urban scene understanding
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9923801-0003-4e7b-b5e7-ee17336f2daf · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond InstructBLIP: towards general-purpose vision-language models with instruction tuning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e14e38a-ef65-48a6-857e-611fd4578a0f · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Objaverse: A universe of annotated 3d objects
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2292a49c-c425-4563-b398-826e48e06cbb · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond PLA: Language-driven open- vocabulary 3d scene understanding
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 86b3923f-4e51-4d43-a509-03b3ccf645af · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Shape anchor guided holistic indoor scene understanding
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7801d887-d0a7-4588-bc59-705e5e17dd87 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond The Llama 3 Herd of Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e214b70-bc06-4223-bca8-bfbf1fec0407 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Data Filtering Networks
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f97cf4f-c0b2-4a96-adc1-fb1c65d52bc5 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond 3D-FUTURE: 3d furniture shape with texture
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 59b10919-b9f0-4428-8f26-eb9a5aa5a895 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Dual attention network for scene segmentation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5bb43f78-c254-4742-99d4-d7c19063c656 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a71aaacb-70b9-43bd-94e3-790785a75993 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af4ddc67-a50b-4506-92b4-f0c4b5a18e8b · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Semantic Abstraction: Open-World 3D Scene Understanding from 2D Vision-Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d386a71c-09b0-4503-b496-18857ad0e4c0 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Scene Graph Reasoning for Visual Question Answering
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64e9cabf-2778-458e-890c-600cd8ad29e0 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Probabilistic future prediction for video scene understanding
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5f9309be-a5d9-43f1-b7ba-ff7243fe9e11 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Mutual Scene Synthesis for Mixed Reality Telepresence
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cd208689-2c21-4649-9265-497b5ecb4787 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Segment any- thing
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ddc1ef5c-9df6-4241-a33f-432e321158ad · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond AI2-THOR: An Interactive 3D Environment for Visual AI
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 186746fb-fbd3-4210-a5b6-6942d03cd0a1 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond TopViewRS: Vision-Language Models as Top-View Spatial Reasoners
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b525e52b-e32f-4c9d-83c1-fc608ca3262b · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond From pixels to graphs: Open-vocabulary scene graph generation with vision-language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2dc21d4a-bb74-48ca-931f-ca17e087a1d2 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Robotic indoor scene captioning from streaming video
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d9940563-1aec-469c-a000-5ae28d1c5688 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Improved baselines with visual instruction tuning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6a4f2041-f045-4fc6-84de-31e7bc6b54b1 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond LLaV A-NeXT: Improved reasoning, ocr, and world knowledge, 2024
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5016f070-9371-487d-9a87-1b836e57649b · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Visual instruction tuning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 67103b3c-282d-47df-9eb3-e90e47be7e44 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83fde7c9-9487-40a5-9421-63b16ca6e1b3 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond GPT-4o System Card
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b16e217-16a5-4831-8f26-e9856c98a318 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond OpenScene: 3d scene understanding with open vocabularies
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ff1de0b6-4ca1-4b15-934c-a74ca9c3ab13 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Seamless scene segmentation
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 472161c7-5c06-48c0-8109-2644c32d1764 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Scene graph refinement network for visual question answering
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4bab1390-faa2-42fa-9cb2-93dbb2cfb7b2 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Recognizing indoor scenes
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4d666cab-f5ab-4b10-8b0e-6ea5aa25ff42 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Learning transferable visual models from natural language supervision
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation faf5f261-8a75-4ab0-997b-8a852e59a898 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Monocular Depth Estimation using Diffusion Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d19b8cda-f938-4146-9ed4-ef835b0d8691 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Structured query- based image retrieval using scene graphs
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0a56fd43-45e3-422f-8b51-80b24bd832c4 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Disentangling orthogonal planes for indoor panoramic room layout estimation with cross-scale distortion awareness
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6d1f5540-7123-44e0-849e-2bfa049b3cf2 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond A benchmark for the evaluation of rgb-d slam systems
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9384d069-a990-41ac-a93c-e8b606b903ce · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Distilled semantics for comprehensive scene under- standing from videos
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cbe49e0e-65f2-453c-9568-01d7a18b06b2 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond No more ambiguity in 360deg room layout via bi-layout estimation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0db297da-263b-4d79-9c68-95d1f754fb7c · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond LLaVA-SG: Leveraging Scene Graphs as Visual Semantic Expression in Vision-Language Models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f4dfcf14-e209-46f1-a679-7315b9f71037 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f0785af-c445-4ec8-9ce8-98939b8685e4 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 653f2056-5c20-4f5f-95ed-7dfcc113768d · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond EmbodiedScan: A holistic multi- modal 3d perception suite towards embodied ai
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 830b18c8-db7a-4e83-a440-22dbe9613375 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond SUN Database: Exploring a large collection of scene categories
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b4cefe0b-812a-4c15-b771-73f57e960be2 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Unified perceptual parsing for scene understanding
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e2b20125-2f40-469b-9ef9-6bf07abe832c · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Depth Anything: Unleashing the power of large-scale unlabeled data
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 850efc0c-36ca-4216-9cbe-4df5737aa17d · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Graph-structured referring expression reasoning in the wild
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 214692b9-1074-408a-867c-0ea8f074b2f4 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond PHYSCENE: Physically interactable 3d scene synthesis for embodied ai
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f4b0342b-d019-4478-acef-ccf8c51d2b37 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond HOLODECK: Language guided generation of 3d embodied ai environments
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 538e56b8-ab74-41d0-8a83-c9bc2864b89b · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Swin3D++: Effective Multi-Source Pretraining for 3D Indoor Scene Understanding
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 520f711c-08e3-4b90-a522-eebfad3189e9 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bd2f766-dd0a-4847-9c83-87ea93607a0f · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Human-aware object placement for visual environment reconstruction
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation aff02606-eb00-435c-8066-80d079a16ba2 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Context prior for scene segmenta- tion
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4cf6d4f0-d78e-4a71-a42e-16ca4901d708 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Agent3D-Zero: An Agent for Zero-shot 3D Understanding
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55d7e85e-47c4-453a-aee3-fc84dadafc55 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond DeepContext: Context-encoding neural pathways for 3d holistic scene understanding
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 80e5cca4-8608-498e-a5d9-2740888eb56a · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond LUMINOUS: Indoor Scene Generation for Embodied AI Challenges
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cd825d64-eb09-4f38-becb-ec1de78d05b4 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Comprehensive image captioning via scene graph decom- position
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 78cbf77e-e712-43bd-8fe9-31b8bc4b1d14 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond select prompt
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 032cc422-851c-4ad1-9918-62d70b60726e · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond For instance, in Figure 11, the relationship [1, support, 4] is considered correct if it is correctly extracted from the JSON file
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9576c886-b533-45db-a771-3a16e590c17e · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond For example, in Figure 11, object 1 has relationships such as [[1, support, 4], [1, support, 5], [1, support, 6], [1, support, 7]]
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d5ccf514-aff7-4fd2-8455-f0d7ed1725d1 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond For example, in Figure 11, there are four layers: the first layer includes 1: 1,2,3, the second layer contains 2: 4,5,6,7,8,9,10, and so on
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a6f765fd-4eae-47f5-b73e-3d9a4c38608e · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond In Figure 11, if an object, such as 1, appears in the JSON, it is considered as accurate
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5bb40f7d-1503-4b41-9471-a2a1b48e9bf4 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond This metric quantifies the accuracy of the positive predictions made by the model
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fcbf0c22-5052-4110-a2d5-e73bf40eed38 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond This metric assesses the model’s ability to identify all relevant instances
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 38784c3b-d74a-4904-bea2-7c632ec29d69 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond This metric is the harmonic mean of Precision and Recall, providing a balanced measure of both metrics
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f1f70c2b-6f71-4e68-be0f-12830c52b7d3 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond object” with a numerical suffix starting from 1. The value of each “object
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f4496e56-7bcf-4663-9aa3-a880d1185d74 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond container
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1704547c-a002-41cb-9f46-1f56fd61050e · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Please consider a desk and its tablecloth as one object
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 473674bc-6de2-4713-98a6-eb3c50fa62e2 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Unresolved cited work
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6f9ca715-f0e9-4f48-af0b-95135e7f648b · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond object1”: {“description
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a9be1a3a-0e8b-4611-8a18-85123cdfde4f · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Unresolved cited work
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3d4e02e5-e3ae-4bff-8c7a-5d14a520ce07 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond {container}
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1fcd0d44-701d-43b9-93cc-8ca40f56b816 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond {container}
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 09c97110-a4c1-4754-9eff-6622d6aa43e2 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond {container}
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f2760584-2be5-4911-8e39-1523483fb605 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond {container}
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 126ed720-09d1-4489-92e1-4deac030f969 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Unresolved cited work
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 959063d5-b067-4721-96c5-390bcb4261b2 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond object1”: {“description
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c5782cdb-7aea-4462-a741-812b7c474756 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond left”, “center/middle
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d5f82e30-2497-4467-a527-abe08eeb01e1 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond In this situation, you still need to select the the suitable bounding box based on the relative position of these three objects
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0bfbdae0-ed21-4bc6-84a7-c1d59798a045 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond reason” and “color
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e2cc38cc-d8b7-436c-8e85-389144bfb296 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond If none of the bounding box meets the description, you should select one randomly
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation dc5bd52b-7543-4f94-902c-d1ebb9eefb49 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond Unresolved cited work
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 048d2161-5125-4b16-bf82-629c487b261f · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond You should select the bounding box and its corresponding color according to the description
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b1a311d3-b5b4-4bde-8146-eb90cb413bd8 · outbound
ROOT: VLM based System for Indoor Scene Understanding and Beyond {description}
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 79d5886f-8c07-4b1b-8754-0e8a348100fd · inbound
Generative Physical AI in Vision: A Survey ROOT: VLM based System for Indoor Scene Understanding and Beyond
Reference 242
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fd3770d-7085-48db-bd4e-98922f244666 · inbound
DreamScene: 3D Gaussian-based End-to-end Text-to-3D Scene Generation ROOT: VLM based System for Indoor Scene Understanding and Beyond
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 73fb5f02-18e2-4dea-acb2-da36c4314e0d · inbound
Hierarchical Evidence-Driven Reasoning for Long Document Understanding ROOT: VLM based System for Indoor Scene Understanding and Beyond
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.