Pith. sign in

Paper Citation Record · LEDGER

ROOT: VLM based System for Indoor Scene Understanding and Beyond

As of 13 August 2026, this Paper Citation Record lists 88 of 88 outbound references and 3 inbound Pith citation observations for arXiv:2411.15714.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15714 v1

Coverage vector

measured 88 of 88 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:02:18.803304Z

measured 91 of 91 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T18:53:00.943821Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T16:18:54.098490Z

Reference resolution

88 of 88 outbound references displayed

  • verified exact3
  • verified fuzzy55
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d6c7504f-9516-4f5f-b774-72047c101214 · outbound

This paper cites GPT-4 Technical Report.

ROOT: VLM based System for Indoor Scene Understanding and Beyond GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:02:18.564198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:02:18.564198Z digest=sha256:82f015a3cf49d9322e042ad267fbe20c3603f244f8aa92522c1769bf295c73dd

Observation c3eaba14-492b-4cdc-ae2d-120fe54ba347 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:02:18.568011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:02:18.568011Z digest=sha256:b80a9770d0efa7dc61e668e41b416b13a2d840ba000476b7f4131a689266c21c

Observation 58a61e6b-a608-41f5-946d-d143b49a35fa · outbound

This paper cites Towards in-context scene understanding.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Towards in-context scene understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T14:02:18.571482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:02:18.571482Z digest=sha256:e4df33312532331efd03f9f9c48b34873f223e6cc145de247bc3d38387cd26ee

Observation 60fd053f-fe81-4fb7-89fa-20945be29804 · outbound

This paper cites MAPLM: A real-world large-scale vision-language benchmark for map and traffic scene under- standing.

ROOT: VLM based System for Indoor Scene Understanding and Beyond MAPLM: A real-world large-scale vision-language benchmark for map and traffic scene under- standing

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T14:02:18.574421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:02:18.574421Z digest=sha256:6c47cab6f907699c08d2c71717fdc0db7a17dd7082894c262572f44ccc7b2988

Observation b117f13b-f66a-4f9d-927c-15e4d3827d26 · outbound

This paper cites SpatialVLM: Endow- ing vision-language models with spatial reasoning capabili- ties.

ROOT: VLM based System for Indoor Scene Understanding and Beyond SpatialVLM: Endow- ing vision-language models with spatial reasoning capabili- ties

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T14:02:18.577277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:02:18.577277Z digest=sha256:30988c5b8986cb18141016e356c589f80866a6d23fc0fc74d22249cfb3e16c15

Observation 9b20ee71-6525-4ef2-b534-c7759cc61d09 · outbound

This paper cites Poly- Diffuse: Polygonal shape reconstruction via guided set diffu- sion models.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Poly- Diffuse: Polygonal shape reconstruction via guided set diffu- sion models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T14:02:18.580966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:02:18.580966Z digest=sha256:1fbc30399435fef8da40c0f0b1add0dcb6770f3c51e81ae8ba2c97dc15b34224

Observation 020dc9fc-4d84-4baa-bf18-e566e74af60d · outbound

This paper cites InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

ROOT: VLM based System for Indoor Scene Understanding and Beyond InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T14:02:18.583838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:02:18.583838Z digest=sha256:b68178148cf77cc503eed545010ade5529d66012ec83f6f6b513a869e4b256a2

Observation f52cd7bf-92f4-4007-965f-8c41948b3f5b · outbound

This paper cites The cityscapes dataset for semantic urban scene understanding.

ROOT: VLM based System for Indoor Scene Understanding and Beyond The cityscapes dataset for semantic urban scene understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T14:02:18.586825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:02:18.586825Z digest=sha256:acb2c6a658ea3a76a99153213b68c7bb1aa60906063b19650fe58dc3ab829102

Observation d9923801-0003-4e7b-b5e7-ee17336f2daf · outbound

This paper cites InstructBLIP: towards general-purpose vision-language models with instruction tuning.

ROOT: VLM based System for Indoor Scene Understanding and Beyond InstructBLIP: towards general-purpose vision-language models with instruction tuning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T14:02:18.589469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:02:18.589469Z digest=sha256:f9302fb005e882f4a35d4d998033407f30c849183d05d2364b9b9e81af1ca13b

Observation 4e14e38a-ef65-48a6-857e-611fd4578a0f · outbound

This paper cites Objaverse: A universe of annotated 3d objects.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Objaverse: A universe of annotated 3d objects

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T14:02:18.592056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:02:18.592056Z digest=sha256:3549da510db7b5b494fc74901da47989911d0858cfa3601124ef67ed3454510b

Observation 2292a49c-c425-4563-b398-826e48e06cbb · outbound

This paper cites PLA: Language-driven open- vocabulary 3d scene understanding.

ROOT: VLM based System for Indoor Scene Understanding and Beyond PLA: Language-driven open- vocabulary 3d scene understanding

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.446113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.594765Z digest=sha256:f201b6359f294626980832f744f0ba6a5e19f08f3284ef17741e7ad0b2a68958

Observation 86b3923f-4e51-4d43-a509-03b3ccf645af · outbound

This paper cites Shape anchor guided holistic indoor scene understanding.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Shape anchor guided holistic indoor scene understanding

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.438257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.597376Z digest=sha256:128e9f742347655fbc57d107be89b95cda98a7593cc2516d9b86c6fbfaee6f99

Observation 7801d887-d0a7-4588-bc59-705e5e17dd87 · outbound

This paper cites The Llama 3 Herd of Models.

ROOT: VLM based System for Indoor Scene Understanding and Beyond The Llama 3 Herd of Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T14:02:18.600014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:02:18.600014Z digest=sha256:edbb4a0c2e035f5f9f9a8956dafd6daeccd94717e8d3d1ab380b91210fa23df9

Observation 5e214b70-bc06-4223-bca8-bfbf1fec0407 · outbound

This paper cites Data Filtering Networks.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Data Filtering Networks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T14:02:18.603007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:02:18.603007Z digest=sha256:7fef74c1fe88725f311c567a4e04380f458750a8227462e44bd9f84d84c44d75

Observation 3f97cf4f-c0b2-4a96-adc1-fb1c65d52bc5 · outbound

This paper cites 3D-FUTURE: 3d furniture shape with texture.

ROOT: VLM based System for Indoor Scene Understanding and Beyond 3D-FUTURE: 3d furniture shape with texture

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.430344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.606099Z digest=sha256:273dda83a9d8760f9f9ba49b1b49b2db8d66e6c2577ae67c6456e8988392d12c

Observation 59b10919-b9f0-4428-8f26-eb9a5aa5a895 · outbound

This paper cites Dual attention network for scene segmentation.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Dual attention network for scene segmentation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.422370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.608595Z digest=sha256:be4aaab0dae104585e15c0b204623167be4932248e2dda0f1573b5b5a3fc619d

Observation 5bb43f78-c254-4742-99d4-d7c19063c656 · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T14:02:18.611357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:02:18.611357Z digest=sha256:b9ed25bfa396df78e36bd707c8ee49ee9c6aafc67f421302fcc516249f1b10f7

Observation a71aaacb-70b9-43bd-94e3-790785a75993 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

ROOT: VLM based System for Indoor Scene Understanding and Beyond ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T14:02:18.614445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:02:18.614445Z digest=sha256:0963dd0cc4cea91f763ab011e96aec546e4e12db715ca4f2eb59b9f9d43aa553

Observation af4ddc67-a50b-4506-92b4-f0c4b5a18e8b · outbound

This paper cites Semantic Abstraction: Open-World 3D Scene Understanding from 2D Vision-Language Models.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Semantic Abstraction: Open-World 3D Scene Understanding from 2D Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T14:02:18.617908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:02:18.617908Z digest=sha256:246fd58906e0cd1c8a87686b04b2fd07f0c8c4f19c6480d8e3ed818c1d6d90dd

Observation d386a71c-09b0-4503-b496-18857ad0e4c0 · outbound

This paper cites Scene Graph Reasoning for Visual Question Answering.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Scene Graph Reasoning for Visual Question Answering

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T14:02:18.620273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:02:18.620273Z digest=sha256:4be0b68643d153fc4b70fb0681b200df6fcddbb7fa47af522b771a99a120efd3

Observation 64e9cabf-2778-458e-890c-600cd8ad29e0 · outbound

This paper cites Probabilistic future prediction for video scene understanding.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Probabilistic future prediction for video scene understanding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.414570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.622764Z digest=sha256:746fcf8f060d800f03120c68e50efc8f292af32ffe0c139149adf4862030e41f

Observation 5f9309be-a5d9-43f1-b7ba-ff7243fe9e11 · outbound

This paper cites Mutual Scene Synthesis for Mixed Reality Telepresence.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Mutual Scene Synthesis for Mixed Reality Telepresence

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-12T14:02:18.924277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.625028Z digest=sha256:ce2e5608ed5adf4e3f615a758a9c454276c8257659ad79c9f0b9d47e0deec38b

Observation cd208689-2c21-4649-9265-497b5ecb4787 · outbound

This paper cites Segment any- thing.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Segment any- thing

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.406400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.627565Z digest=sha256:7b9f02051909a6fdafb9c865042d36e9d3f24b1ff0e8bde4983106fd5cad34f3

Observation ddc1ef5c-9df6-4241-a33f-432e321158ad · outbound

This paper cites AI2-THOR: An Interactive 3D Environment for Visual AI.

ROOT: VLM based System for Indoor Scene Understanding and Beyond AI2-THOR: An Interactive 3D Environment for Visual AI

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T14:02:18.629784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:02:18.629784Z digest=sha256:569abcdec6b10e8e0afd8ee4b670770c4bb11ba03d65208ed0bf5bca38ced307

Observation 186746fb-fbd3-4210-a5b6-6942d03cd0a1 · outbound

This paper cites TopViewRS: Vision-Language Models as Top-View Spatial Reasoners.

ROOT: VLM based System for Indoor Scene Understanding and Beyond TopViewRS: Vision-Language Models as Top-View Spatial Reasoners

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T14:02:18.632789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:02:18.632789Z digest=sha256:f1ad4a47ddf3f1c8531d04b7c49f1160a9a34c0112f0f2443f3f2bc43d1558c1

Observation b525e52b-e32f-4c9d-83c1-fc608ca3262b · outbound

This paper cites From pixels to graphs: Open-vocabulary scene graph generation with vision-language models.

ROOT: VLM based System for Indoor Scene Understanding and Beyond From pixels to graphs: Open-vocabulary scene graph generation with vision-language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.399268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.635912Z digest=sha256:ef8b79b21154756cd19cb1e084c54022fd980361188cfb30057630ae4304189f

Observation 2dc21d4a-bb74-48ca-931f-ca17e087a1d2 · outbound

This paper cites Robotic indoor scene captioning from streaming video.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Robotic indoor scene captioning from streaming video

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.392138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.638751Z digest=sha256:3a702717456a4a292b3498ba95ad7089d1cfe9b801ac7f3bc6a131a21841580c

Observation d9940563-1aec-469c-a000-5ae28d1c5688 · outbound

This paper cites Improved baselines with visual instruction tuning.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Improved baselines with visual instruction tuning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.385091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.641460Z digest=sha256:53405d516b9812640ef28f954e5f87ff04fc69828883a21d23629dcbfb0d7c30

Observation 6a4f2041-f045-4fc6-84de-31e7bc6b54b1 · outbound

This paper cites LLaV A-NeXT: Improved reasoning, ocr, and world knowledge, 2024.

ROOT: VLM based System for Indoor Scene Understanding and Beyond LLaV A-NeXT: Improved reasoning, ocr, and world knowledge, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.377558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.644200Z digest=sha256:70072164348284a302fc655da2e75c85c0b252cae426e0e93ed65bb49a496af7

Observation 5016f070-9371-487d-9a87-1b836e57649b · outbound

This paper cites Visual instruction tuning.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Visual instruction tuning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.369396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.647071Z digest=sha256:cb1da014e0461e12515aaf80d189b15c686f20518398ec0c970d3b5c45351a3c

Observation 67103b3c-282d-47df-9eb3-e90e47be7e44 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T14:02:18.649747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:02:18.649747Z digest=sha256:e34e5d544573bfaa7ee106f08fdec691c8cbcd723e20e4b083565817d3fa7c9d

Observation 83fde7c9-9487-40a5-9421-63b16ca6e1b3 · outbound

This paper cites GPT-4o System Card.

ROOT: VLM based System for Indoor Scene Understanding and Beyond GPT-4o System Card

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T14:02:18.652797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:02:18.652797Z digest=sha256:fd1d79da584b701bbc77d75c2df4618c2e509a0045c0da03b22d21c9f0bb4fd2

Observation 0b16e217-16a5-4831-8f26-e9856c98a318 · outbound

This paper cites OpenScene: 3d scene understanding with open vocabularies.

ROOT: VLM based System for Indoor Scene Understanding and Beyond OpenScene: 3d scene understanding with open vocabularies

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.361358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.655704Z digest=sha256:65eac40db41353f51b223fd01a493e650a6415d482745f36ccb6691c0660ffc7

Observation ff1de0b6-4ca1-4b15-934c-a74ca9c3ab13 · outbound

This paper cites Seamless scene segmentation.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Seamless scene segmentation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.353574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.658502Z digest=sha256:863925dda140b30bf72f027ab12545a37df9ae1360532dc91239f0bf6969778d

Observation 472161c7-5c06-48c0-8109-2644c32d1764 · outbound

This paper cites Scene graph refinement network for visual question answering.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Scene graph refinement network for visual question answering

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.345518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.661492Z digest=sha256:4c79d2412213bb7ce6d00db7df34ad8b767205a33498ecefd16530049f149eb8

Observation 4bab1390-faa2-42fa-9cb2-93dbb2cfb7b2 · outbound

This paper cites Recognizing indoor scenes.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Recognizing indoor scenes

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.337626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.664257Z digest=sha256:b6eb93ced1de2a7f38452a4bff774aa736bc5bd835c295352f601ef897036f67

Observation 4d666cab-f5ab-4b10-8b0e-6ea5aa25ff42 · outbound

This paper cites Learning transferable visual models from natural language supervision.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Learning transferable visual models from natural language supervision

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.329988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.666918Z digest=sha256:b0964d34b9f294e7aa5a40d24dfe47aadfd916c46aef5231f69800af62e71513

Observation faf5f261-8a75-4ab0-997b-8a852e59a898 · outbound

This paper cites Monocular Depth Estimation using Diffusion Models.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Monocular Depth Estimation using Diffusion Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T14:02:18.669520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:02:18.669520Z digest=sha256:c733e0b93f931846987151e919fd33df451a0611b4ab6b661a5956c296c2c138

Observation d19b8cda-f938-4146-9ed4-ef835b0d8691 · outbound

This paper cites Structured query- based image retrieval using scene graphs.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Structured query- based image retrieval using scene graphs

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.322128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.672821Z digest=sha256:6dae760baa1d0135a2d34f13e98667e69262ff01085a98a7e9b518a2cd63c4b1

Observation 0a56fd43-45e3-422f-8b51-80b24bd832c4 · outbound

This paper cites Disentangling orthogonal planes for indoor panoramic room layout estimation with cross-scale distortion awareness.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Disentangling orthogonal planes for indoor panoramic room layout estimation with cross-scale distortion awareness

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.314216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.675554Z digest=sha256:23c628d364b58be9155d49b4e657df4b2039e93cf8f8e83434e53c56e9eafbdf

Observation 6d1f5540-7123-44e0-849e-2bfa049b3cf2 · outbound

This paper cites A benchmark for the evaluation of rgb-d slam systems.

ROOT: VLM based System for Indoor Scene Understanding and Beyond A benchmark for the evaluation of rgb-d slam systems

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.306154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.678290Z digest=sha256:4b52b1d49e65df76d7a937106f6002f6cd596ce90ecec9c43c42ed4d83259710

Observation 9384d069-a990-41ac-a93c-e8b606b903ce · outbound

This paper cites Distilled semantics for comprehensive scene under- standing from videos.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Distilled semantics for comprehensive scene under- standing from videos

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.299127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.680996Z digest=sha256:27f2c9ceb3e05e9612af9c83d1c306a455c8edf0bd9b7563a6d8799d88ae670f

Observation cbe49e0e-65f2-453c-9568-01d7a18b06b2 · outbound

This paper cites No more ambiguity in 360deg room layout via bi-layout estimation.

ROOT: VLM based System for Indoor Scene Understanding and Beyond No more ambiguity in 360deg room layout via bi-layout estimation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.291871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.683807Z digest=sha256:15c6e6a9c2557804786588e891683cb6abdd01b4654e8eeea4a3576ce84b31cc

Observation 0db297da-263b-4d79-9c68-95d1f754fb7c · outbound

This paper cites LLaVA-SG: Leveraging Scene Graphs as Visual Semantic Expression in Vision-Language Models.

ROOT: VLM based System for Indoor Scene Understanding and Beyond LLaVA-SG: Leveraging Scene Graphs as Visual Semantic Expression in Vision-Language Models

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-12T14:02:18.877852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.686551Z digest=sha256:b51e3bdbb8cc36d1e8a0c18f0b737c7403a44989f6498d6d2aa4a706bc056ff1

Observation f4dfcf14-e209-46f1-a679-7315b9f71037 · outbound

This paper cites Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T14:02:18.689585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:02:18.689585Z digest=sha256:a449eab392035bfc6402a4bbf3e5960307c52257c3265bc42815684042b97222

Observation 1f0785af-c445-4ec8-9ce8-98939b8685e4 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T14:02:18.692598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:02:18.692598Z digest=sha256:67bccff15294745f9a34b59c15aca9d49e0cd4a43e65dd036c702a32021444ab

Observation 653f2056-5c20-4f5f-95ed-7dfcc113768d · outbound

This paper cites EmbodiedScan: A holistic multi- modal 3d perception suite towards embodied ai.

ROOT: VLM based System for Indoor Scene Understanding and Beyond EmbodiedScan: A holistic multi- modal 3d perception suite towards embodied ai

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.283685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.695338Z digest=sha256:e8f65b558fda9b290d51687746c3bd68bcf37332f233e48d3213473bb53c4c11

Observation 830b18c8-db7a-4e83-a440-22dbe9613375 · outbound

This paper cites SUN Database: Exploring a large collection of scene categories.

ROOT: VLM based System for Indoor Scene Understanding and Beyond SUN Database: Exploring a large collection of scene categories

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.275756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.697643Z digest=sha256:26a395c5f07b4077eac3be8204aed4d7fc83db34eb4b93c9d6f4e5d189377d15

Observation b4cefe0b-812a-4c15-b771-73f57e960be2 · outbound

This paper cites Unified perceptual parsing for scene understanding.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Unified perceptual parsing for scene understanding

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.267578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.700021Z digest=sha256:4c8fe4a275913ebf799e30f4fc9ed0347970c7c65fa36fb2a1020a533ac21330

Observation e2b20125-2f40-469b-9ef9-6bf07abe832c · outbound

This paper cites Depth Anything: Unleashing the power of large-scale unlabeled data.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Depth Anything: Unleashing the power of large-scale unlabeled data

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.259406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.702439Z digest=sha256:c30e8cd8eaf97ca826a53e353f5f4ea905d26d0ceb02bfdfc01ff090acde3b0c

Observation 850efc0c-36ca-4216-9cbe-4df5737aa17d · outbound

This paper cites Graph-structured referring expression reasoning in the wild.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Graph-structured referring expression reasoning in the wild

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.251461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.704771Z digest=sha256:451decf75f697a2f8c3b204112873d71936f79629065066b02523f3009e19ef6

Observation 214692b9-1074-408a-867c-0ea8f074b2f4 · outbound

This paper cites PHYSCENE: Physically interactable 3d scene synthesis for embodied ai.

ROOT: VLM based System for Indoor Scene Understanding and Beyond PHYSCENE: Physically interactable 3d scene synthesis for embodied ai

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.243136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.706996Z digest=sha256:9a77fa2347355032e63bfdaf3939d1898c07e53a71928c4be239bd373d2f285d

Observation f4b0342b-d019-4478-acef-ccf8c51d2b37 · outbound

This paper cites HOLODECK: Language guided generation of 3d embodied ai environments.

ROOT: VLM based System for Indoor Scene Understanding and Beyond HOLODECK: Language guided generation of 3d embodied ai environments

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.234386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.709129Z digest=sha256:cf0c00b22d1c924915025c89a5d85c377ee4eab999af2274ae327555a2fb3515

Observation 538e56b8-ab74-41d0-8a83-c9bc2864b89b · outbound

This paper cites Swin3D++: Effective Multi-Source Pretraining for 3D Indoor Scene Understanding.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Swin3D++: Effective Multi-Source Pretraining for 3D Indoor Scene Understanding

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T14:02:18.711377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:02:18.711377Z digest=sha256:a51f4bd7ee1a142df94ea12fe5ebbbaa16369b917d7b9515e49d561ca19a18fa

Observation 520f711c-08e3-4b90-a522-eebfad3189e9 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

ROOT: VLM based System for Indoor Scene Understanding and Beyond MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T14:02:18.714146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:02:18.714146Z digest=sha256:f6a60858536ad1e14901f804c6c4dc20fa1c9504ce2a2bb6706789ac829132e6

Observation 3bd2f766-dd0a-4847-9c83-87ea93607a0f · outbound

This paper cites Human-aware object placement for visual environment reconstruction.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Human-aware object placement for visual environment reconstruction

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.225556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.717055Z digest=sha256:fdca49eef71d98425938a8891bbb6dd921d6fd8a13e537a037faa7dfb6ba8dac

Observation aff02606-eb00-435c-8066-80d079a16ba2 · outbound

This paper cites Context prior for scene segmenta- tion.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Context prior for scene segmenta- tion

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.216349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.720332Z digest=sha256:5379c0bb94a4720deb159bfff9b3d6d82f593d04dce09b4c5856ffa9d2024bc3

Observation 4cf6d4f0-d78e-4a71-a42e-16ca4901d708 · outbound

This paper cites Agent3D-Zero: An Agent for Zero-shot 3D Understanding.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Agent3D-Zero: An Agent for Zero-shot 3D Understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T14:02:18.722908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:02:18.722908Z digest=sha256:0130a82d1ab588e6590e3b79044ded45b963e9e48e510940ab39e88734003944

Observation 55d7e85e-47c4-453a-aee3-fc84dadafc55 · outbound

This paper cites DeepContext: Context-encoding neural pathways for 3d holistic scene understanding.

ROOT: VLM based System for Indoor Scene Understanding and Beyond DeepContext: Context-encoding neural pathways for 3d holistic scene understanding

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.209079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.725757Z digest=sha256:2dc41e2d5651de5bdecd8ecb4a08c120f7f95264623cc6ac498acfdb56de41b0

Observation 80e5cca4-8608-498e-a5d9-2740888eb56a · outbound

This paper cites LUMINOUS: Indoor Scene Generation for Embodied AI Challenges.

ROOT: VLM based System for Indoor Scene Understanding and Beyond LUMINOUS: Indoor Scene Generation for Embodied AI Challenges

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-08-12T14:02:18.831038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.728355Z digest=sha256:7a4641e817c823f5582c928966f58c325fb8da6bb5d6ab67fa4dc30ff69aa617

Observation cd825d64-eb09-4f38-becb-ec1de78d05b4 · outbound

This paper cites Comprehensive image captioning via scene graph decom- position.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Comprehensive image captioning via scene graph decom- position

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.201427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.731107Z digest=sha256:0349763332040b18aa512878439d2fb8d88eb96c76ab6bcaac10cd2734c91f64

Observation 78cbf77e-e712-43bd-8fe9-31b8bc4b1d14 · outbound

This paper cites select prompt.

ROOT: VLM based System for Indoor Scene Understanding and Beyond select prompt

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.193275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.733650Z digest=sha256:7980b648fad89b6e2c956abc301dba6345c09949ffe98f0d7b384a5e789a009a

Observation 032cc422-851c-4ad1-9918-62d70b60726e · outbound

This paper cites For instance, in Figure 11, the relationship [1, support, 4] is considered correct if it is correctly extracted from the JSON file.

ROOT: VLM based System for Indoor Scene Understanding and Beyond For instance, in Figure 11, the relationship [1, support, 4] is considered correct if it is correctly extracted from the JSON file

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.185328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.736759Z digest=sha256:1430a268f7930ce97ad677fd7b6f7f7f786e13ea9db35d814da3f4399673d2d3

Observation 9576c886-b533-45db-a771-3a16e590c17e · outbound

This paper cites For example, in Figure 11, object 1 has relationships such as [[1, support, 4], [1, support, 5], [1, support, 6], [1, support, 7]].

ROOT: VLM based System for Indoor Scene Understanding and Beyond For example, in Figure 11, object 1 has relationships such as [[1, support, 4], [1, support, 5], [1, support, 6], [1, support, 7]]

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.177323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.739995Z digest=sha256:1eb11ffb62d0176c002b251407aedc515fca2fc86ec8527ed4640b454d5d479c

Observation d5ccf514-aff7-4fd2-8455-f0d7ed1725d1 · outbound

This paper cites For example, in Figure 11, there are four layers: the first layer includes 1: 1,2,3, the second layer contains 2: 4,5,6,7,8,9,10, and so on.

ROOT: VLM based System for Indoor Scene Understanding and Beyond For example, in Figure 11, there are four layers: the first layer includes 1: 1,2,3, the second layer contains 2: 4,5,6,7,8,9,10, and so on

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.169305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.743083Z digest=sha256:0cf69c555d243b5f1e813e5245a8dc2e9600e3bbbec959226ea55a96125020ad

Observation a6f765fd-4eae-47f5-b73e-3d9a4c38608e · outbound

This paper cites In Figure 11, if an object, such as 1, appears in the JSON, it is considered as accurate.

ROOT: VLM based System for Indoor Scene Understanding and Beyond In Figure 11, if an object, such as 1, appears in the JSON, it is considered as accurate

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.161476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.745814Z digest=sha256:ab9d65c84ea06676ad215815ae6253dc6148b0b4d34a2ac536c31c00d9e0ea17

Observation 5bb40f7d-1503-4b41-9471-a2a1b48e9bf4 · outbound

This paper cites This metric quantifies the accuracy of the positive predictions made by the model.

ROOT: VLM based System for Indoor Scene Understanding and Beyond This metric quantifies the accuracy of the positive predictions made by the model

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.153574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.748686Z digest=sha256:e401a575c930d25ac53c73ddf1409c37c7cd046f51477f816fa5b27ba9321281

Observation fcbf0c22-5052-4110-a2d5-e73bf40eed38 · outbound

This paper cites This metric assesses the model’s ability to identify all relevant instances.

ROOT: VLM based System for Indoor Scene Understanding and Beyond This metric assesses the model’s ability to identify all relevant instances

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.145709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.751197Z digest=sha256:9164f32860b6cfaa4b6de91560fa8f061ee0d9e4ceb8f5d4fd89da6dd5edac25

Observation 38784c3b-d74a-4904-bea2-7c632ec29d69 · outbound

This paper cites This metric is the harmonic mean of Precision and Recall, providing a balanced measure of both metrics.

ROOT: VLM based System for Indoor Scene Understanding and Beyond This metric is the harmonic mean of Precision and Recall, providing a balanced measure of both metrics

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.137776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.754029Z digest=sha256:832668161bba247c6b6b45f3129dd92b26c2864c290ee4d79bb1c670ff4b45db

Observation f1f70c2b-6f71-4e68-be0f-12830c52b7d3 · outbound

This paper cites object” with a numerical suffix starting from 1. The value of each “object.

ROOT: VLM based System for Indoor Scene Understanding and Beyond object” with a numerical suffix starting from 1. The value of each “object

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.129729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.756746Z digest=sha256:968aeec5fa385bf3ca823c2ef66903b5598514382841b86482464815cd6073f4

Observation f4496e56-7bcf-4663-9aa3-a880d1185d74 · outbound

This paper cites container.

ROOT: VLM based System for Indoor Scene Understanding and Beyond container

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.115143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.762355Z digest=sha256:45d5b18a3d02becf6d5fd391507bed6d9644d26f1eae36d4a148de63c4b73241

Observation 1704547c-a002-41cb-9f46-1f56fd61050e · outbound

This paper cites Please consider a desk and its tablecloth as one object.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Please consider a desk and its tablecloth as one object

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.108303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.764937Z digest=sha256:acd42d85be9ccfaf92e911a75ea1568cc127370cf554e78c03097c57db4db6c7

Observation 473674bc-6de2-4713-98a6-eb3c50fa62e2 · outbound

This paper cites an unresolved cited work.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:02:19.100746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.767731Z digest=sha256:98334eaa9f7eb606741035daf60dedef1e406f1e5ed4e1881b3b00f984b1088b

Observation 6f9ca715-f0e9-4f48-af0b-95135e7f648b · outbound

This paper cites object1”: {“description.

ROOT: VLM based System for Indoor Scene Understanding and Beyond object1”: {“description

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.093106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.770412Z digest=sha256:06714035556d2f3818eb6e26b7ed832ce9e79363acc7420303096f97b98338aa

Observation a9be1a3a-0e8b-4611-8a18-85123cdfde4f · outbound

This paper cites an unresolved cited work.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:02:19.122288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.772948Z digest=sha256:e3b32c4882f79e3d7b50ec7a9fca39492b2aee51424d2fae6f2a68cde116d961

Observation 3d4e02e5-e3ae-4bff-8c7a-5d14a520ce07 · outbound

This paper cites {container}.

ROOT: VLM based System for Indoor Scene Understanding and Beyond {container}

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.085336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.775742Z digest=sha256:74fe61ce7fcad1879442e0ae3b4988e897d2442d1c94e9d8d0b2962b8b7affc8

Observation 1fcd0d44-701d-43b9-93cc-8ca40f56b816 · outbound

This paper cites {container}.

ROOT: VLM based System for Indoor Scene Understanding and Beyond {container}

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.077809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.778358Z digest=sha256:2a32c04090583d94d90c5ff96064ace657371f2ef7cb6d4d1455671fd234a59b

Observation 09c97110-a4c1-4754-9eff-6622d6aa43e2 · outbound

This paper cites {container}.

ROOT: VLM based System for Indoor Scene Understanding and Beyond {container}

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.069999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.780957Z digest=sha256:239693973f46a5faf75abfe6b170adfd789a5f0e4585f80e5e448e356febbd77

Observation f2760584-2be5-4911-8e39-1523483fb605 · outbound

This paper cites {container}.

ROOT: VLM based System for Indoor Scene Understanding and Beyond {container}

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.062377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.783483Z digest=sha256:6456724ad8c4b4f01bb2f956ba8b2db88ec4eb7c00777cf82cd8a29047e93c8e

Observation 126ed720-09d1-4489-92e1-4deac030f969 · outbound

This paper cites an unresolved cited work.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:02:19.054369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.785677Z digest=sha256:437252441aa71b1a1fe0dcde8f98bacc880dfd1fd7516c46831a30f4cf594d78

Observation 959063d5-b067-4721-96c5-390bcb4261b2 · outbound

This paper cites object1”: {“description.

ROOT: VLM based System for Indoor Scene Understanding and Beyond object1”: {“description

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.046490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.787785Z digest=sha256:d4205c56b4b1587e17f3d5b907e56510c7a85b8a4092a37c1786b0c63f715c29

Observation c5782cdb-7aea-4462-a741-812b7c474756 · outbound

This paper cites left”, “center/middle.

ROOT: VLM based System for Indoor Scene Understanding and Beyond left”, “center/middle

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.038490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.790078Z digest=sha256:40c6fc623a3ca2c4edc976f797e21f9482fa3a82c771335a36209dcf4d275545

Observation d5f82e30-2497-4467-a527-abe08eeb01e1 · outbound

This paper cites In this situation, you still need to select the the suitable bounding box based on the relative position of these three objects.

ROOT: VLM based System for Indoor Scene Understanding and Beyond In this situation, you still need to select the the suitable bounding box based on the relative position of these three objects

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.029581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.792289Z digest=sha256:d64b667c38c17015c07617483fde1ccb3d89e172f2c3ef644aee11c0f9d8deb6

Observation 0bfbdae0-ed21-4bc6-84a7-c1d59798a045 · outbound

This paper cites reason” and “color.

ROOT: VLM based System for Indoor Scene Understanding and Beyond reason” and “color

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.021518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.794459Z digest=sha256:95f7a1dd0976b3595ce2a0ddb1428bc901089301efc71f2dcb12ff07cfc8641d

Observation e2cc38cc-d8b7-436c-8e85-389144bfb296 · outbound

This paper cites If none of the bounding box meets the description, you should select one randomly.

ROOT: VLM based System for Indoor Scene Understanding and Beyond If none of the bounding box meets the description, you should select one randomly

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:19.013943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.796703Z digest=sha256:87ef8a23e06bfa542b7d186a21df2a580e61784bc9a40f3c885272b104e34c34

Observation dc5bd52b-7543-4f94-902c-d1ebb9eefb49 · outbound

This paper cites an unresolved cited work.

ROOT: VLM based System for Indoor Scene Understanding and Beyond Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:02:19.005959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.798882Z digest=sha256:c334742c003ec8dd2ecc37b57aa3dc0218c2c14f638c5082a2fe87e56becf5ae

Observation 048d2161-5125-4b16-bf82-629c487b261f · outbound

This paper cites You should select the bounding box and its corresponding color according to the description.

ROOT: VLM based System for Indoor Scene Understanding and Beyond You should select the bounding box and its corresponding color according to the description

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:18.997407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.801089Z digest=sha256:2b5b75d2594dae4afc329a431cb260bc11392949327b291d0054ddb398d9cc5d

Observation b1a311d3-b5b4-4bde-8146-eb90cb413bd8 · outbound

This paper cites {description}.

ROOT: VLM based System for Indoor Scene Understanding and Beyond {description}

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:02:18.988634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:02:18.803304Z digest=sha256:9b8a6dcd81a69af1ea922af7090294ca35858ae4a2e48c2b16a5792ad498c09b

Pith citing papers

Observation 79d5886f-8c07-4b1b-8754-0e8a348100fd · inbound

Generative Physical AI in Vision: A Survey cites this paper.

Generative Physical AI in Vision: A Survey ROOT: VLM based System for Indoor Scene Understanding and Beyond

Reference 242

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:00.943821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:53:00.943821Z digest=sha256:418cdf8f1914aebf3ea3a95f0c3e7579e9a198de4f83599522c7578f1e7d0187

Observation 6fd3770d-7085-48db-bd4e-98922f244666 · inbound

DreamScene: 3D Gaussian-based End-to-end Text-to-3D Scene Generation cites this paper.

DreamScene: 3D Gaussian-based End-to-end Text-to-3D Scene Generation ROOT: VLM based System for Indoor Scene Understanding and Beyond

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:18:54.158852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T16:18:48.178757Z digest=sha256:ffa0c33aa9993a25e79e92dd79f02467e92cd75f6fdea7987ec2cef4b8d3aa0f

Observation 73fb5f02-18e2-4dea-acb2-da36c4314e0d · inbound

Hierarchical Evidence-Driven Reasoning for Long Document Understanding cites this paper.

Hierarchical Evidence-Driven Reasoning for Long Document Understanding ROOT: VLM based System for Indoor Scene Understanding and Beyond

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-11T16:13:40.332262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:13:40.332262Z digest=sha256:a7e9c9ae27e088e24f47385cb269d98be063fc82d80a07dac652bce0603184c6