Pith. sign in

Paper Citation Record · LEDGER

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models

As of 21 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 0 inbound Pith citation observations for arXiv:2607.06534.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06534 v2

Coverage vector

measured 74 of 74 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-14T16:00:13.133298Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

74 of 74 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved74
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 18d113dc-df46-4858-aca8-9ee5694ca5bd · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Scanqa: 3d question answering for spatial scene understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:cdc0fa070913150d691ae81f8f24e5ae3a7523d2100ec736e81c8f83f8fa0dd2

Observation 6093b5e2-73ba-4d47-b323-4cff332e5ce9 · outbound

This paper cites Longformer: The Long-Document Transformer.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Longformer: The Long-Document Transformer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:149b31d05ca5721467129f6357acf782a176a3fd45bf2dbe9d8be40a8171d567

Observation a29dff35-013d-472d-a90c-cd5008ecd547 · outbound

This paper cites SAM 3: Segment Anything with Concepts.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models SAM 3: Segment Anything with Concepts

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:bf0ec572635a3b90dc81898b327a666c2e087487e9a9107e1fbf8dc2e5dc9b99

Observation abb6cf32-2618-412d-9753-8afbc5da791a · outbound

This paper cites PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:d4f0b0a5382372ecb3d2f940e96126fef157fbcd014cf90d2232d26b9d4dd898

Observation 8ce3de25-9e8d-4b6b-9482-dbec9ce1ef56 · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Scanrefer: 3d object localization in rgb-d scans using natural language

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:43d7ce6eba27555475d0645aaf87ddcb980fb9f77af20cca2cc437ba89826354

Observation dd593664-6fa5-4c1d-8a68-a5cce473f5eb · outbound

This paper cites Structure-aware transformer for graph representa- tion learning.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Structure-aware transformer for graph representa- tion learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:32063cb2f9081561bdc4e5609b2762c9cfeb81415a234f99cd97584d4bb2e642

Observation c78004cc-3f6d-4142-b167-eb5238b32b78 · outbound

This paper cites Language conditioned spatial relation reasoning for 3d object grounding.Advances in Neural Information Processing Systems, 35:20522–20535, 2022.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Language conditioned spatial relation reasoning for 3d object grounding.Advances in Neural Information Processing Systems, 35:20522–20535, 2022

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:ab7e77d84d0b9970cda7142f035eb8812cd232ca999260307796c52d85555e56

Observation fe8d2d96-fa5a-47ef-ad23-3c8e8a611cee · outbound

This paper cites LL3DA: Visual Interactive Instruction Tuning for Omni-3D Understanding, Reasoning, and Planning.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models LL3DA: Visual Interactive Instruction Tuning for Omni-3D Understanding, Reasoning, and Planning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:ac4bd658fa7b9f201701e89d0815fa1f0750767e538f07b2c51e7ee4bb757e6d

Observation 73cb8e9f-bd05-4e9c-8d28-1d4c82116ddc · outbound

This paper cites V ote2cap-detr++: Decoupling localization and describing for end-to-end 3d dense captioning.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models V ote2cap-detr++: Decoupling localization and describing for end-to-end 3d dense captioning.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:e61dda0606ede478125b3aa209b2ed5edbf9e7a7d41f2c46403ad7711db412c9

Observation 1da281cf-c681-460b-a0eb-f30c1554eb71 · outbound

This paper cites Grounded 3D-LLM with Referent Tokens.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Grounded 3D-LLM with Referent Tokens

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:c3abd7f92594249107674e72bb59573a5a871de8e44ff232e7504fd48694c0c1

Observation f38d90ad-c53b-487d-825e-04ec4db55ef9 · outbound

This paper cites Scan2cap: Context-aware dense captioning in rgb-d scans.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Scan2cap: Context-aware dense captioning in rgb-d scans

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:e27a5cc5da3db90f3e1593b97c3dde20c78ab2967a6e225d27b44298a36b6135

Observation d81dd9fd-3594-4495-a5a1-a2c6ec9a35af · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Generating Long Sequences with Sparse Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:a11ac0c4994f04e769fae88569128cdc69abc31e2fbbf8822fc8a5bb627e02c0

Observation 50d261a5-eb1c-4d9e-905a-9ae7cfe49868 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:226f3d25aec4da0912c5b8ed20e298509dc99b52a709c11d4cf177dcb6074414

Observation 9a347b4b-5fab-47e4-bef3-2f7b7e637d18 · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:045f452e6cf1c96b9ae6e444460187ad84f2c6860f7d7208c261ef9ee12d026e

Observation ff3275c8-3bb9-4a47-91d4-0795ff91d100 · outbound

This paper cites Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:b028ecf39ffb40490808c03160bcd4ac6d89a08ab83dca1c74adda77e2660f95

Observation 672134eb-02e5-47a5-b15c-0193ac4f1b06 · outbound

This paper cites Inductive representation learning on large graphs.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Inductive representation learning on large graphs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:281b5618695c5b7f4428912910163babbd790e14cff18473e3add8641a963fae

Observation 1d88bab1-4d95-41c7-a74d-ed4124779002 · outbound

This paper cites Language- grounded dynamic scene graphs for interactive object search with mobile manipulation.IEEE Robotics and Automation Letters, 9(10):8298–8305, 2024.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Language- grounded dynamic scene graphs for interactive object search with mobile manipulation.IEEE Robotics and Automation Letters, 9(10):8298–8305, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:9bb2d2f2ce4814e947c5fe4d40a08498929fc99338c21b53b2266a320ea53f77

Observation 85949127-6c34-428c-864c-8d78daaf9c7b · outbound

This paper cites 3D-LLM: Injecting the 3D World into Large Language Models.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models 3D-LLM: Injecting the 3D World into Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:070d41cbbe671d3d778f4928b29f8a63c35a81cbf44874a192b26f7b25d2abc9

Observation 85780207-6dc0-4fb8-9a2c-4f35d74597f0 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:6647bccbf307f1d3a890dde9fff91c4ab186a35560b680791b9821021d5f41b9

Observation 6d645eff-9322-418b-bf88-1cc0453a0ee0 · outbound

This paper cites Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:0a989b9160083432cd5dffc7f2034f32a12a163a6af5e711c4fd85f06b5b8439

Observation b4c7db12-90f0-472e-9256-59719d01bcaf · outbound

This paper cites Chat-scene: Bridging 3d scene and large language models with object identifiers.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Chat-scene: Bridging 3d scene and large language models with object identifiers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:1b6d088a44f0fb4fe3b27bc9c144fea483c30cb81631834153d5b3b09a1db1f0

Observation 81b3c3a5-3096-4ab4-8c28-68326279be89 · outbound

This paper cites An Embodied Generalist Agent in 3D World.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models An Embodied Generalist Agent in 3D World

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:55fb347fafbe3eecf1a89183381bb50fc6c3fdf57798ceb417815c454eee769e

Observation 11a95950-a76c-4cc4-8635-79aac8f0637c · outbound

This paper cites Hughes, Y.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Hughes, Y

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:6b080690d8a821fc9cc59f7be0fe67b107da7271e03e2c523928df3191ce53c7

Observation 231d215d-6829-421a-88bd-cd8a3560b065 · outbound

This paper cites Sceneverse: Scaling 3d vision-language learning for grounded scene understanding.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Sceneverse: Scaling 3d vision-language learning for grounded scene understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:a484c238584b06af9e71d56434ec0f54fe806ce0b3294587b8458fef7ff15e10

Observation 6a99e49b-69fd-4bfb-a92f-dc6f19e05656 · outbound

This paper cites Context-aware alignment and mutual masking for 3d-language pre-training.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Context-aware alignment and mutual masking for 3d-language pre-training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:2c6dd8a68ae6d3269c93b81d761719614cd68d0643658460730a2c0566cc7274

Observation 53eae517-2a04-4d87-84dc-ebf983974cde · outbound

This paper cites Robin3D: Improving 3D Large Language Model via Robust Instruction Tuning.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Robin3D: Improving 3D Large Language Model via Robust Instruction Tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:f9e3328a43f13b42b4aa3dbcd29d4b8c235a94917b88f12d1599569bf6d9d9b8

Observation 6c0b34c1-8a66-4b66-82ca-1b0416fd7baa · outbound

This paper cites Chang, and Manolis Savva.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Chang, and Manolis Savva

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:6ce377105eb59dbb94ee6a4204db6aa409096ee74163078017b6cc309e252699

Observation 92a261bf-a0aa-4bd0-af57-212de3b6fb8d · outbound

This paper cites LISA: Reasoning Segmentation via Large Language Model.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models LISA: Reasoning Segmentation via Large Language Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:0352e5df46f70aaab03a1435f46bc3de24566cd4c37b7a958d0287f647a638c2

Observation 1eeda62c-6ff5-44e1-871e-7666a122c453 · outbound

This paper cites Videochat: Chat-centric video understanding.Science China Information Sciences, 68(10): 200102, 2025.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Videochat: Chat-centric video understanding.Science China Information Sciences, 68(10): 200102, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:f5b3b290766e8ea85da577fdd5928bd24d87c517435585e199058e121ce8713d

Observation 2432d7fc-c235-4e75-b994-546c51ec4e3f · outbound

This paper cites Gradformer: Graph Transformer with Exponential Decay.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Gradformer: Graph Transformer with Exponential Decay

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:3afad4b7c4e7d300a88dccbf25ec8bccfbd56fc5dc5d01a5b5a9f3898139482b

Observation 1e7c65c4-845c-4fee-8682-829399bca933 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Improved Baselines with Visual Instruction Tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:724efd873eec4e337a43a9899c02931e01a1aca93b9ea5254dd8c3c2b401adc8

Observation c64a1162-17b2-4e73-bbdb-b13bcb9acf60 · outbound

This paper cites Visual Instruction Tuning.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Visual Instruction Tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:78ca1f0dfafa2099f6b582ce386d40d98bd64d2107f4d049c8a697d13003f50e

Observation ac8571d8-371f-47c4-a69a-c6022e24db79 · outbound

This paper cites an unresolved cited work.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:41a990a30b91577f564d0ab96f5c03202a7c1df9842ffd119e33962f3f808667

Observation 8039ebf8-a282-4033-8622-266947f928a5 · outbound

This paper cites Spatialpin: Enhancing spatial reasoning capabilities of vision-language models through prompting and interacting 3d priors.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Spatialpin: Enhancing spatial reasoning capabilities of vision-language models through prompting and interacting 3d priors

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:8c1ff9615a0db1c61f2d2f40fdc4a2c475aa0ea89b17566d49ef6f3e84619ac4

Observation 86527407-ee36-4d03-88c6-b7e2199b2697 · outbound

This paper cites Coopera: Continual open-ended human-robot assistance.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Coopera: Continual open-ended human-robot assistance

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:8e55d4098ecc197826b54f120cc39b9b55e74ed7cb3b8d493784d7c77e9b9b5c

Observation 5bbd0dd0-64e9-4099-8b3e-3dd7bbbe172b · outbound

This paper cites Cyclevla: Proactive self-correcting vision-language-action models via subtask backtracking and minimum bayes risk decoding.arXiv preprint arXiv:2601.02295, 2026.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Cyclevla: Proactive self-correcting vision-language-action models via subtask backtracking and minimum bayes risk decoding.arXiv preprint arXiv:2601.02295, 2026

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:958b3c787c55b9cb44962c39401130dfa6d035306d9690df92cb34321fbf1818

Observation 6d7833d2-f947-402c-8f2f-9c41c8434e18 · outbound

This paper cites SQA3D: Situated Question Answering in 3D Scenes.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models SQA3D: Situated Question Answering in 3D Scenes

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:910a9ac67c1d5d5a56b73813371388fdd2a94b85ce9f0c058280cd89e32503c4

Observation eab4be75-4893-4ee8-b8e5-8024a106ac54 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models DINOv2: Learning Robust Visual Features without Supervision

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:44c89581bcf5d8d93b93ba56ede2ba69cbe7aef576f3a51a59065af22178c78f

Observation ee8a8ff2-5bec-4c2e-ba9c-559777b0249d · outbound

This paper cites Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:144ccc1c8f1dfa5fdf7a61c70dc9c316589b8947d2e5291bb623d87d2ff2426e

Observation 7da5096e-7bb0-4516-ae76-90ecc597f730 · outbound

This paper cites Virtualhome: Simulating household activities via programs.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Virtualhome: Simulating household activities via programs

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:2fa366a1968ca4f2e1340fe3f69363d594ba305793ccd414c2222e7f2655e9c6

Observation 8b09774e-cfb5-4f95-963d-35b1ef020cd8 · outbound

This paper cites Watch-And-Help: A Challenge for Social Perception and Human-AI Collaboration.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Watch-And-Help: A Challenge for Social Perception and Human-AI Collaboration

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:c4aaac04b7b877ea318500579f73b5ed8d6b2d47207f66b94c8a380c71cf4558

Observation ef7cd8c7-a793-46c6-8dd9-b1459c24fdf4 · outbound

This paper cites Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:9d9175c7603faae51a0278d27a9cfa937cac570881532c88c49d7a51973dfb59

Observation fbb5f7de-c5d8-41ae-84f1-d333c7925267 · outbound

This paper cites GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:35f805fef32113655380f6f6cf243d43233317e3b433bdf856c6c1e16b98ae24

Observation 72b9cd3f-1cee-4d5a-943e-0e7a70f062ad · outbound

This paper cites Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:37a6a56032797791b216ad3aa1b2b505a87fa91ac87e824a8a369dbbff423f65

Observation 98e9f947-7003-48a4-a3df-5864d3af193b · outbound

This paper cites SayPlan: Grounding Large Language Models using 3D Scene Graphs for Scalable Robot Task Planning.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models SayPlan: Grounding Large Language Models using 3D Scene Graphs for Scalable Robot Task Planning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:e7420f76dc6995274fb83b2d48f02127359fa4c0324b2b97986e42c1f300ec9d

Observation 4dceaeec-536c-49bc-8cd8-193fad153315 · outbound

This paper cites Mask3d: Mask transformer for 3d semantic instance segmentation.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Mask3d: Mask transformer for 3d semantic instance segmentation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:7048189bc46b734b8ffa284a38a803c15a2ebc7c429105776c49971f0712a1ae

Observation 6bec8900-f821-4dc3-a3fd-9423f61bbea0 · outbound

This paper cites Qwen3 Technical Report.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Qwen3 Technical Report

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:9e06f43f3eb7658e1d79f5fbfd0bba348ea7fd9687aa7a98cf2b47f854e2d49f

Observation 90eb3ffe-4717-4df7-9f66-134393fe633d · outbound

This paper cites Pts3D-LLM: Studying the Impact of Token Structure for 3D Scene Understanding With Large Language Models.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Pts3D-LLM: Studying the Impact of Token Structure for 3D Scene Understanding With Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:f0d781f432ea5d6861f76cebd04bfe51c875a7392a6d40c45f18a9b64134504c

Observation 9da267cb-c9fd-4ddf-a9d6-fc519d3b5842 · outbound

This paper cites Vl-sat: Visual-linguistic semantics assisted training for 3d semantic scene graph prediction in point cloud.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Vl-sat: Visual-linguistic semantics assisted training for 3d semantic scene graph prediction in point cloud

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:bab2aeed9b27550ae8acc3787a7cf50c323644c1df93b3f7af6e2a6b8fac4e35

Observation 61be677b-120e-4c62-b730-24e8dde14bad · outbound

This paper cites Hier- archical open-vocabulary 3d scene graphs for language-grounded robot navigation.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Hier- archical open-vocabulary 3d scene graphs for language-grounded robot navigation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:da16e03199ae01d9c1de45eaa14c87b47fc5d45ba3abef8c0912a4aca7811736

Observation 80495fc6-c120-4d4b-b23b-ba311b4775aa · outbound

This paper cites CAMON: Cooperative Agents for Multi-Object Navigation with LLM-based Conversations.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models CAMON: Cooperative Agents for Multi-Object Navigation with LLM-based Conversations

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:0f2d2cee6595bc2bb4c66de96e5f66e5c15f0fdf93b3b981ba5d67b67eeccbf1

Observation c6a6da55-e19a-4953-9282-fefb9ee759fa · outbound

This paper cites 3ur-llm: An end-to-end multimodal large language model for 3d scene understanding.IEEE Transactions on Multimedia, 27: 2899–2911, 2025.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models 3ur-llm: An end-to-end multimodal large language model for 3d scene understanding.IEEE Transactions on Multimedia, 27: 2899–2911, 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:0f2266fb2afd5f8166123256b2c3cd7207e755fca3d28243212afdc1530c5fef

Observation 45d0eba0-a298-4438-825f-c87b9e76e99e · outbound

This paper cites Descrip3d: Enhancing large language model-based 3d scene understanding with object-level text descriptions.arXiv preprint arXiv:2507.14555, 2025.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Descrip3d: Enhancing large language model-based 3d scene understanding with object-level text descriptions.arXiv preprint arXiv:2507.14555, 2025

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:24108207f3b485452e9617be02df0c29c3086c5a629172a9bf18860285ef8b16

Observation f0d30aa6-cc46-4dfc-8d64-6b5af58ca6c1 · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember and Recall Spaces.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Thinking in Space: How Multimodal Large Language Models See, Remember and Recall Spaces

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:c29601cb69aae113a27ec3cdaa046ab86db3eb8c410173504ae8028ba5508f78

Observation c6545e75-d4e4-44af-b350-3270f63cf087 · outbound

This paper cites HomeRobot: Open-Vocabulary Mobile Manipulation.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models HomeRobot: Open-Vocabulary Mobile Manipulation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:c97c8379ea9916fdc17c6682c048da23f02657a4c9e4136eff06cbec35330601

Observation 31703ae3-c50c-4801-b584-a2955fc490a8 · outbound

This paper cites Do transformers really perform badly for graph representation?Advances in neural information processing systems, 34:28877–28888, 2021.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Do transformers really perform badly for graph representation?Advances in neural information processing systems, 34:28877–28888, 2021

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:6f5070e0868ee7a998b55abacefbee8d88c985d3ff8e6635539ef980bff384b5

Observation c58a59f4-c85e-4ab0-a1ae-37115813a222 · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:e695bf9a08c7279b49de9ac70f7e3a06dc27a0646eff2530190722de03621dab

Observation d14963bf-45f9-4ee6-b63f-a2a5bf184e64 · outbound

This paper cites Inst3d-lmm: Instance-aware 3d scene understanding with multi-modal instruction tuning.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Inst3d-lmm: Instance-aware 3d scene understanding with multi-modal instruction tuning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:524e586d264685bde8a6d524d14509bb1b85c215e9d7a2afaeb84a10b5af804c

Observation ae538b21-02db-4513-a29b-d949a67e8a23 · outbound

This paper cites Big bird: Transformers for longer sequences.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Big bird: Transformers for longer sequences

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:b3daa16f35dc38cfb2e5ffe6a85a07a4a79cfe4d30561386b3286a8a4e607c8d

Observation e1b60801-1547-4a74-9c20-7ce8b66a2a9a · outbound

This paper cites 3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene Understanding.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models 3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene Understanding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:f25ac5b6ad3a682719b2197caf9672fb548bff95f51396bb3016598e0b38c5a4

Observation 43830ec1-18cc-4a42-8487-d91eaab667d2 · outbound

This paper cites Multi3drefer: Grounding text description to multiple 3d objects.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Multi3drefer: Grounding text description to multiple 3d objects

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:81940a6b0dc351d5077265f9dc5cfb44f4160c52123413c007b8af6b21ace381

Observation 275ac01e-0370-46f5-8b5d-3a42a63baeae · outbound

This paper cites 3dvg-transformer: Relation modeling for visual grounding on point clouds.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models 3dvg-transformer: Relation modeling for visual grounding on point clouds

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:785d8fe9e05fa07fa70d41d554b165054075f25e7408992fe6a0283ea13fa063

Observation 40876b77-b838-49e0-8447-3d26b5d35922 · outbound

This paper cites BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:2cb8a405f3891aa52e73aba463ebe54834024a920e304346b8d0750e797198cb

Observation 1c0e8e9a-66ba-4de5-99df-f09942d2d45f · outbound

This paper cites Lscenellm: Enhancing large 3d scene understanding using adaptive visual preferences.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Lscenellm: Enhancing large 3d scene understanding using adaptive visual preferences

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:4c89a4129747fc61c8fc9ced77fdca9f0ba23f4e841bb102ec8c83a71ec81e8c

Observation 16b844de-e247-4a59-8dd6-893f6907134c · outbound

This paper cites Uni3D: Exploring Unified 3D Representation at Scale.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Uni3D: Exploring Unified 3D Representation at Scale

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:40de082b10e9b8fa0eb065651123421cba6753070359f2bb3b1b61bd7545648b

Observation 246857a5-26a0-4628-8f81-3ffbafb0b847 · outbound

This paper cites 3d-vista: Pre-trained transformer for 3d vision and text alignment.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models 3d-vista: Pre-trained transformer for 3d vision and text alignment

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:8317da19fd787673bdbdb13d8822ed1995127ef97f9b262738c313f2bd2fbb52

Observation 3e550046-13ab-4117-a8a3-a83962f5ac53 · outbound

This paper cites Unifying 3d vision-language understanding via promptable queries.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Unifying 3d vision-language understanding via promptable queries

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:4b3462c7c3fd5abc75bc62f78fa4f1a4c34cf2060d396547549b655116da23d1

Observation a7554b1a-c3ad-448d-b301-5a7916f53a12 · outbound

This paper cites the tablebetweenthe sofaandthe bookshelf,next tothe lamp.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models the tablebetweenthe sofaandthe bookshelf,next tothe lamp

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:e92f94c91385adde10e0fa21d84d233fffd7e7d8b572619397eebfc7760c4e19

Observation 55b82348-49fa-424b-bdc0-1622f54c577a · outbound

This paper cites the chairnext tothe deskthat supports the monitornearthe window.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models the chairnext tothe deskthat supports the monitornearthe window

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:5f95a2873062532b370f0efdcc5b289846a512e699836e0a1f29c14ae5b5771d

Observation 78de21ff-1c22-44d1-a730-f60cecad06a8 · outbound

This paper cites the lampabovethe nightstand, where the nightstandis next tothe bedthat facesthe window.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models the lampabovethe nightstand, where the nightstandis next tothe bedthat facesthe window

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:c0615a7387d9101bbf4ec302f9f3a730bc6dc779336fc6a28c4df76604a2cd33

Observation e3774508-7c74-4d8a-94da-840c12956ed8 · outbound

This paper cites Q:” and ending with “?.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Q:” and ending with “?

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:b5b8450c0a4352fd1c70d4b46c25d5a3bdcaca34a1cd16bb39bfbaeb9537db03

Observation 0cbcf9e3-3703-4a02-89bd-e6366a03efbc · outbound

This paper cites what” or “which object.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models what” or “which object

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:c55f129c253932baccf6560c18d82c533409af770e8de79813e6cb2d627f2918

Observation 696ec95e-af61-488d-b30a-1e7a7a98b28e · outbound

This paper cites an unresolved cited work.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:c0c485e308c7ff0b6498a46673a986d3df88a9cd2a98ca594dbef479df47dbad

Observation a8e61a81-3cdf-479d-a34b-0cf784834890 · outbound

This paper cites In {room}, how many {class} are there?.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models In {room}, how many {class} are there?

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:c38019b4063902c1dab39ec984793df4b00eb9624b85ffe7af3ec2ea0f8b3f09

Pith citing papers

No inbound Pith citation observations are available.