Pith. sign in

Paper Citation Record · LEDGER

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences

As of 17 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 7 inbound Pith citation observations for arXiv:2412.01292.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01292 v2

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:35:37.029873Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:22:03.376214Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T10:49:47.240788Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3fe63a2a-e12e-4997-9012-8109844ba71e · outbound

This paper cites Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:38.070144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T04:35:36.758229Z digest=sha256:642c0675216be48a6908a139f92e1c61309f59ad3f2bcd14342f3eecfa7a438e

Observation cce39ea2-b811-48b5-afb0-58f06112dc11 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Scanqa: 3d question answering for spatial scene understanding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:38.053855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T04:35:36.764332Z digest=sha256:5c978720c138c5276857c38598636d236ab266df0dc705bd7b48aeea00d65aca

Observation 89625eb6-3a0a-47af-b56b-0c07ee467f36 · outbound

This paper cites Qwen Technical Report.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.770448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.770448Z digest=sha256:f4be473c340768d7d151ec68d2afd4b75bea1e590e2c5d5664bbf70439fc4eb0

Observation 0f011e68-62db-40b9-9a65-35802a41da18 · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:38.034213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T04:35:36.775993Z digest=sha256:e11c117cc17db83388180291949375e11f2c7bbb2a9722ae3842b1d4a63f8ab0

Observation 8314fb60-969c-494a-8404-e6af4f4ace65 · outbound

This paper cites ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.781754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.781754Z digest=sha256:57e87680478dbf81d9cc15b69283613b3f86334609e3b969e3f49a95037cb17d

Observation 766b527c-8945-45cc-85b4-3158c4abfe4d · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Scanrefer: 3d object localization in rgb-d scans using natural language

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.787399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.787399Z digest=sha256:15d629506e81f511590076322fff620d7d3bd5975dd604f77d58cd276fc0680b

Observation 99481ef7-81e6-43aa-8a80-6e604f803d84 · outbound

This paper cites $A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences $A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.793861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.793861Z digest=sha256:89f76b2b9cb0bc39b34facd6d201f6302e6a816f306319fc16e71315738d4be7

Observation 0ed7d7af-6a74-46ef-a9c0-42411919e95f · outbound

This paper cites Ll3da: Visual interactive instruction tuning for omni-3d understand- ing reasoning and planning.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Ll3da: Visual interactive instruction tuning for omni-3d understand- ing reasoning and planning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:38.007324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T04:35:36.799342Z digest=sha256:f7a53385aa927d23c9285861a88147ddb2d3dda81cd21cbd688ae2d000c59ba0

Observation 55b59e00-9332-4f1b-85b8-fea3868d9847 · outbound

This paper cites End-to-end 3d dense captioning with vote2cap- detr.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences End-to-end 3d dense captioning with vote2cap- detr

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.988183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T04:35:36.805160Z digest=sha256:123522a0ad1d9892479f8dcce2f17248766423b2337768be077c26bdf31cf2a8

Observation 0084ae39-f534-471a-b858-6b40fbd4d9d1 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.810807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.810807Z digest=sha256:5c534a952554b897619660bd42b2962b2c5cc98ed2eadfc7b57a910ab0024c60

Observation 95bbf4d7-0bfb-4bee-9130-8cbbc9b3daf7 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.816633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.816633Z digest=sha256:917efefb7d1ed00e97217ed0d6ba4b616574ff3815c52dac8fd1d45c8444f4f6

Observation b7b8680e-eed0-4471-94fc-ec625400368c · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.823134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.823134Z digest=sha256:d8011d386f8d9db780bf8c918cbd6149be1c33cea7d603320d06e23d2f1a2816

Observation 831ae781-da79-4ceb-86bc-b97f7354b054 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.830027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.830027Z digest=sha256:248e0fddf53e38c3a5c1520eb364a6be0882705f632e8cd2dcc05d1817453fc8

Observation a4448104-fa27-41da-bf9f-d48299e08c4d · outbound

This paper cites Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.836304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.836304Z digest=sha256:68ea1da05f9041e8ba4991f33a2bb97a94545ee02aa87886a01528a66952d8c8

Observation 9598c877-e40e-4dca-a05f-1c8b005f0efc · outbound

This paper cites 3d-llm: Injecting the 3d world into large language models.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences 3d-llm: Injecting the 3d world into large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.844087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.844087Z digest=sha256:f7088ae6d457de2f1b400237fbb0eb2af3d1a85221226c98a4e165793345ec33

Observation 1a5b9862-ebd9-4560-9457-c18768df34dd · outbound

This paper cites Chat-scene: Bridging 3d scene and large language models with object identifiers, 2024.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Chat-scene: Bridging 3d scene and large language models with object identifiers, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.933127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T04:35:36.852883Z digest=sha256:4cc8e74ff9573eca196a49376a3fc75311ffedbe2496ebdb3b863eb7dcb22790

Observation 27f7e11d-5cba-4ba7-9b8b-aa69f5a921d3 · outbound

This paper cites Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.858454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.858454Z digest=sha256:9aeb88274ac6d7f45c2dc4b031a6db436f189b53ba49e596990eeb535792b5dd

Observation b8e95ab9-4675-4a26-8643-a5b8ab002919 · outbound

This paper cites An Embodied Generalist Agent in 3D World.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences An Embodied Generalist Agent in 3D World

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.865168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.865168Z digest=sha256:f2a57620766028c1bc53c600d896e732ed82a81a83c9d86575ce11ae24831520

Observation dc713f0b-4911-4c64-bd7e-847baf796a85 · outbound

This paper cites VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.871188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.871188Z digest=sha256:65f81d2737cb175ce4d74f6ad33aa5d88c93f52771dc6a3816cc932036be2989

Observation 16476d8f-cf5d-460d-8cbd-6c8c3c118703 · outbound

This paper cites SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.876736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.876736Z digest=sha256:a7691d368a96d0252eb17e7aad1d56a701990860cd68f13f71e1f972094ced5f

Observation ac1239c6-4c5f-47d1-a80f-5c66a1dc739c · outbound

This paper cites Context-aware alignment and mutual masking for 3d-language pre-training.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Context-aware alignment and mutual masking for 3d-language pre-training

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.916311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T04:35:36.882736Z digest=sha256:3ef3a1baef78bbbf5a404c0f69fed4e6477df1733cd50d27e992baed34152975

Observation 91ac9420-3e8a-4900-b510-e51b59bc384a · outbound

This paper cites Beyond the nav-graph: Vision-and-language navigation in continuous environments.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Beyond the nav-graph: Vision-and-language navigation in continuous environments

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.897796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T04:35:36.887832Z digest=sha256:16b64f968f44e61578054e8a89f386d781673898d979e6be276f771dcd7c841f

Observation f90ef089-cb8b-405a-a583-f9d18e022e63 · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.894102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.894102Z digest=sha256:94ddbc12688b8017f4416ed9b1bffd82f8ec356db4a48ab4c12109395576c018

Observation 28576c4a-b2b0-42e9-9c68-fe7aef66df65 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Rouge: A package for automatic evaluation of summaries

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.900360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.900360Z digest=sha256:dcd9cb86c749d7f8ab24a46db01e88c8612019cedfacbfaf79fbab0a11e93f56

Observation e4249674-4ce6-4727-a50f-667fa6c79017 · outbound

This paper cites Visual instruction tuning.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Visual instruction tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.905491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.905491Z digest=sha256:ff42e1be404428898dbd60e3c44bf8a208366ae9288c47f2ae55cf5a7798daa2

Observation 7b1d5d94-7ef6-46d9-b3b1-145ecc5d3287 · outbound

This paper cites Decoupled Weight Decay Regularization.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Decoupled Weight Decay Regularization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.911294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.911294Z digest=sha256:a28dd5a9ab2ff62ed2b85d036c2058fba97905e88f6390d7890923bf7ce77465

Observation 03f47964-be3c-47ca-a0ed-e811e786e0ab · outbound

This paper cites Openscene: 3d scene understanding with open vocabularies.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Openscene: 3d scene understanding with open vocabularies

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.828298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T04:35:36.916675Z digest=sha256:ffca58a4dd60c4bbb80820a8b2cecc614829c9d1b3cf7010c73c4bccdd292c9b

Observation aa0ba61a-c73e-45dd-8cca-f1cafe004031 · outbound

This paper cites Pointnet++: Deep hierarchical feature learning on point sets in a metric space.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Pointnet++: Deep hierarchical feature learning on point sets in a metric space

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.921651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.921651Z digest=sha256:6990f1b1bdfb568d9be1167e38db89d2f2f2b5703197c20623fa4e4636d85d04

Observation 2cf8fe14-0928-4d78-bda2-90b5bf9784aa · outbound

This paper cites Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.926180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.926180Z digest=sha256:2fbdb5ec479afd6f0f50737d2a849a852181355624af023f6de2e13e2007d7e4

Observation d32fbdfe-714f-49aa-8c0e-604f83201fc7 · outbound

This paper cites Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.931272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.931272Z digest=sha256:448648e436bcf9d581ab3a314d6bd9f8983653fd26a90bdd418340538ef5b4a8

Observation 0b67ca78-061a-4d9c-bb91-f968e9452ec5 · outbound

This paper cites Eye movements in iconic visual search.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Eye movements in iconic visual search

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.781151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T04:35:36.937120Z digest=sha256:0a752005e562f4bf5494c10a869fc1001243e128f95f7a5219f96c80c3bd9365

Observation 4a93a63b-c80b-4b5e-9944-eeee7bd933a5 · outbound

This paper cites Mask3d: Mask transformer for 3d semantic instance segmentation.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Mask3d: Mask transformer for 3d semantic instance segmentation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.942835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.942835Z digest=sha256:4a4e4ddba6cc9c1b1ed7dbd042b0594e96e1fdb94afa99e6343816a54cdf13b0

Observation 49ba5686-d19b-4bfd-9813-167bdebcc65a · outbound

This paper cites Indoor scene segmen- tation using a structured light sensor.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Indoor scene segmen- tation using a structured light sensor

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.743530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T04:35:36.947620Z digest=sha256:7a42e8de7c2309a2ce66838188a82b0e6e51d990b946fd5e05d42a6c2296bcdd

Observation 5eec47bb-6fef-436a-9893-525d27d01bd3 · outbound

This paper cites Sun rgb-d: A rgb-d scene understanding benchmark suite.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Sun rgb-d: A rgb-d scene understanding benchmark suite

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.953258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.953258Z digest=sha256:f0812b746007ddba53b8d0c605f74af29ae30cbfd87f1f34568ac65ea82effe0

Observation e3ae5547-944b-4072-b5ab-31fefc9fe72a · outbound

This paper cites Fgprompt: fine-grained goal prompting for image-goal navigation.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Fgprompt: fine-grained goal prompting for image-goal navigation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.704457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T04:35:36.958878Z digest=sha256:fef1b265acb073abbc8c8af14fb7de923bf7f26ee19009838a39d3c3970b5251

Observation 2abc01b9-451d-48fa-9b09-c1b2d4935959 · outbound

This paper cites MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D Priors.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D Priors

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.964112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.964112Z digest=sha256:0e9c6faf5944147e524ee14bbcfe1df6b9c2e620f49042922454d6316f172ee7

Observation d38280e9-3764-4b09-9095-8f114c10ff3e · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences LLaMA: Open and Efficient Foundation Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.969395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.969395Z digest=sha256:456a7eb1748077e4a32291d714fca4cac165ca2aa6c81dcf1000ba7321c09f7b

Observation 0fd286fb-d35a-4bb2-86ca-7f27d0c4447f · outbound

This paper cites A feature-integration theory of attention.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences A feature-integration theory of attention

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.684125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T04:35:36.974290Z digest=sha256:987fdd08a0bf066adbf775cebf01d3e1909b5200067d5f4b39695efac7dc4227

Observation 44949439-52ab-4c0d-89a1-1e7d694004e8 · outbound

This paper cites Cider: Consensus-based image description evalu- ation.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Cider: Consensus-based image description evalu- ation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.666338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T04:35:36.979864Z digest=sha256:d79cce10416c622c891784f5ad668c3d3e40eea7b9bab25d5d6b0278f8d34325

Observation 6070df36-5bfd-4529-b8ac-c1d978427a70 · outbound

This paper cites Rio: 3d object instance re-localization in changing indoor environments.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Rio: 3d object instance re-localization in changing indoor environments

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.649299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T04:35:36.985024Z digest=sha256:f22fcaf43019b7e387dcc897ca26d61d826b13a5c7d91825ce788d9cb2c41bc9

Observation 470c7e11-cbf2-4bdb-b133-de04097cc0c4 · outbound

This paper cites Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.991429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.991429Z digest=sha256:699f424a9ea6dcf3db4f78c5e313a5e3f0e39413ec3216d8de097c3c9ce67a67

Observation 4a82ddc5-e9f0-4d88-8fdc-c8425cae31b4 · outbound

This paper cites OccLLaMA: An Occupancy-Language-Action Generative World Model for Autonomous Driving.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences OccLLaMA: An Occupancy-Language-Action Generative World Model for Autonomous Driving

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.997620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.997620Z digest=sha256:40d9afb1032c8cb73de77d0be75f92c0ce280d8f0923c9ee7ee1290ac62b41ae

Observation bb40708d-d3b5-4e09-aa2b-cb2dba6761f4 · outbound

This paper cites What attributes guide the deployment of visual attention and how do they do it? Nature reviews neuroscience, 5(6):495–501, 2004.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences What attributes guide the deployment of visual attention and how do they do it? Nature reviews neuroscience, 5(6):495–501, 2004

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.629469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T04:35:37.004468Z digest=sha256:b89ffd1c4d56e5b86039e7b3cfdc70644bd3d6cefcd541a41d126fc92a01455a

Observation 6e6ffa06-16de-4985-b0fd-5779beaf94e6 · outbound

This paper cites PointLLM: Empowering Large Language Models to Understand Point Clouds.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences PointLLM: Empowering Large Language Models to Understand Point Clouds

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:37.009769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:37.009769Z digest=sha256:c8cb87280e19fb175e2add5cc4007f6b35f799c6ca022a31cf08abc25c76ff38

Observation 8aca1262-4a18-4809-adfd-c15946878935 · outbound

This paper cites LiDAR-LLM: Exploring the Potential of Large Language Models for 3D LiDAR Understanding.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences LiDAR-LLM: Exploring the Potential of Large Language Models for 3D LiDAR Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:37.014890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:37.014890Z digest=sha256:4a6756acf7d22ffe66754e0c29a306bbe84afb0b15fc5f385f1d396e34e7352d

Observation 6ba662bc-0477-4522-9029-b31ae0f3edb1 · outbound

This paper cites DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:37.020259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:37.020259Z digest=sha256:5469a74ad3250764b3bb0e58b3ce70d0a20a061f57ba41063dcfac6e4e6ecf22

Observation 11dde869-e4c9-4999-a23d-7b200624bf83 · outbound

This paper cites Uni3d: A unified baseline for multi-dataset 3d object detection.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences Uni3d: A unified baseline for multi-dataset 3d object detection

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:35:37.606717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T04:35:37.025291Z digest=sha256:dfcabacce99f41060b78884c5d6078da7152f1d5ce9b62ee421693f6a2d68deb

Observation 83f99ab1-5a8a-4afa-a952-dbb517c0ee36 · outbound

This paper cites 3D-VLA: A 3D Vision-Language-Action Generative World Model.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:37.029873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:37.029873Z digest=sha256:186ad33a193d19bb02b308e036658f884ce849a2fdd78ecdab8f66806c95ad83

Pith citing papers

Observation 7e8b58bb-e232-48c0-9cc4-1cc51e9b19ae · inbound

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models cites this paper.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:03.376214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:22:03.376214Z digest=sha256:b3faddd11a5d54f3a2f4f6947ac51c2e5c371f2089df4e4b08ea0317058686d6

Observation 9faf74b7-9151-42b8-a42e-61024e8edca4 · inbound

LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding cites this paper.

LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:46.765433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:46.765433Z digest=sha256:f3845daa3589baf4cebd53f32c93cb694b01b7902e9a1ef4f95d67136c84fdb9

Observation 91e291e6-f4df-43a5-88c0-2cb0ff3dd88f · inbound

A Large-Scale Referring Remote Sensing Image Segmentation Dataset and Benchmark cites this paper.

A Large-Scale Referring Remote Sensing Image Segmentation Dataset and Benchmark LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:01:49.176786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:01:49.176786Z digest=sha256:e84e715e771209cc56b349de0e51610aa13ad789cf23ddc802cab3fcce2f05bf

Observation 537570de-39ae-44c7-b1c1-77720a093e87 · inbound

Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs cites this paper.

Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T10:25:48.086502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:25:48.086502Z digest=sha256:ef5000fdc299044a04065a21040559a812872c6c6ea9058a87abcfc8fc09a532

Observation 5587a349-bc45-4921-9800-e501f3c2a3f7 · inbound

GaussianVLM: Scene-centric 3D Vision-Language Models using Language-aligned Gaussian Splats for Embodied Reasoning and Beyond cites this paper.

GaussianVLM: Scene-centric 3D Vision-Language Models using Language-aligned Gaussian Splats for Embodied Reasoning and Beyond LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T21:11:41.554989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:11:41.554989Z digest=sha256:d56731673b33200908d174b65239847c1275ed05624fbd22c9bde7782a8f80c8

Observation 966be097-444d-475d-91d4-6aabacd26c1e · inbound

3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding cites this paper.

3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-08-06T10:49:47.249706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T10:49:47.135209Z digest=sha256:bb0ed2835ac54ae72c2b3317a043284fa839ad1667353a70941c6c4f9228ed53

Observation c02b4ead-fde1-4b93-92b0-417ce4d8e435 · inbound

Nav-R1: Reasoning and Navigation in Embodied Scenes cites this paper.

Nav-R1: Reasoning and Navigation in Embodied Scenes LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T17:31:43.089996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:31:43.089996Z digest=sha256:b8f31bd668b78a86c2c54abcc952f93f2c36a0bd4de92fd32c2456aa0801129c