Pith. sign in

Paper Citation Record · LEDGER

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations

As of 18 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 7 inbound Pith citation observations for arXiv:2506.17545.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17545 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:34:54.732094Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:00:00.845845Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T21:16:14.446824Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ed8e7168-1445-4b8f-b31a-47ef08ff9f0a · outbound

This paper cites Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:35:01.552537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:34:50.085085Z digest=sha256:9bcc7671f4bbf019db0e85d5e5b5c3e6a55e82f78464ccc90f96fb27eea51520

Observation 9226bca7-4524-49b8-8091-469a56ba4367 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Scanqa: 3d question answering for spatial scene understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:50.187260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:50.187260Z digest=sha256:9dcc4bcfee77b4da99f75067268448cbd103a158d648e518bbda5fb71df70e47

Observation 4ad1c548-4f4c-4ed3-a81e-26e7e230146d · outbound

This paper cites Qwen2.5-VL Technical Report.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:50.293925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:50.293925Z digest=sha256:592205082db8f403202c04308db3d40b9331265b0de027978ccbe4f3368a65d4

Observation 3061aad6-af57-490d-8a9f-02616bb75303 · outbound

This paper cites ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:50.424528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:50.424528Z digest=sha256:16f65fa1e69777d9fe6a8ca77c8bca16a4469b4fa9613623d4d1474036ed1000

Observation a866976b-438a-45c9-a23b-79a47a285988 · outbound

This paper cites Do as i can, not as i say: Grounding language in robotic affordances.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Do as i can, not as i say: Grounding language in robotic affordances

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:50.502146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:50.502146Z digest=sha256:7c2ec8a225757d0ff79bed84dec76d643354dd80cc61fcc17c11e94662545cda

Observation 8886deb5-0959-48ae-b47a-f5854c681229 · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Scanrefer: 3d object localization in rgb-d scans using natural language

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:50.590561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:50.590561Z digest=sha256:fc43734583eca58d9388f30b3deb33857bacb670cd73c62f77d161b55005fc0e

Observation 962879af-d356-4703-98e2-e3a1c2052ec4 · outbound

This paper cites R1-v: Reinforcing super generalization ability in vision-language models with less than $3.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations R1-v: Reinforcing super generalization ability in vision-language models with less than $3

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:50.674777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:50.674777Z digest=sha256:d4a5f3faf780a8f5bb6642b5fb9300917d6aeebc942227d665f65c4201c9ba8b

Observation 279ff78f-65b3-41ff-b81e-d483adeb9d9a · outbound

This paper cites Language conditioned spatial relation reasoning for 3d object grounding.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Language conditioned spatial relation reasoning for 3d object grounding

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:35:01.194834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:34:50.786284Z digest=sha256:b8cf1fa1b23dba0a058750638cfee74b3ddc9dbcfea977c9b181a28381a4a406

Observation d0ab54f3-d442-4292-b066-44c289b2dfec · outbound

This paper cites End-to-end 3d dense captioning with vote2cap-detr.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations End-to-end 3d dense captioning with vote2cap-detr

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:35:00.957226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:34:50.864776Z digest=sha256:7b002f118b26a78987fa3026eb389049cddd4c5c18fa8245cb7ea46a674bda33

Observation 285d4ae2-ca5a-45aa-8e85-414c5362e4b1 · outbound

This paper cites V ote2cap-detr++: Decoupling localization and describing for end-to-end 3d dense captioning.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations V ote2cap-detr++: Decoupling localization and describing for end-to-end 3d dense captioning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:35:00.706023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:34:51.009696Z digest=sha256:e2f10755fc6ff8844c45e1e08485a711810f75a6042bcee9f936e3f7c119f021

Observation 28a4d20a-9564-41cb-88a2-23ecabea99b9 · outbound

This paper cites Scan2cap: Context-aware dense captioning in rgb-d scans.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Scan2cap: Context-aware dense captioning in rgb-d scans

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:35:00.402504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:34:51.112310Z digest=sha256:5670cfda6277f3570117751447a21a3f63f2fd41af059be6089fbb097490eb69

Observation 28484630-a5d9-4b57-a6a9-8d67a508401c · outbound

This paper cites Functionality understanding and segmentation in 3d scenes.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Functionality understanding and segmentation in 3d scenes

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:35:00.222262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:34:51.198878Z digest=sha256:5a5d9a46f70bdaa75885b2e17cfc5f8c7ceb13256f351e98cc339cde634efce9

Observation 7e5f8c23-6c1c-48b2-b91e-11f77dea73ad · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:59.955079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:34:51.302373Z digest=sha256:7eaaa9cdfe59592c797ca6507e7dcf54a79941cc6ad9ab4b19d6e96e9734577e

Observation b253733e-2e3f-400e-8a32-56f2b7b5b772 · outbound

This paper cites Scenefun3d: fine-grained functionality and affordance understanding in 3d scenes.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Scenefun3d: fine-grained functionality and affordance understanding in 3d scenes

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:59.666719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:34:51.373750Z digest=sha256:faf4c7b51eeae2f111969facfecdf34ae02b042464b83a40809b6758f6657323

Observation 7acd6dc5-915f-4e15-b313-13cd14a70da8 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:51.432004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:51.432004Z digest=sha256:d05cf063e99aa64df7d31409e2955e628348fa3bffeb0d7e88c7f9b623a85abc

Observation d14b9ef8-5e7f-4d64-b888-4815955d8d59 · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:51.467823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:51.467823Z digest=sha256:9ccdc38f5435197bbd10760a681ac249e032403c59548180f63e39b0238d110e

Observation 7f54d23d-d774-4890-8f84-618e83889bcc · outbound

This paper cites The ecological approach to visual perception: classic edition.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations The ecological approach to visual perception: classic edition

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:51.538593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:51.538593Z digest=sha256:8bc8db8ebcad00809635d817f339c338504e3e9daffe16546186c21fc89f37b1

Observation f0353bd0-68cd-4d0e-b8aa-ee09a817f793 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:51.580536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:51.580536Z digest=sha256:931fb6d5450df55e039b6607c76b505e15a71068e63ce2f7c1f228c8a4a3fbef

Observation 005ced98-3c95-48b2-9856-325d9f985566 · outbound

This paper cites Viewrefer: Grasp the multi-view knowledge for 3d visual grounding.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Viewrefer: Grasp the multi-view knowledge for 3d visual grounding

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:59.393195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:34:51.630120Z digest=sha256:62b120a9fe93e4dbe8d4dbf65af7b6a987c3e50245e844555efffd2df320fa74

Observation b9d52e38-0094-444f-aa1f-65c4f97ce632 · outbound

This paper cites Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:51.681819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:51.681819Z digest=sha256:e7745a0199ca8fdc7d7d098f9cff77191a6ba69789698455f8d76e6626394752

Observation 14668d63-bd0c-49fa-be77-36c52ef52330 · outbound

This paper cites Chat-scene: Bridging 3d scene and large language models with object identifiers.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Chat-scene: Bridging 3d scene and large language models with object identifiers

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:59.186224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:34:51.737382Z digest=sha256:7b94d88f51a984555f1b20fce4a635aafa71031f37462a4eb1a59e0153a83832

Observation 8c1c1e4c-3e24-4bd9-b9eb-e8941c7b684e · outbound

This paper cites An Embodied Generalist Agent in 3D World.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations An Embodied Generalist Agent in 3D World

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:51.827382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:51.827382Z digest=sha256:7b38fe1b08a80a17ad1b1bfd2c6f53f4c85e0e6e46be6cf509eac18bc1c1f758

Observation 353c30ed-8014-4f78-8788-fbacb505715b · outbound

This paper cites Text-guided graph neural networks for referring 3d instance segmentation.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Text-guided graph neural networks for referring 3d instance segmentation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:58.944257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:34:51.886696Z digest=sha256:61d4e25d29def84da95a11a95dd6078f732ed18e8e760f494d71944b9fb1367d

Observation fbc57c2c-1005-4aa4-8900-2cdb4a8f0385 · outbound

This paper cites Multi-view transformer for 3d visual grounding.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Multi-view transformer for 3d visual grounding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:58.705701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:34:51.957621Z digest=sha256:84ffc9709797a81bd5352a0bbf0be50ec2147358b918be8f415e45c8cc59d7bf

Observation 1e59bc04-336c-47d1-8f1d-34864794b21f · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:52.047794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:52.047794Z digest=sha256:4dae83d361aa0d4e4a0e1b1c188e8016fc1e4231ce46e4cb4ab068457326e379

Observation 4c7ebe5b-73ec-46c8-b30a-254cba943b27 · outbound

This paper cites Bottom up top down detection transformers for language grounding in images and point clouds.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Bottom up top down detection transformers for language grounding in images and point clouds

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:58.471374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:34:52.121287Z digest=sha256:63cf8feafe7148651321eb3e4e3181fe498a0eee23fad35e422252ff7ef6e368

Observation 8b4f9729-5c84-4207-bc3a-ddc83b247789 · outbound

This paper cites Lerf: Language embedded radiance fields.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Lerf: Language embedded radiance fields

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:52.196735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:52.196735Z digest=sha256:8715e7a35166e08ae19e5cb0ae071f1c0861c1ec25117f358d21a00889efddf5

Observation 5ff9199b-effe-478f-8564-e0160dc4a7d3 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations VideoChat: Chat-Centric Video Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:52.248369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:52.248369Z digest=sha256:a9ceb23d33885c6d55596a78f6bb262070205d895f298f0224d00c02e71bfdf6

Observation e234c58c-8045-4310-aedb-429afb0f45c3 · outbound

This paper cites SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:52.340466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:52.340466Z digest=sha256:c62cfd346e7638a22078c5af0da236336f5fc088e84d93b1dc8688ebc18d617d

Observation 5228ce6a-ba5c-47d6-9e27-6a5406917599 · outbound

This paper cites Sqa3d: Situated question answering in 3d scenes.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Sqa3d: Situated question answering in 3d scenes

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:52.404835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:52.404835Z digest=sha256:3770704d4cae5fe4b75ecbb0b80e13c6a815d2a493f33fe8d2cbae703a3c6ed4

Observation 7ecf0d1f-e033-4411-999d-0edd71d31830 · outbound

This paper cites Introducing openai o1.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Introducing openai o1

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:58.064915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:34:52.525748Z digest=sha256:416481be16d0fd243a8447610790eb9500c23851b8f4d4d943f8ff0014237379

Observation 1ce32298-8785-4299-8788-d676de5a47d2 · outbound

This paper cites Openscene: 3d scene understanding with open vocabularies.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Openscene: 3d scene understanding with open vocabularies

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:57.587518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:34:52.613220Z digest=sha256:f3d8c63acc9963b3a276648efea4b59d5ae8d061a6c04be3abc2be20b75bf0ce

Observation e405d8fc-4ca3-4d86-bc3f-00c087ee3ed9 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Learning transferable visual models from natural language supervision

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:57.324756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:34:52.768733Z digest=sha256:0cc5f08c5d2f3282541252d566c41939b11ea72a4c5e5b76ea31a088e13ad7a7

Observation 69fe2126-b29c-4af0-95a6-dd0306ff9b4e · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations SAM 2: Segment Anything in Images and Videos

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:53.004776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:53.004776Z digest=sha256:8de7f13c633444b0f2ece981b7e1e038343b3e54398df116895ad0f78b046e3a

Observation c3fa0c00-fe7f-4026-abb8-a8b4ab348688 · outbound

This paper cites High-resolution image synthesis with latent diffusion models, 2021.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations High-resolution image synthesis with latent diffusion models, 2021

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:53.075943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:53.075943Z digest=sha256:33fa31e49984e4b3879a9e4f45856f38e2e6114b89e236b03fc494dc3335383e

Observation 82a2a810-cd69-4659-bd7e-d28eb2c4c698 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Proximal Policy Optimization Algorithms

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:53.147630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:53.147630Z digest=sha256:51dced3aba3772a4f61433ae4988b672f03b75ebba8432575f59097a595d40af

Observation 4bd2fa6a-d9fb-41b5-a809-f01b28e6e8d0 · outbound

This paper cites Mask3d: Mask transformer for 3d semantic instance segmentation.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Mask3d: Mask transformer for 3d semantic instance segmentation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:57.078831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:34:53.210278Z digest=sha256:4b590b5dcf76d9a005ba8a2df1234032d650e9d8cf007409e900e95511cd8bd6

Observation 45df17a8-1db0-49b3-bf58-476513600938 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:53.304750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:53.304750Z digest=sha256:2643a8cfff4b061537f6e7807a6b107c609a0402886d58d11d05f44398526c49

Observation b2563195-10da-4789-9fb6-4383212f9d6a · outbound

This paper cites Chatgpt for robotics: Design principles and model abilities.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Chatgpt for robotics: Design principles and model abilities

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:56.868146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:34:53.418030Z digest=sha256:30a10354a23703dd844160cd1cea0cb71abba1c11819ad9bbbf8317f334de41b

Observation 67ca3aea-21d9-42f0-82e8-0d874e7f886f · outbound

This paper cites Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:53.529692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:53.529692Z digest=sha256:787ddd09736ba882551ed8874c36e010bce7ee7a4537f0904f06e1da702d6585

Observation 78530adf-24b4-4652-b6b5-b232e77c5410 · outbound

This paper cites Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:53.624919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:53.624919Z digest=sha256:039fc50ecd555b4a3a2a071f1e303dd08fce05bff308abc1f43b3c14152d1d85

Observation 6bf266ad-34ce-4b37-a064-63a11da1e28c · outbound

This paper cites Distilling coarse-to-fine semantic matching knowledge for weakly supervised 3d visual grounding.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Distilling coarse-to-fine semantic matching knowledge for weakly supervised 3d visual grounding

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:56.650264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:34:53.702757Z digest=sha256:61af7463fa6739c0959e8f1ea258cb0cce3fc6a8d5160816075438672b30af84

Observation f3770be4-4d4c-46c4-ab61-fa2d0697575d · outbound

This paper cites Pointllm: Empower- ing large language models to understand point clouds.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Pointllm: Empower- ing large language models to understand point clouds

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:56.416286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:34:53.845949Z digest=sha256:1eb17bca0dbb760a34a4e82ee6bf80f3fba37c5a2c56e9a9100d9f5da66684a9

Observation 05da5af7-46a8-4a2d-af28-13db6388021c · outbound

This paper cites Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual referring.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual referring

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:56.111190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:34:53.974587Z digest=sha256:4c25560784d7df5b85ccc8d060852f2a4f77a3b3c6c94d6fc07fc992872b4fe5

Observation 7a2cf311-ede6-42a4-a5a4-ce6053d4b62d · outbound

This paper cites Visual programming for zero-shot open-vocabulary 3d visual grounding.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Visual programming for zero-shot open-vocabulary 3d visual grounding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:54.128620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:54.128620Z digest=sha256:bca2df4bfb69b5bf1c91a9e303f62ad66ab877e3638ae69c4e18603a5727c7b8

Observation 1a39f9cf-49bf-44fc-bc8c-6d39f64c550d · outbound

This paper cites Empowering large language models with 3d situation awareness.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Empowering large language models with 3d situation awareness

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:55.935976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:34:54.272841Z digest=sha256:6de53a43345ee89d3bf606e0b098c90aed06fe27ade6613804ef975b07fe6824

Observation 494884c1-1a1c-49cf-a2f6-49713e22119b · outbound

This paper cites 3dvg-transformer: Relation modeling for visual grounding on point clouds.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations 3dvg-transformer: Relation modeling for visual grounding on point clouds

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:55.760801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:34:54.384753Z digest=sha256:14d7bc5d09f93149c683cd707ecb8d20466cf53d0b2136a29f37b51c05fc113d

Observation 3df51210-40c0-4bee-ae88-a9c4dfca88e0 · outbound

This paper cites BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:54.466311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:54.466311Z digest=sha256:ebe7259c60de1bc3e65931343652de02b6e945627361e70a03cfcc334c88532e

Observation ca3eb5f5-d067-4075-b7e1-c7514176b17b · outbound

This paper cites Uni3D: Exploring Unified 3D Representation at Scale.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Uni3D: Exploring Unified 3D Representation at Scale

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:54.555746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:54.555746Z digest=sha256:91037214df272eaf5f5d36314374db61960c69bda94d9a65ab700b2b61561b4b

Observation 64d34c28-84ba-4862-84ad-17a14bb5ecd8 · outbound

This paper cites 3d-vista: Pre-trained transformer for 3d vision and text alignment.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations 3d-vista: Pre-trained transformer for 3d vision and text alignment

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:55.621691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:34:54.649230Z digest=sha256:2870a207dbed0058fe379095322259e27bc075e1f790e92283b96042487bb53b

Observation ad927a9d-0c9f-4e92-80ff-8f16ffaec360 · outbound

This paper cites [EVENT]" in the video, determine the precise time period of the occurrence of the object. Provide the start and end times (in seconds, precise to one decimal place) in the format.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations [EVENT]" in the video, determine the precise time period of the occurrence of the object. Provide the start and end times (in seconds, precise to one decimal place) in the format

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:55.426439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:34:54.732094Z digest=sha256:b74d59247b66cc48b12d97e1206ee8923760c67ecf015513ecfe9b8be826f343

Pith citing papers

Observation a03740dc-991f-46e0-b652-48ad2b8c165e · inbound

3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding cites this paper.

3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T10:49:47.113836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:49:47.113836Z digest=sha256:9055cd6884b83d0e957c0e87e820157d66118c441c5582d3532367a24c8cf904

Observation dc1d0c90-7740-4d73-9f6f-43a60c62986c · inbound

What if? Emulative Simulation with World Models for Situated Reasoning cites this paper.

What if? Emulative Simulation with World Models for Situated Reasoning Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations

Reference 115

Resolution
unresolved
no resolver link, observed 2026-07-15T13:51:30.008232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:51:30.008232Z digest=sha256:058cb4de503dcb522f28c7274a446ba41587400e3592338d17d7c940c5f29f89

Observation b6c9fc21-e333-4f3e-a41d-595aa2598a86 · inbound

Token Warping Helps MLLMs Look from Nearby Viewpoints cites this paper.

Token Warping Helps MLLMs Look from Nearby Viewpoints Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations

Reference 117

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:08:17.348876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T21:07:55.062113Z digest=sha256:287771b630c46ebe5e1c3cafadab610fce983891d7cffe3e1c6dff9f303c7298

Observation f454b41e-fca4-4ae5-8299-3dba8be8881e · inbound

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models cites this paper.

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:28:04.506945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-20T05:27:30.938311Z digest=sha256:efc0eb82af89968c1233f51ad73b467e44152ae9592bc0f26e4bf0c90829b0fe

Observation 8b7f3447-d5a3-469b-a3dd-700edaba72ba · inbound

Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMs cites this paper.

Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMs Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:16:14.448373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T17:14:48.328093Z digest=sha256:180a1cd9bfda7964407576eea60d52c37f7f77a6524d199e4eff9e8cfcdd5269

Observation 5e2d5bd9-7531-4e87-b024-96a063370336 · inbound

Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMs cites this paper.

Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMs Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:34:38.162717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T11:26:47.680406Z digest=sha256:da03cdb0043f064899e3bd826cf816cdce560ce8094fe8df54469999cd6d8f5e

Observation 445c9bd0-66c0-4450-b783-e35f4a0881a3 · inbound

ThinkAfford: Affordance-Centric Reasoning for Fine-Grained 3D Grounding in Cluttered Scenes cites this paper.

ThinkAfford: Affordance-Centric Reasoning for Fine-Grained 3D Grounding in Cluttered Scenes Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T13:00:00.845845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:00:00.845845Z digest=sha256:4c53f1d8e331e4d1361aaf5b534761994c66e9697f74ba31a048c30d3449ddb2