Pith. sign in

Paper Citation Record · LEDGER

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations

As of 9 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 6 inbound Pith citation observations for arXiv:2506.17545.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17545 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:34:54.732094Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T10:49:47.113836Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T21:16:14.446824Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ed8e7168-1445-4b8f-b31a-47ef08ff9f0a · outbound

This paper cites Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:35:01.552537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:50.085085Z digest=sha256:bbaebe7830954c9ac2693396c295b4c7bde73e8fde32e210756191473819a2bc

Observation 9226bca7-4524-49b8-8091-469a56ba4367 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Scanqa: 3d question answering for spatial scene understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:50.187260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:50.187260Z digest=sha256:b60a8a553c57ee24e96d69bd004782511b243710c1c72563db6bcf2a3ddba60c

Observation 4ad1c548-4f4c-4ed3-a81e-26e7e230146d · outbound

This paper cites Qwen2.5-VL Technical Report.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:50.293925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:50.293925Z digest=sha256:7d96e01ae515cb301757d86339290873f88ec99a49ef37a465a8f5075d992bae

Observation 3061aad6-af57-490d-8a9f-02616bb75303 · outbound

This paper cites ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:50.424528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:50.424528Z digest=sha256:bc195d4a148ea84161cd0cce818cdbb84a1955c1ca12886ce03690fc0ef86d29

Observation a866976b-438a-45c9-a23b-79a47a285988 · outbound

This paper cites Do as i can, not as i say: Grounding language in robotic affordances.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Do as i can, not as i say: Grounding language in robotic affordances

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:50.502146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:50.502146Z digest=sha256:262ff8e3c6911835e8c7b45761caad293c32120e01e0d06859cd8cfa8d6449d0

Observation 8886deb5-0959-48ae-b47a-f5854c681229 · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Scanrefer: 3d object localization in rgb-d scans using natural language

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:50.590561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:50.590561Z digest=sha256:2a53102ac2510d9cc7a533b479d9398e385128ec3470699db59d30b662208611

Observation 962879af-d356-4703-98e2-e3a1c2052ec4 · outbound

This paper cites R1-v: Reinforcing super generalization ability in vision-language models with less than $3.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations R1-v: Reinforcing super generalization ability in vision-language models with less than $3

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:50.674777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:50.674777Z digest=sha256:81c7b2ab3c423dad0c5ac03eb1017f12274d17d261cc9d01497f8688f5c5d6c1

Observation 279ff78f-65b3-41ff-b81e-d483adeb9d9a · outbound

This paper cites Language conditioned spatial relation reasoning for 3d object grounding.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Language conditioned spatial relation reasoning for 3d object grounding

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:35:01.194834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:50.786284Z digest=sha256:a8e364c5d1489830a411b1c40b51dce52e00345ef40cdf4a1faadf0748917a64

Observation d0ab54f3-d442-4292-b066-44c289b2dfec · outbound

This paper cites End-to-end 3d dense captioning with vote2cap-detr.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations End-to-end 3d dense captioning with vote2cap-detr

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:35:00.957226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:50.864776Z digest=sha256:a80c1b02e2f1a2c985b703c48760a8e749b761b9a4f9626210f4b306c3258854

Observation 285d4ae2-ca5a-45aa-8e85-414c5362e4b1 · outbound

This paper cites V ote2cap-detr++: Decoupling localization and describing for end-to-end 3d dense captioning.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations V ote2cap-detr++: Decoupling localization and describing for end-to-end 3d dense captioning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:35:00.706023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:51.009696Z digest=sha256:4db3edd40dd077b38523f21abce4bd7a2e7b510d7f14ce0770003da8d5294822

Observation 28a4d20a-9564-41cb-88a2-23ecabea99b9 · outbound

This paper cites Scan2cap: Context-aware dense captioning in rgb-d scans.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Scan2cap: Context-aware dense captioning in rgb-d scans

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:35:00.402504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:51.112310Z digest=sha256:ed1f82f164d4aa012d9eada905491d532ca7fbc993a95073ce16b8e6c8201f16

Observation 28484630-a5d9-4b57-a6a9-8d67a508401c · outbound

This paper cites Functionality understanding and segmentation in 3d scenes.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Functionality understanding and segmentation in 3d scenes

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:35:00.222262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:51.198878Z digest=sha256:197e398bf45d54475817a3f22458ad7ba9875856ae2e0f5dbeeab9fa087b39b4

Observation 7e5f8c23-6c1c-48b2-b91e-11f77dea73ad · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:59.955079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:51.302373Z digest=sha256:ffb2480d1dff8b6761b53f14ed92fbb50a669dcfc636011f3743b41dc4e49d15

Observation b253733e-2e3f-400e-8a32-56f2b7b5b772 · outbound

This paper cites Scenefun3d: fine-grained functionality and affordance understanding in 3d scenes.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Scenefun3d: fine-grained functionality and affordance understanding in 3d scenes

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:59.666719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:51.373750Z digest=sha256:66fc2b25cf7d3178a3a9ff0c1013bf2613549b252baf605b76a223e83fc62c83

Observation 7acd6dc5-915f-4e15-b313-13cd14a70da8 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:51.432004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:51.432004Z digest=sha256:fb6b2bab3cf020afe0e242bff30d2eb1cdf10865e38ef922c940f2b8f5ce5281

Observation d14b9ef8-5e7f-4d64-b888-4815955d8d59 · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:51.467823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:51.467823Z digest=sha256:c49a547876114233d86cc6cc0e1ac2fea110043dd433f48c2c384e2622512e3f

Observation 7f54d23d-d774-4890-8f84-618e83889bcc · outbound

This paper cites The ecological approach to visual perception: classic edition.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations The ecological approach to visual perception: classic edition

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:51.538593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:51.538593Z digest=sha256:fd4f4cc5695ecbea171b529806b3bf35aceba1c72f6c7b7b20f72c66097972bf

Observation f0353bd0-68cd-4d0e-b8aa-ee09a817f793 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:51.580536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:51.580536Z digest=sha256:a90d31ce4047d5a0963b857b2db9f6f468a7328447e2d5217dc29fb709bdf52f

Observation 005ced98-3c95-48b2-9856-325d9f985566 · outbound

This paper cites Viewrefer: Grasp the multi-view knowledge for 3d visual grounding.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Viewrefer: Grasp the multi-view knowledge for 3d visual grounding

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:59.393195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:51.630120Z digest=sha256:dfdb925b0f1bfcd9f61bdceec4639d3c06b0b8beea982a7abb8d543481a2d68b

Observation b9d52e38-0094-444f-aa1f-65c4f97ce632 · outbound

This paper cites Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:51.681819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:51.681819Z digest=sha256:02a0a143f52ce1e0b89dcd5c676b122fc629202c77652b240673b842a0143fef

Observation 14668d63-bd0c-49fa-be77-36c52ef52330 · outbound

This paper cites Chat-scene: Bridging 3d scene and large language models with object identifiers.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Chat-scene: Bridging 3d scene and large language models with object identifiers

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:59.186224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:51.737382Z digest=sha256:2903bcbb244c7ba6b73c12ed311339a78f065ab66e5c66799d6ffde88c280437

Observation 8c1c1e4c-3e24-4bd9-b9eb-e8941c7b684e · outbound

This paper cites An Embodied Generalist Agent in 3D World.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations An Embodied Generalist Agent in 3D World

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:51.827382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:51.827382Z digest=sha256:747ee1db5ed7df34e62702c645adf4ac08b8fe5b9a304a83e055f09acc5f35aa

Observation 353c30ed-8014-4f78-8788-fbacb505715b · outbound

This paper cites Text-guided graph neural networks for referring 3d instance segmentation.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Text-guided graph neural networks for referring 3d instance segmentation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:58.944257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:51.886696Z digest=sha256:373aca557b699e86f52c3b6472d49511adea43f4acb89a06c6762af57425faab

Observation fbc57c2c-1005-4aa4-8900-2cdb4a8f0385 · outbound

This paper cites Multi-view transformer for 3d visual grounding.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Multi-view transformer for 3d visual grounding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:58.705701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:51.957621Z digest=sha256:425f3352820cf1e2062eadeac35bdc0c1e77a0a9deb3d7ec36928b34c9c7b8ab

Observation 1e59bc04-336c-47d1-8f1d-34864794b21f · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:52.047794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:52.047794Z digest=sha256:2e8f5b9214d5bba54782f0473d9112cdf01383d6de14fb9ed2d5d74af26d9afc

Observation 4c7ebe5b-73ec-46c8-b30a-254cba943b27 · outbound

This paper cites Bottom up top down detection transformers for language grounding in images and point clouds.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Bottom up top down detection transformers for language grounding in images and point clouds

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:58.471374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:52.121287Z digest=sha256:fe17a84bd49bfc3150be87fb9a079dbc906e4ad10fc808fbe36268bb0c15cbb6

Observation 8b4f9729-5c84-4207-bc3a-ddc83b247789 · outbound

This paper cites Lerf: Language embedded radiance fields.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Lerf: Language embedded radiance fields

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:52.196735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:52.196735Z digest=sha256:824ed0899ec9f116630a540feec6bcb9b503d3cdb4870c99f3001817268ce4b4

Observation 5ff9199b-effe-478f-8564-e0160dc4a7d3 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations VideoChat: Chat-Centric Video Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:52.248369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:52.248369Z digest=sha256:689fe528533d05ba53cff98f21c3f49a676a3b9b33633fc3969fa04c82e023dc

Observation e234c58c-8045-4310-aedb-429afb0f45c3 · outbound

This paper cites SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:52.340466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:52.340466Z digest=sha256:6759231222717c1cd57568cf2e0ab779818f1ab79273774807f1a438b6dc3613

Observation 5228ce6a-ba5c-47d6-9e27-6a5406917599 · outbound

This paper cites Sqa3d: Situated question answering in 3d scenes.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Sqa3d: Situated question answering in 3d scenes

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:52.404835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:52.404835Z digest=sha256:725980f81716f0debfef73d1012401a06577cb4f0250949604db83031b4f1f90

Observation 7ecf0d1f-e033-4411-999d-0edd71d31830 · outbound

This paper cites Introducing openai o1.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Introducing openai o1

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:58.064915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:52.525748Z digest=sha256:97edae312720eb7065357e405645d363a06f3c6b026baa655c2474c6f3ae39f8

Observation 1ce32298-8785-4299-8788-d676de5a47d2 · outbound

This paper cites Openscene: 3d scene understanding with open vocabularies.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Openscene: 3d scene understanding with open vocabularies

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:57.587518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:52.613220Z digest=sha256:c06fe4b267af77ea7505cad221af90c84644b68e61ae338ee4ced063be1f8384

Observation e405d8fc-4ca3-4d86-bc3f-00c087ee3ed9 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Learning transferable visual models from natural language supervision

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:57.324756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:52.768733Z digest=sha256:01dd245de94c0c131b4427a0df3c118815b52e884416401ad862825116a72a26

Observation 69fe2126-b29c-4af0-95a6-dd0306ff9b4e · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations SAM 2: Segment Anything in Images and Videos

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:53.004776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:53.004776Z digest=sha256:b0f4a83c05f3b73abe521269504c28c19dba20c40b0680274e3b9a3d3ad4d907

Observation c3fa0c00-fe7f-4026-abb8-a8b4ab348688 · outbound

This paper cites High-resolution image synthesis with latent diffusion models, 2021.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations High-resolution image synthesis with latent diffusion models, 2021

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:53.075943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:53.075943Z digest=sha256:1df5347d3de7cb6ab2a79a2b118315f3ad8e4d5ccfda8590cf65599ff9780d73

Observation 82a2a810-cd69-4659-bd7e-d28eb2c4c698 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Proximal Policy Optimization Algorithms

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:53.147630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:53.147630Z digest=sha256:23bf16173c0279a66d96a5b4b95876bd83dccc8f3df92cd4a6e3da81a6f17fa1

Observation 4bd2fa6a-d9fb-41b5-a809-f01b28e6e8d0 · outbound

This paper cites Mask3d: Mask transformer for 3d semantic instance segmentation.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Mask3d: Mask transformer for 3d semantic instance segmentation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:57.078831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:53.210278Z digest=sha256:5ad632a81164fe8a87c90fed01689ff301b745eb688b696c6d80cbed7dd07a86

Observation 45df17a8-1db0-49b3-bf58-476513600938 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:53.304750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:53.304750Z digest=sha256:4dde0b354ff132ec7f160d7f97b938ad68ccbcd26db7eee6a92b055fa783b9c5

Observation b2563195-10da-4789-9fb6-4383212f9d6a · outbound

This paper cites Chatgpt for robotics: Design principles and model abilities.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Chatgpt for robotics: Design principles and model abilities

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:56.868146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:53.418030Z digest=sha256:87ef19ba851958eade7e8b2ac30491f66ac0bbd9471ebebe5825bfb89fe2ee67

Observation 67ca3aea-21d9-42f0-82e8-0d874e7f886f · outbound

This paper cites Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:53.529692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:53.529692Z digest=sha256:d1b265931a300d81ec3c45ba5716f53de84c8bbc03cde326c76ba946ea1c11e5

Observation 78530adf-24b4-4652-b6b5-b232e77c5410 · outbound

This paper cites Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:53.624919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:53.624919Z digest=sha256:da84b6cadd4697edf215f7604466f989899879314cc9aa4369b520159146c0c0

Observation 6bf266ad-34ce-4b37-a064-63a11da1e28c · outbound

This paper cites Distilling coarse-to-fine semantic matching knowledge for weakly supervised 3d visual grounding.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Distilling coarse-to-fine semantic matching knowledge for weakly supervised 3d visual grounding

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:56.650264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:53.702757Z digest=sha256:20cc89f32cfbaf351b230ddf660614da9fce02c5eb9feb2d211ac48eb8e934b9

Observation f3770be4-4d4c-46c4-ab61-fa2d0697575d · outbound

This paper cites Pointllm: Empower- ing large language models to understand point clouds.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Pointllm: Empower- ing large language models to understand point clouds

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:56.416286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:53.845949Z digest=sha256:4e11ffe9fff6b42f7ab05d6de30e62de4e76d0552a7441cbb5bcf71e0618af4b

Observation 05da5af7-46a8-4a2d-af28-13db6388021c · outbound

This paper cites Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual referring.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual referring

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:56.111190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:53.974587Z digest=sha256:a9bcb838df11be5c991a12adeab55502c03c6bcd6ed56bcb824d9430592d16cd

Observation 7a2cf311-ede6-42a4-a5a4-ce6053d4b62d · outbound

This paper cites Visual programming for zero-shot open-vocabulary 3d visual grounding.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Visual programming for zero-shot open-vocabulary 3d visual grounding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:54.128620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:54.128620Z digest=sha256:00986ac0d5e5a243e33b1bc45a8d0b1be9b7a7397d504c5120f3b14b4e6a6ede

Observation 1a39f9cf-49bf-44fc-bc8c-6d39f64c550d · outbound

This paper cites Empowering large language models with 3d situation awareness.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Empowering large language models with 3d situation awareness

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:55.935976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:54.272841Z digest=sha256:e516bb4d5cc2f1514dab9e1eb88c17a69f1d6cb3278f6ac343824789601f82b0

Observation 494884c1-1a1c-49cf-a2f6-49713e22119b · outbound

This paper cites 3dvg-transformer: Relation modeling for visual grounding on point clouds.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations 3dvg-transformer: Relation modeling for visual grounding on point clouds

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:55.760801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:54.384753Z digest=sha256:cb0d4f695f8d57aad4b77c25f519d2c0b7ed0966e2b819868935192e786b0ee9

Observation 3df51210-40c0-4bee-ae88-a9c4dfca88e0 · outbound

This paper cites BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:54.466311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:54.466311Z digest=sha256:d36e3749ccc2d5c7e76bc75b28f625caf9aa840ee930fb3bd715fcf8e1073aeb

Observation ca3eb5f5-d067-4075-b7e1-c7514176b17b · outbound

This paper cites Uni3D: Exploring Unified 3D Representation at Scale.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Uni3D: Exploring Unified 3D Representation at Scale

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:54.555746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:54.555746Z digest=sha256:049ba80f926f3c40ebb2d9e8a88225a2b4d8cd81cf421a4e25c6820c1022bbe9

Observation 64d34c28-84ba-4862-84ad-17a14bb5ecd8 · outbound

This paper cites 3d-vista: Pre-trained transformer for 3d vision and text alignment.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations 3d-vista: Pre-trained transformer for 3d vision and text alignment

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:55.621691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:54.649230Z digest=sha256:a97aaabd2e5f6306189bb8ba6a9316be5cdd651c9b750e045a62800a9f16d19f

Observation ad927a9d-0c9f-4e92-80ff-8f16ffaec360 · outbound

This paper cites [EVENT]" in the video, determine the precise time period of the occurrence of the object. Provide the start and end times (in seconds, precise to one decimal place) in the format.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations [EVENT]" in the video, determine the precise time period of the occurrence of the object. Provide the start and end times (in seconds, precise to one decimal place) in the format

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:55.426439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:54.732094Z digest=sha256:0c2025e506ab60dcb9233bdbf3aed2db9c2fd2da1420528e59ac7a05f1d2c27e

Pith citing papers

Observation a03740dc-991f-46e0-b652-48ad2b8c165e · inbound

3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding cites this paper.

3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T10:49:47.113836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:49:47.113836Z digest=sha256:8d55da5ba73757f19579b0523ef37d44a8314979003556b9dab134049be1dc19

Observation dc1d0c90-7740-4d73-9f6f-43a60c62986c · inbound

What if? Emulative Simulation with World Models for Situated Reasoning cites this paper.

What if? Emulative Simulation with World Models for Situated Reasoning Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations

Reference 115

Resolution
unresolved
no resolver link, observed 2026-07-15T13:51:30.008232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:51:30.008232Z digest=sha256:f590e6824f05b7950bd2ce0f4d4151b52676810de6e1ad61486d008b67d2c5ab

Observation b6c9fc21-e333-4f3e-a41d-595aa2598a86 · inbound

Token Warping Helps MLLMs Look from Nearby Viewpoints cites this paper.

Token Warping Helps MLLMs Look from Nearby Viewpoints Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations

Reference 117

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:08:17.348876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T21:07:55.062113Z digest=sha256:34f5b7ef5b1e2721e2204acba67f2ac457e1dbb686d165f13c9ff8049d4dbc87

Observation f454b41e-fca4-4ae5-8299-3dba8be8881e · inbound

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models cites this paper.

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:28:04.506945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T05:27:30.938311Z digest=sha256:d20755117c27a86f5a4fec5b399fe42bf6e5fefee1b166f8c1fa0269adba2c5e

Observation 8b7f3447-d5a3-469b-a3dd-700edaba72ba · inbound

Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMs cites this paper.

Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMs Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:16:14.448373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T17:14:48.328093Z digest=sha256:9714c9a04fca5fe44139e10c66e40c5071c60e64d21e8ba742b1fb53b92d9b98

Observation 5e2d5bd9-7531-4e87-b024-96a063370336 · inbound

Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMs cites this paper.

Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMs Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:34:38.162717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T11:26:47.680406Z digest=sha256:66e4658c0564bd82269eda06c5de4d6e100f420ee6a300199eb90ae4f7c65077