Pith. sign in

Paper Citation Record · LEDGER

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition

As of 8 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2507.11261.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.11261 v2

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:18:14.922481Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy46
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 99a5221d-5176-43af-abae-2f73a0eb4af0 · outbound

This paper cites Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:23.475431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:10.932060Z digest=sha256:86b4a853c6c57b0909fabf5c12008183f6d5a168bb841890aeff481743756ad8

Observation f975c64a-485f-4d78-9f8d-dc42fbe8e48a · outbound

This paper cites Cot3dref: Chain-of-thoughts data-efficient 3d visual grounding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Cot3dref: Chain-of-thoughts data-efficient 3d visual grounding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:23.222477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:11.018378Z digest=sha256:8d016c66f526171557fe16d5a5a6adca0b995d744140f96235dd4c64f3f2d663

Observation f4e61f26-db46-490f-aba8-0eb1df005f13 · outbound

This paper cites Visual question answering from another perspective: Clevr mental rotation tests.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Visual question answering from another perspective: Clevr mental rotation tests

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:22.964487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:11.096372Z digest=sha256:05f71a81a6a77cdf8c4daf70801db620f7badce461bc552b22dd0b8c0740fdcd

Observation d4dc29a6-0669-4bdd-976d-eb467977906c · outbound

This paper cites Assertiveness-based agent communica- tion for a personalized medicine on medical imaging diag- nosis.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Assertiveness-based agent communica- tion for a personalized medicine on medical imaging diag- nosis

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:22.729024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:11.162571Z digest=sha256:2d96d031de9f1d5e26da3c7dec6f4ba2176aac2c993c5a44aaa0f3262aa828b0

Observation a17f72e6-33c7-48fc-a038-94af86b7efcf · outbound

This paper cites Mikasa: Multi-key-anchor & scene-aware transformer for 3d visual grounding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Mikasa: Multi-key-anchor & scene-aware transformer for 3d visual grounding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:22.578516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:11.232514Z digest=sha256:a8fdb3d008fe5f2858ebaef043a5185c4c3b80eb3a084686d68d2f0c00abaa05

Observation 526e06e0-2a1c-4bec-b757-a07af60d434d · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Scanrefer: 3d object localization in rgb-d scans using natural language

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:22.419009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:11.272693Z digest=sha256:b17444424703dd83773c2eac4909d95caaf045c33ff0b8920e9ecc31cdba98fc

Observation 59de1d01-e895-4fa6-baea-6d5f6edca71a · outbound

This paper cites SCJD: Sparse Correlation and Joint Distillation for Efficient 3D Human Pose Estimation.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition SCJD: Sparse Correlation and Joint Distillation for Efficient 3D Human Pose Estimation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:11.315194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:11.315194Z digest=sha256:501097c4be844802cbc66f14fb568487bc292f1f7b428eaa51b6ed2387956ec7

Observation eccb3eba-6be2-470e-bbeb-bf59790be25d · outbound

This paper cites Talk2bev: Language-enhanced bird’s- eye view maps for autonomous driving.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Talk2bev: Language-enhanced bird’s- eye view maps for autonomous driving

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:22.228420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:11.372304Z digest=sha256:c303875cb0f70e710f8edf2b54a338797526b0b3453d7c5bb3ca247d82b1265a

Observation 1dc3fd23-afd9-4983-abd1-963c1e72f9a5 · outbound

This paper cites Drive as you speak: Enabling human-like interac- tion with large language models in autonomous vehicles.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Drive as you speak: Enabling human-like interac- tion with large language models in autonomous vehicles

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:22.060228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:11.459556Z digest=sha256:05651bd41f761fb8d3441c34226aa4c103f043d80f93b9bcc30b5796b70a0a8d

Observation 1c2babb7-5748-4946-a129-a39f62ca7b64 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:11.526902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:11.526902Z digest=sha256:18c32d318c3a5b03b5e1b73a30959ca8d63e5be67b5a42f594733058a47338e4

Observation 0ffb4a29-cc63-41ed-8328-6490f340850e · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:11.558820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:11.558820Z digest=sha256:7d2a8eeb05525c4f67a6be37dadad702a401beb1dda088766226857b8f621d64

Observation ade96116-986d-4715-878b-0a4040aa684d · outbound

This paper cites Scenegenie: Scene graph guided diffusion models for image synthesis.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Scenegenie: Scene graph guided diffusion models for image synthesis

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:21.859773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:11.609396Z digest=sha256:e91376f75e259014dffce8cc93741f3a69a7dd762ac172bcb7122ced1e0e2cca

Observation 94dd748e-1e5b-4db3-abdf-d3114273a179 · outbound

This paper cites Dense reinforce- ment learning for safety validation of autonomous vehicles.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Dense reinforce- ment learning for safety validation of autonomous vehicles

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:21.686087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:11.725233Z digest=sha256:d3f870f3abe50f7adb3ff2160bcf24eb85d7a1a344130b46ca0a782df41191c7

Observation af1fb14e-7516-443a-8274-73fcc2fbc43c · outbound

This paper cites Viewinfer3d: 3d visual ground- ing based on embodied viewpoint inference.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Viewinfer3d: 3d visual ground- ing based on embodied viewpoint inference

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:21.525736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:11.814980Z digest=sha256:6e3be3b46139fb86e6c33a6aac1582bada2abcfc36f849490a03f6f5ff0ff477

Observation e05dd4c4-9cf5-4bd0-8f03-d88fee1038db · outbound

This paper cites Viewrefer: Grasp the multi-view knowledge for 3d visual grounding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Viewrefer: Grasp the multi-view knowledge for 3d visual grounding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:21.382853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:11.879963Z digest=sha256:c6ad4424b21e929cb95d90645c0b7c3a045838cab58445b3572750fa21fa9ef9

Observation 10802696-30ec-4432-802d-8d1ad2a2a2ba · outbound

This paper cites Transrefer3d: Entity-and- relation aware transformer for fine-grained 3d visual ground- ing.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Transrefer3d: Entity-and- relation aware transformer for fine-grained 3d visual ground- ing

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:21.222781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:11.960630Z digest=sha256:1dafe776ef1af81a0fe4877c007ba7daee5f12c649577073ad9317af40615180

Observation 871a281a-822f-48d2-b87e-f017dd6f71da · outbound

This paper cites 3d-llm: In- jecting the 3d world into large language models.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition 3d-llm: In- jecting the 3d world into large language models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:12.036499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:12.036499Z digest=sha256:9f0f6ead71fdeb1a7816d0b31f122569d86c88ee84b446080dfe865915bb13d8

Observation be0bd385-f4ef-4432-8a79-37595d71b9e9 · outbound

This paper cites Multi- view transformer for 3d visual grounding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Multi- view transformer for 3d visual grounding

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:21.108849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:12.169841Z digest=sha256:37837999c6c0827e7aa2a0b0bef4c59c8b19b279f7d112c94b3cf0101e86742c

Observation fe61d437-8966-45a3-b4d8-9ffb6c95ed08 · outbound

This paper cites Structure-clip: Towards scene graph knowledge to enhance multi-modal structured repre- sentations.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Structure-clip: Towards scene graph knowledge to enhance multi-modal structured repre- sentations

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:20.956537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:12.226913Z digest=sha256:f6d34a020928f34141114b27b3117170d06426d3acc851ce6d4620f2984555da

Observation 1447d34a-0364-45f7-9836-3487acfcbaec · outbound

This paper cites Nan-detr: noising multi-anchor makes detr better for object detection.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Nan-detr: noising multi-anchor makes detr better for object detection

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:20.759454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:12.302592Z digest=sha256:fb9e65d42087f3205d1e9ab5820a59f945b0fffeb97f2187acd3435ec641a203

Observation c7234e13-59d2-4e77-a0e8-722b6bfeb5ee · outbound

This paper cites Bottom up top down detection transform- ers for language grounding in images and point clouds.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Bottom up top down detection transform- ers for language grounding in images and point clouds

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:20.616504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:12.337407Z digest=sha256:bcae07ba7c3aec96e2b357bd577c23e7d336b7f895bb0341a1b11b35980b6109

Observation 20737413-972d-41be-8b38-9cf6e05427b4 · outbound

This paper cites Sita: Struc- turally imperceptible and transferable adversarial attacks for stylized image generation.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Sita: Struc- turally imperceptible and transferable adversarial attacks for stylized image generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:20.474489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:12.420336Z digest=sha256:2802c6edaad8f57022546ca39801ab5c342f99b2709df80a07188155d688e0f1

Observation 0a65a579-0f29-4f63-8d2d-aea7735b381e · outbound

This paper cites Is-ggt: Itera- tive scene graph generation with generative transformers.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Is-ggt: Itera- tive scene graph generation with generative transformers

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:20.356956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:12.464267Z digest=sha256:bd29197af1dd26eda3767b264d29863bd5b70278c80281f70c59dce084a955d2

Observation 379d5fb8-6792-49bd-994f-0a4a458fb316 · outbound

This paper cites Panogen: Text-conditioned panoramic environment generation for vision-and-language navigation.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Panogen: Text-conditioned panoramic environment generation for vision-and-language navigation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:20.153322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:12.567526Z digest=sha256:871bb6b31fd095f664d05c95f7e1c407f5b798020cb46709905f8ecb3b532c9a

Observation 10e9e04c-bdb5-43e1-9a72-f80606081a36 · outbound

This paper cites Delving into in- visible semantics for generalized one-shot neural human ren- dering.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Delving into in- visible semantics for generalized one-shot neural human ren- dering

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:20.017204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:12.610425Z digest=sha256:7eed6ff5077e9e650011c0766758d226dcc20e4c7a5d8890d61d80881ec3bb04

Observation eac386b2-18cb-4782-833e-b7e25bd217ad · outbound

This paper cites Multi- modal situated reasoning in 3d scenes.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Multi- modal situated reasoning in 3d scenes

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:19.881172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:12.672393Z digest=sha256:a33487f4cfe1a6abd3aa7ba8abcb9182d05489c153032c5514090e16b8b2bb5e

Observation 92964dcf-8d01-48f2-8a93-fcf256401bc5 · outbound

This paper cites DeepSeek-V3 Technical Report.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition DeepSeek-V3 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:12.730320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:12.730320Z digest=sha256:d188c69c1e7539a93a7fbb2f3c3ee6d007f3831c096a6e4c9bc2cd8e99e8f547

Observation 69934a9d-b77c-4eec-bdd4-0c747b923a4b · outbound

This paper cites Rotation-adaptive point cloud domain generalization via intricate orientation learning.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Rotation-adaptive point cloud domain generalization via intricate orientation learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:19.793645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:12.792481Z digest=sha256:32c6f96e8f043ca5b808d7c69da3b2a344b77278e25dbe5fb47ae04caebcd32e

Observation 8d336ad7-7f9f-40a3-919c-4fe67cba887e · outbound

This paper cites Decoupled Weight Decay Regularization.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Decoupled Weight Decay Regularization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:12.856045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:12.856045Z digest=sha256:16d4f84f2907bcf826feb5cbbfe3f56fc13266f34895a6c8f9ae4c1edae16aab

Observation f2e1b9c6-b79e-4826-b41e-d6d679641ab6 · outbound

This paper cites Mmscan: A multi-modal 3d scene dataset with hierarchical grounded language annota- tions.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Mmscan: A multi-modal 3d scene dataset with hierarchical grounded language annota- tions

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:19.641153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:12.907511Z digest=sha256:0c05336753ab836b5dbd081a2fcd2e2acec77b7d07181361ed42a989b4b895db

Observation 6cb05e79-39f8-48c6-83b4-0778956ae5e7 · outbound

This paper cites Situa- tional awareness matters in 3d vision language reasoning.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Situa- tional awareness matters in 3d vision language reasoning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:19.530994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:13.002554Z digest=sha256:62db68e5364becf11356b9f55b5c4dbb7e998b333593536058bf9d6ce09f6524

Observation b1d8da55-3518-489d-841a-e30ffcb33ad9 · outbound

This paper cites Textrank: Bringing order into text.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Textrank: Bringing order into text

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:19.415622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:13.049968Z digest=sha256:e449ea0a1362a4ee6a7cfef8431ef39590487bbc815972e3a50f4dc37c7c9b9a

Observation 7baa5367-9431-4ebf-9dab-b0e614a4469c · outbound

This paper cites Gaussian prompter: Link- ing 2d prompts for 3d gaussian segmentation.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Gaussian prompter: Link- ing 2d prompts for 3d gaussian segmentation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:19.242942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:13.117883Z digest=sha256:b406f4868a2c3c81b73435be01f9efa541ae79d2aa164b24969aa4341842b5b6

Observation c0cb8c5d-fc43-440c-8122-b81e486fd0af · outbound

This paper cites An approach to gener- ate a caption for an image collection using scene graph gen- eration.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition An approach to gener- ate a caption for an image collection using scene graph gen- eration

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:19.100794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:13.165411Z digest=sha256:cf3d27efd1faaaabffba7095a319ad91dc3c3068cb7850400b5979171394e191

Observation 9d956151-940c-4880-8a83-724d8a8d0321 · outbound

This paper cites Pointnet++: Deep hierarchical feature learning on point sets in a metric space.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Pointnet++: Deep hierarchical feature learning on point sets in a metric space

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:13.209464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:13.209464Z digest=sha256:cdceb5a6b7d8f87ce418efc79698e2d2fc8ddd32a76147cdceaa55a057991a6b

Observation 7aca918f-4f06-4bbe-8c3a-8ccfbe288b3a · outbound

This paper cites Languagerefer: Spatial-language model for 3d visual grounding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Languagerefer: Spatial-language model for 3d visual grounding

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:18.899318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:13.264346Z digest=sha256:a48cf5b768d3824881f59233e9d810243015b01c5399403d6fc76d6a1f01ec52

Observation 0dc6fb11-c63d-4a44-8000-0772f5480f73 · outbound

This paper cites Aware visual grounding in 3d scenes.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Aware visual grounding in 3d scenes

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:18.591055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:13.302172Z digest=sha256:fcdeb9d17eaa3ad48560e8a52dd06a6f49da510eb11660b7580e25ef99cd2605

Observation c8ca42ed-25ea-4e09-bf48-939444c1df02 · outbound

This paper cites Multi-Agent Collaboration: Harnessing the Power of Intelligent LLM Agents.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Multi-Agent Collaboration: Harnessing the Power of Intelligent LLM Agents

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:13.353401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:13.353401Z digest=sha256:48ed40c130ab4b2e5b8e878e14d9f2994e5491aa764a48c4578842fdbc066c85

Observation 1478cb02-15b0-46d6-92fe-72defd0eb2b3 · outbound

This paper cites Attention is all you need.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Attention is all you need

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:13.395000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:13.395000Z digest=sha256:729c022189212ecb2458bff00d551743e75998a5a0aa747a729f638df422bf41

Observation 3fb44b95-ce78-4a3d-a504-3ebc3c94a86f · outbound

This paper cites OpenChat: Advancing Open-source Language Models with Mixed-Quality Data.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition OpenChat: Advancing Open-source Language Models with Mixed-Quality Data

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:13.463357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:13.463357Z digest=sha256:125e388d9a0650f441a838652e4bb0f676f18422abec9f854dde698aa2fa6560

Observation 4abf7b5a-2bff-4eba-b2a7-666fe6aa4823 · outbound

This paper cites Gˆ 3-lq: Marry- ing hyperbolic alignment with explicit semantic-geometric modeling for 3d visual grounding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Gˆ 3-lq: Marry- ing hyperbolic alignment with explicit semantic-geometric modeling for 3d visual grounding

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:18.360406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:13.530305Z digest=sha256:2bf5b966d08c45c1a2812584d212abef4d1be95bf44f52b7b70a955c3d0296a7

Observation c3063547-1f1d-4b65-bd7c-d838fb358d06 · outbound

This paper cites Eda: Explicit text-decoupling and dense alignment for 3d visual grounding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Eda: Explicit text-decoupling and dense alignment for 3d visual grounding

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:18.114493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:13.577268Z digest=sha256:ae2f025dc47e34feb6ba54bb94b5f7cf63e25bcd1841bff652c0885d3b341d6a

Observation 4f8a1019-7534-49a9-91ef-647eee7f552d · outbound

This paper cites Multi- scale flow-based occluding effect and content separation for cartoon animations.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Multi- scale flow-based occluding effect and content separation for cartoon animations

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:17.875702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:13.656275Z digest=sha256:37c98a934326f4b584e56953f0bd747c108c6a668bcc6608d656051584eeeb92

Observation dc991a12-da9e-4040-aea2-80f2cd5f3ce3 · outbound

This paper cites Multi-attribute interactions matter for 3d visual grounding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Multi-attribute interactions matter for 3d visual grounding

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:17.663023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:13.741285Z digest=sha256:d0f3ecbb70611277647319252290eaf7aebeffe52ebb15cc1c7587e55a26c722

Observation 4f1a0181-73b4-4da4-9ebb-d7a7579bb0f1 · outbound

This paper cites Learning with unreliability: Fast few-shot voxel radiance fields with relative geometric consistency.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Learning with unreliability: Fast few-shot voxel radiance fields with relative geometric consistency

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:17.436049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:13.829101Z digest=sha256:ab46814b521e648ddb0e79027ee520a76765344dd42473fdbf774f5f2a58ebef

Observation 8d5e569b-edb4-48de-9a4f-f24c278d71fa · outbound

This paper cites Qwen2.5 Technical Report.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Qwen2.5 Technical Report

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:13.861692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:13.861692Z digest=sha256:fc3b4881f3b814f5086b6c66c1e475322a20c4ec2c29f2ae7f967b29fee9129d

Observation dee07287-1d11-4c9a-8419-d1521f62ed39 · outbound

This paper cites G2face: High-fidelity reversible face anonymization via gen- erative and geometric priors.IEEE Transactions on Informa- tion Forensics and Security, 2024.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition G2face: High-fidelity reversible face anonymization via gen- erative and geometric priors.IEEE Transactions on Informa- tion Forensics and Security, 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:17.218042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:13.971665Z digest=sha256:e9d4cb9a1347f45a00abb47f00f9ccd78fc208b524a7ccb38c834f5f4ef397e6

Observation 410dc413-a061-453c-afe6-3e74887c004c · outbound

This paper cites Sat: 2d semantics assisted training for 3d visual grounding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Sat: 2d semantics assisted training for 3d visual grounding

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:17.017863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:14.062628Z digest=sha256:1993c7634440ae096503555f7c0dd879812894b310c066ad7b51298ae99281ea

Observation c51e1a3b-11e9-4557-aa96-26c9f73191b1 · outbound

This paper cites AppAgent: Multimodal Agents as Smartphone Users.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition AppAgent: Multimodal Agents as Smartphone Users

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:14.165736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:14.165736Z digest=sha256:b43c602272ffe7ce711e9a128c10565930657b59e4e532b43cae9acbf824e74b

Observation 7285a952-e166-495d-96f7-3e13a0c27a45 · outbound

This paper cites Visually-prompted language model for fine-grained scene graph generation in an open world.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Visually-prompted language model for fine-grained scene graph generation in an open world

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:16.729026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:14.205944Z digest=sha256:1dd6eb109a55e6f77cb83fd2c4c5c6a2b377eaf75d4bc3eceed9956e9d34d42d

Observation 50924109-6dab-4f3e-b74f-a1aaabe18993 · outbound

This paper cites Visual programming for zero-shot open-vocabulary 3d visual grounding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Visual programming for zero-shot open-vocabulary 3d visual grounding

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:16.560839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:14.295905Z digest=sha256:5c23d5793716afff44a33428ff48b6013d8d51480f9c40283e68dc0861e3007c

Observation 0d2a8417-3822-4740-b6da-d5ad400d68f9 · outbound

This paper cites Multi3drefer: Grounding text description to multiple 3d ob- jects.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Multi3drefer: Grounding text description to multiple 3d ob- jects

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:16.325888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:14.401402Z digest=sha256:fc43fca35e650f1b1c40e0f74d561b57e238f3de32b19880fb9da6885b050c5e

Observation d7249bfb-ba41-47a9-91ba-e54441ddf6c2 · outbound

This paper cites Towards clip-driven language-free 3d visual grounding via 2d-3d relational en- hancement and consistency.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Towards clip-driven language-free 3d visual grounding via 2d-3d relational en- hancement and consistency

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:16.133141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:14.470853Z digest=sha256:ef348eab9f97e83c5b5d5fd3b86c9d975464cb5eeaef3ad676f2d20b0b38ae73

Observation 4fe4caae-2393-4b23-ac9e-4e82e6bb8ea6 · outbound

This paper cites 3dvg- transformer: Relation modeling for visual grounding on point clouds.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition 3dvg- transformer: Relation modeling for visual grounding on point clouds

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:15.958457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:14.561773Z digest=sha256:645a7e217022120f0b42243b15557932cacf53df74dae3025230eb69c42f000f

Observation b2d39ab7-26b5-4deb-b54a-4aa8bcd4585e · outbound

This paper cites Recdreamer: Consistent text-to-3d generation via uniform score distillation.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Recdreamer: Consistent text-to-3d generation via uniform score distillation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:15.769685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:14.688627Z digest=sha256:125145bbe66b0ad98eeffc0f3c453117f0387fdfb73e1220e013f6fa48b98d6e

Observation 0f8264bd-dc26-47cd-bb8e-d9ba862e5af3 · outbound

This paper cites Learning an interpretable stylized subspace for 3d-aware animatable artforms.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Learning an interpretable stylized subspace for 3d-aware animatable artforms

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:15.560384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:14.768209Z digest=sha256:0aab2c6c8b7ecb6d7aad4a8c5960522c42d661721cf5b63a13c9297d6076422b

Observation 24fadd34-555f-417a-b47b-443df573e8d0 · outbound

This paper cites Navgpt: Explicit reasoning in vision-and-language navigation with large lan- guage models.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Navgpt: Explicit reasoning in vision-and-language navigation with large lan- guage models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:15.365406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:14.861449Z digest=sha256:924906f212ba1377a2190800191c269bbe9bce6c9e73c3052ca0ecc86d498dde

Observation 39697c49-7872-4fd2-bc57-0ee034b57ec5 · outbound

This paper cites Unifying 3d vision-language understanding via prompt- able queries.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Unifying 3d vision-language understanding via prompt- able queries

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:15.154024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:18:14.922481Z digest=sha256:e6b3b603b7a0f56d31f163cca15af44079b66540428e10b4000339fbec90897f

Pith citing papers

No inbound Pith citation observations are available.