Pith. sign in

Paper Citation Record · LEDGER

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition

As of 16 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2507.11261.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.11261 v2

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:18:14.922481Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy46
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 99a5221d-5176-43af-abae-2f73a0eb4af0 · outbound

This paper cites Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:23.475431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:10.932060Z digest=sha256:e94cef3a4034c4a158bafc61f84f020d02390e8a98ad0757b076f4006d3eb500

Observation f975c64a-485f-4d78-9f8d-dc42fbe8e48a · outbound

This paper cites Cot3dref: Chain-of-thoughts data-efficient 3d visual grounding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Cot3dref: Chain-of-thoughts data-efficient 3d visual grounding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:23.222477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:11.018378Z digest=sha256:5bcae93e3f7e80ddd0608ea8f1f7776674134786fe261cdd269712beab333656

Observation f4e61f26-db46-490f-aba8-0eb1df005f13 · outbound

This paper cites Visual question answering from another perspective: Clevr mental rotation tests.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Visual question answering from another perspective: Clevr mental rotation tests

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:22.964487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:11.096372Z digest=sha256:79f961efa1020c8ca7a3527c33a7553c065b485cbf3ab1e38ae000e85451d5c8

Observation d4dc29a6-0669-4bdd-976d-eb467977906c · outbound

This paper cites Assertiveness-based agent communica- tion for a personalized medicine on medical imaging diag- nosis.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Assertiveness-based agent communica- tion for a personalized medicine on medical imaging diag- nosis

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:22.729024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:11.162571Z digest=sha256:2fd484aa7d749de0b1c0d4204518f151a428669a4fd75658238404b433d04b72

Observation a17f72e6-33c7-48fc-a038-94af86b7efcf · outbound

This paper cites Mikasa: Multi-key-anchor & scene-aware transformer for 3d visual grounding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Mikasa: Multi-key-anchor & scene-aware transformer for 3d visual grounding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:22.578516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:11.232514Z digest=sha256:471d3abb15704b9eaf7354c4bf1827ff0fb7abb3f585fe56dcc5dfe0a0919045

Observation 526e06e0-2a1c-4bec-b757-a07af60d434d · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Scanrefer: 3d object localization in rgb-d scans using natural language

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:22.419009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:11.272693Z digest=sha256:cf4b8c1371a25e762d1af1efc5cafa828c8544d33100a9ba72016e5482938bf1

Observation 59de1d01-e895-4fa6-baea-6d5f6edca71a · outbound

This paper cites SCJD: Sparse Correlation and Joint Distillation for Efficient 3D Human Pose Estimation.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition SCJD: Sparse Correlation and Joint Distillation for Efficient 3D Human Pose Estimation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:11.315194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:11.315194Z digest=sha256:7f36f6b306dd14bfc1b8348150ec3833768aeb404b0c615e9fd501c8d1a20622

Observation eccb3eba-6be2-470e-bbeb-bf59790be25d · outbound

This paper cites Talk2bev: Language-enhanced bird’s- eye view maps for autonomous driving.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Talk2bev: Language-enhanced bird’s- eye view maps for autonomous driving

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:22.228420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:11.372304Z digest=sha256:600f181268d0b6dbfa0a4aa03a8157b8694e54a7d064642c0ed7c1b0960ff6fe

Observation 1dc3fd23-afd9-4983-abd1-963c1e72f9a5 · outbound

This paper cites Drive as you speak: Enabling human-like interac- tion with large language models in autonomous vehicles.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Drive as you speak: Enabling human-like interac- tion with large language models in autonomous vehicles

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:22.060228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:11.459556Z digest=sha256:a00698f74126fc288b571962e5046331fd5f6c6cdce6047063fabb8b2cd09f5d

Observation 1c2babb7-5748-4946-a129-a39f62ca7b64 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:11.526902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:11.526902Z digest=sha256:d2b0115ed5e66d6d23a39ee0c6d38ba574b164d57261321fed31d78d819d39c7

Observation 0ffb4a29-cc63-41ed-8328-6490f340850e · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:11.558820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:11.558820Z digest=sha256:3f66bb72fd1d29dc5d2f3a36b9d5549f95d7a0fe72fc5133797bc50faa0f3d14

Observation ade96116-986d-4715-878b-0a4040aa684d · outbound

This paper cites Scenegenie: Scene graph guided diffusion models for image synthesis.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Scenegenie: Scene graph guided diffusion models for image synthesis

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:21.859773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:11.609396Z digest=sha256:36893a819d2c28c4b7b57c94da3a0aab4fcb5fc4a54a72330e03db40c69ce155

Observation 94dd748e-1e5b-4db3-abdf-d3114273a179 · outbound

This paper cites Dense reinforce- ment learning for safety validation of autonomous vehicles.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Dense reinforce- ment learning for safety validation of autonomous vehicles

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:21.686087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:11.725233Z digest=sha256:447f9199369899559ed799df6f1c6536e862345a9ac4087697cdcf67ae191ded

Observation af1fb14e-7516-443a-8274-73fcc2fbc43c · outbound

This paper cites Viewinfer3d: 3d visual ground- ing based on embodied viewpoint inference.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Viewinfer3d: 3d visual ground- ing based on embodied viewpoint inference

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:21.525736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:11.814980Z digest=sha256:5a89cd6f426c2539666437e0ed2fb1ee4b425c571e83f16e9777edfa994ba91b

Observation e05dd4c4-9cf5-4bd0-8f03-d88fee1038db · outbound

This paper cites Viewrefer: Grasp the multi-view knowledge for 3d visual grounding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Viewrefer: Grasp the multi-view knowledge for 3d visual grounding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:21.382853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:11.879963Z digest=sha256:12f5f2ab999e9883b9f0008a654cade0993aa0cbb51128b01f008e13d9ca1124

Observation 10802696-30ec-4432-802d-8d1ad2a2a2ba · outbound

This paper cites Transrefer3d: Entity-and- relation aware transformer for fine-grained 3d visual ground- ing.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Transrefer3d: Entity-and- relation aware transformer for fine-grained 3d visual ground- ing

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:21.222781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:11.960630Z digest=sha256:d25097e6ac968f360a4bd1557595015536f45f6c7161f441d916342e43819fc8

Observation 871a281a-822f-48d2-b87e-f017dd6f71da · outbound

This paper cites 3d-llm: In- jecting the 3d world into large language models.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition 3d-llm: In- jecting the 3d world into large language models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:12.036499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:12.036499Z digest=sha256:73dae0e4e6b95bf627885991d56b1311ed5c0799eddd29ba1e0908ab83cfbe27

Observation be0bd385-f4ef-4432-8a79-37595d71b9e9 · outbound

This paper cites Multi- view transformer for 3d visual grounding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Multi- view transformer for 3d visual grounding

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:21.108849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:12.169841Z digest=sha256:cf87d7e893a163021252835b2944102dc151aecf1a334c156802efdd92b6518d

Observation fe61d437-8966-45a3-b4d8-9ffb6c95ed08 · outbound

This paper cites Structure-clip: Towards scene graph knowledge to enhance multi-modal structured repre- sentations.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Structure-clip: Towards scene graph knowledge to enhance multi-modal structured repre- sentations

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:20.956537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:12.226913Z digest=sha256:2a8ff1557e77c13ffdf26fa13a424f8695b8645d561cd6a52a8b5d198cb4b013

Observation 1447d34a-0364-45f7-9836-3487acfcbaec · outbound

This paper cites Nan-detr: noising multi-anchor makes detr better for object detection.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Nan-detr: noising multi-anchor makes detr better for object detection

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:20.759454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:12.302592Z digest=sha256:a6fe4e819a901ef6a2d6fa50ca29d747a8814f869d341ba85d10a96272b2453a

Observation c7234e13-59d2-4e77-a0e8-722b6bfeb5ee · outbound

This paper cites Bottom up top down detection transform- ers for language grounding in images and point clouds.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Bottom up top down detection transform- ers for language grounding in images and point clouds

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:20.616504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:12.337407Z digest=sha256:32fe2667c3a05a2baa693e69b91b8bf0063d318b4140f74555e36537f038aaa1

Observation 20737413-972d-41be-8b38-9cf6e05427b4 · outbound

This paper cites Sita: Struc- turally imperceptible and transferable adversarial attacks for stylized image generation.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Sita: Struc- turally imperceptible and transferable adversarial attacks for stylized image generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:20.474489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:12.420336Z digest=sha256:f03e214d7423b1a198847c59777c55b6a6f4aa1af8b11966d53a5e1dd0bdb2a4

Observation 0a65a579-0f29-4f63-8d2d-aea7735b381e · outbound

This paper cites Is-ggt: Itera- tive scene graph generation with generative transformers.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Is-ggt: Itera- tive scene graph generation with generative transformers

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:20.356956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:12.464267Z digest=sha256:8429a2695d27f0259b0e450d7e8069bef605c5756a5dbd6d09f308b5733510e6

Observation 379d5fb8-6792-49bd-994f-0a4a458fb316 · outbound

This paper cites Panogen: Text-conditioned panoramic environment generation for vision-and-language navigation.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Panogen: Text-conditioned panoramic environment generation for vision-and-language navigation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:20.153322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:12.567526Z digest=sha256:9cfc014851b9ae48a4578c992913989415ba2c801b5cb97d0be008f169ddeb6d

Observation 10e9e04c-bdb5-43e1-9a72-f80606081a36 · outbound

This paper cites Delving into in- visible semantics for generalized one-shot neural human ren- dering.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Delving into in- visible semantics for generalized one-shot neural human ren- dering

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:20.017204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:12.610425Z digest=sha256:f8ba18012ee9e3431e7ffb9abdc746253750b719276c1a6e4dcd1f397439cb47

Observation eac386b2-18cb-4782-833e-b7e25bd217ad · outbound

This paper cites Multi- modal situated reasoning in 3d scenes.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Multi- modal situated reasoning in 3d scenes

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:19.881172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:12.672393Z digest=sha256:3baadd068ad24faeb72f93c07df6d58350df2ad7d9554d019facc3e72b0f5e93

Observation 92964dcf-8d01-48f2-8a93-fcf256401bc5 · outbound

This paper cites DeepSeek-V3 Technical Report.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition DeepSeek-V3 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:12.730320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:12.730320Z digest=sha256:1c8a5068bc5f338e8a458189c5d904e1b667ab6656d03d456f26181cb3610797

Observation 69934a9d-b77c-4eec-bdd4-0c747b923a4b · outbound

This paper cites Rotation-adaptive point cloud domain generalization via intricate orientation learning.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Rotation-adaptive point cloud domain generalization via intricate orientation learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:19.793645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:12.792481Z digest=sha256:a533606ec644857eae9901cf4304f0ab380e55a699729036c99c9e670e059d43

Observation 8d336ad7-7f9f-40a3-919c-4fe67cba887e · outbound

This paper cites Decoupled Weight Decay Regularization.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Decoupled Weight Decay Regularization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:12.856045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:12.856045Z digest=sha256:928e6e2fee169e5a422a340e6c203141a82d847eb69c9f43b988a6e9bac55330

Observation f2e1b9c6-b79e-4826-b41e-d6d679641ab6 · outbound

This paper cites Mmscan: A multi-modal 3d scene dataset with hierarchical grounded language annota- tions.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Mmscan: A multi-modal 3d scene dataset with hierarchical grounded language annota- tions

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:19.641153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:12.907511Z digest=sha256:e80568b69d8854eb96080144362fd98f89f4a331e5f79f0ddc60b1c951a2d8cf

Observation 6cb05e79-39f8-48c6-83b4-0778956ae5e7 · outbound

This paper cites Situa- tional awareness matters in 3d vision language reasoning.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Situa- tional awareness matters in 3d vision language reasoning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:19.530994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:13.002554Z digest=sha256:52ca2a8eb1527f63502e5ef25aad362da0f481db75e59127b7aaa180b6694a3a

Observation b1d8da55-3518-489d-841a-e30ffcb33ad9 · outbound

This paper cites Textrank: Bringing order into text.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Textrank: Bringing order into text

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:19.415622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:13.049968Z digest=sha256:d6b0b265c8f7931a93db21ce7f2e326ce50a415d3f448ab2ca3b88df7f3e8467

Observation 7baa5367-9431-4ebf-9dab-b0e614a4469c · outbound

This paper cites Gaussian prompter: Link- ing 2d prompts for 3d gaussian segmentation.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Gaussian prompter: Link- ing 2d prompts for 3d gaussian segmentation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:19.242942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:13.117883Z digest=sha256:8e63a64462f298822c643f688e471e172a59511c7fccdcc1486a470871e6a756

Observation c0cb8c5d-fc43-440c-8122-b81e486fd0af · outbound

This paper cites An approach to gener- ate a caption for an image collection using scene graph gen- eration.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition An approach to gener- ate a caption for an image collection using scene graph gen- eration

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:19.100794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:13.165411Z digest=sha256:289fa4943903d78b426a60f7916d287ff2fcd8a079cb1389e497407bbd70f9bb

Observation 9d956151-940c-4880-8a83-724d8a8d0321 · outbound

This paper cites Pointnet++: Deep hierarchical feature learning on point sets in a metric space.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Pointnet++: Deep hierarchical feature learning on point sets in a metric space

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:13.209464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:13.209464Z digest=sha256:f959bbd2937afad256ac3a46b0a633874f4533d8deeceeff37eb94c1172900ef

Observation 7aca918f-4f06-4bbe-8c3a-8ccfbe288b3a · outbound

This paper cites Languagerefer: Spatial-language model for 3d visual grounding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Languagerefer: Spatial-language model for 3d visual grounding

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:18.899318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:13.264346Z digest=sha256:039219da15765687057c692e63b8d2b8ad40e49e37c946cb140dbc1f6341a154

Observation 0dc6fb11-c63d-4a44-8000-0772f5480f73 · outbound

This paper cites Aware visual grounding in 3d scenes.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Aware visual grounding in 3d scenes

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:18.591055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:13.302172Z digest=sha256:ae00b93a12ec87cb90ebb5cbdffb81bbebf19a43a4c14ced76a768ab34c57f8a

Observation c8ca42ed-25ea-4e09-bf48-939444c1df02 · outbound

This paper cites Multi-Agent Collaboration: Harnessing the Power of Intelligent LLM Agents.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Multi-Agent Collaboration: Harnessing the Power of Intelligent LLM Agents

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:13.353401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:13.353401Z digest=sha256:ee4dac01a009448be557fac5346b4e90d2b06a28287f3aee0a133457841ebff5

Observation 1478cb02-15b0-46d6-92fe-72defd0eb2b3 · outbound

This paper cites Attention is all you need.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Attention is all you need

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:13.395000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:13.395000Z digest=sha256:cf2333801db4298d896a888fb59e3197d6515ad1c4d3acacb12529699955f5fe

Observation 3fb44b95-ce78-4a3d-a504-3ebc3c94a86f · outbound

This paper cites OpenChat: Advancing Open-source Language Models with Mixed-Quality Data.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition OpenChat: Advancing Open-source Language Models with Mixed-Quality Data

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:13.463357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:13.463357Z digest=sha256:a6a7698824056bbbc45f638d9e227aa8559004945bc47dcc5bc406825cfec5a5

Observation 4abf7b5a-2bff-4eba-b2a7-666fe6aa4823 · outbound

This paper cites Gˆ 3-lq: Marry- ing hyperbolic alignment with explicit semantic-geometric modeling for 3d visual grounding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Gˆ 3-lq: Marry- ing hyperbolic alignment with explicit semantic-geometric modeling for 3d visual grounding

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:18.360406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:13.530305Z digest=sha256:a4d6c7e69a116f09086c4d94426a39897a7ee8468c28290431cd8e17019a4ca9

Observation c3063547-1f1d-4b65-bd7c-d838fb358d06 · outbound

This paper cites Eda: Explicit text-decoupling and dense alignment for 3d visual grounding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Eda: Explicit text-decoupling and dense alignment for 3d visual grounding

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:18.114493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:13.577268Z digest=sha256:ef19c6e65b933c27d9ebe78b9361ed5defa2bed2ad8368cdacf353fd0004a185

Observation 4f8a1019-7534-49a9-91ef-647eee7f552d · outbound

This paper cites Multi- scale flow-based occluding effect and content separation for cartoon animations.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Multi- scale flow-based occluding effect and content separation for cartoon animations

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:17.875702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:13.656275Z digest=sha256:e8caf7c0a653e3ef8314a6fb1131fc59d6fefba58239ee7f57d5ee684363b50f

Observation dc991a12-da9e-4040-aea2-80f2cd5f3ce3 · outbound

This paper cites Multi-attribute interactions matter for 3d visual grounding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Multi-attribute interactions matter for 3d visual grounding

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:17.663023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:13.741285Z digest=sha256:c945ccd150d38549ad13865dd85cf49d92893547c2182cc3f4d6dcac93deb3c6

Observation 4f1a0181-73b4-4da4-9ebb-d7a7579bb0f1 · outbound

This paper cites Learning with unreliability: Fast few-shot voxel radiance fields with relative geometric consistency.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Learning with unreliability: Fast few-shot voxel radiance fields with relative geometric consistency

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:17.436049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:13.829101Z digest=sha256:dea3ea46b22b7346fb681b27a5f13db2ead58a0036102f3b5cc9231d5510a3b4

Observation 8d5e569b-edb4-48de-9a4f-f24c278d71fa · outbound

This paper cites Qwen2.5 Technical Report.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Qwen2.5 Technical Report

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:13.861692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:13.861692Z digest=sha256:23147daf19e7380405078f479bf744e4c1c1c032da381fa5fa08c64bce6eb823

Observation dee07287-1d11-4c9a-8419-d1521f62ed39 · outbound

This paper cites G2face: High-fidelity reversible face anonymization via gen- erative and geometric priors.IEEE Transactions on Informa- tion Forensics and Security, 2024.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition G2face: High-fidelity reversible face anonymization via gen- erative and geometric priors.IEEE Transactions on Informa- tion Forensics and Security, 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:17.218042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:13.971665Z digest=sha256:4b5ae8f299acc2267fd4e18b855f4e46e12e8b81d23fd15210b4dd7a875d8990

Observation 410dc413-a061-453c-afe6-3e74887c004c · outbound

This paper cites Sat: 2d semantics assisted training for 3d visual grounding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Sat: 2d semantics assisted training for 3d visual grounding

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:17.017863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:14.062628Z digest=sha256:86ffa61b7590bc4a5be6bdde2dcccc32aece98bc3a7e05a8e4581a4d0c2f769e

Observation c51e1a3b-11e9-4557-aa96-26c9f73191b1 · outbound

This paper cites AppAgent: Multimodal Agents as Smartphone Users.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition AppAgent: Multimodal Agents as Smartphone Users

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:14.165736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:14.165736Z digest=sha256:b4e8cf580f11a6a242f2b138dea17bba50c111a447f6ac4601f2e100e63cbc3c

Observation 7285a952-e166-495d-96f7-3e13a0c27a45 · outbound

This paper cites Visually-prompted language model for fine-grained scene graph generation in an open world.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Visually-prompted language model for fine-grained scene graph generation in an open world

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:16.729026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:14.205944Z digest=sha256:1027147a9c379574eda6db9599ede6f3c0a17ffce51542acfb85d080660753f3

Observation 50924109-6dab-4f3e-b74f-a1aaabe18993 · outbound

This paper cites Visual programming for zero-shot open-vocabulary 3d visual grounding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Visual programming for zero-shot open-vocabulary 3d visual grounding

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:16.560839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:14.295905Z digest=sha256:fc387d26d74a21117369f21a4692aea8b52cf8781f3dc50e8156d3a9aa2db609

Observation 0d2a8417-3822-4740-b6da-d5ad400d68f9 · outbound

This paper cites Multi3drefer: Grounding text description to multiple 3d ob- jects.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Multi3drefer: Grounding text description to multiple 3d ob- jects

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:16.325888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:14.401402Z digest=sha256:6b9f1a2482c3c7b745bc2fcf45d9bc99d6049473a3fbd4566aea8aa05d075c03

Observation d7249bfb-ba41-47a9-91ba-e54441ddf6c2 · outbound

This paper cites Towards clip-driven language-free 3d visual grounding via 2d-3d relational en- hancement and consistency.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Towards clip-driven language-free 3d visual grounding via 2d-3d relational en- hancement and consistency

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:16.133141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:14.470853Z digest=sha256:c0f18081eebba192fa2afad85508b76ae97f932622860295f0496270270667bc

Observation 4fe4caae-2393-4b23-ac9e-4e82e6bb8ea6 · outbound

This paper cites 3dvg- transformer: Relation modeling for visual grounding on point clouds.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition 3dvg- transformer: Relation modeling for visual grounding on point clouds

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:15.958457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:14.561773Z digest=sha256:205726301c3c54df136ba9c194ab2422548ae3203e9b78764ee4b332dcc9d22a

Observation b2d39ab7-26b5-4deb-b54a-4aa8bcd4585e · outbound

This paper cites Recdreamer: Consistent text-to-3d generation via uniform score distillation.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Recdreamer: Consistent text-to-3d generation via uniform score distillation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:15.769685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:14.688627Z digest=sha256:0bbcab1bd243c4f1634e925e6becc4d6acabde82a4a55a6de83fa165c7c9108e

Observation 0f8264bd-dc26-47cd-bb8e-d9ba862e5af3 · outbound

This paper cites Learning an interpretable stylized subspace for 3d-aware animatable artforms.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Learning an interpretable stylized subspace for 3d-aware animatable artforms

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:15.560384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:14.768209Z digest=sha256:5d3a3025e39a63e1d5e477783fb8886e9362220d0918f35ab6581e8ccea9b90a

Observation 24fadd34-555f-417a-b47b-443df573e8d0 · outbound

This paper cites Navgpt: Explicit reasoning in vision-and-language navigation with large lan- guage models.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Navgpt: Explicit reasoning in vision-and-language navigation with large lan- guage models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:15.365406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:14.861449Z digest=sha256:862f40c30c2c069a2ce88ed9c66620fe04a09ffb2bf667337c3bb3f4bdacdd05

Observation 39697c49-7872-4fd2-bc57-0ee034b57ec5 · outbound

This paper cites Unifying 3d vision-language understanding via prompt- able queries.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Unifying 3d vision-language understanding via prompt- able queries

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:15.154024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:18:14.922481Z digest=sha256:cb1f7308edfb2f0bf82d494ef25e2f6061c7cc9bbdf62d4e51adaa6a2abcd9ff

Pith citing papers

No inbound Pith citation observations are available.