Pith. sign in

Paper Citation Record · LEDGER

3D Scene Graph Guided Vision-Language Pre-training

As of 18 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2411.18666.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.18666 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:13:36.729314Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

63 of 63 outbound references displayed

  • verified exact2
  • verified fuzzy45
  • unresolved15
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d904d31c-8ce4-418e-9f95-364ff0aa5bd3 · outbound

This paper cites Referit3D: Neural listeners for fine-grained 3D object identification in real-world scenes.

3D Scene Graph Guided Vision-Language Pre-training Referit3D: Neural listeners for fine-grained 3D object identification in real-world scenes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.941437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.067631Z digest=sha256:10d4723dfeb0ecfd3b4cfc7a54418a5f2e6d90ea67a0715ad9d0852bb17337a8

Observation be7fb37e-cab9-46f4-9462-15f2ac409ff0 · outbound

This paper cites Scanqa: 3D question answering for spatial scene understanding.

3D Scene Graph Guided Vision-Language Pre-training Scanqa: 3D question answering for spatial scene understanding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.911304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.075948Z digest=sha256:422edcc2334c02126a0a39520c71a63965ad8df494b21edb19d8f8e8160ba076

Observation 9256b3a2-a057-4bea-9edd-a959b4aeb8bf · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments.

3D Scene Graph Guided Vision-Language Pre-training Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.879255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.090937Z digest=sha256:747e67d0eeded616e9dde088ea1233b7eb1c65570776d78ac38de0ec9951f77b

Observation e1a35ee1-bbbc-4f12-ad66-ecba3a64f95b · outbound

This paper cites 3DJCG: A unified framework for joint dense captioning and visual grounding on 3D point clouds.

3D Scene Graph Guided Vision-Language Pre-training 3DJCG: A unified framework for joint dense captioning and visual grounding on 3D point clouds

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.855501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.098706Z digest=sha256:423edff19a8f4b40e30dbf1fef514a600c7358284fa258c3992503be7dd8427d

Observation 54ee0b72-a096-4a8e-a3a1-09e805387cb3 · outbound

This paper cites Scanrefer: 3D object localization in RGB-D scans using natural language.

3D Scene Graph Guided Vision-Language Pre-training Scanrefer: 3D object localization in RGB-D scans using natural language

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.828887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.105899Z digest=sha256:b08d0d29d542c7e6367bb8bfb9c2c64efe0422756223bae2a0e5234ee0bfab68

Observation e0553d85-eb06-498a-a401-b7fe7060acf6 · outbound

This paper cites UniT3D: A Unified Transformer for 3D Dense Captioning and Visual Grounding.

3D Scene Graph Guided Vision-Language Pre-training UniT3D: A Unified Transformer for 3D Dense Captioning and Visual Grounding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T11:13:36.119765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:13:36.119765Z digest=sha256:d82ea9e347b1cd80d4a9056a8add674ad6294d019fb406fcb577339e981fef2a

Observation 482814b2-9443-4391-b56c-45ebf84c4f28 · outbound

This paper cites D 3net: A unified speaker-listener architec- ture for 3D dense captioning and visual grounding.

3D Scene Graph Guided Vision-Language Pre-training D 3net: A unified speaker-listener architec- ture for 3D dense captioning and visual grounding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.787228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.132167Z digest=sha256:d0cb15050bcf2bdbbed459e39a963658bb921795e0d49e9c2b042108a81be0bf

Observation cdc9c32c-7958-482a-8728-3b504033a28e · outbound

This paper cites Language Conditioned Spatial Relation Reasoning for 3D Object Grounding.

3D Scene Graph Guided Vision-Language Pre-training Language Conditioned Spatial Relation Reasoning for 3D Object Grounding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T11:13:36.142026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:13:36.142026Z digest=sha256:4285a26e27e2a0bc2cc09c57bc32d8c502fa6bfdf5a5375136cdb2eeaf995a93

Observation 645089ad-5419-4f2b-8c07-12326b1c8259 · outbound

This paper cites End-to-End 3D Dense Captioning with Vote2Cap-DETR.

3D Scene Graph Guided Vision-Language Pre-training End-to-End 3D Dense Captioning with Vote2Cap-DETR

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-12T11:13:37.175474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.148543Z digest=sha256:7a55fc5f65d3c99e62b0c0d21eeb2e1abfb60dde94352cfb534085aa607ac7a1

Observation 1eae4423-c0f7-440f-8270-d92e8eede3d9 · outbound

This paper cites Simclr: A simple framework for contrastive learning of visual representations.

3D Scene Graph Guided Vision-Language Pre-training Simclr: A simple framework for contrastive learning of visual representations

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.748738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.164142Z digest=sha256:6c3b1780004c9becc79d9746958e76afc0f9f0c003b64c2531de482b0c44717b

Observation 9c04a5c1-edd0-450b-9aab-bb5b3a831654 · outbound

This paper cites Scan2cap: Context-aware dense captioning in RGB- D scans.

3D Scene Graph Guided Vision-Language Pre-training Scan2cap: Context-aware dense captioning in RGB- D scans

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.717928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.177460Z digest=sha256:86b92c089add04569c5fd2b6c9435a15569424ce7c271059f1d716dacde306cb

Observation 06d9a5d4-d9d6-4a16-a87b-242937f7b853 · outbound

This paper cites Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling.

3D Scene Graph Guided Vision-Language Pre-training Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T11:13:36.186265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:13:36.186265Z digest=sha256:d6f6dcb80c9bb4ca64dee6f8604b8efd5ec81ff47a1f77989e8084849547607f

Observation 2ca348ef-f9d4-4bee-a0fb-a2895c547056 · outbound

This paper cites Scannet: Richly-annotated 3D reconstructions of indoor scenes.

3D Scene Graph Guided Vision-Language Pre-training Scannet: Richly-annotated 3D reconstructions of indoor scenes

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.688778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.192424Z digest=sha256:250fe69b44d721ff5eed78a8f14f0b12fa3432d27e92bfb5ef6f97923c84ab44

Observation 6148387c-a619-43d0-b9ce-9d56e1a5d581 · outbound

This paper cites Multi-CLIP: Contrastive Vision-Language Pre-training for Question Answering tasks in 3D Scenes.

3D Scene Graph Guided Vision-Language Pre-training Multi-CLIP: Contrastive Vision-Language Pre-training for Question Answering tasks in 3D Scenes

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T11:13:36.201254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:13:36.201254Z digest=sha256:b87838a316bac05abee1015ab2b9611bafd38c91611edc990603831b1a2f62e5

Observation 6e037ebd-0304-4fea-be66-fe391167025e · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

3D Scene Graph Guided Vision-Language Pre-training BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T11:13:36.208235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:13:36.208235Z digest=sha256:a7e14db4fc692b19be85de61427060afd30beec60d1c1e64642e77986b4b4056

Observation 46c911d3-3f83-4605-b491-89e16ab45a50 · outbound

This paper cites Multi-modal align- ment using representation codebook.

3D Scene Graph Guided Vision-Language Pre-training Multi-modal align- ment using representation codebook

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.661247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.216194Z digest=sha256:1099ecd68c81a68b62eb9be263515a61cf02b40df824c7c643f01932621e5203

Observation 35ecd526-1ac5-4141-9a44-49c7922836c0 · outbound

This paper cites Free-form description guided 3D visual graph network for object grounding in point cloud.

3D Scene Graph Guided Vision-Language Pre-training Free-form description guided 3D visual graph network for object grounding in point cloud

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.636225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.227686Z digest=sha256:760411a45f66431dd15140f79d5ece2d16f65b71eb41b252530deb1ab8e0eb24

Observation 89e366f2-9e98-47b9-922b-6467d7e3e23c · outbound

This paper cites ViewRefer: Grasp the Multi-view Knowledge for 3D Visual Grounding with GPT and Prototype Guidance.

3D Scene Graph Guided Vision-Language Pre-training ViewRefer: Grasp the Multi-view Knowledge for 3D Visual Grounding with GPT and Prototype Guidance

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T11:13:36.245100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:13:36.245100Z digest=sha256:1a18667d0c74045b2fc533cc4b2f9153b5b7eb6ed2557cfcfb5d9759bb2665a6

Observation 0a7836a0-ad93-4bc4-9554-f6f8d85d3038 · outbound

This paper cites Masked autoencoders are scalable vision learners.

3D Scene Graph Guided Vision-Language Pre-training Masked autoencoders are scalable vision learners

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T11:13:36.264686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:13:36.264686Z digest=sha256:8597029a8fefcf0c8d7829bd285238bf407a92909233454bc8b7b07a771d4942

Observation 3b84bbb6-3d71-4fbc-b32b-defdf2e8dc1d · outbound

This paper cites Text-guided graph neural networks for refer- ring 3D instance segmentation.

3D Scene Graph Guided Vision-Language Pre-training Text-guided graph neural networks for refer- ring 3D instance segmentation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.559829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.275054Z digest=sha256:7df56ba43b23cb55200c43c8510dfa4e51c5bcd1eab972ef5ea18dd43b3e9625

Observation d703c403-646a-40e8-b776-484f3a84121a · outbound

This paper cites Multi- view transformer for 3D visual grounding.

3D Scene Graph Guided Vision-Language Pre-training Multi- view transformer for 3D visual grounding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.519069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.290301Z digest=sha256:0fe4647f05fa0d31f3d7c4e0903252b91096b0bc12c060519408fe80e2c14f45

Observation dc79bff9-a612-41e8-806b-6d6231baab9e · outbound

This paper cites Clip2point: Transfer clip to point cloud classifica- tion with image-depth pre-training.

3D Scene Graph Guided Vision-Language Pre-training Clip2point: Transfer clip to point cloud classifica- tion with image-depth pre-training

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.488174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.299677Z digest=sha256:11229c95da87e78ec5a88ba3a2840f8ebfe3058a3ee72436aacbdf7b3ee96704

Observation 80f76758-69ec-477c-a06e-b42c95b9799a · outbound

This paper cites Bottom up top down detection transform- ers for language grounding in images and point clouds.

3D Scene Graph Guided Vision-Language Pre-training Bottom up top down detection transform- ers for language grounding in images and point clouds

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.454472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.307834Z digest=sha256:2891e67531a8af19f3aa45c69c7cfa5968626bdbd035370f9a9d4593f639e3aa

Observation 3d2c4fcc-275b-4085-85de-12e4066ecb49 · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

3D Scene Graph Guided Vision-Language Pre-training Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.414822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.316408Z digest=sha256:19adf2f19995b38cc92ab641881f9c5adb25c947ae4159bb6426d72dc1a4dee5

Observation d5ce4dcf-1918-435f-a60f-b404bfceaabc · outbound

This paper cites More: Multi-order relation mining for dense captioning in 3D scenes.

3D Scene Graph Guided Vision-Language Pre-training More: Multi-order relation mining for dense captioning in 3D scenes

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.384467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.329853Z digest=sha256:1ce9ae5aabf72dd800378cc816025ac4d13896a8b0d006710fde935e0ff5aaf7

Observation f806a9a1-4cc2-4bae-b76e-e3f58be63352 · outbound

This paper cites Context-aware alignment and mutual masking for 3D-language pre-training.

3D Scene Graph Guided Vision-Language Pre-training Context-aware alignment and mutual masking for 3D-language pre-training

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.317360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.363821Z digest=sha256:74ab976e8a4b680c90b129e5bba91f89ba942670801e8d60476d8c89a6102544

Observation d952bcfe-9492-4d7d-8cd3-870011cc3b07 · outbound

This paper cites Lang3DSG: Language-based contrastive pre-training for 3D Scene Graph prediction.

3D Scene Graph Guided Vision-Language Pre-training Lang3DSG: Language-based contrastive pre-training for 3D Scene Graph prediction

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-12T11:13:36.965742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.375827Z digest=sha256:55356d1768891cf390a3fb55544ed9a58cf655ab65c1d0d957128ed095a181d8

Observation 93031e7c-7f48-4ca6-8c0a-6c4c58a0c5b5 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

3D Scene Graph Guided Vision-Language Pre-training Rouge: A package for automatic evaluation of summaries

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T11:13:36.386569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:13:36.386569Z digest=sha256:d91bb998c3725493523f1c6257317445789feb72d027d959f45b88b8e9887393

Observation a96efaeb-f091-42ea-8ad6-0bf0c82b7355 · outbound

This paper cites 3D-SPS: Single- stage 3D visual grounding via referred point progressive se- lection.

3D Scene Graph Guided Vision-Language Pre-training 3D-SPS: Single- stage 3D visual grounding via referred point progressive se- lection

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.252235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.393540Z digest=sha256:f10f2e3b73ef7153ce0995a6473210b263bddc9c0c485cd8c825a507653901ed

Observation daad1910-59d8-48ae-aed0-8cba1f351a42 · outbound

This paper cites Sgformer: Semantic graph transformer for point cloud-based 3D scene graph generation.

3D Scene Graph Guided Vision-Language Pre-training Sgformer: Semantic graph transformer for point cloud-based 3D scene graph generation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.228351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.403331Z digest=sha256:42e28c613baef7f7495e5a1c134263fcb44087c9ea2ffbe6032b693daf23bf25

Observation f53f8afe-8753-45ca-955c-3c2fe18e140f · outbound

This paper cites Heterogeneous graph learning for scene graph prediction in 3d point clouds.

3D Scene Graph Guided Vision-Language Pre-training Heterogeneous graph learning for scene graph prediction in 3d point clouds

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.187707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.409952Z digest=sha256:e9501a35fc2a1df9a1b5c4683595a683c8b6f20410e3cbe15453028815b087b0

Observation f2fef0cf-69cc-43bf-ba2f-1c2b0ffa5488 · outbound

This paper cites Complete 3D relationships extraction modality align- ment network for 3D dense captioning.

3D Scene Graph Guided Vision-Language Pre-training Complete 3D relationships extraction modality align- ment network for 3D dense captioning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.161343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.419632Z digest=sha256:cf54899dffacc4bcbec263c97d342805eb18b034294c6b7ad23857416f41a8ea

Observation 90d7c137-08f4-455e-8419-9246f79f905e · outbound

This paper cites An end-to- end transformer model for 3d object detection.

3D Scene Graph Guided Vision-Language Pre-training An end-to- end transformer model for 3d object detection

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.126738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.427835Z digest=sha256:f282389c4d974f8ee0d41fef8d12e46d827169b82f14fbf9744d47fa1ab8069c

Observation 2f28da01-60c4-4736-bdc0-7dbb32a1651d · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

3D Scene Graph Guided Vision-Language Pre-training Bleu: a method for automatic evaluation of machine translation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.099631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.433566Z digest=sha256:d55168fd7c469146e2e0f30ed15db023e4bfe9d7111b5bc027b80af5d803d779

Observation ec56fba2-5f36-4b63-bb68-0ddc2c83501c · outbound

This paper cites Clip-guided vision-language pre-training for question answering in 3D scenes.

3D Scene Graph Guided Vision-Language Pre-training Clip-guided vision-language pre-training for question answering in 3D scenes

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.069425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.439474Z digest=sha256:e741c7a2e18c885b128deb13bdfa2ba3eee41d627ffe6d5c9dcf86c7086313c5

Observation 742e42e4-e2b4-48bf-abdc-1e281c67cc51 · outbound

This paper cites Pytorch: An im- perative style, high-performance deep learning library.

3D Scene Graph Guided Vision-Language Pre-training Pytorch: An im- perative style, high-performance deep learning library

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.031479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.452029Z digest=sha256:4c20bdd2b93f40ea474b5997b17f56da5b5467cc219d2778dd82a451ce948f35

Observation 945f0afd-6094-4cd5-80b9-154dc8648ea0 · outbound

This paper cites Glove: Global vectors for word representation.

3D Scene Graph Guided Vision-Language Pre-training Glove: Global vectors for word representation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.995167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.468658Z digest=sha256:e454ecf42eb4c806461db8f71d2b45c9d6fa308756ba90a87307811671c83b1c

Observation 2e0cbfa8-531e-4385-8f5f-79eb676a4e87 · outbound

This paper cites PointNet: Deep learning on point sets for 3D classification and segmentation.

3D Scene Graph Guided Vision-Language Pre-training PointNet: Deep learning on point sets for 3D classification and segmentation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.968465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.477527Z digest=sha256:cfb77899419444f9c287197af1d44deacdef6f6f771b3c6f6d657b34f73e0f5a

Observation 8238dd07-accd-4a70-972b-1ffa67a82a56 · outbound

This paper cites PointNet++: Deep hierarchical feature learning on point sets in a metric space.

3D Scene Graph Guided Vision-Language Pre-training PointNet++: Deep hierarchical feature learning on point sets in a metric space

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.936699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.486896Z digest=sha256:e30e58e50a1be56860bd61eabc8b5783c157e08d97f10d258357f4de54669113

Observation a65fc4c1-120f-4427-8298-6b69357c270f · outbound

This paper cites Qi, Or Litany, Kaiming He, and Leonidas J.

3D Scene Graph Guided Vision-Language Pre-training Qi, Or Litany, Kaiming He, and Leonidas J

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.888364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.495368Z digest=sha256:976b8a4ddeb34ed5b14af745cd0bb8a15b3594824a5fb95f4bccd0b22617dd67

Observation 60169241-a941-4d23-95e7-6676998543c8 · outbound

This paper cites Improving language understanding by gen- erative pre-training.

3D Scene Graph Guided Vision-Language Pre-training Improving language understanding by gen- erative pre-training

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T11:13:36.508319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:13:36.508319Z digest=sha256:0704fa2b859e18c2955f737fccecbe82f81fb77be5a58228d56683f0571f7133

Observation beb3ef1a-8106-4f14-9b6b-4425e5b13a0f · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

3D Scene Graph Guided Vision-Language Pre-training Learn- ing transferable visual models from natural language super- vision

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.830923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.521458Z digest=sha256:f76d67061d55b875b6c428c8bb10e0ff39f4e2c5dfcd6bf318a6185ec700b0a4

Observation 9c700dbd-8fea-4789-aab8-76fcf5bc6ca6 · outbound

This paper cites Mask3D: Mask trans- former for 3D semantic instance segmentation.

3D Scene Graph Guided Vision-Language Pre-training Mask3D: Mask trans- former for 3D semantic instance segmentation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.801646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.536173Z digest=sha256:255bce1e7baa82c142fe1eb25841918e8028fa299cee844a9e10621ca7a81a39

Observation f32d1a60-7929-4f51-a895-82975aa858dc · outbound

This paper cites VL-BERT: Pre-training of Generic Visual-Linguistic Representations.

3D Scene Graph Guided Vision-Language Pre-training VL-BERT: Pre-training of Generic Visual-Linguistic Representations

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T11:13:36.545186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:13:36.545186Z digest=sha256:4edbd057f43035de964c1f2b8d00f8403f00241af46cbdbd3b3af1b2e0825e51

Observation 4bbadab6-a8a5-4c20-ae13-f7d8e9a51c01 · outbound

This paper cites Cider: Consensus-based image description evalua- tion.

3D Scene Graph Guided Vision-Language Pre-training Cider: Consensus-based image description evalua- tion

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T11:13:36.554591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:13:36.554591Z digest=sha256:98b600fa8748aaec2df4db59a92a454eff433da1c4f2b90d33181599da24382a

Observation e85005a0-7a23-434d-9c1a-0e7bf4e854a9 · outbound

This paper cites Learning 3D semantic scene graphs from 3D indoor reconstructions.

3D Scene Graph Guided Vision-Language Pre-training Learning 3D semantic scene graphs from 3D indoor reconstructions

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.746357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.565038Z digest=sha256:e0ca24af02e14348c96e113e65ee07a998c1ab778b822ae1837702f44818bd7e

Observation 592bafb9-5d43-4280-9f7b-3e9d4e1d3d66 · outbound

This paper cites Spatiality-guided Transformer for 3D Dense Captioning on Point Clouds.

3D Scene Graph Guided Vision-Language Pre-training Spatiality-guided Transformer for 3D Dense Captioning on Point Clouds

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T11:13:36.575391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:13:36.575391Z digest=sha256:f9a65ab20875c5822bdb81d45af13f4d7acb5f0037b3cb47fd2869372461eb0e

Observation 9f91b276-cfd1-4b6a-8521-a2fbc03c2abf · outbound

This paper cites Vl-sat: visual-linguistic semantics assisted training for 3D semantic scene graph prediction in point cloud.

3D Scene Graph Guided Vision-Language Pre-training Vl-sat: visual-linguistic semantics assisted training for 3D semantic scene graph prediction in point cloud

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.715574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.587490Z digest=sha256:7fe720ebef257a5f092a310142b558dca47de000f6e0f26ff5f1b0e4a8e6c46b

Observation 2336401d-5b95-4b6b-95f0-c4171c7dabad · outbound

This paper cites EDA: Explicit Text-Decoupling and Dense Alignment for 3D Visual Grounding.

3D Scene Graph Guided Vision-Language Pre-training EDA: Explicit Text-Decoupling and Dense Alignment for 3D Visual Grounding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T11:13:36.597094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:13:36.597094Z digest=sha256:473948a9f1fc9eb66aa3ef043b692c3ccceac151ba331d68c165e24a67862047

Observation 81e19616-fc70-4b3a-a2f2-5ca0e21e3f8d · outbound

This paper cites Vision-language pre-training with triple contrastive learning.

3D Scene Graph Guided Vision-Language Pre-training Vision-language pre-training with triple contrastive learning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.685199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.605670Z digest=sha256:b47165c9140ff79e69f753b30a8a6a3c23c967d2bb1386045d76c5c63888d480

Observation 0827a075-662e-4da7-8d6e-f29b3d33139a · outbound

This paper cites Sat: 2D semantics assisted training for 3D visual grounding.

3D Scene Graph Guided Vision-Language Pre-training Sat: 2D semantics assisted training for 3D visual grounding

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.652600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.617166Z digest=sha256:33fffd9f7eb08850f7fa0edf4d012cd393c25bcb31203c241861545037d524b8

Observation d319560b-93fa-44d1-95ff-4b01c1fe1c08 · outbound

This paper cites Deep modular co-attention networks for visual question an- swering.

3D Scene Graph Guided Vision-Language Pre-training Deep modular co-attention networks for visual question an- swering

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.624968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.623795Z digest=sha256:ef024dfe3d3ec92475d4832458c93fe3ed1c1683697e8a55e44c5dc21568f27f

Observation f8d960d7-c1ee-41cf-8919-c7a6636a089c · outbound

This paper cites Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual refer- ring.

3D Scene Graph Guided Vision-Language Pre-training Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual refer- ring

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.593057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.631511Z digest=sha256:056154a69883ffcb1d24c945e4970d83d9774d60fb936802236b7376c37ddc49

Observation 81c79412-2899-4c42-93f5-6d4138ddf27d · outbound

This paper cites X-trans2cap: Cross-modal knowledge transfer using transformer for 3D dense caption- ing.

3D Scene Graph Guided Vision-Language Pre-training X-trans2cap: Cross-modal knowledge transfer using transformer for 3D dense caption- ing

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.554651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.639579Z digest=sha256:aec54272e919e93322e550317fcd1bcbddda313b60a6fa4a55d8b1fc76a33cf0

Observation 018b2df1-9eb0-4195-bf05-63c3cbfa76a9 · outbound

This paper cites Exploiting edge-oriented reasoning for 3D point-based scene graph analysis.

3D Scene Graph Guided Vision-Language Pre-training Exploiting edge-oriented reasoning for 3D point-based scene graph analysis

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.522722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.652096Z digest=sha256:5a312e1bbc29001c8c268a29a4bedd61f397603e8a6d182f577d10c610fd2a16

Observation 3f647af2-0f9f-46d1-bbea-603f5937e669 · outbound

This paper cites Pointclip: Point cloud understanding by clip.

3D Scene Graph Guided Vision-Language Pre-training Pointclip: Point cloud understanding by clip

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T11:13:36.670122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:13:36.670122Z digest=sha256:4214c72a7c954768b92df134d347d4fac5a32d7b64bc3a5f267d7ba6829f20a9

Observation 667b6665-a768-4f13-8b0c-b287402c671c · outbound

This paper cites Vision-language pre-training with object con- trastive learning for 3D scene understanding.

3D Scene Graph Guided Vision-Language Pre-training Vision-language pre-training with object con- trastive learning for 3D scene understanding

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.450236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.681011Z digest=sha256:2c54b54e9261ad7bdb3244bcc12e95eb36e5b94f76a6ae47c0929346e5198081

Observation c6572169-f869-47a6-b384-75e071c1701d · outbound

This paper cites 3DVG- Transformer: Relation modeling for visual grounding on point clouds.

3D Scene Graph Guided Vision-Language Pre-training 3DVG- Transformer: Relation modeling for visual grounding on point clouds

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.421331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.687868Z digest=sha256:32d4b8710b48f5a35d78dd9e7fcb7058c03bc4ec600fcd2802d814dd285600bd

Observation 83364dbb-1be1-4fd3-a251-2fbb110d3df6 · outbound

This paper cites Towards explainable 3D grounded visual question answer- ing: A new benchmark and strong baseline.

3D Scene Graph Guided Vision-Language Pre-training Towards explainable 3D grounded visual question answer- ing: A new benchmark and strong baseline

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.382388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.700903Z digest=sha256:1b257f7b05b9f05f6e7a0cca8f0ab54f009ba00a73e8fa9fc11b1d1ea0145356

Observation effb4dd5-c393-4efc-b84e-89a983b5065d · outbound

This paper cites Contextual Modeling for 3D Dense Captioning on Point Clouds.

3D Scene Graph Guided Vision-Language Pre-training Contextual Modeling for 3D Dense Captioning on Point Clouds

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T11:13:36.711134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:13:36.711134Z digest=sha256:3c2e76d616f34e50460ec4dfd034d7bdb929d5d025e39de16521e4109ead75a5

Observation 9ee57161-b2fc-41df-826c-c257492387c4 · outbound

This paper cites Point- clip v2: Prompting clip and gpt for powerful 3D open-world learning.

3D Scene Graph Guided Vision-Language Pre-training Point- clip v2: Prompting clip and gpt for powerful 3D open-world learning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.343863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.722077Z digest=sha256:1d4bc51c203345c5612710998f69cc098850f997cb20e2a1093cb87002c7deab

Observation 7db195e7-147b-4288-8c5a-e034a489808c · outbound

This paper cites 3d-vista: Pre-trained transformer for 3D vision and text alignment.

3D Scene Graph Guided Vision-Language Pre-training 3d-vista: Pre-trained transformer for 3D vision and text alignment

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.313045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.729314Z digest=sha256:b6f86a8b004d3899e2dc3e91a34e14245174319cccc312c0133628020eaafe1c

Observation 6251aefd-8e42-42d6-aaf2-bfdafc7408d0 · outbound

This paper cites an unresolved cited work.

3D Scene Graph Guided Vision-Language Pre-training Unresolved cited work

Reference 545

Resolution
parse uncertain
no resolver link, observed 2026-08-12T11:13:36.344522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:13:36.344522Z digest=sha256:033a0af483663cd43328483c04dda98770ab57bde2b982872a2fbc987e007bfe

Pith citing papers

No inbound Pith citation observations are available.