Pith. sign in

Paper Citation Record · LEDGER

3D Scene Graph Guided Vision-Language Pre-training

As of 18 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2411.18666.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.18666 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:13:36.729314Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

63 of 63 outbound references displayed

  • verified exact2
  • verified fuzzy45
  • unresolved15
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d904d31c-8ce4-418e-9f95-364ff0aa5bd3 · outbound

This paper cites Referit3D: Neural listeners for fine-grained 3D object identification in real-world scenes.

3D Scene Graph Guided Vision-Language Pre-training Referit3D: Neural listeners for fine-grained 3D object identification in real-world scenes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.941437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.067631Z digest=sha256:a61aed95415204596be14b02e46207202af338a035949372c1a16e87199c7fda

Observation be7fb37e-cab9-46f4-9462-15f2ac409ff0 · outbound

This paper cites Scanqa: 3D question answering for spatial scene understanding.

3D Scene Graph Guided Vision-Language Pre-training Scanqa: 3D question answering for spatial scene understanding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.911304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.075948Z digest=sha256:3817e83ab369fb1712edb107916895e183c977d50f58b2b5701fa5459affdc90

Observation 9256b3a2-a057-4bea-9edd-a959b4aeb8bf · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments.

3D Scene Graph Guided Vision-Language Pre-training Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.879255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.090937Z digest=sha256:ac7c414c50364b5734cce85c06d1957f9b33de682784b81804dee975bf8020f0

Observation e1a35ee1-bbbc-4f12-ad66-ecba3a64f95b · outbound

This paper cites 3DJCG: A unified framework for joint dense captioning and visual grounding on 3D point clouds.

3D Scene Graph Guided Vision-Language Pre-training 3DJCG: A unified framework for joint dense captioning and visual grounding on 3D point clouds

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.855501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.098706Z digest=sha256:0a6d5552773294d40e5dd3dc9ca475e2e084e5297fe8293f85fbda5ef530a776

Observation 54ee0b72-a096-4a8e-a3a1-09e805387cb3 · outbound

This paper cites Scanrefer: 3D object localization in RGB-D scans using natural language.

3D Scene Graph Guided Vision-Language Pre-training Scanrefer: 3D object localization in RGB-D scans using natural language

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.828887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.105899Z digest=sha256:d4b60385f3cfc33591e4f216f3c400cc362c175e3ac43b3a11f7fc9b0bf7ae27

Observation e0553d85-eb06-498a-a401-b7fe7060acf6 · outbound

This paper cites UniT3D: A Unified Transformer for 3D Dense Captioning and Visual Grounding.

3D Scene Graph Guided Vision-Language Pre-training UniT3D: A Unified Transformer for 3D Dense Captioning and Visual Grounding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T11:13:36.119765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:13:36.119765Z digest=sha256:91b8ce2d079672ea7db67a0ffa4f4f3a1ff8da65b1a315899145bb861f4a4a58

Observation 482814b2-9443-4391-b56c-45ebf84c4f28 · outbound

This paper cites D 3net: A unified speaker-listener architec- ture for 3D dense captioning and visual grounding.

3D Scene Graph Guided Vision-Language Pre-training D 3net: A unified speaker-listener architec- ture for 3D dense captioning and visual grounding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.787228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.132167Z digest=sha256:1003e4dd638419d679a26cdcf40b3928ec060d1a2732e2a23e7767433b4b99a1

Observation cdc9c32c-7958-482a-8728-3b504033a28e · outbound

This paper cites Language Conditioned Spatial Relation Reasoning for 3D Object Grounding.

3D Scene Graph Guided Vision-Language Pre-training Language Conditioned Spatial Relation Reasoning for 3D Object Grounding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T11:13:36.142026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:13:36.142026Z digest=sha256:0f02fbc1abc4595eb72f73b5d4e53e9b9d3f87697180fc1226ee81b77a6ec8c7

Observation 645089ad-5419-4f2b-8c07-12326b1c8259 · outbound

This paper cites End-to-End 3D Dense Captioning with Vote2Cap-DETR.

3D Scene Graph Guided Vision-Language Pre-training End-to-End 3D Dense Captioning with Vote2Cap-DETR

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-12T11:13:37.175474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.148543Z digest=sha256:45791d507c6840f2f1258d85fdeb701f6725557276249bae16b04949ee984e49

Observation 1eae4423-c0f7-440f-8270-d92e8eede3d9 · outbound

This paper cites Simclr: A simple framework for contrastive learning of visual representations.

3D Scene Graph Guided Vision-Language Pre-training Simclr: A simple framework for contrastive learning of visual representations

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.748738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.164142Z digest=sha256:a795d94d4cb28289e23fe39ac333394b16b1e655b474667a711144979d1039a0

Observation 9c04a5c1-edd0-450b-9aab-bb5b3a831654 · outbound

This paper cites Scan2cap: Context-aware dense captioning in RGB- D scans.

3D Scene Graph Guided Vision-Language Pre-training Scan2cap: Context-aware dense captioning in RGB- D scans

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.717928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.177460Z digest=sha256:1bf077c17182a4ac071aa1557a5df4ff680309e1f2170e52aa05e8aba7c2262f

Observation 06d9a5d4-d9d6-4a16-a87b-242937f7b853 · outbound

This paper cites Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling.

3D Scene Graph Guided Vision-Language Pre-training Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T11:13:36.186265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:13:36.186265Z digest=sha256:675552f0ae99e84ceb52fc2e84fedaef59de50e79a88636bd0d2d6cb04fe6282

Observation 2ca348ef-f9d4-4bee-a0fb-a2895c547056 · outbound

This paper cites Scannet: Richly-annotated 3D reconstructions of indoor scenes.

3D Scene Graph Guided Vision-Language Pre-training Scannet: Richly-annotated 3D reconstructions of indoor scenes

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.688778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.192424Z digest=sha256:2205da78be04b03c8533d402dd79793f9ca2fc0d8477ebff6a7f70b2ed175bf8

Observation 6148387c-a619-43d0-b9ce-9d56e1a5d581 · outbound

This paper cites Multi-CLIP: Contrastive Vision-Language Pre-training for Question Answering tasks in 3D Scenes.

3D Scene Graph Guided Vision-Language Pre-training Multi-CLIP: Contrastive Vision-Language Pre-training for Question Answering tasks in 3D Scenes

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T11:13:36.201254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:13:36.201254Z digest=sha256:0aa9bcafc459ddcc9248b7e4f46511e6f89f3e5fdf26629177c220d8ac498218

Observation 6e037ebd-0304-4fea-be66-fe391167025e · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

3D Scene Graph Guided Vision-Language Pre-training BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T11:13:36.208235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:13:36.208235Z digest=sha256:b5d2acdae127db58b18b457ef2ab3dd6ff9cb5309a4600962afe21536799b292

Observation 46c911d3-3f83-4605-b491-89e16ab45a50 · outbound

This paper cites Multi-modal align- ment using representation codebook.

3D Scene Graph Guided Vision-Language Pre-training Multi-modal align- ment using representation codebook

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.661247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.216194Z digest=sha256:f5096b89f7bfcd83246bd894833617653fb988ddd8657442b2f0eb73debc6bdc

Observation 35ecd526-1ac5-4141-9a44-49c7922836c0 · outbound

This paper cites Free-form description guided 3D visual graph network for object grounding in point cloud.

3D Scene Graph Guided Vision-Language Pre-training Free-form description guided 3D visual graph network for object grounding in point cloud

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.636225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.227686Z digest=sha256:897b6d32b939e180c6292d42fc1917560bbfd9628dc6ed02fb3d53ad090b6810

Observation 89e366f2-9e98-47b9-922b-6467d7e3e23c · outbound

This paper cites ViewRefer: Grasp the Multi-view Knowledge for 3D Visual Grounding with GPT and Prototype Guidance.

3D Scene Graph Guided Vision-Language Pre-training ViewRefer: Grasp the Multi-view Knowledge for 3D Visual Grounding with GPT and Prototype Guidance

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T11:13:36.245100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:13:36.245100Z digest=sha256:0adeca4f75e1140eb5633eee4570e49970d2ce2371f5a5d92feacdf55f931de3

Observation 0a7836a0-ad93-4bc4-9554-f6f8d85d3038 · outbound

This paper cites Masked autoencoders are scalable vision learners.

3D Scene Graph Guided Vision-Language Pre-training Masked autoencoders are scalable vision learners

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T11:13:36.264686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:13:36.264686Z digest=sha256:b03347bd3cbe8f3ff649abf72084bfc538d717ee12f95f1a6f2d2cb533f2b7e6

Observation 3b84bbb6-3d71-4fbc-b32b-defdf2e8dc1d · outbound

This paper cites Text-guided graph neural networks for refer- ring 3D instance segmentation.

3D Scene Graph Guided Vision-Language Pre-training Text-guided graph neural networks for refer- ring 3D instance segmentation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.559829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.275054Z digest=sha256:7839bce58506e075dbf330e7b97e273f9f031df2fdc4733943728eba86b9c2c7

Observation d703c403-646a-40e8-b776-484f3a84121a · outbound

This paper cites Multi- view transformer for 3D visual grounding.

3D Scene Graph Guided Vision-Language Pre-training Multi- view transformer for 3D visual grounding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.519069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.290301Z digest=sha256:39c77c039810a9d8a0f7577e347e568df11671e0cc258e3033b2d390ad3034cb

Observation dc79bff9-a612-41e8-806b-6d6231baab9e · outbound

This paper cites Clip2point: Transfer clip to point cloud classifica- tion with image-depth pre-training.

3D Scene Graph Guided Vision-Language Pre-training Clip2point: Transfer clip to point cloud classifica- tion with image-depth pre-training

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.488174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.299677Z digest=sha256:acad5c2c57e7506fe8fe824a3c1ec11dc5adadb524591137fae5dda3e5ef28b1

Observation 80f76758-69ec-477c-a06e-b42c95b9799a · outbound

This paper cites Bottom up top down detection transform- ers for language grounding in images and point clouds.

3D Scene Graph Guided Vision-Language Pre-training Bottom up top down detection transform- ers for language grounding in images and point clouds

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.454472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.307834Z digest=sha256:04e31768c2001feaa2e841674ed32f5c4c1517c729651f74356acd480e69ede6

Observation 3d2c4fcc-275b-4085-85de-12e4066ecb49 · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

3D Scene Graph Guided Vision-Language Pre-training Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.414822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.316408Z digest=sha256:19a2ee86a16a7de938b9890ea58356512d65e8cb1390cd3afd76ae5ccd49a0f5

Observation d5ce4dcf-1918-435f-a60f-b404bfceaabc · outbound

This paper cites More: Multi-order relation mining for dense captioning in 3D scenes.

3D Scene Graph Guided Vision-Language Pre-training More: Multi-order relation mining for dense captioning in 3D scenes

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.384467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.329853Z digest=sha256:4cbf1cb33da9fcd010de2aa7138f48daaf74b1ffa5dd188825daa46483be54d3

Observation f806a9a1-4cc2-4bae-b76e-e3f58be63352 · outbound

This paper cites Context-aware alignment and mutual masking for 3D-language pre-training.

3D Scene Graph Guided Vision-Language Pre-training Context-aware alignment and mutual masking for 3D-language pre-training

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.317360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.363821Z digest=sha256:23894e99e143179f22970ba37835f84aac32e804cbabc5de0a5130325190739d

Observation d952bcfe-9492-4d7d-8cd3-870011cc3b07 · outbound

This paper cites Lang3DSG: Language-based contrastive pre-training for 3D Scene Graph prediction.

3D Scene Graph Guided Vision-Language Pre-training Lang3DSG: Language-based contrastive pre-training for 3D Scene Graph prediction

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-12T11:13:36.965742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.375827Z digest=sha256:94ece84a62f05b40aa4210ef1f70426737c5edbab2f57d836fcc709d3a9fbd4a

Observation 93031e7c-7f48-4ca6-8c0a-6c4c58a0c5b5 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

3D Scene Graph Guided Vision-Language Pre-training Rouge: A package for automatic evaluation of summaries

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T11:13:36.386569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:13:36.386569Z digest=sha256:d5a39d84aeac5f57ddbfc05a467c972d9682659125838d7522f17d7bca0a146d

Observation a96efaeb-f091-42ea-8ad6-0bf0c82b7355 · outbound

This paper cites 3D-SPS: Single- stage 3D visual grounding via referred point progressive se- lection.

3D Scene Graph Guided Vision-Language Pre-training 3D-SPS: Single- stage 3D visual grounding via referred point progressive se- lection

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.252235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.393540Z digest=sha256:ddae7cb00fc62046a2151086fc7e40589b324c773268d80c9ad4e2b7be88f135

Observation daad1910-59d8-48ae-aed0-8cba1f351a42 · outbound

This paper cites Sgformer: Semantic graph transformer for point cloud-based 3D scene graph generation.

3D Scene Graph Guided Vision-Language Pre-training Sgformer: Semantic graph transformer for point cloud-based 3D scene graph generation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.228351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.403331Z digest=sha256:3db9c770100ab0fa9b5b725f0c69801012f74f75725e674b67f05fa9481a92fe

Observation f53f8afe-8753-45ca-955c-3c2fe18e140f · outbound

This paper cites Heterogeneous graph learning for scene graph prediction in 3d point clouds.

3D Scene Graph Guided Vision-Language Pre-training Heterogeneous graph learning for scene graph prediction in 3d point clouds

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.187707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.409952Z digest=sha256:23d3918171431cf5bf5cb74c70d90027f2bde959f67e9740d6dc743a8023c66a

Observation f2fef0cf-69cc-43bf-ba2f-1c2b0ffa5488 · outbound

This paper cites Complete 3D relationships extraction modality align- ment network for 3D dense captioning.

3D Scene Graph Guided Vision-Language Pre-training Complete 3D relationships extraction modality align- ment network for 3D dense captioning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.161343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.419632Z digest=sha256:cc56798638a87b68dec03931e77f23f7d7bfc31c3760efe6bd7cdf23cefc5fa1

Observation 90d7c137-08f4-455e-8419-9246f79f905e · outbound

This paper cites An end-to- end transformer model for 3d object detection.

3D Scene Graph Guided Vision-Language Pre-training An end-to- end transformer model for 3d object detection

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.126738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.427835Z digest=sha256:362b4cf88f48e08d466d1457c83b8a2d846f9814edf85390dc1e7f0a967d4b87

Observation 2f28da01-60c4-4736-bdc0-7dbb32a1651d · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

3D Scene Graph Guided Vision-Language Pre-training Bleu: a method for automatic evaluation of machine translation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.099631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.433566Z digest=sha256:6ee130e648506f7fc05228fdad9e7a94fac640cff9e851d876bdcb3d4444c855

Observation ec56fba2-5f36-4b63-bb68-0ddc2c83501c · outbound

This paper cites Clip-guided vision-language pre-training for question answering in 3D scenes.

3D Scene Graph Guided Vision-Language Pre-training Clip-guided vision-language pre-training for question answering in 3D scenes

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.069425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.439474Z digest=sha256:0675265037be5d94e9e17becfc98aa2dcec152b833e910cabe52a3b622a2fe7b

Observation 742e42e4-e2b4-48bf-abdc-1e281c67cc51 · outbound

This paper cites Pytorch: An im- perative style, high-performance deep learning library.

3D Scene Graph Guided Vision-Language Pre-training Pytorch: An im- perative style, high-performance deep learning library

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:38.031479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.452029Z digest=sha256:bcd3da130c524c6550c9f760609e6e973c60128008c035d7a0d0058322c496ea

Observation 945f0afd-6094-4cd5-80b9-154dc8648ea0 · outbound

This paper cites Glove: Global vectors for word representation.

3D Scene Graph Guided Vision-Language Pre-training Glove: Global vectors for word representation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.995167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.468658Z digest=sha256:d327d2443d66fccd4f533ca6fdb7f49462d63e9e5c80d879ad390ad501b49738

Observation 2e0cbfa8-531e-4385-8f5f-79eb676a4e87 · outbound

This paper cites PointNet: Deep learning on point sets for 3D classification and segmentation.

3D Scene Graph Guided Vision-Language Pre-training PointNet: Deep learning on point sets for 3D classification and segmentation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.968465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.477527Z digest=sha256:f23554ea3cf3f6b1e9c36c90c086ca16a66645ffd3723bc066d504dc1c8ba421

Observation 8238dd07-accd-4a70-972b-1ffa67a82a56 · outbound

This paper cites PointNet++: Deep hierarchical feature learning on point sets in a metric space.

3D Scene Graph Guided Vision-Language Pre-training PointNet++: Deep hierarchical feature learning on point sets in a metric space

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.936699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.486896Z digest=sha256:98998fbe837ce49a4c20eaf76d825db711264da3c60103c954bedcdc05e2483f

Observation a65fc4c1-120f-4427-8298-6b69357c270f · outbound

This paper cites Qi, Or Litany, Kaiming He, and Leonidas J.

3D Scene Graph Guided Vision-Language Pre-training Qi, Or Litany, Kaiming He, and Leonidas J

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.888364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.495368Z digest=sha256:b659c65f3671d03ccce4731175287c4454e52425706525bf7e2e6797d0b8b970

Observation 60169241-a941-4d23-95e7-6676998543c8 · outbound

This paper cites Improving language understanding by gen- erative pre-training.

3D Scene Graph Guided Vision-Language Pre-training Improving language understanding by gen- erative pre-training

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T11:13:36.508319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:13:36.508319Z digest=sha256:12f1c9af4a32696fc29d2ab90b32e31c3c072f8e3cc794ff45825c5f46482d00

Observation beb3ef1a-8106-4f14-9b6b-4425e5b13a0f · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

3D Scene Graph Guided Vision-Language Pre-training Learn- ing transferable visual models from natural language super- vision

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.830923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.521458Z digest=sha256:c5f07685b778900605dec285ef552613eb864248dc102b04eef1694e494d9928

Observation 9c700dbd-8fea-4789-aab8-76fcf5bc6ca6 · outbound

This paper cites Mask3D: Mask trans- former for 3D semantic instance segmentation.

3D Scene Graph Guided Vision-Language Pre-training Mask3D: Mask trans- former for 3D semantic instance segmentation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.801646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.536173Z digest=sha256:6bba9390fe0c4ef8a18706bd389c6b92d0a5966e5e867796177f0e98d21c92f3

Observation f32d1a60-7929-4f51-a895-82975aa858dc · outbound

This paper cites VL-BERT: Pre-training of Generic Visual-Linguistic Representations.

3D Scene Graph Guided Vision-Language Pre-training VL-BERT: Pre-training of Generic Visual-Linguistic Representations

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T11:13:36.545186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:13:36.545186Z digest=sha256:28ef02819a70c1f3a07f8f57ef1bad10b52a4687dcfdefe6850c276467aabb2c

Observation 4bbadab6-a8a5-4c20-ae13-f7d8e9a51c01 · outbound

This paper cites Cider: Consensus-based image description evalua- tion.

3D Scene Graph Guided Vision-Language Pre-training Cider: Consensus-based image description evalua- tion

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T11:13:36.554591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:13:36.554591Z digest=sha256:486c0f03e9511fc3febc06e86dbaa43272339913011aafe051a60f062e1d6bb1

Observation e85005a0-7a23-434d-9c1a-0e7bf4e854a9 · outbound

This paper cites Learning 3D semantic scene graphs from 3D indoor reconstructions.

3D Scene Graph Guided Vision-Language Pre-training Learning 3D semantic scene graphs from 3D indoor reconstructions

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.746357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.565038Z digest=sha256:87a672d9fee22e163936457cd473af39b63c7c516889b34c0cd5a0d91657797e

Observation 592bafb9-5d43-4280-9f7b-3e9d4e1d3d66 · outbound

This paper cites Spatiality-guided Transformer for 3D Dense Captioning on Point Clouds.

3D Scene Graph Guided Vision-Language Pre-training Spatiality-guided Transformer for 3D Dense Captioning on Point Clouds

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T11:13:36.575391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:13:36.575391Z digest=sha256:b8b3a327c952efe20fb4149d887ed31993d97efe33f82a74c957488db0b93dd4

Observation 9f91b276-cfd1-4b6a-8521-a2fbc03c2abf · outbound

This paper cites Vl-sat: visual-linguistic semantics assisted training for 3D semantic scene graph prediction in point cloud.

3D Scene Graph Guided Vision-Language Pre-training Vl-sat: visual-linguistic semantics assisted training for 3D semantic scene graph prediction in point cloud

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.715574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.587490Z digest=sha256:291bec7c84864b8acadfebe3dd9a91fb0918aac1ce002e8f019a7ca38205769f

Observation 2336401d-5b95-4b6b-95f0-c4171c7dabad · outbound

This paper cites EDA: Explicit Text-Decoupling and Dense Alignment for 3D Visual Grounding.

3D Scene Graph Guided Vision-Language Pre-training EDA: Explicit Text-Decoupling and Dense Alignment for 3D Visual Grounding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T11:13:36.597094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:13:36.597094Z digest=sha256:9dba317a1b32a32d6036dd234c8640d245635c8484e93674c90005ddbeb91a4c

Observation 81e19616-fc70-4b3a-a2f2-5ca0e21e3f8d · outbound

This paper cites Vision-language pre-training with triple contrastive learning.

3D Scene Graph Guided Vision-Language Pre-training Vision-language pre-training with triple contrastive learning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.685199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.605670Z digest=sha256:e72adfb255ed8bf196f5d4ca8663fe28274b27f13d02a13cab27bb3ecde86351

Observation 0827a075-662e-4da7-8d6e-f29b3d33139a · outbound

This paper cites Sat: 2D semantics assisted training for 3D visual grounding.

3D Scene Graph Guided Vision-Language Pre-training Sat: 2D semantics assisted training for 3D visual grounding

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.652600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.617166Z digest=sha256:15dc95c6cdbe4277934f1a336d44f16f1d995f29b7f967795b5324b6838f7735

Observation d319560b-93fa-44d1-95ff-4b01c1fe1c08 · outbound

This paper cites Deep modular co-attention networks for visual question an- swering.

3D Scene Graph Guided Vision-Language Pre-training Deep modular co-attention networks for visual question an- swering

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.624968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.623795Z digest=sha256:6f856e1cd359564c183fd1266b52455f09d489af85ba12127f1772d6052bab50

Observation f8d960d7-c1ee-41cf-8919-c7a6636a089c · outbound

This paper cites Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual refer- ring.

3D Scene Graph Guided Vision-Language Pre-training Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual refer- ring

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.593057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.631511Z digest=sha256:e4fb7c799043c67b1b8430a80fb3400f63e1d53f68cb910deb9b245dd045992f

Observation 81c79412-2899-4c42-93f5-6d4138ddf27d · outbound

This paper cites X-trans2cap: Cross-modal knowledge transfer using transformer for 3D dense caption- ing.

3D Scene Graph Guided Vision-Language Pre-training X-trans2cap: Cross-modal knowledge transfer using transformer for 3D dense caption- ing

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.554651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.639579Z digest=sha256:a96b5b946d474d875b2b8c3425762c493a215394d3ca295adf72260645be5285

Observation 018b2df1-9eb0-4195-bf05-63c3cbfa76a9 · outbound

This paper cites Exploiting edge-oriented reasoning for 3D point-based scene graph analysis.

3D Scene Graph Guided Vision-Language Pre-training Exploiting edge-oriented reasoning for 3D point-based scene graph analysis

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.522722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.652096Z digest=sha256:0cd3f1fb2f3bb9f3e0ed95780127a179016fbe33852dae051d294c8f0fe7e59d

Observation 3f647af2-0f9f-46d1-bbea-603f5937e669 · outbound

This paper cites Pointclip: Point cloud understanding by clip.

3D Scene Graph Guided Vision-Language Pre-training Pointclip: Point cloud understanding by clip

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T11:13:36.670122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:13:36.670122Z digest=sha256:699791f00203fbd6e11015efdfaa85b87cb55daa94d3536dd8831dfd1c5c8fc0

Observation 667b6665-a768-4f13-8b0c-b287402c671c · outbound

This paper cites Vision-language pre-training with object con- trastive learning for 3D scene understanding.

3D Scene Graph Guided Vision-Language Pre-training Vision-language pre-training with object con- trastive learning for 3D scene understanding

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.450236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.681011Z digest=sha256:867da8e9bef910711bfdb370319ef4db9ccca6b7a814f9b8e09cbc19ec8ebdb7

Observation c6572169-f869-47a6-b384-75e071c1701d · outbound

This paper cites 3DVG- Transformer: Relation modeling for visual grounding on point clouds.

3D Scene Graph Guided Vision-Language Pre-training 3DVG- Transformer: Relation modeling for visual grounding on point clouds

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.421331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.687868Z digest=sha256:4bee71f91785f6d29c14c932059482e57320e452d0365a717ff62df33908a219

Observation 83364dbb-1be1-4fd3-a251-2fbb110d3df6 · outbound

This paper cites Towards explainable 3D grounded visual question answer- ing: A new benchmark and strong baseline.

3D Scene Graph Guided Vision-Language Pre-training Towards explainable 3D grounded visual question answer- ing: A new benchmark and strong baseline

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.382388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.700903Z digest=sha256:7a0eb9b76e60779fe05e9b7f850609bcba9bf56a82d0f99dee57e60f437d1aef

Observation effb4dd5-c393-4efc-b84e-89a983b5065d · outbound

This paper cites Contextual Modeling for 3D Dense Captioning on Point Clouds.

3D Scene Graph Guided Vision-Language Pre-training Contextual Modeling for 3D Dense Captioning on Point Clouds

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T11:13:36.711134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:13:36.711134Z digest=sha256:a8a5ac48fe55a64a193ef887ac7992729947cc2c53004a872a851887938dfa05

Observation 9ee57161-b2fc-41df-826c-c257492387c4 · outbound

This paper cites Point- clip v2: Prompting clip and gpt for powerful 3D open-world learning.

3D Scene Graph Guided Vision-Language Pre-training Point- clip v2: Prompting clip and gpt for powerful 3D open-world learning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.343863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.722077Z digest=sha256:5d22c5394d1f188dc7b124a6293edd671192e9e76600806e0b81e03c02996374

Observation 7db195e7-147b-4288-8c5a-e034a489808c · outbound

This paper cites 3d-vista: Pre-trained transformer for 3D vision and text alignment.

3D Scene Graph Guided Vision-Language Pre-training 3d-vista: Pre-trained transformer for 3D vision and text alignment

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:13:37.313045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:13:36.729314Z digest=sha256:1cee801886e14a50f71d7e91e9518d660468051f8b57af4650a700e33deaa52e

Observation 6251aefd-8e42-42d6-aaf2-bfdafc7408d0 · outbound

This paper cites an unresolved cited work.

3D Scene Graph Guided Vision-Language Pre-training Unresolved cited work

Reference 545

Resolution
parse uncertain
no resolver link, observed 2026-08-12T11:13:36.344522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:13:36.344522Z digest=sha256:54ea69ce0758c59fd252924d897308f4d0c4d79ca01690838ad6ada1d5e6089e

Pith citing papers

No inbound Pith citation observations are available.