Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T11:13:36.729314Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2411.18666.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T11:13:36.729314Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
63 of 63 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d904d31c-8ce4-418e-9f95-364ff0aa5bd3 · outbound
3D Scene Graph Guided Vision-Language Pre-training Referit3D: Neural listeners for fine-grained 3D object identification in real-world scenes
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation be7fb37e-cab9-46f4-9462-15f2ac409ff0 · outbound
3D Scene Graph Guided Vision-Language Pre-training Scanqa: 3D question answering for spatial scene understanding
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9256b3a2-a057-4bea-9edd-a959b4aeb8bf · outbound
3D Scene Graph Guided Vision-Language Pre-training Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e1a35ee1-bbbc-4f12-ad66-ecba3a64f95b · outbound
3D Scene Graph Guided Vision-Language Pre-training 3DJCG: A unified framework for joint dense captioning and visual grounding on 3D point clouds
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 54ee0b72-a096-4a8e-a3a1-09e805387cb3 · outbound
3D Scene Graph Guided Vision-Language Pre-training Scanrefer: 3D object localization in RGB-D scans using natural language
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e0553d85-eb06-498a-a401-b7fe7060acf6 · outbound
3D Scene Graph Guided Vision-Language Pre-training UniT3D: A Unified Transformer for 3D Dense Captioning and Visual Grounding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 482814b2-9443-4391-b56c-45ebf84c4f28 · outbound
3D Scene Graph Guided Vision-Language Pre-training D 3net: A unified speaker-listener architec- ture for 3D dense captioning and visual grounding
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cdc9c32c-7958-482a-8728-3b504033a28e · outbound
3D Scene Graph Guided Vision-Language Pre-training Language Conditioned Spatial Relation Reasoning for 3D Object Grounding
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 645089ad-5419-4f2b-8c07-12326b1c8259 · outbound
3D Scene Graph Guided Vision-Language Pre-training End-to-End 3D Dense Captioning with Vote2Cap-DETR
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1eae4423-c0f7-440f-8270-d92e8eede3d9 · outbound
3D Scene Graph Guided Vision-Language Pre-training Simclr: A simple framework for contrastive learning of visual representations
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9c04a5c1-edd0-450b-9aab-bb5b3a831654 · outbound
3D Scene Graph Guided Vision-Language Pre-training Scan2cap: Context-aware dense captioning in RGB- D scans
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 06d9a5d4-d9d6-4a16-a87b-242937f7b853 · outbound
3D Scene Graph Guided Vision-Language Pre-training Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ca348ef-f9d4-4bee-a0fb-a2895c547056 · outbound
3D Scene Graph Guided Vision-Language Pre-training Scannet: Richly-annotated 3D reconstructions of indoor scenes
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6148387c-a619-43d0-b9ce-9d56e1a5d581 · outbound
3D Scene Graph Guided Vision-Language Pre-training Multi-CLIP: Contrastive Vision-Language Pre-training for Question Answering tasks in 3D Scenes
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e037ebd-0304-4fea-be66-fe391167025e · outbound
3D Scene Graph Guided Vision-Language Pre-training BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46c911d3-3f83-4605-b491-89e16ab45a50 · outbound
3D Scene Graph Guided Vision-Language Pre-training Multi-modal align- ment using representation codebook
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 35ecd526-1ac5-4141-9a44-49c7922836c0 · outbound
3D Scene Graph Guided Vision-Language Pre-training Free-form description guided 3D visual graph network for object grounding in point cloud
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 89e366f2-9e98-47b9-922b-6467d7e3e23c · outbound
3D Scene Graph Guided Vision-Language Pre-training ViewRefer: Grasp the Multi-view Knowledge for 3D Visual Grounding with GPT and Prototype Guidance
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a7836a0-ad93-4bc4-9554-f6f8d85d3038 · outbound
3D Scene Graph Guided Vision-Language Pre-training Masked autoencoders are scalable vision learners
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b84bbb6-3d71-4fbc-b32b-defdf2e8dc1d · outbound
3D Scene Graph Guided Vision-Language Pre-training Text-guided graph neural networks for refer- ring 3D instance segmentation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d703c403-646a-40e8-b776-484f3a84121a · outbound
3D Scene Graph Guided Vision-Language Pre-training Multi- view transformer for 3D visual grounding
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation dc79bff9-a612-41e8-806b-6d6231baab9e · outbound
3D Scene Graph Guided Vision-Language Pre-training Clip2point: Transfer clip to point cloud classifica- tion with image-depth pre-training
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 80f76758-69ec-477c-a06e-b42c95b9799a · outbound
3D Scene Graph Guided Vision-Language Pre-training Bottom up top down detection transform- ers for language grounding in images and point clouds
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3d2c4fcc-275b-4085-85de-12e4066ecb49 · outbound
3D Scene Graph Guided Vision-Language Pre-training Scaling up visual and vision-language representa- tion learning with noisy text supervision
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d5ce4dcf-1918-435f-a60f-b404bfceaabc · outbound
3D Scene Graph Guided Vision-Language Pre-training More: Multi-order relation mining for dense captioning in 3D scenes
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f806a9a1-4cc2-4bae-b76e-e3f58be63352 · outbound
3D Scene Graph Guided Vision-Language Pre-training Context-aware alignment and mutual masking for 3D-language pre-training
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d952bcfe-9492-4d7d-8cd3-870011cc3b07 · outbound
3D Scene Graph Guided Vision-Language Pre-training Lang3DSG: Language-based contrastive pre-training for 3D Scene Graph prediction
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 93031e7c-7f48-4ca6-8c0a-6c4c58a0c5b5 · outbound
3D Scene Graph Guided Vision-Language Pre-training Rouge: A package for automatic evaluation of summaries
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a96efaeb-f091-42ea-8ad6-0bf0c82b7355 · outbound
3D Scene Graph Guided Vision-Language Pre-training 3D-SPS: Single- stage 3D visual grounding via referred point progressive se- lection
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation daad1910-59d8-48ae-aed0-8cba1f351a42 · outbound
3D Scene Graph Guided Vision-Language Pre-training Sgformer: Semantic graph transformer for point cloud-based 3D scene graph generation
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f53f8afe-8753-45ca-955c-3c2fe18e140f · outbound
3D Scene Graph Guided Vision-Language Pre-training Heterogeneous graph learning for scene graph prediction in 3d point clouds
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f2fef0cf-69cc-43bf-ba2f-1c2b0ffa5488 · outbound
3D Scene Graph Guided Vision-Language Pre-training Complete 3D relationships extraction modality align- ment network for 3D dense captioning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 90d7c137-08f4-455e-8419-9246f79f905e · outbound
3D Scene Graph Guided Vision-Language Pre-training An end-to- end transformer model for 3d object detection
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2f28da01-60c4-4736-bdc0-7dbb32a1651d · outbound
3D Scene Graph Guided Vision-Language Pre-training Bleu: a method for automatic evaluation of machine translation
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ec56fba2-5f36-4b63-bb68-0ddc2c83501c · outbound
3D Scene Graph Guided Vision-Language Pre-training Clip-guided vision-language pre-training for question answering in 3D scenes
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 742e42e4-e2b4-48bf-abdc-1e281c67cc51 · outbound
3D Scene Graph Guided Vision-Language Pre-training Pytorch: An im- perative style, high-performance deep learning library
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 945f0afd-6094-4cd5-80b9-154dc8648ea0 · outbound
3D Scene Graph Guided Vision-Language Pre-training Glove: Global vectors for word representation
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2e0cbfa8-531e-4385-8f5f-79eb676a4e87 · outbound
3D Scene Graph Guided Vision-Language Pre-training PointNet: Deep learning on point sets for 3D classification and segmentation
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8238dd07-accd-4a70-972b-1ffa67a82a56 · outbound
3D Scene Graph Guided Vision-Language Pre-training PointNet++: Deep hierarchical feature learning on point sets in a metric space
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a65fc4c1-120f-4427-8298-6b69357c270f · outbound
3D Scene Graph Guided Vision-Language Pre-training Qi, Or Litany, Kaiming He, and Leonidas J
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 60169241-a941-4d23-95e7-6676998543c8 · outbound
3D Scene Graph Guided Vision-Language Pre-training Improving language understanding by gen- erative pre-training
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation beb3ef1a-8106-4f14-9b6b-4425e5b13a0f · outbound
3D Scene Graph Guided Vision-Language Pre-training Learn- ing transferable visual models from natural language super- vision
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9c700dbd-8fea-4789-aab8-76fcf5bc6ca6 · outbound
3D Scene Graph Guided Vision-Language Pre-training Mask3D: Mask trans- former for 3D semantic instance segmentation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f32d1a60-7929-4f51-a895-82975aa858dc · outbound
3D Scene Graph Guided Vision-Language Pre-training VL-BERT: Pre-training of Generic Visual-Linguistic Representations
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bbadab6-a8a5-4c20-ae13-f7d8e9a51c01 · outbound
3D Scene Graph Guided Vision-Language Pre-training Cider: Consensus-based image description evalua- tion
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e85005a0-7a23-434d-9c1a-0e7bf4e854a9 · outbound
3D Scene Graph Guided Vision-Language Pre-training Learning 3D semantic scene graphs from 3D indoor reconstructions
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 592bafb9-5d43-4280-9f7b-3e9d4e1d3d66 · outbound
3D Scene Graph Guided Vision-Language Pre-training Spatiality-guided Transformer for 3D Dense Captioning on Point Clouds
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f91b276-cfd1-4b6a-8521-a2fbc03c2abf · outbound
3D Scene Graph Guided Vision-Language Pre-training Vl-sat: visual-linguistic semantics assisted training for 3D semantic scene graph prediction in point cloud
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2336401d-5b95-4b6b-95f0-c4171c7dabad · outbound
3D Scene Graph Guided Vision-Language Pre-training EDA: Explicit Text-Decoupling and Dense Alignment for 3D Visual Grounding
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81e19616-fc70-4b3a-a2f2-5ca0e21e3f8d · outbound
3D Scene Graph Guided Vision-Language Pre-training Vision-language pre-training with triple contrastive learning
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0827a075-662e-4da7-8d6e-f29b3d33139a · outbound
3D Scene Graph Guided Vision-Language Pre-training Sat: 2D semantics assisted training for 3D visual grounding
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d319560b-93fa-44d1-95ff-4b01c1fe1c08 · outbound
3D Scene Graph Guided Vision-Language Pre-training Deep modular co-attention networks for visual question an- swering
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f8d960d7-c1ee-41cf-8919-c7a6636a089c · outbound
3D Scene Graph Guided Vision-Language Pre-training Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual refer- ring
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 81c79412-2899-4c42-93f5-6d4138ddf27d · outbound
3D Scene Graph Guided Vision-Language Pre-training X-trans2cap: Cross-modal knowledge transfer using transformer for 3D dense caption- ing
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 018b2df1-9eb0-4195-bf05-63c3cbfa76a9 · outbound
3D Scene Graph Guided Vision-Language Pre-training Exploiting edge-oriented reasoning for 3D point-based scene graph analysis
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3f647af2-0f9f-46d1-bbea-603f5937e669 · outbound
3D Scene Graph Guided Vision-Language Pre-training Pointclip: Point cloud understanding by clip
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 667b6665-a768-4f13-8b0c-b287402c671c · outbound
3D Scene Graph Guided Vision-Language Pre-training Vision-language pre-training with object con- trastive learning for 3D scene understanding
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c6572169-f869-47a6-b384-75e071c1701d · outbound
3D Scene Graph Guided Vision-Language Pre-training 3DVG- Transformer: Relation modeling for visual grounding on point clouds
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 83364dbb-1be1-4fd3-a251-2fbb110d3df6 · outbound
3D Scene Graph Guided Vision-Language Pre-training Towards explainable 3D grounded visual question answer- ing: A new benchmark and strong baseline
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation effb4dd5-c393-4efc-b84e-89a983b5065d · outbound
3D Scene Graph Guided Vision-Language Pre-training Contextual Modeling for 3D Dense Captioning on Point Clouds
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ee57161-b2fc-41df-826c-c257492387c4 · outbound
3D Scene Graph Guided Vision-Language Pre-training Point- clip v2: Prompting clip and gpt for powerful 3D open-world learning
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7db195e7-147b-4288-8c5a-e034a489808c · outbound
3D Scene Graph Guided Vision-Language Pre-training 3d-vista: Pre-trained transformer for 3D vision and text alignment
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6251aefd-8e42-42d6-aaf2-bfdafc7408d0 · outbound
3D Scene Graph Guided Vision-Language Pre-training Unresolved cited work
Reference 545
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.