Pith. sign in

Paper Citation Record · LEDGER

A Review of 3D Object Detection with Vision-Language Models

As of 18 August 2026, this Paper Citation Record lists 7 of 7 outbound references and 1 inbound Pith citation observation for arXiv:2504.18738.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.18738 v1

Coverage vector

measured 7 of 7 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:13:25.257063Z

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:34:07.995899Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-16T12:16:17.039197Z

Reference resolution

7 of 7 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 98957a5e-5342-4994-8ea7-0436f0597901 · outbound

This paper cites Leveraging VLM-Based Pipelines to Annotate 3D Objects.

A Review of 3D Object Detection with Vision-Language Models Leveraging VLM-Based Pipelines to Annotate 3D Objects

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T10:13:25.741575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:13:25.139862Z digest=sha256:9eebb3ad80fcf4b71ba24d6bafaa985e071c63b4e17527f344e8c18e58299998

Observation 33ac26f8-7c1c-4483-ba56-11a9eb3bf281 · outbound

This paper cites SparseVoxFormer: Sparse Voxel-based Transformer for Multi-modal 3D Object Detection.

A Review of 3D Object Detection with Vision-Language Models SparseVoxFormer: Sparse Voxel-based Transformer for Multi-modal 3D Object Detection

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T10:13:25.627225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:13:25.173316Z digest=sha256:b98b8297c9bc2e1ebde8cb892089c1290d48a3fe2ec329a314b8f2e06769c740

Observation 7018c0b3-d611-47e9-8421-b7f103a98985 · outbound

This paper cites Towards Robust and Secure Embodied AI: A Survey on Vulnerabilities and Attacks.

A Review of 3D Object Detection with Vision-Language Models Towards Robust and Secure Embodied AI: A Survey on Vulnerabilities and Attacks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T10:13:25.257063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:13:25.257063Z digest=sha256:619fa62c9862f15e06e5bee76120d8185afd12133705385b8c83fbad44c3828d

Observation 8febbffd-7988-46e7-b6ce-d4a41c792e9f · outbound

This paper cites OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference.

A Review of 3D Object Detection with Vision-Language Models OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference

Reference 2022

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T10:13:26.010057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:13:25.113330Z digest=sha256:fa435af9bfdd560b4ef178cc7b38c75469cc6a42c2d31111fbe51a13fe503f5b

Observation f2f5b1d2-5aff-40db-a7c9-fc4a66b70aa8 · outbound

This paper cites Instruct 3D-to-3D: Text Instruction Guided 3D-to-3D conversion.

A Review of 3D Object Detection with Vision-Language Models Instruct 3D-to-3D: Text Instruction Guided 3D-to-3D conversion

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T10:13:25.147697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:13:25.147697Z digest=sha256:d6d621d4720ed8129e3e9bc3d1ad71c612a20bb826aa8c942cce9de5ae902591

Observation 0abec588-bf00-4e50-ae31-19e61f2e006e · outbound

This paper cites M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models.

A Review of 3D Object Detection with Vision-Language Models M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T10:13:25.100455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:13:25.100455Z digest=sha256:eb89adca9e26117edf6e64b1101b1af3ba658df2a8a4e62f5e03102ad80d233c

Observation ba40995c-8a85-4c7d-bded-77e27378dee3 · outbound

This paper cites OV-SCAN: Semantically Consistent Alignment for Novel Object Discovery in Open-Vocabulary 3D Object Detection.

A Review of 3D Object Detection with Vision-Language Models OV-SCAN: Semantically Consistent Alignment for Novel Object Discovery in Open-Vocabulary 3D Object Detection

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-16T10:13:25.128583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:13:25.128583Z digest=sha256:f196d11978614f452843f2dae3ddb824e0d7646cf3b76323ed25239b10f65f22

Pith citing papers

Observation f52befdf-3477-40ff-8c9b-f9d2922e3eb1 · inbound

Plant Disease Detection through Multimodal Large Language Models and Convolutional Neural Networks cites this paper.

Plant Disease Detection through Multimodal Large Language Models and Convolutional Neural Networks A Review of 3D Object Detection with Vision-Language Models

Reference 102

Resolution
verified exact
local_arxiv, observed 2026-08-16T05:34:08.064561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:34:07.995899Z digest=sha256:2641b1b30c6d459e14d40298c5c1492af227eeeaa77fc48057a9ff6ce01df948