Pith. sign in

Paper Citation Record · LEDGER

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes

As of 8 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2508.16812.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.16812 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:13:06.935892Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact11
  • verified fuzzy32
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 840f8725-d2b5-4625-9829-401e7c8370a8 · outbound

This paper cites TransFusion: Robust LiDAR-camera fusion for 3D object detection with transformers.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes TransFusion: Robust LiDAR-camera fusion for 3D object detection with transformers

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.347860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:01.437631Z digest=sha256:b83203581032245a2f9d740c27e81d1ce03c446de56b42fe6d1d1fc88d5b183d

Observation d9c232d4-c6b1-404e-8f17-5e6da800d8ae · outbound

This paper cites Is Space-Time Attention All You Need for Video Understanding?.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Is Space-Time Attention All You Need for Video Understanding?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:01.481907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:01.481907Z digest=sha256:b2562560e041fb196bad95f4ec221482581703318465d923030e42f7935f4aa5

Observation 0323529f-d181-4656-9e98-f1df76763a63 · outbound

This paper cites Lang, Sourabh V ora, V enice Erin Liong, Qiang Xu, Anush Krishnan, Y u Pan, Giancarlo Baldan, and Oscar Beijbom.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Lang, Sourabh V ora, V enice Erin Liong, Qiang Xu, Anush Krishnan, Y u Pan, Giancarlo Baldan, and Oscar Beijbom

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.331370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:01.550840Z digest=sha256:0a014fbb196d80ce4d6d2e2f367e0f62a4b5020572508f0f207e754255707976

Observation 92bfde35-7860-4701-81bd-53daddc5e63f · outbound

This paper cites CoDA: Collaborative novel box discovery and cross-modal alignment for open-vocabulary 3D object detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes CoDA: Collaborative novel box discovery and cross-modal alignment for open-vocabulary 3D object detection

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.315917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:01.619614Z digest=sha256:7b90f0bd0110290e1165160d815fbb78e1194f8a9dc20f99f993422c696a961c

Observation 5e6d3ffb-ca94-414c-9c2d-5ae815fbcfda · outbound

This paper cites Collaborative Novel Object Discovery and Box-Guided Cross-Modal Alignment for Open-Vocabulary 3D Object Detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Collaborative Novel Object Discovery and Box-Guided Cross-Modal Alignment for Open-Vocabulary 3D Object Detection

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:08.748655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:01.666088Z digest=sha256:4fd64a63cd896b7c417ce51791962d37b8cd8dc0db3a101c3f67f047397a564e

Observation b85c14d1-2749-4e90-94ee-c905dc863425 · outbound

This paper cites Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:01.758342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:01.758342Z digest=sha256:531530b9b5ce9499a519b602d0ae2b11423ace0b51cd1ebb51dbfa302f3bfc4d

Observation 103b1be2-881b-4092-a9de-cdf760126ccd · outbound

This paper cites Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language Model.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language Model

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:08.568522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:01.824856Z digest=sha256:6b50f48cb075848588330f22657873604c45f3daef8e7c693ddfbc73c13ab78e

Observation fcc17589-f298-479f-9021-265f579e2956 · outbound

This paper cites Find n' Propagate: Open-Vocabulary 3D Object Detection in Urban Environments.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Find n' Propagate: Open-Vocabulary 3D Object Detection in Urban Environments

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:08.337877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:01.870710Z digest=sha256:78698b20977db88e21d77258b85a3747fb0e3238fddf40dd71998f4d5372e293

Observation b9c0e4d5-992e-4efb-9ce4-80e17e4a012e · outbound

This paper cites Fully sparse 3D object detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Fully sparse 3D object detection

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.297804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:01.958733Z digest=sha256:43f63863c5f23e464231b97caccc7277d47aed23dea19b04abfdd1aa06497bb2

Observation 65a7e81d-22fa-4f21-9b72-55d7e8009575 · outbound

This paper cites Multi-modal transformer for video retrieval.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Multi-modal transformer for video retrieval

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.281744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:02.024764Z digest=sha256:f8b1b539c202bf3e2ffd3dfa00647050e9a4e5a2350c1db81938ad7c0ac8d220

Observation ff00328c-a284-450d-a55a-0bed8daf3e52 · outbound

This paper cites Are we ready for autonomous driving? the KITTI vision benchmark suite.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Are we ready for autonomous driving? the KITTI vision benchmark suite

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:02.093300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:02.093300Z digest=sha256:28aa010dd41a3f0a5f16dfa9e00bdb8a7f17b6ba50c51c7eac63d99a4cd1eda3

Observation 5d3e40f7-cc5b-4a84-9b7c-20e5284fab6a · outbound

This paper cites ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:02.141605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:02.141605Z digest=sha256:96db4ce1d2182536e277068d2fab4cee0d4be3e370d708bc02e861c8cb76d444

Observation e667cdb0-2ac7-45a1-b463-6f4ee791ac75 · outbound

This paper cites Open-vocabulary object detection via vision and language knowledge distillation.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Open-vocabulary object detection via vision and language knowledge distillation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.265554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:02.230640Z digest=sha256:2c25ac11fa1ab5320d0b013aa203fb4bc1a695fc2cb63bc10720a139d08a9328

Observation 4daed349-6ad7-46ce-8ed6-7562da3577c9 · outbound

This paper cites OneLLM: One framework to align all modalities with language.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes OneLLM: One framework to align all modalities with language

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.250070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:02.295762Z digest=sha256:19189bc5cc85ffbccb6023a978639312ef65f5c64c058bbacfaf37ca9838b2b6

Observation 1feac009-2953-4277-953f-eb5175876ced · outbound

This paper cites Jones, and Vishal M.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Jones, and Vishal M

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.231011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:02.472014Z digest=sha256:1ad13fed46aabe408f862c616f505ad5b56ab613aa98b643f141ad140b23c103

Observation a42b1dbc-f053-4791-b64e-845761745580 · outbound

This paper cites Jones, and Vishal M.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Jones, and Vishal M

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.214938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:02.563889Z digest=sha256:eff0fd9579f4166fff2c41a1c47bed489622f981d601401b4cdbbe2057b6b270

Observation eb6ac23e-a5bd-4605-8d95-8fc262d9962e · outbound

This paper cites Long short-term memory.Neural Comput., 9 (8):1735–1780, November 1997.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Long short-term memory.Neural Comput., 9 (8):1735–1780, November 1997

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:02.655417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:02.655417Z digest=sha256:4d727e49c5a5e3ee5a897ee9f0d7efefb53b76c66cc5b0b606a89d15e79a3bab

Observation a27b0254-1bb4-49a8-990b-1cdea1a248cb · outbound

This paper cites ConceptFusion: Open-set Multimodal 3D Mapping.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes ConceptFusion: Open-set Multimodal 3D Mapping

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:02.731888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:02.731888Z digest=sha256:6034da1d9aa792eaf2b71c1a004fbe2fabe5bdae21d1fbcbfc281b5eee167bd0

Observation 1e0ae063-aee7-4e35-80d7-5e7f021646cb · outbound

This paper cites Action genome: Actions as composition of spatio-temporal scene graphs.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Action genome: Actions as composition of spatio-temporal scene graphs

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.198930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:02.811170Z digest=sha256:ecba798a0867858ff1ab9a91e57e6381e1c8ad10049b69503b5a2efe8ad4e619

Observation 81b7e8c8-d47a-4b5b-980c-41e71feb883e · outbound

This paper cites PF3Det: A prompted foundation feature assisted visual LiDAR 3D detector.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes PF3Det: A prompted foundation feature assisted visual LiDAR 3D detector

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.182991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:02.874237Z digest=sha256:294b8852684d3387775c5f5c2cbb31d568ea591b46ad10265cc692291b15d7ec

Observation c734e3c7-3a44-4ebe-b8d5-b10c179ca67a · outbound

This paper cites Grounded language-image pre-training.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Grounded language-image pre-training

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.165658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:02.980304Z digest=sha256:eca687c49e5c858c8f3db7c689a6fa00a9770b70ef0144f060c4f3fad0f8a7c0

Observation f5844c12-2d4d-467f-b559-f74bc885388d · outbound

This paper cites OpenShape: Scaling up 3D shape representation towards open-world understanding.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes OpenShape: Scaling up 3D shape representation towards open-world understanding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.146762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:03.053531Z digest=sha256:79a0c20e6c5a254624c10fc30740307426791b760c4f702d601df1cf194dd5ef

Observation 47339047-0ceb-48aa-be97-d11fc8eb48e5 · outbound

This paper cites Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.127565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:03.160071Z digest=sha256:05a9bb0423da075371482caa690bccc62057c94faf879b28b5edbf9560c46746

Observation a1648445-19ac-4617-ac79-9ebce9000c08 · outbound

This paper cites Open-vocabulary point-cloud object detection without 3D annotation.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Open-vocabulary point-cloud object detection without 3D annotation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.111280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:03.291623Z digest=sha256:d58ee0f7d3542b6b2926d810fa774d6bd50771006c47e2a347ec12a6daab5da8

Observation 7f458115-3bbd-44f3-ab37-2168437dc4ee · outbound

This paper cites An End-to-End Transformer Model for 3D Object Detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes An End-to-End Transformer Model for 3D Object Detection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:03.369392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:03.369392Z digest=sha256:ce9a9b03d8255f82352bdf958aca425ccc737c2cca33708b85f4eb20745fda92

Observation b0c79d08-7717-4a79-a1ee-1dd63bd53f09 · outbound

This paper cites Modeling temporal structure of decomposable motion segments for activity classification.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Modeling temporal structure of decomposable motion segments for activity classification

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.094901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:03.471688Z digest=sha256:be5def73fcd5a8c272b7d2c3680fae27e0e761e1780c3d1c53f5e3d1fd8f54e8

Observation dabae4fa-8546-48ee-8e6e-dc06fcded168 · outbound

This paper cites PyT orch: An imperative style, high-performance deep learning library.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes PyT orch: An imperative style, high-performance deep learning library

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.076037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:03.583519Z digest=sha256:afd615b5a745df6711349fa544876eeb4b646083c349d471cd68a75a618dc0ea

Observation 1bd61506-50b3-4391-9cd5-498424ec4d70 · outbound

This paper cites OpenScene: 3D scene understanding with open vocabularies.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes OpenScene: 3D scene understanding with open vocabularies

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.058826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:03.686010Z digest=sha256:6d535f2d738bceba2f5246631df87ed7d0deaa024c0f37e5f401dc7dcbe6e17e

Observation 166f1343-26ab-447c-9b6d-14a5b9b70c69 · outbound

This paper cites Qi, Hao Su, Kaichun Mo, and Leonidas J.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Qi, Hao Su, Kaichun Mo, and Leonidas J

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.042963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:03.788184Z digest=sha256:a45b5ed87b52e837b8bbfa7963f66404e22e99a9ca2ad165f5b391fbd8d2688a

Observation 2f536505-b225-487c-a519-1c9ef552fcc6 · outbound

This paper cites Frustum PointNets for 3D Object Detection from RGB-D Data.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Frustum PointNets for 3D Object Detection from RGB-D Data

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:03.886207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:03.886207Z digest=sha256:7462e9d331392c6cee77a74bccba5f0175bcc49dec179ce0b7317c38cf645016

Observation 0303b28b-8ef5-4362-897b-8707f7336768 · outbound

This paper cites Deep Hough Voting for 3D Object Detection in Point Clouds.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Deep Hough Voting for 3D Object Detection in Point Clouds

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:08.084109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:03.987522Z digest=sha256:83cfeac74c1cc401201a107020dda98963c27b5eca522fdd96fa13e13bfffcfe

Observation 6d60b02a-280e-4cea-8201-22b4b411d242 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Learning Transferable Visual Models From Natural Language Supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:04.056417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:04.056417Z digest=sha256:bf5ae73d39076348dfa1b233bbc7bec91c1b66c1422c126d497533cb7b826971

Observation 70bcaecc-0c36-41b1-9450-c0d54195480f · outbound

This paper cites PointRCNN: 3D Object Proposal Generation and Detection from Point Cloud.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes PointRCNN: 3D Object Proposal Generation and Detection from Point Cloud

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:08.034757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:04.157255Z digest=sha256:f33cfb75c1bfb09bebd25d959492238bc97dcc5a074954cc554d33c490781da5

Observation cd9525d2-b4da-4dea-9189-e38ee5217e14 · outbound

This paper cites PV -RCNN: Point-voxel feature set abstraction for 3D object detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes PV -RCNN: Point-voxel feature set abstraction for 3D object detection

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.021635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:04.256401Z digest=sha256:a3cbd103c8d53a363d344ec7df0cc95e44e042589e44427b8e1e4d43680136b9

Observation 75bf4426-3b0a-45c4-98d4-8e737839456c · outbound

This paper cites VideoBERT: A joint model for video and language representation learning.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes VideoBERT: A joint model for video and language representation learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.005712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:04.338428Z digest=sha256:0b715024e9d06cd59dd4195248c2441d23975a8b6fa76737690140b55bd45da5

Observation 9536a125-4495-40da-9be9-6671aa89b29a · outbound

This paper cites Scalability in perception for autonomous driving: W aymo open dataset, 2020.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Scalability in perception for autonomous driving: W aymo open dataset, 2020

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.986709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:04.558032Z digest=sha256:de17be4ae08183e314df428d1d28468eb5903c4dc858b44d91c5d6f193b1bdaf

Observation ca3ff166-92f1-495c-ac5c-aa2b4d7e419f · outbound

This paper cites Learning spatiotemporal features with 3D convolutional networks.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Learning spatiotemporal features with 3D convolutional networks

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.956623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:04.735453Z digest=sha256:3753c2b7074d8f4be691d92f973ebb24569e7fdef1b39fa33eaa4a1ae156551a

Observation a3914318-82fa-4337-b98e-878381f55c9f · outbound

This paper cites DSVT: Dynamic sparse voxel transformer with rotated sets.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes DSVT: Dynamic sparse voxel transformer with rotated sets

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.935947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:04.934308Z digest=sha256:1dfd1896a4436b55ce5a97479346d68247faea9ff89cfe93684ba0368ec7b236

Observation 76edf698-e339-4015-b383-1b7c9cfd170c · outbound

This paper cites Hierarchical open-vocabulary 3D scene graphs for language-grounded robot navigation.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Hierarchical open-vocabulary 3D scene graphs for language-grounded robot navigation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:05.358512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:05.358512Z digest=sha256:40c14c3e1e78b2c38ad3867d56c9fad8ac7f24742a939896b7a2faecb2875c39

Observation 5f9d2efa-63c2-4004-8447-bccd567f0af8 · outbound

This paper cites Argoverse 2: Next generation datasets for self-driving perception and forecasting.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Argoverse 2: Next generation datasets for self-driving perception and forecasting

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.914171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:05.456330Z digest=sha256:6194aee8b7b671ab8c37b2d17ba3f8acdb2a8138ae305ce290b13bd2bbac95cc

Observation 6e6596de-4638-4dfd-90ba-f7288e8a3ee8 · outbound

This paper cites Transformation- equivariant 3D object detection for autonomous driving.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Transformation- equivariant 3D object detection for autonomous driving

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.895793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:05.555846Z digest=sha256:3179166b2c545f241d6ffa8b3bbd2a92adfc7fc571c43d7f8c4eeb95a5436bfc

Observation 74593d27-35fd-4cd5-be04-5504b25e89aa · outbound

This paper cites Towards Open Vocabulary Learning: A Survey.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Towards Open Vocabulary Learning: A Survey

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:08.006377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:05.638687Z digest=sha256:d8cd86505e36b02f5ffc96963ebbc8bd9da8f9c029879a639124c3b8cc4f2b33

Observation 7e732a7a-114b-4d1f-b7be-a104fda0f275 · outbound

This paper cites FusionViT: Hierarchical 3D Object Detection via LiDAR-Camera Vision Transformer Fusion.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes FusionViT: Hierarchical 3D Object Detection via LiDAR-Camera Vision Transformer Fusion

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:05.763277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:05.763277Z digest=sha256:090268bb221278f19f67ada99f3069bfc15cc5a30a341dd3f851b9f8bcfae0f7

Observation d4098634-ae30-4531-82c2-b551b1eddb60 · outbound

This paper cites 3DifFusionDet: Diffusion model for 3D object detection with robust LiDAR-camera fusion.arXiv preprint arXiv:2311.0374, 2023.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes 3DifFusionDet: Diffusion model for 3D object detection with robust LiDAR-camera fusion.arXiv preprint arXiv:2311.0374, 2023

Reference 46

Resolution
verified exact
raw_fallback, observed 2026-08-05T17:13:07.953455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:05.930584Z digest=sha256:03d4fb66f6fbb81c92f6dd1bc0a5fb1b9b7baab4d148fe2211092567c4778565

Observation 6c5b56a5-94f0-4a35-af10-38ab6e6203b0 · outbound

This paper cites EffiPerception: an Efficient Framework for Various Perception Tasks.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes EffiPerception: an Efficient Framework for Various Perception Tasks

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:07.702417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:06.036014Z digest=sha256:94939e53154ee1ae9fcacb10530c886e2248d8db151f2e605ea98a1456ca55bf

Observation 91b533a6-5f6e-4b49-86d4-4fb4a3360e09 · outbound

This paper cites Graph R-CNN for Scene Graph Generation.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Graph R-CNN for Scene Graph Generation

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:07.519380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:06.116837Z digest=sha256:0411b4e933555589b7d4b354803f6c2802664a4910510249520baada5a534be2

Observation d37adc28-e3b3-44ff-ba67-1b6365d0cabd · outbound

This paper cites Open-vocabulary DETR with conditional matching.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Open-vocabulary DETR with conditional matching

Reference 49

Resolution
verified exact
doi, observed 2026-08-05T17:13:07.148949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:06.198543Z digest=sha256:50aa8ab425a655a8db503145afc8b2812499dc15ce0ade07d2699b1a9c607ed1

Observation 32c79f97-0f4c-4abb-9c36-4cfb103414ca · outbound

This paper cites Open-Vocabulary Object Detection Using Captions.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Open-Vocabulary Object Detection Using Captions

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:07.337789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:06.300876Z digest=sha256:8e44a2f477f7731465b8cb79fd862681a3a96d98b7b2f34fd76b612e21f576db

Observation f02b52ce-72b7-4aed-9ed4-519b2f9548b2 · outbound

This paper cites FM-OV3D: Foundation model-based cross-modal knowledge blending for open-vocabulary 3D detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes FM-OV3D: Foundation model-based cross-modal knowledge blending for open-vocabulary 3D detection

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.878591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:06.391908Z digest=sha256:e5527ae8c51debe0fcbb3d9014753794e77769f8b7474a48a695a9198fc34b53

Observation 0ffd2678-aa19-4326-be4b-169e30a6c3a8 · outbound

This paper cites OpenSight: A simple open-vocabulary framework for LiDAR-based object detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes OpenSight: A simple open-vocabulary framework for LiDAR-based object detection

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.860512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:06.478801Z digest=sha256:383f968a1009a89db7aded29c316a516c039fee80d63200603e3e9057f9177c6

Observation c38fb163-4d80-43bf-ba6d-f93513cef3f2 · outbound

This paper cites G, Anastasis Stathopoulos, Manmohan Chandraker, and Dimitris Metaxas.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes G, Anastasis Stathopoulos, Manmohan Chandraker, and Dimitris Metaxas

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.839413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:06.561654Z digest=sha256:e86ff15b98eccfd6d6f599fb6f4867c2a1c72476269b755e340a2b2aa6b1f44b

Observation 8fd75d1e-f795-4927-90fd-bd099936e1de · outbound

This paper cites OcTr: Octree-based transformer for 3D object detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes OcTr: Octree-based transformer for 3D object detection

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.816793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:06.650142Z digest=sha256:b682cd41992a326b8b21a0a183cd04030048cdbb0f4b196a49bbe4f0fb85b259

Observation e84cafae-c49f-4a74-b668-7d794db8d9da · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes CogVLM: Visual Expert for Pretrained Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:06.743796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:06.743796Z digest=sha256:8300044cb1a50af57b53ce469b8abdea5058616c867eff7a8ac1d1eb8593a474

Observation af50caab-0370-43a9-b0b5-631aea67a62b · outbound

This paper cites PointCLIP V2: Prompting CLIP and GPT for powerful 3D open-world learning.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes PointCLIP V2: Prompting CLIP and GPT for powerful 3D open-world learning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.799818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:06.857653Z digest=sha256:e1ea9d6004091abc5c46e38963959adbbd1e53f5a1c3e3ed0727a07cce7bb0cb

Observation f4e88895-2b6d-422c-95bc-46b5a19585ec · outbound

This paper cites Then, the model continues to be trained for 20 epochs.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Then, the model continues to be trained for 20 epochs

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.784314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:13:06.935892Z digest=sha256:21efb0ed266db95697fd81f0db45602914ad30776a0b67ba6458e40b750db8d2

Observation 3a331ea2-4e45-46c3-b90b-5d7b8e65f565 · outbound

This paper cites OneLLM: One Framework to Align All Modalities with Language.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes OneLLM: One Framework to Align All Modalities with Language

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:02.378508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:02.378508Z digest=sha256:51c4e687bf4ae4155a5c6299a815732b0769872d47077c39cc2430ac44bef8bf

Pith citing papers

No inbound Pith citation observations are available.