Pith. sign in

Paper Citation Record · LEDGER

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes

As of 20 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2508.16812.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.16812 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:13:06.935892Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact11
  • verified fuzzy32
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 840f8725-d2b5-4625-9829-401e7c8370a8 · outbound

This paper cites TransFusion: Robust LiDAR-camera fusion for 3D object detection with transformers.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes TransFusion: Robust LiDAR-camera fusion for 3D object detection with transformers

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.347860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:01.437631Z digest=sha256:b6aafc7a1783e700497ebbfef9ef8133de70bb4daf33fad6adeeb20b80caad10

Observation d9c232d4-c6b1-404e-8f17-5e6da800d8ae · outbound

This paper cites Is Space-Time Attention All You Need for Video Understanding?.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Is Space-Time Attention All You Need for Video Understanding?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:01.481907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:01.481907Z digest=sha256:86a05b095b75d693dc9f78b1bcc8da8ed67ea340f3303611ea282683619b5b52

Observation 0323529f-d181-4656-9e98-f1df76763a63 · outbound

This paper cites Lang, Sourabh V ora, V enice Erin Liong, Qiang Xu, Anush Krishnan, Y u Pan, Giancarlo Baldan, and Oscar Beijbom.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Lang, Sourabh V ora, V enice Erin Liong, Qiang Xu, Anush Krishnan, Y u Pan, Giancarlo Baldan, and Oscar Beijbom

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.331370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:01.550840Z digest=sha256:5fd32fde7744dedbf37b3dcc55abbc291874e49466c3ab01978babded20bd1e9

Observation 92bfde35-7860-4701-81bd-53daddc5e63f · outbound

This paper cites CoDA: Collaborative novel box discovery and cross-modal alignment for open-vocabulary 3D object detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes CoDA: Collaborative novel box discovery and cross-modal alignment for open-vocabulary 3D object detection

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.315917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:01.619614Z digest=sha256:1e1f2da6c3b5a61dca851a4cfcb4e94b81c60d4d5b8d54c39dbfc353cc2b76be

Observation 5e6d3ffb-ca94-414c-9c2d-5ae815fbcfda · outbound

This paper cites Collaborative Novel Object Discovery and Box-Guided Cross-Modal Alignment for Open-Vocabulary 3D Object Detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Collaborative Novel Object Discovery and Box-Guided Cross-Modal Alignment for Open-Vocabulary 3D Object Detection

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:08.748655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:01.666088Z digest=sha256:99c10a4b11687f6b3eab2581679f7cf21a639357a51e29d709e8c566ba5815e1

Observation b85c14d1-2749-4e90-94ee-c905dc863425 · outbound

This paper cites Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:01.758342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:01.758342Z digest=sha256:fd5a2853450b550484ae4c95a5800db4b043b57368175da61b42d9b038d6d329

Observation 103b1be2-881b-4092-a9de-cdf760126ccd · outbound

This paper cites Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language Model.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language Model

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:08.568522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:01.824856Z digest=sha256:192e09dac1eb28cb2f980a26b9be514ce05b37db607f0ed2069613bd89ffd7f8

Observation fcc17589-f298-479f-9021-265f579e2956 · outbound

This paper cites Find n' Propagate: Open-Vocabulary 3D Object Detection in Urban Environments.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Find n' Propagate: Open-Vocabulary 3D Object Detection in Urban Environments

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:08.337877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:01.870710Z digest=sha256:36e50f4298abcb0ea31fd4a8d95923f301fe72b16b1ec59d55f58a48ae4745f5

Observation b9c0e4d5-992e-4efb-9ce4-80e17e4a012e · outbound

This paper cites Fully sparse 3D object detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Fully sparse 3D object detection

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.297804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:01.958733Z digest=sha256:d85a4001102499b8b8e0749036ced6dd76ddcf40a7baa4081adf92d9ffbcf86c

Observation 65a7e81d-22fa-4f21-9b72-55d7e8009575 · outbound

This paper cites Multi-modal transformer for video retrieval.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Multi-modal transformer for video retrieval

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.281744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:02.024764Z digest=sha256:f6de7ec57b045b53f88fce02f2701d8c41d683dad2c40aeea311e32a986975d2

Observation ff00328c-a284-450d-a55a-0bed8daf3e52 · outbound

This paper cites Are we ready for autonomous driving? the KITTI vision benchmark suite.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Are we ready for autonomous driving? the KITTI vision benchmark suite

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:02.093300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:02.093300Z digest=sha256:72ebe55761dfd8a99412c93805d9ac7065d8772725b97aaf0ea1224f3a6162e6

Observation 5d3e40f7-cc5b-4a84-9b7c-20e5284fab6a · outbound

This paper cites ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:02.141605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:02.141605Z digest=sha256:e06131f08507917914123834ae8345f779a9e379093a0bd387515654d8ecab02

Observation e667cdb0-2ac7-45a1-b463-6f4ee791ac75 · outbound

This paper cites Open-vocabulary object detection via vision and language knowledge distillation.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Open-vocabulary object detection via vision and language knowledge distillation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.265554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:02.230640Z digest=sha256:c74b1381db3eac125f578cc6a8edb08412c1f89e056c8d9e69e258ed79af3af8

Observation 4daed349-6ad7-46ce-8ed6-7562da3577c9 · outbound

This paper cites OneLLM: One framework to align all modalities with language.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes OneLLM: One framework to align all modalities with language

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.250070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:02.295762Z digest=sha256:92dd6c2ab7ead8b2c1bba2fe9dd165ac5442b466209d6069c238a7910178c7d0

Observation 1feac009-2953-4277-953f-eb5175876ced · outbound

This paper cites Jones, and Vishal M.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Jones, and Vishal M

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.231011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:02.472014Z digest=sha256:f1c42068dac813a5f64cfaaf757fda4b15fb59024d7802261438276b4e059dea

Observation a42b1dbc-f053-4791-b64e-845761745580 · outbound

This paper cites Jones, and Vishal M.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Jones, and Vishal M

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.214938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:02.563889Z digest=sha256:886a46937133081b653707b869e996046e3f0883b39bea1ba10d77eef8b04112

Observation eb6ac23e-a5bd-4605-8d95-8fc262d9962e · outbound

This paper cites Long short-term memory.Neural Comput., 9 (8):1735–1780, November 1997.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Long short-term memory.Neural Comput., 9 (8):1735–1780, November 1997

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:02.655417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:02.655417Z digest=sha256:959118e73f431eb413724faead3e8df2ca9a9d9f1598b1fe5bdd5f5f31517bb5

Observation a27b0254-1bb4-49a8-990b-1cdea1a248cb · outbound

This paper cites ConceptFusion: Open-set Multimodal 3D Mapping.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes ConceptFusion: Open-set Multimodal 3D Mapping

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:02.731888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:02.731888Z digest=sha256:52a41f46989425d73c97b8d0545527041f0c32a3ba4477567d399e4cc4e96049

Observation 1e0ae063-aee7-4e35-80d7-5e7f021646cb · outbound

This paper cites Action genome: Actions as composition of spatio-temporal scene graphs.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Action genome: Actions as composition of spatio-temporal scene graphs

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.198930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:02.811170Z digest=sha256:4175ef66bc15f960fb4a5c951850753ad291941a3224e2441c6b3df07075dcb4

Observation 81b7e8c8-d47a-4b5b-980c-41e71feb883e · outbound

This paper cites PF3Det: A prompted foundation feature assisted visual LiDAR 3D detector.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes PF3Det: A prompted foundation feature assisted visual LiDAR 3D detector

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.182991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:02.874237Z digest=sha256:d372008376596df3ffb6d4191291b06a8cca680f3d4020df9b016c1e49a802b2

Observation c734e3c7-3a44-4ebe-b8d5-b10c179ca67a · outbound

This paper cites Grounded language-image pre-training.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Grounded language-image pre-training

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.165658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:02.980304Z digest=sha256:d6d88bbe98c97244ae167fe948a76b249dda42d6316196ed55cee31afa220eac

Observation f5844c12-2d4d-467f-b559-f74bc885388d · outbound

This paper cites OpenShape: Scaling up 3D shape representation towards open-world understanding.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes OpenShape: Scaling up 3D shape representation towards open-world understanding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.146762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:03.053531Z digest=sha256:c2a0f2f99fda886a5522e5d7a7b4861e31dbc9e67526a1bc4159ea3ec3a9bdab

Observation 47339047-0ceb-48aa-be97-d11fc8eb48e5 · outbound

This paper cites Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.127565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:03.160071Z digest=sha256:7eb8bf4eee0aa6df41385da1f5752143de14804a85ea35414d7da69affcb7d81

Observation a1648445-19ac-4617-ac79-9ebce9000c08 · outbound

This paper cites Open-vocabulary point-cloud object detection without 3D annotation.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Open-vocabulary point-cloud object detection without 3D annotation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.111280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:03.291623Z digest=sha256:222afaef54504c9571d81fd396c8ea50eaa5205d7229d0942e9f6ed7c9c19a8f

Observation 7f458115-3bbd-44f3-ab37-2168437dc4ee · outbound

This paper cites An End-to-End Transformer Model for 3D Object Detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes An End-to-End Transformer Model for 3D Object Detection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:03.369392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:03.369392Z digest=sha256:9e64571511f01e6ab3b3dc0c5d39981170f8ab598ba7fe194ab735692f2e16cf

Observation b0c79d08-7717-4a79-a1ee-1dd63bd53f09 · outbound

This paper cites Modeling temporal structure of decomposable motion segments for activity classification.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Modeling temporal structure of decomposable motion segments for activity classification

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.094901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:03.471688Z digest=sha256:1e5f7a57a9c1d2146ced143ca48b2c62769874c55108d869016bad30555d793c

Observation dabae4fa-8546-48ee-8e6e-dc06fcded168 · outbound

This paper cites PyT orch: An imperative style, high-performance deep learning library.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes PyT orch: An imperative style, high-performance deep learning library

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.076037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:03.583519Z digest=sha256:cdefa05af3b69b34bcba08b7f9dcf806b7838ac3434d60f5dc93819c097cb7b0

Observation 1bd61506-50b3-4391-9cd5-498424ec4d70 · outbound

This paper cites OpenScene: 3D scene understanding with open vocabularies.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes OpenScene: 3D scene understanding with open vocabularies

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.058826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:03.686010Z digest=sha256:2c79da44e8118a5467ce47c1ab8cae01a9b97b448922ef57578828a243f9940b

Observation 166f1343-26ab-447c-9b6d-14a5b9b70c69 · outbound

This paper cites Qi, Hao Su, Kaichun Mo, and Leonidas J.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Qi, Hao Su, Kaichun Mo, and Leonidas J

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.042963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:03.788184Z digest=sha256:18737c57076dbf395ebbff4979f1c54700b6813bc66e21b7e45e8712a5801f5b

Observation 2f536505-b225-487c-a519-1c9ef552fcc6 · outbound

This paper cites Frustum PointNets for 3D Object Detection from RGB-D Data.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Frustum PointNets for 3D Object Detection from RGB-D Data

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:03.886207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:03.886207Z digest=sha256:3e9ea3b955a80cc1e669f351ac6d622ded2e5bb151beb9a1599f1394dd70d4a7

Observation 0303b28b-8ef5-4362-897b-8707f7336768 · outbound

This paper cites Deep Hough Voting for 3D Object Detection in Point Clouds.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Deep Hough Voting for 3D Object Detection in Point Clouds

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:08.084109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:03.987522Z digest=sha256:bc7854378007a340024853ec6ab2a38cac6cb53675c6dfd0a5a2b3b3d5a03584

Observation 6d60b02a-280e-4cea-8201-22b4b411d242 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Learning Transferable Visual Models From Natural Language Supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:04.056417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:04.056417Z digest=sha256:ded2f31a9ec300261d84f8e9c68aad7ad893e1f61590460cac168daa3e10d9f9

Observation 70bcaecc-0c36-41b1-9450-c0d54195480f · outbound

This paper cites PointRCNN: 3D Object Proposal Generation and Detection from Point Cloud.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes PointRCNN: 3D Object Proposal Generation and Detection from Point Cloud

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:08.034757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:04.157255Z digest=sha256:b8386b69b1cef8157297c2a19edf0c413bede5b880f6126e6515f2374a901b25

Observation cd9525d2-b4da-4dea-9189-e38ee5217e14 · outbound

This paper cites PV -RCNN: Point-voxel feature set abstraction for 3D object detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes PV -RCNN: Point-voxel feature set abstraction for 3D object detection

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.021635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:04.256401Z digest=sha256:4dabc309d6ed94af6c624896602f13a4c36cb0dfa12f16405aa30722143e0878

Observation 75bf4426-3b0a-45c4-98d4-8e737839456c · outbound

This paper cites VideoBERT: A joint model for video and language representation learning.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes VideoBERT: A joint model for video and language representation learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.005712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:04.338428Z digest=sha256:289e6da0ed07cdd04c2dde591c69ccd6b6047f64283626f723cac63a401bc1a7

Observation 9536a125-4495-40da-9be9-6671aa89b29a · outbound

This paper cites Scalability in perception for autonomous driving: W aymo open dataset, 2020.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Scalability in perception for autonomous driving: W aymo open dataset, 2020

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.986709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:04.558032Z digest=sha256:1544edfbb9f673bcefb3334935da68726de743487f8b0326404fbdb12b93db72

Observation ca3ff166-92f1-495c-ac5c-aa2b4d7e419f · outbound

This paper cites Learning spatiotemporal features with 3D convolutional networks.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Learning spatiotemporal features with 3D convolutional networks

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.956623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:04.735453Z digest=sha256:5d506651ae26ec0d868075b34c520530f1af534fb6e36cdeba4e9c88c5f5c9ef

Observation a3914318-82fa-4337-b98e-878381f55c9f · outbound

This paper cites DSVT: Dynamic sparse voxel transformer with rotated sets.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes DSVT: Dynamic sparse voxel transformer with rotated sets

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.935947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:04.934308Z digest=sha256:ab9322e1aab29c0da605389888e6cedd901a040be896fe8552cda062d8a84c27

Observation 76edf698-e339-4015-b383-1b7c9cfd170c · outbound

This paper cites Hierarchical open-vocabulary 3D scene graphs for language-grounded robot navigation.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Hierarchical open-vocabulary 3D scene graphs for language-grounded robot navigation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:05.358512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:05.358512Z digest=sha256:515b48b9f0c782b790c72f4d75338e9366e6cfd4b847cd1862527e14af0137f2

Observation 5f9d2efa-63c2-4004-8447-bccd567f0af8 · outbound

This paper cites Argoverse 2: Next generation datasets for self-driving perception and forecasting.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Argoverse 2: Next generation datasets for self-driving perception and forecasting

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.914171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:05.456330Z digest=sha256:d32da8891f9d01b1b45093045e87c8a754e56b0c4309ce5db8f8948c730a6ab7

Observation 6e6596de-4638-4dfd-90ba-f7288e8a3ee8 · outbound

This paper cites Transformation- equivariant 3D object detection for autonomous driving.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Transformation- equivariant 3D object detection for autonomous driving

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.895793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:05.555846Z digest=sha256:29fb1a4afa850bad91aa9af96a90b7fa126bbccb03a4e3767cf3a664ba31a034

Observation 74593d27-35fd-4cd5-be04-5504b25e89aa · outbound

This paper cites Towards Open Vocabulary Learning: A Survey.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Towards Open Vocabulary Learning: A Survey

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:08.006377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:05.638687Z digest=sha256:9057920729e179039aa2d37207b11c4cba1cf382149e47e47a9606d06aae9ebf

Observation 7e732a7a-114b-4d1f-b7be-a104fda0f275 · outbound

This paper cites FusionViT: Hierarchical 3D Object Detection via LiDAR-Camera Vision Transformer Fusion.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes FusionViT: Hierarchical 3D Object Detection via LiDAR-Camera Vision Transformer Fusion

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:05.763277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:05.763277Z digest=sha256:35d648e289526d457b1b3d9f60610795b72f6621e2eb7ff65f4ea2652f941173

Observation d4098634-ae30-4531-82c2-b551b1eddb60 · outbound

This paper cites 3DifFusionDet: Diffusion model for 3D object detection with robust LiDAR-camera fusion.arXiv preprint arXiv:2311.0374, 2023.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes 3DifFusionDet: Diffusion model for 3D object detection with robust LiDAR-camera fusion.arXiv preprint arXiv:2311.0374, 2023

Reference 46

Resolution
verified exact
raw_fallback, observed 2026-08-05T17:13:07.953455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:05.930584Z digest=sha256:74685168ab50566eb2945dc7744e11165efac4aa12974524cdfe0f973f8846a9

Observation 6c5b56a5-94f0-4a35-af10-38ab6e6203b0 · outbound

This paper cites EffiPerception: an Efficient Framework for Various Perception Tasks.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes EffiPerception: an Efficient Framework for Various Perception Tasks

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:07.702417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:06.036014Z digest=sha256:09f0e4ab262215058233478bc0c285c37d18c336395319b75db88b302948004e

Observation 91b533a6-5f6e-4b49-86d4-4fb4a3360e09 · outbound

This paper cites Graph R-CNN for Scene Graph Generation.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Graph R-CNN for Scene Graph Generation

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:07.519380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:06.116837Z digest=sha256:8e7846bc2ed31dff92a8d4138c47fdce695afc370c684abe9a3a9e3c2008f3bb

Observation d37adc28-e3b3-44ff-ba67-1b6365d0cabd · outbound

This paper cites Open-vocabulary DETR with conditional matching.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Open-vocabulary DETR with conditional matching

Reference 49

Resolution
verified exact
doi, observed 2026-08-05T17:13:07.148949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:06.198543Z digest=sha256:0c5397d5701fed98cae72a179e86e3d872c5bd85ffcb970501a1e7bdfea82561

Observation 32c79f97-0f4c-4abb-9c36-4cfb103414ca · outbound

This paper cites Open-Vocabulary Object Detection Using Captions.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Open-Vocabulary Object Detection Using Captions

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:07.337789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:06.300876Z digest=sha256:979c60264c42277dcc22272cb8aedcbf534b71cd6c7373e14ecd9fbbdf0cafdf

Observation f02b52ce-72b7-4aed-9ed4-519b2f9548b2 · outbound

This paper cites FM-OV3D: Foundation model-based cross-modal knowledge blending for open-vocabulary 3D detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes FM-OV3D: Foundation model-based cross-modal knowledge blending for open-vocabulary 3D detection

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.878591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:06.391908Z digest=sha256:6e6ff9d3b139715ccaa83fce2b44e3cd2ba6e5193042249cc2276d60d3b3fc52

Observation 0ffd2678-aa19-4326-be4b-169e30a6c3a8 · outbound

This paper cites OpenSight: A simple open-vocabulary framework for LiDAR-based object detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes OpenSight: A simple open-vocabulary framework for LiDAR-based object detection

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.860512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:06.478801Z digest=sha256:d897c802f75084cf38d5e8e7e2e86925a8c4b34dbb39bb58d68af637d13a8a16

Observation c38fb163-4d80-43bf-ba6d-f93513cef3f2 · outbound

This paper cites G, Anastasis Stathopoulos, Manmohan Chandraker, and Dimitris Metaxas.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes G, Anastasis Stathopoulos, Manmohan Chandraker, and Dimitris Metaxas

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.839413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:06.561654Z digest=sha256:9aa7dcca1a01145a0687b4d947714a0b79b2b7669f9384ff0a84801720ad69db

Observation 8fd75d1e-f795-4927-90fd-bd099936e1de · outbound

This paper cites OcTr: Octree-based transformer for 3D object detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes OcTr: Octree-based transformer for 3D object detection

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.816793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:06.650142Z digest=sha256:2135f9d9f33f5a21f952cb7170f92cfd88f299c00829460841d55fb719c84319

Observation e84cafae-c49f-4a74-b668-7d794db8d9da · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes CogVLM: Visual Expert for Pretrained Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:06.743796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:06.743796Z digest=sha256:5370d1a2a4a4b9577b05fb366ef32699f2049fd3c474631121325a97a84c5447

Observation af50caab-0370-43a9-b0b5-631aea67a62b · outbound

This paper cites PointCLIP V2: Prompting CLIP and GPT for powerful 3D open-world learning.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes PointCLIP V2: Prompting CLIP and GPT for powerful 3D open-world learning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.799818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:06.857653Z digest=sha256:cae5063ec928cfc6c390f1a343085dc5b0870e27f1d48f0a9996d56c69703eee

Observation f4e88895-2b6d-422c-95bc-46b5a19585ec · outbound

This paper cites Then, the model continues to be trained for 20 epochs.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Then, the model continues to be trained for 20 epochs

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.784314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T17:13:06.935892Z digest=sha256:8bea821e89df6d071ecd0df88a0fee4e2c031356d1926d18f6583960ddfc78f6

Observation 3a331ea2-4e45-46c3-b90b-5d7b8e65f565 · outbound

This paper cites OneLLM: One Framework to Align All Modalities with Language.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes OneLLM: One Framework to Align All Modalities with Language

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:02.378508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:02.378508Z digest=sha256:c95a87c8334cfb14d8b2b6447acc1e4786b04000ebbf11dcbf4b1c69dfa5c836

Pith citing papers

No inbound Pith citation observations are available.