Pith. sign in

Paper Citation Record · LEDGER

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet

As of 9 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2607.06097.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06097 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-08T16:39:53.448383Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact5
  • verified fuzzy49
  • unresolved0
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 770cf900-7afa-4733-9f33-39959fc5b2bf · outbound

This paper cites Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.068846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:703d1c195931aae12ca6a8384039796fb008e04ba5f5f563b01903b7c4146dbe

Observation 2c946741-9a0c-4648-be0f-92d6f0977cbb · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.070455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:98b2fea4cf138b764ebd949986488f2790d8d1426818d598c09e880159034c1c

Observation 5a8dc24a-39b3-4a2e-9951-97b85b7a79be · outbound

This paper cites 3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet 3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.093792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:f41593fd3ec2b9d043f0140989689dee4d7a7279a1ea2bf2b8ae162ec282fa05

Observation 7995818b-976a-4a7b-92f5-cb1da9b0de65 · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Scanrefer: 3d object localization in rgb-d scans using natural language

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.092047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:9612ae0fd5ed286f3c0aef408f5e6bf0f7556502fa48c9434daa9c6dff014b0f

Observation 968b226e-7f3d-4e69-ad5d-54b99e90cf27 · outbound

This paper cites D 3 net: A unified speaker-listener architecture for 3d dense captioning and visual grounding.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet D 3 net: A unified speaker-listener architecture for 3d dense captioning and visual grounding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.063322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:78a6c3dc722fe8b7095c28a84e775387cc5a3f9b0dfc8b82c4c39060cd6c7039

Observation 857daf50-1891-4fdf-ba81-64f5b0b19714 · outbound

This paper cites End-to-end 3d dense captioning with vote2cap-detr.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet End-to-end 3d dense captioning with vote2cap-detr

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.067137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:37ce92cbac4c5c3831c89818ae3706e7768db552998c9607deba5fcd753efcf0

Observation b9b56170-c9d7-4ce9-b064-c85786464f7f · outbound

This paper cites V ote2cap-detr++: Decoupling localization and describing for end-to-end 3d dense captioning, 2023.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet V ote2cap-detr++: Decoupling localization and describing for end-to-end 3d dense captioning, 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.070675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:771714cd476e6eac8facb8ccda761f56070ddde3fe57ddb414db2f5651bc1f59

Observation a393026d-9d76-4cfc-929c-073404d8e129 · outbound

This paper cites Segment and Select: Vision-Language Segmentation in 3D Scenarios.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Segment and Select: Vision-Language Segmentation in 3D Scenarios

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-08T16:45:07.676993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:87a00dcb6d16e5d70f16635290917a0135e75985f2e2aed05973766e96a69bfc

Observation fa03f646-0c25-40ff-8fbb-f3f4c1cec50f · outbound

This paper cites Uniter: Universal image-text representation learning.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Uniter: Universal image-text representation learning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.095248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:f40a6d176342716b416d1ac881fd946108bd8ee58b6021f2623dc2b2acef0740

Observation 46e1c952-63e7-4e4e-ac5e-040a46854904 · outbound

This paper cites Scan2cap: Context-aware dense captioning in rgb- d scans.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Scan2cap: Context-aware dense captioning in rgb- d scans

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.055428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:b6e5705f77689cf1b26bc0e48039e31a144a6ffdf9ee37ad99ae9c9e3c3d17cc

Observation d517a76e-1d56-4441-8de4-e4494312825f · outbound

This paper cites Unit3d: A unified trans- former for 3d dense captioning and visual grounding.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Unit3d: A unified trans- former for 3d dense captioning and visual grounding

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.051526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:24a0ba6b6403bb0d2727f1c713a2071892a2edf662e8532a6b62ef29b61c2bac

Observation 69dc074a-e70a-4b58-b9a8-68ae88779ae0 · outbound

This paper cites Back-tracing representative points for voting- based 3d object detection in point clouds.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Back-tracing representative points for voting- based 3d object detection in point clouds

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.053576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:84b37e4aebfcff2aebf07dd7a7093cfa3dd4b61cae1ab25471d542d6db4d442e

Observation 6f8b0d8d-a725-405a-ba53-9a6c8f64b670 · outbound

This paper cites 4d spatio-temporal convnets: Minkowski convolutional neural networks, 2019.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet 4d spatio-temporal convnets: Minkowski convolutional neural networks, 2019

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.061281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:da0c1d90cd6da7fccf3e18df895805c93a9c469b4b51b76f5c75aea259106112

Observation bb682023-94e7-4dd0-8de3-e105e7967eae · outbound

This paper cites Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.082660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:bd686d2ee392cd4c0b3736982aca4539c034ca20fe3ee9cd05ed75310fe2ad74

Observation 0c37d686-f28e-44b4-88c5-f3e78b0c0fe3 · outbound

This paper cites V otenet: A deep learning label fusion method for multi-atlas segmenta- tion.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet V otenet: A deep learning label fusion method for multi-atlas segmenta- tion

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.046074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:aa820955c940be768312e5bf34045c226bfc918fda05382060a5b58badd069a0

Observation 6c6aa963-c870-4022-8526-6430dc607748 · outbound

This paper cites An empirical study of training end-to-end vision-and-language transformers.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet An empirical study of training end-to-end vision-and-language transformers

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.049835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:8e84462d4b7694a8f75796a87e2b41eb88522dd969f6baa97b7f277d4bd8d6a9

Observation 5ca690bb-f449-47eb-ad4d-3b06f20a5d27 · outbound

This paper cites Learning lightweight lane detection CNNs by self atten- tion distillation.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Learning lightweight lane detection CNNs by self atten- tion distillation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.016672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:bf33e9ba03affc7ae0722e707d7027aa9497ec2b56e6a98409834fb4785614d5

Observation 5d5521b8-bc05-4580-8616-3beeca30e485 · outbound

This paper cites Point-to-Voxel Knowledge Distillation for Li- DAR Semantic Segmentation.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Point-to-Voxel Knowledge Distillation for Li- DAR Semantic Segmentation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.025710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:edd685b033cbe04e68d4a42d9d5d6af61dd0bec2d29417e002085db148d21e35

Observation bb32d021-743e-4184-a67d-c1db2e4ac1fe · outbound

This paper cites Scaling up vision-language pre-training for image captioning.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Scaling up vision-language pre-training for image captioning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.015541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:85f55decf3f56c63b88f6e099925d4712138a5187386daee8f8dc8a2812f265f

Observation c82a5b9a-1dae-4fe3-b802-998d7dd0a511 · outbound

This paper cites Nerf-det++: Incorporat- ing semantic cues and perspective-aware depth supervision for indoor multi-view 3d detection.IEEE Transactions on Image Processing, 2025.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Nerf-det++: Incorporat- ing semantic cues and perspective-aware depth supervision for indoor multi-view 3d detection.IEEE Transactions on Image Processing, 2025

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.007184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:0571ee2862cd781adac0ae71a741eaa154b28c899776791c774215e9383b69aa

Observation 49ebd448-0871-44d6-a417-5ec01004c8ca · outbound

This paper cites Perturb, predict & para- phrase: Semi-supervised learning using noisy student for im- age captioning.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Perturb, predict & para- phrase: Semi-supervised learning using noisy student for im- age captioning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.028092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:520cb90e6acef22c7524f2c47c70b7a31ea855cea23c84809b1ca868bc362799

Observation 2ba7fa23-cc38-42ff-bc3a-dc6805539583 · outbound

This paper cites Recurrent fusion network for image captioning.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Recurrent fusion network for image captioning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.042013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:4f04ce74f74c014ba40b26012ecb569b836402b43fe8e6e8318abbb72d61ac0e

Observation ace88239-e12d-494b-aa21-3a6de33b8bda · outbound

This paper cites More: Multi-order relation mining for dense captioning in 3d scenes.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet More: Multi-order relation mining for dense captioning in 3d scenes

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.084358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:de7d3cbfeaed2a5f79978b0b024cd4234de883b147772b468e499b0f25059ab0

Observation 51eada91-8947-4a2a-9920-1917e97ea4c0 · outbound

This paper cites Context-aware alignment and mutual masking for 3d- language pre-training.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Context-aware alignment and mutual masking for 3d- language pre-training

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.103381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:7e06bd964e59bc8aec6c5a2ebd3f4e28ba67e6b50dda699d7c27fa3594a09512

Observation dc1c7fd1-f2a1-42b7-8d6c-70c35142d45f · outbound

This paper cites DLIP: Distilling Language-Image Pre-training.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet DLIP: Distilling Language-Image Pre-training

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-08T16:45:07.686511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:11d8f36538a5cf4310798fcab6fed74dd4e3720ba27d6daacdb7bc375dc890b9

Observation ac293cb5-3f5a-4f4e-a4b5-6ada96e398f3 · outbound

This paper cites Align before fuse: Vision and language representation learn- ing with momentum distillation.Advances in neural infor- mation processing systems, 34:9694–9705.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Align before fuse: Vision and language representation learn- ing with momentum distillation.Advances in neural infor- mation processing systems, 34:9694–9705

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.097444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:6b142746ab6d7aab976f0124e7d04f765ff29bc66246f943931577e0e8689451

Observation ac034ae4-5770-439d-bba8-0701697233c1 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.088568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:e2541626f113cf6a240efa3393fee9dcded7d1f238f5bd7e5105d3f8d14d142a

Observation fe1ff3a3-27a5-44c5-9619-369381b78827 · outbound

This paper cites Oscar: Object-semantics aligned pre-training for vision-language tasks.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Oscar: Object-semantics aligned pre-training for vision-language tasks

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.090362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:fac40a7962585d9c9c7c4286123986939efb8b5f1ef7dad703d6a72d9e6ab477

Observation b2ab76ca-055c-4786-be4d-185cb75fef55 · outbound

This paper cites MoE3D: Mixture of Experts Meets Multi-Modal 3D Understanding.arXiv 2025.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet MoE3D: Mixture of Experts Meets Multi-Modal 3D Understanding.arXiv 2025

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-08T16:45:07.682981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:4c8af67902f1f2bdb54787cd626f7e68b7bad45c3d4a62574b234f79594058a2

Observation 6db7c481-5339-4174-8308-920c24d3d103 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Rouge: A package for automatic evaluation of summaries

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.095433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:1b9516509b473623db638f307991d0029c87458e68e09c0f959d8e494bb589ec

Observation 646f077b-536c-43b1-bd6e-da1d1cc4596b · outbound

This paper cites Complete 3d relationships extraction modality align- ment network for 3d dense captioning.IEEE Transactions on Visualization and Computer Graphics, 2023.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Complete 3d relationships extraction modality align- ment network for 3d dense captioning.IEEE Transactions on Visualization and Computer Graphics, 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.105503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:bd01b7275e5a7afd566310fc207722b46df1a296f484800788a5f213162f0116

Observation 7b7d8d4d-9552-4b08-9b80-34fc72754ab8 · outbound

This paper cites An end-to- end transformer model for 3d object detection.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet An end-to- end transformer model for 3d object detection

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.085182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:215beaa62709a6120408492ff4b27d2a7a5a08443a48b9bd6b1d1199adaeb5d2

Observation 9132598e-e2ba-42d3-a77c-d75faeadf435 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Bleu: a method for automatic evaluation of machine translation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.099619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:e4e440b0637e3eebdd0932f082a90ac031f40f721471c6c850cf916ac1521e58

Observation b6075cc7-32ed-41c5-b104-b6b5ddb31d35 · outbound

This paper cites Pointnet: Deep learning on point sets for 3d classification and segmentation.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Pointnet: Deep learning on point sets for 3d classification and segmentation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.101499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:95880d62cd490b42750496041d64cb7493b4727a31c5d04c8e5d25630dd58ae1

Observation e65cc2fb-4474-4c2f-aa3e-7e2f27e373c5 · outbound

This paper cites Pointnet++: Deep hierarchical feature learning on point sets in a metric space.Advances in neural information processing systems, 30, 2017.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Pointnet++: Deep hierarchical feature learning on point sets in a metric space.Advances in neural information processing systems, 30, 2017

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.079181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:b696601a9d0a0df671eb8c05614606daa04209deba8218736a8ab2e5ced48dcd

Observation d9b72fbd-de5f-4b21-b9e9-0ab36724dd4b · outbound

This paper cites Qi, Or Litany, Kaiming He, and Leonidas J.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Qi, Or Litany, Kaiming He, and Leonidas J

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.081669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:4f1d7631045da64349730806abbc555b72ed6da52c3ad6e3b3cb6c3b3ea93d38

Observation 06a05bf8-35a4-4a7c-9472-b64c4e205fd7 · outbound

This paper cites Rennie, Etienne Marcheret, Youssef Mroueh, Jarret Ross, and Vaibhava Goel.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Rennie, Etienne Marcheret, Youssef Mroueh, Jarret Ross, and Vaibhava Goel

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.083499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:164cf33a6edd0fd34ec99e66fec5f2188a7076cfdde3d3d72cfb04a580c43e78

Observation f23972b7-ab42-4b7d-bb5f-85c097badd53 · outbound

This paper cites Octnet: Learning deep 3d representations at high resolutions.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Octnet: Learning deep 3d representations at high resolutions

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.091236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:a6f51e4fdac9c0c4bafefbe42cb9cb346c81aea6dbd85c2f6e874216ec4c10c7

Observation 892a1f4d-15d4-49a2-bad3-af402d4079a5 · outbound

This paper cites Fcaf3d: Fully convolutional anchor-free 3d object detection.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Fcaf3d: Fully convolutional anchor-free 3d object detection

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.072951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:f5b16ab7d74283a89b3f948c5971d253f87e6735a87f0d483014fe9ff8c67bcd

Observation 76db27ec-5603-4ef8-8737-61444b85fab7 · outbound

This paper cites an unresolved cited work.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Unresolved cited work

Reference 40

Resolution
parse uncertain
raw_fallback, observed 2026-07-08T16:45:08.066933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:d88c7cbb2ca2de07aaa151efd5fc344f9bbb7f6c18187c76c7b0cc14f36e5641

Observation e438d67d-24ca-44e6-8b77-6e4ecf96f82e · outbound

This paper cites VL-BERT: Pre-training of Generic Visual-Linguistic Representations.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet VL-BERT: Pre-training of Generic Visual-Linguistic Representations

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-08T16:45:07.679500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:715351ac0850706a014e0561049f0abdc09738850ae0d54ca0f8b15e1aeb7f3d

Observation dfcecb9e-7e53-4638-b14d-87e9fff93cdb · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Lawrence Zitnick, and Devi Parikh

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.089264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:2cef3e53345a585874a55dbc8149cc9fe0d9642c296cb27bf643d7d3220e1352

Observation 65cc72ba-d572-4462-9755-2aaa13a88a01 · outbound

This paper cites Cagroup3d: Class- aware grouping for 3d object detection on point clouds.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Cagroup3d: Class- aware grouping for 3d object detection on point clouds

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.059134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:a161334933a650d93415ebf4ea2bab316ebbc3a19511b6a338fc2865fcbf4153

Observation 843ef6ab-92be-4618-9218-1b0916d1afc5 · outbound

This paper cites Rbgnet: Ray-based grouping for 3d object detection.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Rbgnet: Ray-based grouping for 3d object detection

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.055231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:059c38799e5ac544dfffb1ab33a488890f80d6e0bdbb60a4a73900f95bee8f6a

Observation 4f20f37f-a401-4d89-90ef-215406087962 · outbound

This paper cites Spatiality-guided Transformer for 3D Dense Captioning on Point Clouds.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Spatiality-guided Transformer for 3D Dense Captioning on Point Clouds

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-07-08T16:45:07.674452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:58e5ddb0d92c94492c5c014f8ac9dad72372d3f9eecdce9e0bf8ef787c2ad883

Observation 473c21da-dfb2-4285-b3c8-2732efcefc86 · outbound

This paper cites Point transformer v2: Grouped vector atten- tion and partition-based pooling, 2022.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Point transformer v2: Grouped vector atten- tion and partition-based pooling, 2022

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.057219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:adb54fab57b2d8abda66f87fdbab60d221cd83d579a08e86ee04896195fb1f0b

Observation 3eef39bf-dfb0-4a07-ab9d-3a66fcc82be4 · outbound

This paper cites Taseg: Temporal aggregation network for lidar semantic segmentation.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Taseg: Temporal aggregation network for lidar semantic segmentation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.046852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:1c05242261df2054a717e77e9be8fb2ac91ecff7f4c6d3502d3bcc33e79f2929

Observation 3e809ed1-5a40-47b4-aa28-94c4dbec1535 · outbound

This paper cites Second: Sparsely embed- ded convolutional detection.Sensors, 18(10):3337.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Second: Sparsely embed- ded convolutional detection.Sensors, 18(10):3337

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.107399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:2c9c33121cc3f4f24bdcf3f01c684a61e194450bbff3c6ca409abb604b9cb0de

Observation 126695a3-1139-4a12-902b-90a0e4eac2bf · outbound

This paper cites Vision-language pre-training with triple contrastive learning.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Vision-language pre-training with triple contrastive learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.053448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:92d2639ce3be2cbf44ce117f502d5292b3316174e120b2907fce5a0b3c2d9c8d

Observation 9b088032-702b-491d-b486-7351219afa1b · outbound

This paper cites Swin3d: A pretrained transformer backbone for 3d indoor scene understanding, 2023.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Swin3d: A pretrained transformer backbone for 3d indoor scene understanding, 2023

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.068658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:c003875731439b315e6427f542eca3742270078cde75b97de309320c739560e0

Observation 20835cc4-818b-4f34-9c78-5bd570f1c944 · outbound

This paper cites X-trans2cap: Cross- modal knowledge transfer using transformer for 3d dense captioning.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet X-trans2cap: Cross- modal knowledge transfer using transformer for 3d dense captioning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.099452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:583ad18729261e563771cf50b1b4e3b40a71138001ac279a6f9ffafd94079d4a

Observation 6557e455-29be-498b-a661-e0e011d43a58 · outbound

This paper cites H3dnet: 3d object detection using hybrid geometric primi- tives.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet H3dnet: 3d object detection using hybrid geometric primi- tives

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.023270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:4c7d15e718e38b7e7de8b4bb502f5ee90f3c3f30149d5916ab45fccd2a29975e

Observation fb825369-20a4-4810-b10a-8977896acd63 · outbound

This paper cites Contextual modeling for 3d dense captioning on point clouds, 2022.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Contextual modeling for 3d dense captioning on point clouds, 2022

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.025537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:d484584a2b16f8a69972115c2f63c0c042eb3083f4eed9cc00583780bf056c4d

Observation 01df9144-44bd-4d80-8a35-2f8247dfedd1 · outbound

This paper cites V oxelnet: End-to-end learning for point cloud based 3d object detection.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet V oxelnet: End-to-end learning for point cloud based 3d object detection

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.014418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:c86d0bc4d383baadbc2e8b392fa1874d7ec2efacdf3d2a98cc76a7c8273b7b08

Observation 9b2aa759-6b9f-4e44-bba0-7200dcde6860 · outbound

This paper cites 3d-vista: Pre-trained transformer for 3d vision and text alignment.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet 3d-vista: Pre-trained transformer for 3d vision and text alignment

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.012034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:72f5dda164d09ae7dff5de53e82910c7ea64a6e4d636575c5a9c883a03886ceb

Pith citing papers

No inbound Pith citation observations are available.