Pith. sign in

Paper Citation Record · LEDGER

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet

As of 10 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2607.06097.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06097 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-08T16:39:53.448383Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact5
  • verified fuzzy49
  • unresolved0
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 770cf900-7afa-4733-9f33-39959fc5b2bf · outbound

This paper cites Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.068846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:2ae83ca125928f9721e1b99332f79ec2365879e6574bac16d0f4c35fa0194ab3

Observation 2c946741-9a0c-4648-be0f-92d6f0977cbb · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.070455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:8bbed8da6f28568dc4284f3e3dd52365970a55b82aac83f90fd0e6589c5f4d66

Observation 5a8dc24a-39b3-4a2e-9951-97b85b7a79be · outbound

This paper cites 3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet 3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.093792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:c25463e0187d40068395920626e0c9ec62d6fc6600073f14e571e7086b85fe57

Observation 7995818b-976a-4a7b-92f5-cb1da9b0de65 · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Scanrefer: 3d object localization in rgb-d scans using natural language

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.092047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:b1260377aae1122a8d0e9bf4b047fa9edf7e7046197589478a065946a74c2550

Observation 968b226e-7f3d-4e69-ad5d-54b99e90cf27 · outbound

This paper cites D 3 net: A unified speaker-listener architecture for 3d dense captioning and visual grounding.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet D 3 net: A unified speaker-listener architecture for 3d dense captioning and visual grounding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.063322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:c9fd493e91ccfb701a6dea03f6bfbb6c31c6f12cc851a2b139ec0c3cea2b069e

Observation 857daf50-1891-4fdf-ba81-64f5b0b19714 · outbound

This paper cites End-to-end 3d dense captioning with vote2cap-detr.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet End-to-end 3d dense captioning with vote2cap-detr

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.067137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:fcfe793bf82ed5a392e6ba717a77224109c84c460b77da8e7dccd3dda679f8b3

Observation b9b56170-c9d7-4ce9-b064-c85786464f7f · outbound

This paper cites V ote2cap-detr++: Decoupling localization and describing for end-to-end 3d dense captioning, 2023.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet V ote2cap-detr++: Decoupling localization and describing for end-to-end 3d dense captioning, 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.070675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:a7945f2f5965abd83105fc4fa5f5d9707cb88343506cb3fddf9171cb5b54cda4

Observation a393026d-9d76-4cfc-929c-073404d8e129 · outbound

This paper cites Segment and Select: Vision-Language Segmentation in 3D Scenarios.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Segment and Select: Vision-Language Segmentation in 3D Scenarios

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-08T16:45:07.676993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:be3a8e3cb10f39dbd34a1a18599e1aa775e0fdac33fed52052d8f5e726ab6e59

Observation fa03f646-0c25-40ff-8fbb-f3f4c1cec50f · outbound

This paper cites Uniter: Universal image-text representation learning.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Uniter: Universal image-text representation learning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.095248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:2586cb444f5d80b39e885e3063b060fea6cb082760478800d953224ff5d19189

Observation 46e1c952-63e7-4e4e-ac5e-040a46854904 · outbound

This paper cites Scan2cap: Context-aware dense captioning in rgb- d scans.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Scan2cap: Context-aware dense captioning in rgb- d scans

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.055428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:d8179a9b85c8fb938e4893245b507f2c9138ef78609b7a2b9c31bfd43f21c8ea

Observation d517a76e-1d56-4441-8de4-e4494312825f · outbound

This paper cites Unit3d: A unified trans- former for 3d dense captioning and visual grounding.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Unit3d: A unified trans- former for 3d dense captioning and visual grounding

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.051526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:c47d0c02571b62daee0baeef8698e9284af0770fd29834eb26ec272f1be28d82

Observation 69dc074a-e70a-4b58-b9a8-68ae88779ae0 · outbound

This paper cites Back-tracing representative points for voting- based 3d object detection in point clouds.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Back-tracing representative points for voting- based 3d object detection in point clouds

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.053576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:29ca329e652987020265aa5705d2cf24dbec25480d3fdf36cfc1ad3f1275aaa5

Observation 6f8b0d8d-a725-405a-ba53-9a6c8f64b670 · outbound

This paper cites 4d spatio-temporal convnets: Minkowski convolutional neural networks, 2019.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet 4d spatio-temporal convnets: Minkowski convolutional neural networks, 2019

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.061281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:a4ff31453ae45f68bd193582e67074bcf32a257d2a2d570bc74f858dded33994

Observation bb682023-94e7-4dd0-8de3-e105e7967eae · outbound

This paper cites Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.082660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:f8a16e4178b3542c381b8b57f43c336d1d107e525544d5461d56b6bbe4f921ef

Observation 0c37d686-f28e-44b4-88c5-f3e78b0c0fe3 · outbound

This paper cites V otenet: A deep learning label fusion method for multi-atlas segmenta- tion.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet V otenet: A deep learning label fusion method for multi-atlas segmenta- tion

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.046074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:f069cfa3dc226bea65b4a623057242cc19506f3b0722971d0d4eb9b68c20d975

Observation 6c6aa963-c870-4022-8526-6430dc607748 · outbound

This paper cites An empirical study of training end-to-end vision-and-language transformers.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet An empirical study of training end-to-end vision-and-language transformers

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.049835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:1d14e9d5a7e5b4a9d12c19044985278260cc0aad0e77ce4b840ca622e17a596c

Observation 5ca690bb-f449-47eb-ad4d-3b06f20a5d27 · outbound

This paper cites Learning lightweight lane detection CNNs by self atten- tion distillation.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Learning lightweight lane detection CNNs by self atten- tion distillation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.016672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:f975f36c32bd2604ae6a6ce4e8a6bc67e2cb699f2a22be8985795b46ad2a3cda

Observation 5d5521b8-bc05-4580-8616-3beeca30e485 · outbound

This paper cites Point-to-Voxel Knowledge Distillation for Li- DAR Semantic Segmentation.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Point-to-Voxel Knowledge Distillation for Li- DAR Semantic Segmentation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.025710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:ba6f02458ee1acd5c0361e59fe659637cab7f4e1347c48dc32313f4a7ccd6206

Observation bb32d021-743e-4184-a67d-c1db2e4ac1fe · outbound

This paper cites Scaling up vision-language pre-training for image captioning.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Scaling up vision-language pre-training for image captioning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.015541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:9761d2400c0bdfb8f562051613d35e60de22a28a8ad7ee745d85a029d246e863

Observation c82a5b9a-1dae-4fe3-b802-998d7dd0a511 · outbound

This paper cites Nerf-det++: Incorporat- ing semantic cues and perspective-aware depth supervision for indoor multi-view 3d detection.IEEE Transactions on Image Processing, 2025.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Nerf-det++: Incorporat- ing semantic cues and perspective-aware depth supervision for indoor multi-view 3d detection.IEEE Transactions on Image Processing, 2025

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.007184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:fca7ac2d9479d39f077097dbc0ecc96cf8c95b9468612d8e88cbc610615d7c39

Observation 49ebd448-0871-44d6-a417-5ec01004c8ca · outbound

This paper cites Perturb, predict & para- phrase: Semi-supervised learning using noisy student for im- age captioning.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Perturb, predict & para- phrase: Semi-supervised learning using noisy student for im- age captioning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.028092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:cb329f0761732695366b3331662c5ba4269f1c97d5dd7859aa6257406df1fc95

Observation 2ba7fa23-cc38-42ff-bc3a-dc6805539583 · outbound

This paper cites Recurrent fusion network for image captioning.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Recurrent fusion network for image captioning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.042013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:98018649736148bae68470a8ac6dfa0c9c97ee43cfcc7a648b375bc8933321b0

Observation ace88239-e12d-494b-aa21-3a6de33b8bda · outbound

This paper cites More: Multi-order relation mining for dense captioning in 3d scenes.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet More: Multi-order relation mining for dense captioning in 3d scenes

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.084358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:2b14f07ce3b082beab39afe438f5ca30f15130ee6b7459cc0591e71acc43dc90

Observation 51eada91-8947-4a2a-9920-1917e97ea4c0 · outbound

This paper cites Context-aware alignment and mutual masking for 3d- language pre-training.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Context-aware alignment and mutual masking for 3d- language pre-training

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.103381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:092addabf82fe8792b1e49de1c0db684238f04f065e9fdd589f29d00198b1d90

Observation dc1c7fd1-f2a1-42b7-8d6c-70c35142d45f · outbound

This paper cites DLIP: Distilling Language-Image Pre-training.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet DLIP: Distilling Language-Image Pre-training

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-08T16:45:07.686511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:8956ba1a03f88578880cd2089babcbe41377935fa5345c5519fa577d9f85eda3

Observation ac293cb5-3f5a-4f4e-a4b5-6ada96e398f3 · outbound

This paper cites Align before fuse: Vision and language representation learn- ing with momentum distillation.Advances in neural infor- mation processing systems, 34:9694–9705.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Align before fuse: Vision and language representation learn- ing with momentum distillation.Advances in neural infor- mation processing systems, 34:9694–9705

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.097444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:b6b5a5ad65405d477318807c2093a06b8b76d28fe4ab5a88828b55448b33b28e

Observation ac034ae4-5770-439d-bba8-0701697233c1 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.088568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:f79f4477b6d0b95e840cb163ba582524127149160fb9553ad294d37d6b80cac9

Observation fe1ff3a3-27a5-44c5-9619-369381b78827 · outbound

This paper cites Oscar: Object-semantics aligned pre-training for vision-language tasks.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Oscar: Object-semantics aligned pre-training for vision-language tasks

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.090362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:0e472c3a6f227604081d940d2728a2c0ed634d6d87ebbdecf5c38b7d4113bac8

Observation b2ab76ca-055c-4786-be4d-185cb75fef55 · outbound

This paper cites MoE3D: Mixture of Experts Meets Multi-Modal 3D Understanding.arXiv 2025.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet MoE3D: Mixture of Experts Meets Multi-Modal 3D Understanding.arXiv 2025

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-08T16:45:07.682981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:265121a4352fae8097e66c9ae4e10cce99079c40e39e2d5a20f6bcf3e658063e

Observation 6db7c481-5339-4174-8308-920c24d3d103 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Rouge: A package for automatic evaluation of summaries

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.095433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:f4f8c9cb27b44b0ac46aa95cfe0fc81444ffb6689a671254a292a62bba33693c

Observation 646f077b-536c-43b1-bd6e-da1d1cc4596b · outbound

This paper cites Complete 3d relationships extraction modality align- ment network for 3d dense captioning.IEEE Transactions on Visualization and Computer Graphics, 2023.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Complete 3d relationships extraction modality align- ment network for 3d dense captioning.IEEE Transactions on Visualization and Computer Graphics, 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.105503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:c2fae00ab8f2de153c5ae6af4a16c90bc153b0eeee74edba49374fc9826f065a

Observation 7b7d8d4d-9552-4b08-9b80-34fc72754ab8 · outbound

This paper cites An end-to- end transformer model for 3d object detection.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet An end-to- end transformer model for 3d object detection

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.085182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:5d857d881600e85fa9824f5e8ca7018523226fd6264d787506872f7dacac04e0

Observation 9132598e-e2ba-42d3-a77c-d75faeadf435 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Bleu: a method for automatic evaluation of machine translation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.099619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:48f032ee0f5a9123077697673dbd21ddaf22190d3cd647adc3aefa1fc7a3c79c

Observation b6075cc7-32ed-41c5-b104-b6b5ddb31d35 · outbound

This paper cites Pointnet: Deep learning on point sets for 3d classification and segmentation.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Pointnet: Deep learning on point sets for 3d classification and segmentation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.101499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:a8001412ef6ec0aa04fc1a72b2fcd587104724fff82b33dc2744626bce56f59d

Observation e65cc2fb-4474-4c2f-aa3e-7e2f27e373c5 · outbound

This paper cites Pointnet++: Deep hierarchical feature learning on point sets in a metric space.Advances in neural information processing systems, 30, 2017.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Pointnet++: Deep hierarchical feature learning on point sets in a metric space.Advances in neural information processing systems, 30, 2017

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.079181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:2e2b75b6693ffa3a3a417f19cc7e0a809eddc03b3aba842d704629cdc49cdc49

Observation d9b72fbd-de5f-4b21-b9e9-0ab36724dd4b · outbound

This paper cites Qi, Or Litany, Kaiming He, and Leonidas J.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Qi, Or Litany, Kaiming He, and Leonidas J

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.081669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:52572c0f4e944ed7cdbb1ae50bdbd23236ff9a3c8dfff3ddda1a2cf86f1ab41e

Observation 06a05bf8-35a4-4a7c-9472-b64c4e205fd7 · outbound

This paper cites Rennie, Etienne Marcheret, Youssef Mroueh, Jarret Ross, and Vaibhava Goel.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Rennie, Etienne Marcheret, Youssef Mroueh, Jarret Ross, and Vaibhava Goel

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.083499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:132ea6e31efeb31cef31a534307e2d97f0a5cb5b1f674553f08e34e58ac92000

Observation f23972b7-ab42-4b7d-bb5f-85c097badd53 · outbound

This paper cites Octnet: Learning deep 3d representations at high resolutions.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Octnet: Learning deep 3d representations at high resolutions

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.091236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:1099b89f1c981b70f8e1fad4e74dafc699fc24b4b480b0e411b017495c353057

Observation 892a1f4d-15d4-49a2-bad3-af402d4079a5 · outbound

This paper cites Fcaf3d: Fully convolutional anchor-free 3d object detection.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Fcaf3d: Fully convolutional anchor-free 3d object detection

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.072951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:e2142b3e774004adb18566b8a2880ee8f58de1aa4363556aac87aee2e0069fc1

Observation 76db27ec-5603-4ef8-8737-61444b85fab7 · outbound

This paper cites an unresolved cited work.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Unresolved cited work

Reference 40

Resolution
parse uncertain
raw_fallback, observed 2026-07-08T16:45:08.066933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:08c145e917a9ac216e507988e8e5b9e3b0cdd857d9556283f633ec91d6bb6385

Observation e438d67d-24ca-44e6-8b77-6e4ecf96f82e · outbound

This paper cites VL-BERT: Pre-training of Generic Visual-Linguistic Representations.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet VL-BERT: Pre-training of Generic Visual-Linguistic Representations

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-08T16:45:07.679500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:66c07148a662bc4d61e1fe650b932c91fc9ea71f54d7125856c7e64932f04bd3

Observation dfcecb9e-7e53-4638-b14d-87e9fff93cdb · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Lawrence Zitnick, and Devi Parikh

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.089264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:dc63cee448d89a74863ad214e014a5cd527819ab8fd433b8403aeebda02f1d9b

Observation 65cc72ba-d572-4462-9755-2aaa13a88a01 · outbound

This paper cites Cagroup3d: Class- aware grouping for 3d object detection on point clouds.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Cagroup3d: Class- aware grouping for 3d object detection on point clouds

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.059134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:e0aed3ac877daaaf3a125fbf3906e1695e6d7abcc97882ad33970400db1bfbdb

Observation 843ef6ab-92be-4618-9218-1b0916d1afc5 · outbound

This paper cites Rbgnet: Ray-based grouping for 3d object detection.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Rbgnet: Ray-based grouping for 3d object detection

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.055231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:30d0500acce7c5f514203d789b94f0406a5097d7d5b2cb4175e11e3359f47b08

Observation 4f20f37f-a401-4d89-90ef-215406087962 · outbound

This paper cites Spatiality-guided Transformer for 3D Dense Captioning on Point Clouds.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Spatiality-guided Transformer for 3D Dense Captioning on Point Clouds

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-07-08T16:45:07.674452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:c37a8cb6dfd5e2e90163cb3ff6d1171df28e6a919fd53de7a0280082aa5710ce

Observation 473c21da-dfb2-4285-b3c8-2732efcefc86 · outbound

This paper cites Point transformer v2: Grouped vector atten- tion and partition-based pooling, 2022.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Point transformer v2: Grouped vector atten- tion and partition-based pooling, 2022

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.057219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:82f498566e4246d025dfd93365d900f63815219815257f0ba6bdbfa7c46ecdc3

Observation 3eef39bf-dfb0-4a07-ab9d-3a66fcc82be4 · outbound

This paper cites Taseg: Temporal aggregation network for lidar semantic segmentation.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Taseg: Temporal aggregation network for lidar semantic segmentation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.046852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:4f4ba397fb9a4d5413b3244aee73779ec9bc255cbf4ef90576b6fe6028020a02

Observation 3e809ed1-5a40-47b4-aa28-94c4dbec1535 · outbound

This paper cites Second: Sparsely embed- ded convolutional detection.Sensors, 18(10):3337.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Second: Sparsely embed- ded convolutional detection.Sensors, 18(10):3337

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.107399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:ca75ca0f06060ffc237f3cc231fab077425e7db8f4480ae95a8a5302323b33b1

Observation 126695a3-1139-4a12-902b-90a0e4eac2bf · outbound

This paper cites Vision-language pre-training with triple contrastive learning.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Vision-language pre-training with triple contrastive learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.053448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:d8c989d4093905377d312421a34462ad9f9aec64628828a2810b89db21acb4bf

Observation 9b088032-702b-491d-b486-7351219afa1b · outbound

This paper cites Swin3d: A pretrained transformer backbone for 3d indoor scene understanding, 2023.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Swin3d: A pretrained transformer backbone for 3d indoor scene understanding, 2023

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.068658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:f30ee2f0c1c9e0dd1ff20ae2a70c655f8ec1b22331192f7bb83f94789f4f8041

Observation 20835cc4-818b-4f34-9c78-5bd570f1c944 · outbound

This paper cites X-trans2cap: Cross- modal knowledge transfer using transformer for 3d dense captioning.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet X-trans2cap: Cross- modal knowledge transfer using transformer for 3d dense captioning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.099452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:2f82746f44762be72bd03672ba2b6091a6489d08e243b6b7122ec0e2ad9c5f08

Observation 6557e455-29be-498b-a661-e0e011d43a58 · outbound

This paper cites H3dnet: 3d object detection using hybrid geometric primi- tives.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet H3dnet: 3d object detection using hybrid geometric primi- tives

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.023270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:496b2cbb76ad1c642affa04c3a38c18d10984c2257d0e80dd5abb81fb28dc5a3

Observation fb825369-20a4-4810-b10a-8977896acd63 · outbound

This paper cites Contextual modeling for 3d dense captioning on point clouds, 2022.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet Contextual modeling for 3d dense captioning on point clouds, 2022

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.025537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:25510b3ef58b2031b41c7b9f647ecc4a8596030cebad31b0ccad2a2e7e31f9df

Observation 01df9144-44bd-4d80-8a35-2f8247dfedd1 · outbound

This paper cites V oxelnet: End-to-end learning for point cloud based 3d object detection.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet V oxelnet: End-to-end learning for point cloud based 3d object detection

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.014418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:43d73b1d74f115b92a4a5c1a3b8d12a882c1f424bf8029ea0de8cd579c282098

Observation 9b2aa759-6b9f-4e44-bba0-7200dcde6860 · outbound

This paper cites 3d-vista: Pre-trained transformer for 3d vision and text alignment.

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet 3d-vista: Pre-trained transformer for 3d vision and text alignment

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:45:08.012034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T16:39:53.448383Z digest=sha256:fc085c7e468edcc0a74e8f8cd1855fb14367969336be6a26025054578682e193

Pith citing papers

No inbound Pith citation observations are available.