Pith. sign in

Paper Citation Record · LEDGER

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks

As of 8 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 1 inbound Pith citation observation for arXiv:2505.16594.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16594 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:02:31.527586Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T18:06:33.139310Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T05:30:58.338283Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact4
  • verified fuzzy13
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6a669fc7-60d9-4fb6-934c-512dd9d7dfc8 · outbound

This paper cites MaMMUT: A Simple Architecture for Joint Learning for MultiModal Tasks.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks MaMMUT: A Simple Architecture for Joint Learning for MultiModal Tasks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:28.003967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:28.003967Z digest=sha256:ddfac1d122550658a597310bd9bcca40e77905f7f5cd04d81a2fea54d0011a00

Observation b6247258-8cda-4e7a-a7c6-2fc1cfed444c · outbound

This paper cites VLAB: Enhancing Video Language Pre-training by Feature Adapting and Blending.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks VLAB: Enhancing Video Language Pre-training by Feature Adapting and Blending

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:28.118849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:28.118849Z digest=sha256:7df9d7571410a5cdd60c739a8709ba2e09fa08ab13f8b5f6bba0d19fbfc460fa

Observation 3b89894d-f481-4bf4-bdc6-2a0216f92b91 · outbound

This paper cites Valor: Vision-audio-language omni-perception pretraining model and dataset,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Valor: Vision-audio-language omni-perception pretraining model and dataset,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:34.415561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:28.210380Z digest=sha256:40db3ee007ac2e453b3027d5cd96d1a01370433106ef5311f59b7810d0e62ebe

Observation e6e64717-ff25-49de-8e4f-328ddcff60d6 · outbound

This paper cites Self-supervised Spatio-temporal Representation Learning for Videos by Predicting Motion and Appearance Statistics.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Self-supervised Spatio-temporal Representation Learning for Videos by Predicting Motion and Appearance Statistics

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:02:32.305348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:28.384506Z digest=sha256:d6eda5ceee3de3d19f3a5b2dc0bd12f1c25215e3cda9b45aadab60af1099c99e

Observation 4f7c5dd5-252b-4021-9d20-722220d3cc6c · outbound

This paper cites Unsupervised Pre-Training of Image Features on Non-Curated Data.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Unsupervised Pre-Training of Image Features on Non-Curated Data

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:02:32.192245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:28.503146Z digest=sha256:5b2d11e13323699ee225ec29ae1c2efb2782480696271eae7cc2323ea2e72440

Observation b997e172-681e-4e16-924d-19d87ae7e4c3 · outbound

This paper cites Melanoma thick- ness prediction based on convolutional neural network with vgg-19 model transfer learning,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Melanoma thick- ness prediction based on convolutional neural network with vgg-19 model transfer learning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:34.280853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:28.606666Z digest=sha256:62ca003debec198c5a18a97448a6eba40c00fa3790413d9321957d65e1f6234f

Observation 9e3a94ff-32d1-4c6f-bc57-f825a6be1c7f · outbound

This paper cites Pre-training on grayscale imagenet improves medical image classification,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Pre-training on grayscale imagenet improves medical image classification,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:34.149539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:28.716523Z digest=sha256:fb5dc60402f0097ba50f6210a8ac439f80573f45a9626aff8a812b57bb08fbe6

Observation 76007362-10d8-40c1-a4fc-dd3890e65c39 · outbound

This paper cites Low-Rank HOCA: Efficient High-Order Cross-Modal Attention for Video Captioning.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Low-Rank HOCA: Efficient High-Order Cross-Modal Attention for Video Captioning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:02:32.020124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:28.782114Z digest=sha256:b80119026fb5ad69b5a23751653f78a0ac4d149abb9a8feae7b86cae4d9489d5

Observation 2f3fedb6-ae83-4424-bcc7-ae7b5f82b368 · outbound

This paper cites Sensor-augmented egocentric- video captioning with dynamic modal attention,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Sensor-augmented egocentric- video captioning with dynamic modal attention,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:33.986897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:28.879984Z digest=sha256:8201e3b8d701ee1d3d35a20dd96fae14b925ebffe85f1f822736c0b7a2975e3c

Observation b5cdd1d4-1dca-4046-ab49-b6c24b18fef1 · outbound

This paper cites Boosting Video Captioning with Dynamic Loss Network.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Boosting Video Captioning with Dynamic Loss Network

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:02:31.865243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:28.972803Z digest=sha256:9caf6207f34f1b28d6b91af9030df8138a2b95e3a7b112df0d26fc5d65de72e7

Observation 5634ddb2-30e4-4d4d-8e6d-29b3333a1cc4 · outbound

This paper cites Panoptic segmentation,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Panoptic segmentation,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.036353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.036353Z digest=sha256:098691ff961573194b60ca747d9c0d035a55c8f17199913d7ee4ba16342b3061

Observation 949b1838-94dc-43d9-b8fa-5857cc82c437 · outbound

This paper cites An empirical study of context in object detection,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks An empirical study of context in object detection,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:33.817392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:29.104499Z digest=sha256:e87ca8723be71c95dfaa8981551d908f3abd5b30f3dcbbcd1531027410e9814e

Observation afa08809-6553-456f-99cb-ef07545707bb · outbound

This paper cites End-to-end object detection with transformers,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks End-to-end object detection with transformers,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.171195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.171195Z digest=sha256:cc16a0238e1a010db1de856c0ebc0b80397d61eb7784fb388e8473ca555a175b

Observation aa1a084a-51b0-4e57-94fe-f26759cc1cfd · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.260811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.260811Z digest=sha256:f414e1bfbda46513d8015af8ef0a96c1052ea5d7ecbd12c939ebf6ec7be4ca83

Observation b85f535a-1e35-4405-ac04-20df1c5043b4 · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.351108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.351108Z digest=sha256:872c9052b31d36d03802e8428566be57e2f8e74dc9a54f3027de042ff5f60e5c

Observation a9878bc4-bf6a-46ea-a9c2-b4267b312608 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.424508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.424508Z digest=sha256:bc759d3d5e9e2438c66a9eccc4fd0668224f302bcaf2b8715776685331d2f98f

Observation 9225c940-9e52-4dae-8ea7-c98fe7d9f569 · outbound

This paper cites Swinbert: End-to-end transformers with sparse attention for video captioning,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Swinbert: End-to-end transformers with sparse attention for video captioning,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:33.688236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:29.503052Z digest=sha256:80df03a2c876f020f56dc8013c70be7b28bfafea6859be07d9c114fe40ea65c4

Observation 2e89f73c-4b20-447d-bb17-e1b31044843a · outbound

This paper cites Lingoqa: Visual question answering for autonomous driving.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Lingoqa: Visual question answering for autonomous driving

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:33.524482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:29.555237Z digest=sha256:33973b5e7348116213ad75e537c208a391c49a7bcf2840ae1d1bc0457800b9c3

Observation a9ad56f4-133c-4ee6-bf79-03698936ac9c · outbound

This paper cites Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.616411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.616411Z digest=sha256:6fd31e04c8d28e8afa0ce245b479e22ea001fd4e7a0bba5cde27dd9aaff47bfb

Observation 925d6cbf-a016-4395-9a47-23a0f46b9d60 · outbound

This paper cites Nuscenes-mqa: Integrated evaluation of captions and qa for autonomous driving datasets using markup annotations,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Nuscenes-mqa: Integrated evaluation of captions and qa for autonomous driving datasets using markup annotations,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:33.264972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:29.651915Z digest=sha256:2cc738df61140f3a1aa08af6f7cde25d5dee1cd6476cf9ebadec46f751c120be

Observation 6b766470-5748-44f5-891c-013d71e87bed · outbound

This paper cites Language Prompt for Autonomous Driving.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Language Prompt for Autonomous Driving

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.686171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.686171Z digest=sha256:b72a69f0b39d33c63d0044a4418fbaf0aa5d62edae2c486c22e2096f6257e8e1

Observation 25fab8d1-523d-4e0e-9df9-39006acd7f1e · outbound

This paper cites Covla: Comprehensive vision-language-action dataset for autonomous driving,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Covla: Comprehensive vision-language-action dataset for autonomous driving,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.741400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.741400Z digest=sha256:bc8490c4e0d9bcbdf3a25a48e6fafdfd0c819e08833286a721ef7c47967b6a45

Observation 2b7733b6-b7d0-4bc5-83eb-ee539c7f94d8 · outbound

This paper cites nuscenes: A multimodal dataset for autonomous driving,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks nuscenes: A multimodal dataset for autonomous driving,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.802695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.802695Z digest=sha256:7db1698bdebe5284f560bee2ed01947c1bde623b6f903a4dbf84424e3ee6d517

Observation db60efbf-4e4e-445d-92e1-340dcad7f9fc · outbound

This paper cites Scalability in perception for autonomous driving: Waymo open dataset,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Scalability in perception for autonomous driving: Waymo open dataset,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.858783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.858783Z digest=sha256:64bf280529f97ae00665d59b34cdea97c7bb137647c9c62c3ccad5d85931c618

Observation a54ae552-4df9-4c06-bf31-587fc7e0a5b6 · outbound

This paper cites Kitti-360: A novel dataset and bench- marks for urban scene understanding in 2d and 3d,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Kitti-360: A novel dataset and bench- marks for urban scene understanding in 2d and 3d,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.930690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.930690Z digest=sha256:c8220c9e7d344b337c8094c5c9680fdabfa7c59e3aaa4e67af349182721a2fd8

Observation afa8f4a0-e1a6-430d-9fb5-a6d3c098cadd · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.989811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.989811Z digest=sha256:944f7dcd1194fa9e6085032a18f370464b7899376b88c3602f69092048acfd28

Observation 2dee0491-bba6-46c1-a5e7-7ca30aeaa4df · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Swin transformer: Hierarchical vision transformer using shifted windows,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:30.026643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:30.026643Z digest=sha256:42458feedd4c3abb9243595f8cb04961fed8dfe640b9b5d01115525e4e7479ae

Observation 45e652a9-39a0-4b43-9bc1-48b7e076c746 · outbound

This paper cites Training data-efficient image transformers & distillation through attention,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Training data-efficient image transformers & distillation through attention,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:30.122872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:30.122872Z digest=sha256:70a00bdaa03e3e7569463b7b8d0b865427e49ca8984fc303ac3f273305474d00

Observation 58587ce4-673f-4b34-bd10-5ab51d599d13 · outbound

This paper cites Is space-time attention all you need for video understanding?.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Is space-time attention all you need for video understanding?

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:33.026550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:30.219968Z digest=sha256:0599fafb08298637f40e73866fc85d11aa563ed2f23374eea26e11c678a331ea

Observation 07fedb53-97ba-4763-a910-ed65e8bc7cb7 · outbound

This paper cites Vivit: A video vision transformer,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Vivit: A video vision transformer,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:30.284228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:30.284228Z digest=sha256:efd415eef8523d1ddc8a34b1c0a57b473e8a8f50578cb7d8129cf8d0b78095dc

Observation dd8bcb5f-5428-46ce-b46f-026b0068a37b · outbound

This paper cites Tuber: Tubelet transformer for video action detection,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Tuber: Tubelet transformer for video action detection,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:32.884734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:30.350333Z digest=sha256:15bc78f24657455f60681b47da8b9b51160d4aef3a56bdecf5f2819d92bfd7c1

Observation 035c7db7-0b19-416e-906e-ce7a1a86a862 · outbound

This paper cites End-to-end video instance segmentation with transformers,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks End-to-end video instance segmentation with transformers,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:30.423962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:30.423962Z digest=sha256:02de4db9d7e530e38e12679ed437ca343d478f046fee1b29269bf06acc6de197

Observation b5ad08d6-56d4-4d43-95bf-2dd2207e1773 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:30.614682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:30.614682Z digest=sha256:213f331b15f66d34beef167661f0b8973c3117b1a5c4a87833adc33e01d9538f

Observation 27b94720-d27d-4e83-a7e7-a7abde96db36 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:30.732657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:30.732657Z digest=sha256:431f2be886d12e3664937787e601043a396663abc8560222a324b6ddb848efa7

Observation 8cb8e2b0-b50f-42e5-94a6-69dda43a7026 · outbound

This paper cites Adapt: Action-aware driving caption transformer,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Adapt: Action-aware driving caption transformer,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:30.839297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:30.839297Z digest=sha256:dca47f9201c8db6737cb9220a57d4ab2c69351bce754bb23f289f1ead45c9d5f

Observation 464de7ee-7eec-4f00-a9b4-cda0ef90f22f · outbound

This paper cites Textual explanations for self-driving vehicles,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Textual explanations for self-driving vehicles,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:31.005990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:31.005990Z digest=sha256:5034a761a5d620dddb9d5da071d818b99763cacb73d40ad5abdf392a99fca6c5

Observation 8a76d122-1a03-4998-9597-ac78e3d31d83 · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:31.081961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:31.081961Z digest=sha256:361f774deea843ea1c0bbf0a3ba14787cca69b15886d28f3e00ea70d7e072f30

Observation 980049bf-5867-40d3-a61d-4ba25b3636a3 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Learning transferable visual models from natural language supervision,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:31.163340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:31.163340Z digest=sha256:13408e1f6b2b231e39219cdabfc57cd3f6346e6d5ad09d80c1d96d542effee6d

Observation b9b456bb-b459-4b12-83a8-43367164c6c5 · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:31.265644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:31.265644Z digest=sha256:26abd7cac818daa24bb78fd543c0db15bbaf317c9896a5fa8bc5ab63715439b2

Observation 1cb494fc-6a8a-4209-83c1-d2fbd34d1cd9 · outbound

This paper cites Can masking back- ground and object reduce static bias for zero-shot action recognition?.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Can masking back- ground and object reduce static bias for zero-shot action recognition?

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:32.723513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:31.322555Z digest=sha256:56da711aa0fe937e55339c04e0cb84d8ee61fafa022b426f3b6044c408027654

Observation 9facd20a-5781-42ff-b0e5-8da5a3718d0e · outbound

This paper cites Another efficient algorithm for convex hulls in two dimensions,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Another efficient algorithm for convex hulls in two dimensions,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:32.612712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:31.379805Z digest=sha256:279f340773a4e05defdf6468db0efde0c5a1631558ba19de49c71a8f88078ab3

Observation b2b4fe5d-c54d-45b9-9f43-7a3d6cf6d9dc · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Bleu: a method for automatic evaluation of machine translation,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:32.472030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:31.527586Z digest=sha256:25add930a817a2264ae9e68e7f4a55a83ff0ebac5fbd5c3f783df78b578b7658

Observation a349a548-3736-43a6-b28e-36458a90196b · outbound

This paper cites VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:28.305093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:28.305093Z digest=sha256:9dffe7a44772ba222da1808f22220786dadbe04d56b234860039f185f10b33af

Pith citing papers

Observation 11b69413-3366-4aa4-99e1-a5a6081f42eb · inbound

InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding cites this paper.

InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding Temporal Object Captioning for Street Scene Videos from LiDAR Tracks

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:30:58.341882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:06:33.139310Z digest=sha256:a40634d20424ed6192a3dbd7a904961d69d31fe9b0eaa4b28cbf1b6008fd72d0