Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:02:31.527586Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 1 inbound Pith citation observation for arXiv:2505.16594.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:02:31.527586Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-10T18:06:33.139310Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T05:30:58.338283Z
43 of 43 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6a669fc7-60d9-4fb6-934c-512dd9d7dfc8 · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks MaMMUT: A Simple Architecture for Joint Learning for MultiModal Tasks
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6247258-8cda-4e7a-a7c6-2fc1cfed444c · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks VLAB: Enhancing Video Language Pre-training by Feature Adapting and Blending
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b89894d-f481-4bf4-bdc6-2a0216f92b91 · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Valor: Vision-audio-language omni-perception pretraining model and dataset,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e6e64717-ff25-49de-8e4f-328ddcff60d6 · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Self-supervised Spatio-temporal Representation Learning for Videos by Predicting Motion and Appearance Statistics
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4f7c5dd5-252b-4021-9d20-722220d3cc6c · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Unsupervised Pre-Training of Image Features on Non-Curated Data
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b997e172-681e-4e16-924d-19d87ae7e4c3 · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Melanoma thick- ness prediction based on convolutional neural network with vgg-19 model transfer learning,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9e3a94ff-32d1-4c6f-bc57-f825a6be1c7f · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Pre-training on grayscale imagenet improves medical image classification,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 76007362-10d8-40c1-a4fc-dd3890e65c39 · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Low-Rank HOCA: Efficient High-Order Cross-Modal Attention for Video Captioning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2f3fedb6-ae83-4424-bcc7-ae7b5f82b368 · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Sensor-augmented egocentric- video captioning with dynamic modal attention,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b5cdd1d4-1dca-4046-ab49-b6c24b18fef1 · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Boosting Video Captioning with Dynamic Loss Network
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5634ddb2-30e4-4d4d-8e6d-29b3333a1cc4 · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Panoptic segmentation,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 949b1838-94dc-43d9-b8fa-5857cc82c437 · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks An empirical study of context in object detection,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation afa08809-6553-456f-99cb-ef07545707bb · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks End-to-end object detection with transformers,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa1a084a-51b0-4e57-94fe-f26759cc1cfd · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b85f535a-1e35-4405-ac04-20df1c5043b4 · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9878bc4-bf6a-46ea-a9c2-b4267b312608 · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9225c940-9e52-4dae-8ea7-c98fe7d9f569 · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Swinbert: End-to-end transformers with sparse attention for video captioning,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2e89f73c-4b20-447d-bb17-e1b31044843a · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Lingoqa: Visual question answering for autonomous driving
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a9ad56f4-133c-4ee6-bf79-03698936ac9c · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 925d6cbf-a016-4395-9a47-23a0f46b9d60 · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Nuscenes-mqa: Integrated evaluation of captions and qa for autonomous driving datasets using markup annotations,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6b766470-5748-44f5-891c-013d71e87bed · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Language Prompt for Autonomous Driving
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25fab8d1-523d-4e0e-9df9-39006acd7f1e · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Covla: Comprehensive vision-language-action dataset for autonomous driving,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b7733b6-b7d0-4bc5-83eb-ee539c7f94d8 · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks nuscenes: A multimodal dataset for autonomous driving,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db60efbf-4e4e-445d-92e1-340dcad7f9fc · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Scalability in perception for autonomous driving: Waymo open dataset,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a54ae552-4df9-4c06-bf31-587fc7e0a5b6 · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Kitti-360: A novel dataset and bench- marks for urban scene understanding in 2d and 3d,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afa8f4a0-e1a6-430d-9fb5-a6d3c098cadd · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dee0491-bba6-46c1-a5e7-7ca30aeaa4df · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Swin transformer: Hierarchical vision transformer using shifted windows,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45e652a9-39a0-4b43-9bc1-48b7e076c746 · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Training data-efficient image transformers & distillation through attention,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58587ce4-673f-4b34-bd10-5ab51d599d13 · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Is space-time attention all you need for video understanding?
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 07fedb53-97ba-4763-a910-ed65e8bc7cb7 · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Vivit: A video vision transformer,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd8bcb5f-5428-46ce-b46f-026b0068a37b · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Tuber: Tubelet transformer for video action detection,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 035c7db7-0b19-416e-906e-ce7a1a86a862 · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks End-to-end video instance segmentation with transformers,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5ad08d6-56d4-4d43-95bf-2dd2207e1773 · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27b94720-d27d-4e83-a7e7-a7abde96db36 · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cb8e2b0-b50f-42e5-94a6-69dda43a7026 · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Adapt: Action-aware driving caption transformer,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 464de7ee-7eec-4f00-a9b4-cda0ef90f22f · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Textual explanations for self-driving vehicles,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a76d122-1a03-4998-9597-ac78e3d31d83 · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 980049bf-5867-40d3-a61d-4ba25b3636a3 · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Learning transferable visual models from natural language supervision,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9b456bb-b459-4b12-83a8-43367164c6c5 · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Very Deep Convolutional Networks for Large-Scale Image Recognition
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cb494fc-6a8a-4209-83c1-d2fbd34d1cd9 · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Can masking back- ground and object reduce static bias for zero-shot action recognition?
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9facd20a-5781-42ff-b0e5-8da5a3718d0e · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Another efficient algorithm for convex hulls in two dimensions,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b2b4fe5d-c54d-45b9-9f43-7a3d6cf6d9dc · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Bleu: a method for automatic evaluation of machine translation,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a349a548-3736-43a6-b28e-36458a90196b · outbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11b69413-3366-4aa4-99e1-a5a6081f42eb · inbound
InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding Temporal Object Captioning for Street Scene Videos from LiDAR Tracks
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.