Pith. sign in

Paper Citation Record · LEDGER

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks

As of 20 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 1 inbound Pith citation observation for arXiv:2505.16594.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16594 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:02:31.527586Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T18:06:33.139310Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T05:30:58.338283Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact4
  • verified fuzzy13
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6a669fc7-60d9-4fb6-934c-512dd9d7dfc8 · outbound

This paper cites MaMMUT: A Simple Architecture for Joint Learning for MultiModal Tasks.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks MaMMUT: A Simple Architecture for Joint Learning for MultiModal Tasks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:28.003967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:28.003967Z digest=sha256:8ac5947d142b34b66fa0fa8a7c05f08a053b3e14186abd2fad800d6450ce61eb

Observation b6247258-8cda-4e7a-a7c6-2fc1cfed444c · outbound

This paper cites VLAB: Enhancing Video Language Pre-training by Feature Adapting and Blending.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks VLAB: Enhancing Video Language Pre-training by Feature Adapting and Blending

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:28.118849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:28.118849Z digest=sha256:900b0c97b47e070a7f895cf7afa9a0b3e8c1ff228327616cdf655570832371f7

Observation 3b89894d-f481-4bf4-bdc6-2a0216f92b91 · outbound

This paper cites Valor: Vision-audio-language omni-perception pretraining model and dataset,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Valor: Vision-audio-language omni-perception pretraining model and dataset,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:34.415561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:28.210380Z digest=sha256:73125e78ac128001c1362a4084c4526fcf0db1cfa0a6756d42067d7b977a99a3

Observation e6e64717-ff25-49de-8e4f-328ddcff60d6 · outbound

This paper cites Self-supervised Spatio-temporal Representation Learning for Videos by Predicting Motion and Appearance Statistics.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Self-supervised Spatio-temporal Representation Learning for Videos by Predicting Motion and Appearance Statistics

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:02:32.305348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:28.384506Z digest=sha256:7fd0c2ace16658fc7c60f2d334c9fe39f40828d6446ff7657c21e6035ff519df

Observation 4f7c5dd5-252b-4021-9d20-722220d3cc6c · outbound

This paper cites Unsupervised Pre-Training of Image Features on Non-Curated Data.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Unsupervised Pre-Training of Image Features on Non-Curated Data

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:02:32.192245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:28.503146Z digest=sha256:cec9b36fe7586e4507c48bcc3315be93a873d60a2b29606660e2fe664dbc4e1c

Observation b997e172-681e-4e16-924d-19d87ae7e4c3 · outbound

This paper cites Melanoma thick- ness prediction based on convolutional neural network with vgg-19 model transfer learning,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Melanoma thick- ness prediction based on convolutional neural network with vgg-19 model transfer learning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:34.280853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:28.606666Z digest=sha256:c3371daf0ff817afbdfa437831bba34cacf568fc2087d89baf27bea10dc70cb2

Observation 9e3a94ff-32d1-4c6f-bc57-f825a6be1c7f · outbound

This paper cites Pre-training on grayscale imagenet improves medical image classification,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Pre-training on grayscale imagenet improves medical image classification,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:34.149539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:28.716523Z digest=sha256:5dee775a4dbcc03d74478e21070a8079b2908670757ce5ae61d838d2b43d9b27

Observation 76007362-10d8-40c1-a4fc-dd3890e65c39 · outbound

This paper cites Low-Rank HOCA: Efficient High-Order Cross-Modal Attention for Video Captioning.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Low-Rank HOCA: Efficient High-Order Cross-Modal Attention for Video Captioning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:02:32.020124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:28.782114Z digest=sha256:41c9605e0d5411e346287350b0b835330af18347769bc42cd5a61bb8694fdd05

Observation 2f3fedb6-ae83-4424-bcc7-ae7b5f82b368 · outbound

This paper cites Sensor-augmented egocentric- video captioning with dynamic modal attention,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Sensor-augmented egocentric- video captioning with dynamic modal attention,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:33.986897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:28.879984Z digest=sha256:c50e71b2fc73d1c144425883910a1620ce58d3fb71c66ad3b6d45947a8a1189a

Observation b5cdd1d4-1dca-4046-ab49-b6c24b18fef1 · outbound

This paper cites Boosting Video Captioning with Dynamic Loss Network.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Boosting Video Captioning with Dynamic Loss Network

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:02:31.865243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:28.972803Z digest=sha256:4e176aa6a6ea79691a0a5cd8a181764ba8e7d574f7d6f37efce0266067b217e4

Observation 5634ddb2-30e4-4d4d-8e6d-29b3333a1cc4 · outbound

This paper cites Panoptic segmentation,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Panoptic segmentation,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.036353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.036353Z digest=sha256:65ddf156f4de663d8ad20d2eaa0e42a9ef400f9ed72e9119c1bba5a84f7a5bec

Observation 949b1838-94dc-43d9-b8fa-5857cc82c437 · outbound

This paper cites An empirical study of context in object detection,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks An empirical study of context in object detection,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:33.817392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:29.104499Z digest=sha256:ec04ab2d1b1b3d1e82e5a3d49a68da5488f527fe6554994e91623d996fc8673d

Observation afa08809-6553-456f-99cb-ef07545707bb · outbound

This paper cites End-to-end object detection with transformers,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks End-to-end object detection with transformers,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.171195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.171195Z digest=sha256:acd2166b078ff0b84ec6305a678cc5e42e09a9a6068150cecfb8906144cc8f8d

Observation aa1a084a-51b0-4e57-94fe-f26759cc1cfd · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.260811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.260811Z digest=sha256:306feba19fee70b10d68154c9e4ec481b8b834dbb1321a8578b61fb01f877121

Observation b85f535a-1e35-4405-ac04-20df1c5043b4 · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.351108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.351108Z digest=sha256:eb8234be789989264424c8cb3b35ab973a772f08ac504dd101290e45821d8b00

Observation a9878bc4-bf6a-46ea-a9c2-b4267b312608 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.424508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.424508Z digest=sha256:4a0208c7f33897da1dfd376ef482a7803f72375e2fa95021ff07789c47669514

Observation 9225c940-9e52-4dae-8ea7-c98fe7d9f569 · outbound

This paper cites Swinbert: End-to-end transformers with sparse attention for video captioning,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Swinbert: End-to-end transformers with sparse attention for video captioning,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:33.688236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:29.503052Z digest=sha256:8b9e072a201d1b82698dda95a0add9925bd458e0d6e38d85be41f93d25602967

Observation 2e89f73c-4b20-447d-bb17-e1b31044843a · outbound

This paper cites Lingoqa: Visual question answering for autonomous driving.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Lingoqa: Visual question answering for autonomous driving

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:33.524482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:29.555237Z digest=sha256:a6d3839281ec902870b7c457e85da3805ffaecc85c7537b8ce9fea553f890451

Observation a9ad56f4-133c-4ee6-bf79-03698936ac9c · outbound

This paper cites Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.616411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.616411Z digest=sha256:9f456d54ee0c520972d0a3d932dd399d2a89240b629f6c46b01a5668f2229eb6

Observation 925d6cbf-a016-4395-9a47-23a0f46b9d60 · outbound

This paper cites Nuscenes-mqa: Integrated evaluation of captions and qa for autonomous driving datasets using markup annotations,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Nuscenes-mqa: Integrated evaluation of captions and qa for autonomous driving datasets using markup annotations,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:33.264972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:29.651915Z digest=sha256:734d0eb800e9079983d1fae352cdad899281a77bae033843dfb11da485b09ae3

Observation 6b766470-5748-44f5-891c-013d71e87bed · outbound

This paper cites Language Prompt for Autonomous Driving.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Language Prompt for Autonomous Driving

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.686171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.686171Z digest=sha256:2c75a103c06186f5320417a979d709e7aff8c21deb5a38f6e13c9bcb23ed5d9e

Observation 25fab8d1-523d-4e0e-9df9-39006acd7f1e · outbound

This paper cites Covla: Comprehensive vision-language-action dataset for autonomous driving,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Covla: Comprehensive vision-language-action dataset for autonomous driving,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.741400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.741400Z digest=sha256:3c06b70f494c6f86c29447a47ff3ae6d9c1adb5ab61ff352cff10c9f75936f87

Observation 2b7733b6-b7d0-4bc5-83eb-ee539c7f94d8 · outbound

This paper cites nuscenes: A multimodal dataset for autonomous driving,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks nuscenes: A multimodal dataset for autonomous driving,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.802695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.802695Z digest=sha256:c60ab878eabbb195ba63f6f33239b67e25badac028fd84ab7aff02f284fe3d01

Observation db60efbf-4e4e-445d-92e1-340dcad7f9fc · outbound

This paper cites Scalability in perception for autonomous driving: Waymo open dataset,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Scalability in perception for autonomous driving: Waymo open dataset,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.858783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.858783Z digest=sha256:7a2ee93b25c781f2d05afb547b1bdfb0a166b28744265944750caf0b36c43988

Observation a54ae552-4df9-4c06-bf31-587fc7e0a5b6 · outbound

This paper cites Kitti-360: A novel dataset and bench- marks for urban scene understanding in 2d and 3d,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Kitti-360: A novel dataset and bench- marks for urban scene understanding in 2d and 3d,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.930690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.930690Z digest=sha256:02d78828e20960878ff1808687b1cd685c1610cf72502601bfed00b83fb227e0

Observation afa8f4a0-e1a6-430d-9fb5-a6d3c098cadd · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.989811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.989811Z digest=sha256:3bfcb511a54c16a5a7e3573aca3cc5b0f1e363a6c7a5c96ca9a5d50c5d89f4e8

Observation 2dee0491-bba6-46c1-a5e7-7ca30aeaa4df · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Swin transformer: Hierarchical vision transformer using shifted windows,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:30.026643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:30.026643Z digest=sha256:e6a79f3052b51124610191e54cb51a09fadd6e5f1df1f13dbe2f4461a6220218

Observation 45e652a9-39a0-4b43-9bc1-48b7e076c746 · outbound

This paper cites Training data-efficient image transformers & distillation through attention,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Training data-efficient image transformers & distillation through attention,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:30.122872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:30.122872Z digest=sha256:107822a789131871925758be5623908403781432ddc4d3e978ae993b7bab30e7

Observation 58587ce4-673f-4b34-bd10-5ab51d599d13 · outbound

This paper cites Is space-time attention all you need for video understanding?.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Is space-time attention all you need for video understanding?

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:33.026550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:30.219968Z digest=sha256:0836176dca5d65cb71c2757b4ddd01a6dd72e1fde74a7f38e1f31478f1fb5aef

Observation 07fedb53-97ba-4763-a910-ed65e8bc7cb7 · outbound

This paper cites Vivit: A video vision transformer,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Vivit: A video vision transformer,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:30.284228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:30.284228Z digest=sha256:ad7bf456edd3f34706b11dd78c57a6024aae7fa3e75547522df403e51800da83

Observation dd8bcb5f-5428-46ce-b46f-026b0068a37b · outbound

This paper cites Tuber: Tubelet transformer for video action detection,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Tuber: Tubelet transformer for video action detection,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:32.884734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:30.350333Z digest=sha256:ae0b815a289a3a66bab3097c16f68153ddd1768b668b8fe632e7f2f2f4f6c76b

Observation 035c7db7-0b19-416e-906e-ce7a1a86a862 · outbound

This paper cites End-to-end video instance segmentation with transformers,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks End-to-end video instance segmentation with transformers,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:30.423962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:30.423962Z digest=sha256:aed065579feff93e524d25cad18294ee6fb48477de5a0a3979eb17e709dc4965

Observation b5ad08d6-56d4-4d43-95bf-2dd2207e1773 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:30.614682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:30.614682Z digest=sha256:4361fd9a43118d7d4f01ab3c623ed64bfaed8d4f0fdc1eb63ceb47a8f810cf44

Observation 27b94720-d27d-4e83-a7e7-a7abde96db36 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:30.732657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:30.732657Z digest=sha256:7759e8ee95dbfacf2b74cd0eb44a8f316bc4a6973db762905486c1f234390c6f

Observation 8cb8e2b0-b50f-42e5-94a6-69dda43a7026 · outbound

This paper cites Adapt: Action-aware driving caption transformer,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Adapt: Action-aware driving caption transformer,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:30.839297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:30.839297Z digest=sha256:a99396d60600f4f9c4006a987a6efe0743c70e27aa99aa1f757f9611c28261fc

Observation 464de7ee-7eec-4f00-a9b4-cda0ef90f22f · outbound

This paper cites Textual explanations for self-driving vehicles,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Textual explanations for self-driving vehicles,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:31.005990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:31.005990Z digest=sha256:0a1b7b8986298121fbd5a87ecc5d4f748eb194838e6a53acb7a91b20debeaa09

Observation 8a76d122-1a03-4998-9597-ac78e3d31d83 · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:31.081961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:31.081961Z digest=sha256:9e7f4f7ba7b45971709872bd68ba437befc2d7a1189618fd7176509bff3930e0

Observation 980049bf-5867-40d3-a61d-4ba25b3636a3 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Learning transferable visual models from natural language supervision,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:31.163340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:31.163340Z digest=sha256:012d368ac61339af250aa9c1a5fe1e7532ca4b113a7d70fac0461a4e6b6bfeee

Observation b9b456bb-b459-4b12-83a8-43367164c6c5 · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:31.265644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:31.265644Z digest=sha256:d9aa0e99cface04f47e64a1eece7576e1a08f131222768d323ad15d9cd1d9574

Observation 1cb494fc-6a8a-4209-83c1-d2fbd34d1cd9 · outbound

This paper cites Can masking back- ground and object reduce static bias for zero-shot action recognition?.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Can masking back- ground and object reduce static bias for zero-shot action recognition?

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:32.723513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:31.322555Z digest=sha256:8746f870b373d834ceac891c59ffe3fa3cd9403f8c99683877386120fa515931

Observation 9facd20a-5781-42ff-b0e5-8da5a3718d0e · outbound

This paper cites Another efficient algorithm for convex hulls in two dimensions,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Another efficient algorithm for convex hulls in two dimensions,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:32.612712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:31.379805Z digest=sha256:757f365bb1e0e2c48fc61dc61d05b098271277c9bf8d5292c5aef04934ef5eeb

Observation b2b4fe5d-c54d-45b9-9f43-7a3d6cf6d9dc · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Bleu: a method for automatic evaluation of machine translation,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:32.472030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:31.527586Z digest=sha256:1d59ce3f7339dc5402647598dd33c1aabaa6aa0244b2f905e6f85541ca719b4a

Observation a349a548-3736-43a6-b28e-36458a90196b · outbound

This paper cites VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:28.305093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:28.305093Z digest=sha256:5310280e5943fa893a305a867eaffdc3a3645b3e8c13f0592fd88ba3ece59bf3

Pith citing papers

Observation 11b69413-3366-4aa4-99e1-a5a6081f42eb · inbound

InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding cites this paper.

InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding Temporal Object Captioning for Street Scene Videos from LiDAR Tracks

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:30:58.341882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T18:06:33.139310Z digest=sha256:61c83351d53f2f2128a233acadeeb7409bdeba5e929475903bea723a2fadeed5