Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T17:09:23.706315Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2512.10607.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T17:09:23.706315Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
44 of 44 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation aef58132-70af-47e3-981d-a8f1dd0a24c8 · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3de8bd12-c4a3-4ae6-81dd-9e3040ac3c31 · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Frozen in time: A joint video and image encoder for end-to-end retrieval
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57223da2-1323-42c2-b54b-7aedaac20fb0 · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Object segmentation by long term analysis of point trajectories
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ec8f9ad-ab2d-4ca3-88b0-90f5975c6255 · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Mevis: A large-scale bench- mark for video segmentation with motion expressions.arXiv preprint, 2024
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb279774-a2a2-4d49-aa4d-0317a17c0c6d · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Mevis: A large-scale benchmark for video segmentation with motion expressions
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e836d57-eb99-428a-a673-fae38e6452d4 · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Tap-vid: A benchmark for tracking any point in a video
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cde1e77-91b2-4942-a4be-392db3c47f80 · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Tapir: Tracking any point with per-frame initialization and temporal refinement
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f6755db-4ee8-4051-a93c-1a896198f5bc · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Step- former: Self-supervised step discovery and localization in instructional videos
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02fd8f7d-495d-4c5b-b962-a2541a8c1cd3 · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Context-guided spatio-temporal video grounding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae883612-6ff8-4ec3-b5a6-fdbacb368aa1 · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Harley, Zhaoyuan Fang, and Katerina Fragkiadaki
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79d85d07-c7bc-4506-8df5-80a4a1ea4892 · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Pips++: Improved tracking through occlusions via extended point trajectories
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3f971f9-358c-4fc3-a93d-d3026f16ce28 · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation A better use of audio-visual cues: Dense video captioning with bi-modal transformer
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d83ecfe-48f0-4669-aeb5-e8d9310fb367 · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation VideoRAG: Retrieval-Augmented Generation over Video Corpus
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63cadd34-5091-46d2-8b09-0636e0f1acb5 · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Embracing consistency: A one-stage approach for spatio- temporal video grounding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ca2e823-4436-4236-9394-08565064ccda · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Co- tracker: It is better to track together
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e7ef5ab-bb89-4ce2-9340-b5cbe8dfb9ba · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Co- tracker3: Simpler and better point tracking by pseudo- labelling real videos
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bc5818c-8fd0-417e-a7a2-4a303e1c170d · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Cospal: Co-optimizing spatio-temporal context prompting and adapt- ing for weakly supervised video grounding.arXiv preprint,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acf569f3-09ba-4e5c-a3b5-bd2fb4392fd2 · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Unsupervised object discovery and track- ing in video collections
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f3e8864-aa41-4109-aa52-be2b78164700 · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning.Neu- rocomputing, 508:293–304, 2022
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03c7596a-8e0d-48e1-a2aa-b3c9189148fa · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation X-clip: End-to-end multi-grained con- trastive learning for video-text retrieval
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbcefd51-4e0f-4550-a9b3-952260870a1f · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Delta: Dense efficient long-range 3d tracking for any video
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07ff6088-b3f4-4858-840a-d34b39a37d13 · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Unsupervised discovery of actions in in- structional videos
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfcbe662-b9b0-4172-8b78-064f432908d2 · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Learn- ing transferable visual models from natural language super- vision
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8830128-56a8-407e-8ec4-7b20d4a3f0b6 · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Two-stream con- volutional networks for action recognition in videos
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3caa1e21-e743-41b5-b7b7-64487c2b8549 · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Human-centric spatio-temporal video grounding with visual transformers
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72058787-6fce-4a60-be68-312eb18d98d0 · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Human-centric spatio-temporal video grounding with visual transformers
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb08ca90-92cd-4658-afde-b968f42db816 · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Repre- sentation learning with contrastive predictive coding, 2018
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 646b2a06-6c64-4872-a89d-e77122168216 · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Action recognition with trajectories
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0984b6a3-f540-4e40-80e9-ffdbb441b3ec · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation ActionCLIP: A New Paradigm for Video Action Recognition
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d453055-c304-4433-aee9-73bb93a9b9ea · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Tracking everything everywhere all at once
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c865365e-747b-4104-9df0-304c5e1e3458 · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation End-to-end dense video captioning with parallel decoding
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67248719-9360-4fdf-8aea-7097a183c0e3 · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Language as queries for referring video object segmen- tation
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b85600e1-13f3-4eb1-a4bc-5b4987b4f7fb · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Spatialtracker: Tracking any 2d pixels in 3d space
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5163c027-4ad1-48c0-b1ad-548dd93d88b1 · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation SpatialTrackerV2: 3D Point Tracking Made Easy
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 636d8ca8-ba1e-4cdc-b5af-8e847c87a78a · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Videoclip: Contrastive pre-training for zero-shot video-text understanding
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da50673d-f83d-4fe4-851b-4db90ae7451c · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Universal instance perception as object discovery and retrieval
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b7764ea-dabf-4c23-a7fc-61399e6be65f · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Tubedetr: Spatio-temporal video ground- ing with transformers
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52dbcf17-2ab9-434c-be6a-afbed23b942a · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a810d00c-526e-4f9d-85c1-081f80f0a3a3 · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Tapip3d: Tracking any point in persistent 3d geome- try.arXiv preprint arXiv:2504.14717, 2025
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d26d584-b579-49ae-b32f-8708aba588bc · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Where does it exist: Spatio-temporal video grounding for multi-form sentences
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08531455-bcdd-4cd6-ba43-80920be9c8e3 · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Video- text prompting for weakly supervised spatio-temporal video grounding
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3da8936-48f8-4637-bf56-0c9cca9f15ca · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Unsupervised learning from video to detect foreground objects in single images
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ef7c3d8-a217-4a02-bc1c-cef4b6e0600f · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation TAPNext: Tracking Any Point (TAP) as Next Token Prediction
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cb55178-5194-4abe-b175-d239f12c1ff6 · outbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation Dense video object captioning from disjoint super- vision
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.