Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T07:07:44.009137Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2607.21576.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T07:07:44.009137Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fcf5aec1-d427-4934-92b1-88dbb6c38a08 · outbound
Self-Supervised Learning of Structured Dynamics from Videos V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 645df39c-c401-4de0-81ef-3b53daab1ac7 · outbound
Self-Supervised Learning of Structured Dynamics from Videos Back to the Features: DINO as a Foundation for Video World Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e57671f3-34fe-4506-9642-2ad1495ef298 · outbound
Self-Supervised Learning of Structured Dynamics from Videos Revisiting feature prediction for learning visual representations from video.Transactions on Machine Learning Research, 2024
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff4ba9dd-6950-4f73-9f34-ac9508a804ce · outbound
Self-Supervised Learning of Structured Dynamics from Videos VFMF: World modeling by forecasting vision foundation model features.arXiv preprint arXiv:2512.11225,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de75f00c-62f0-4962-9043-dd50da99aee3 · outbound
Self-Supervised Learning of Structured Dynamics from Videos TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e902e301-623b-4c03-9de9-fc397afdd936 · outbound
Self-Supervised Learning of Structured Dynamics from Videos A Short Note on the Kinetics-700 Human Action Dataset
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b3a923d-c9c3-4986-9f24-86606157a485 · outbound
Self-Supervised Learning of Structured Dynamics from Videos Scaling 4D Representations
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe9778c6-5fa5-4553-aa9b-0415daf664ba · outbound
Self-Supervised Learning of Structured Dynamics from Videos Vision transformers need registers
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 672b183b-0146-4b90-b179-08ac7841accb · outbound
Self-Supervised Learning of Structured Dynamics from Videos Probing the 3d awareness of visual foundation models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c1e3832-5293-4856-a12b-ae26f0fa47d4 · outbound
Self-Supervised Learning of Structured Dynamics from Videos Scalable pre-training of large autoregressive image models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f62d957-313a-43c0-bc5f-819ab38c4242 · outbound
Self-Supervised Learning of Structured Dynamics from Videos Multimodal autoregressive pre-training of large vision encoders
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 594d648b-9a2c-4095-a313-23e54a26b8d7 · outbound
Self-Supervised Learning of Structured Dynamics from Videos Learning latent action world models in the wild.arXiv preprint arXiv:2601.05230,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21ed614d-83ad-45e1-ad25-fb18370dd025 · outbound
Self-Supervised Learning of Structured Dynamics from Videos something something
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a5c7dea-afcc-4dae-959e-3bfa5d651e49 · outbound
Self-Supervised Learning of Structured Dynamics from Videos Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f873e337-e9d3-47b8-bf4b-d10e18b844d4 · outbound
Self-Supervised Learning of Structured Dynamics from Videos Siamese masked autoencoders.Advances in Neural Information Processing Systems, 36:40676–40693, 2023
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 547a2166-ea33-4e0a-bb48-c3eee0d8b9d3 · outbound
Self-Supervised Learning of Structured Dynamics from Videos Masked autoencoders are scalable vision learners
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22fc5f9e-72c4-4c09-b583-80c4ac259aef · outbound
Self-Supervised Learning of Structured Dynamics from Videos Rotary position embedding for vision transformer
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1d0eed5-b78e-4a04-97f1-8097b2ed3d86 · outbound
Self-Supervised Learning of Structured Dynamics from Videos VGGT4D: Mining motion cues in visual geometry transformers for 4d scene reconstruction.arXiv preprint arXiv:2511.19971, 2025
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8425ea1-c34c-4e13-bdcd-5392d11716e8 · outbound
Self-Supervised Learning of Structured Dynamics from Videos A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3c5fe62-51b3-41a6-b595-f5cc7558a7ac · outbound
Self-Supervised Learning of Structured Dynamics from Videos Depth Anything 3: Recovering the Visual Space from Any Views
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28d0b606-de80-41da-bdf2-c07efc56d5bc · outbound
Self-Supervised Learning of Structured Dynamics from Videos Towards Understanding Camera Motions in Any Video
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f03464ab-f014-4252-9354-1a64f1527cb4 · outbound
Self-Supervised Learning of Structured Dynamics from Videos DL3DV-10k: A large-scale scene dataset for deep learning-based 3d vision
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43ae3062-fec4-4b39-99d4-c39099bfd6d5 · outbound
Self-Supervised Learning of Structured Dynamics from Videos 4RC: 4D Reconstruction via Conditional Querying Anytime and Anywhere
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a833b1c-186f-4aab-974b-2a97f257b8d4 · outbound
Self-Supervised Learning of Structured Dynamics from Videos V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34d14f7c-97fd-43c8-ac36-5edf4ac34380 · outbound
Self-Supervised Learning of Structured Dynamics from Videos DINOv2: Learning robust visual features without supervision.Transactions on Machine Learning Research, 2023
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 182abe4a-1f14-4b48-82db-e75561513901 · outbound
Self-Supervised Learning of Structured Dynamics from Videos The 2017 DAVIS Challenge on Video Object Segmentation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4583a188-4775-43e9-a355-cbe4adf0db60 · outbound
Self-Supervised Learning of Structured Dynamics from Videos Time does tell: Self-supervised time-tuning of dense image representations
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 450b75fa-98fa-4a39-89df-87ec8183382e · outbound
Self-Supervised Learning of Structured Dynamics from Videos MoSiC: Optimal-transport motion trajectory for dense self- supervised learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa113fe4-b7ac-4da2-bf39-3462434d4f0a · outbound
Self-Supervised Learning of Structured Dynamics from Videos DINOv3
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ea85e0e-05bb-4463-92be-8ee667b13647 · outbound
Self-Supervised Learning of Structured Dynamics from Videos Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2645d77d-e259-461a-a17a-7ac3cebee984 · outbound
Self-Supervised Learning of Structured Dynamics from Videos VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training.Advances in neural information processing systems, 35:10078–10093, 2022
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a851265-df0a-4b5a-9b3a-ae8248a7c565 · outbound
Self-Supervised Learning of Structured Dynamics from Videos SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5aa556d7-3cee-41b0-8ff0-63f37f49b2c4 · outbound
Self-Supervised Learning of Structured Dynamics from Videos Is imagenet worth 1 video? learning strong image encoders from 1 long unlabelled video
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99103198-8a40-45ba-9adf-1fe3682195b8 · outbound
Self-Supervised Learning of Structured Dynamics from Videos PooDLe: Pooled and dense self-supervised learning from naturalistic videos
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 709e7f6f-3816-4c21-9a39-cb847bbe7d88 · outbound
Self-Supervised Learning of Structured Dynamics from Videos VGGT: Visual geometry grounded transformer
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c973aec4-f6bb-4ab7-8b38-b14ba09dcfab · outbound
Self-Supervised Learning of Structured Dynamics from Videos VideoMAE v2: Scaling video masked autoencoders with dual masking
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89c50610-fdc6-4295-9893-a0331c937ea6 · outbound
Self-Supervised Learning of Structured Dynamics from Videos Continuous 3d perception model with persistent state
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58664195-27a9-44fd-b628-84dbfd619214 · outbound
Self-Supervised Learning of Structured Dynamics from Videos DUSt3R: Geometric 3d vision made easy
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd6afa90-1371-427b-b1c3-e89398834c6e · outbound
Self-Supervised Learning of Structured Dynamics from Videos $\pi^3$: Permutation-Equivariant Visual Geometry Learning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fde9c463-4331-40e4-ab44-ae8814e8eb1e · outbound
Self-Supervised Learning of Structured Dynamics from Videos CroCo: Self-supervised pre-training for 3d vision tasks by cross-view completion.Advances in Neural Information Processing Systems, 35:3502–3516, 2022
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 892cc710-3da3-4e36-81e7-21a4b860634c · outbound
Self-Supervised Learning of Structured Dynamics from Videos YouTube-VOS: Sequence-to-sequence video object segmentation
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35b55820-2081-4f19-8147-b4dd6e62b70f · outbound
Self-Supervised Learning of Structured Dynamics from Videos YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8f125c4-0e5e-4172-b6aa-e76d34d9bae5 · outbound
Self-Supervised Learning of Structured Dynamics from Videos In pursuit of pixel supervision for visual pre-training.arXiv preprint arXiv:2512.15715, 2025
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f04e3e93-a621-4f88-bbe7-a42d78715838 · outbound
Self-Supervised Learning of Structured Dynamics from Videos Sigmoid loss for language image pre-training
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e869fc90-2e3a-43d9-89d8-815c453154a3 · outbound
Self-Supervised Learning of Structured Dynamics from Videos MonST3R: A simple approach for estimating geometry in the presence of motion
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87e6ebbf-0bee-495f-86a2-ceff86ce66f4 · outbound
Self-Supervised Learning of Structured Dynamics from Videos DINO-WM: World models on pre-trained visual features enable zero-shot planning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfebb87d-d8e2-43b6-85dc-94f29411fda9 · outbound
Self-Supervised Learning of Structured Dynamics from Videos Recurrent video masked autoencoders
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.