Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-10T00:16:43.190961Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2607.06856.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-10T00:16:43.190961Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
23 of 23 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4226e5dd-79a6-49e8-90fb-2d5b018c1d64 · outbound
Gen4U: Unifying Video Generation and Understanding via Diffusion V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 05d12c6b-149e-43d2-8139-2d3df8eb2573 · outbound
Gen4U: Unifying Video Generation and Understanding via Diffusion PaliGemma: A versatile 3B VLM for transfer
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation afb13ef3-a1bd-48c4-9abc-a824187c2672 · outbound
Gen4U: Unifying Video Generation and Understanding via Diffusion Walk in the cloud: Learning curves for point clouds shape analysis, pp
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a3458577-a4a7-4f4c-a660-a63981166db6 · outbound
Gen4U: Unifying Video Generation and Understanding via Diffusion Scaling 4D Representations
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9620d0e2-5efe-4e86-8c78-32f82fb9e2fa · outbound
Gen4U: Unifying Video Generation and Understanding via Diffusion Microsoft COCO Captions: Data Collection and Evaluation Server
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fe557c13-b1d6-4c7e-8957-92ac7a91ba7f · outbound
Gen4U: Unifying Video Generation and Understanding via Diffusion Whatever next? Predictive brains, situated agents, and the future of cognitive science , volume =
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 21174190-d36d-49e8-842e-c936c346c4fc · outbound
Gen4U: Unifying Video Generation and Understanding via Diffusion Accessed: 2026-04-17
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ae41bede-a22c-4637-8305-daca36e05b63 · outbound
Gen4U: Unifying Video Generation and Understanding via Diffusion Gemma 2: Improving Open Language Models at a Practical Size
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 476340bb-1b15-4b50-8511-7afbd7835261 · outbound
Gen4U: Unifying Video Generation and Understanding via Diffusion doi: 10.1038/s41586-025-08744-2
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9485af90-32a1-488e-815d-c0a10451c24f · outbound
Gen4U: Unifying Video Generation and Understanding via Diffusion Adam: A Method for Stochastic Optimization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 54b15358-3c4a-47c0-b59b-da5c8b931c46 · outbound
Gen4U: Unifying Video Generation and Understanding via Diffusion V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1b609cff-0a19-4121-9a35-6161312c2e27 · outbound
Gen4U: Unifying Video Generation and Understanding via Diffusion A simple recipe for contrastively pre-training video-first encoders beyond 16 frames.2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14386–14397,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c9737ae7-86b5-4b4c-bbf7-8795e08d310c · outbound
Gen4U: Unifying Video Generation and Understanding via Diffusion Walk in the cloud: Learning curves for point clouds shape analysis, pp
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2247903f-0719-48cc-8eaf-f2e9a1063933 · outbound
Gen4U: Unifying Video Generation and Understanding via Diffusion PaliGemma 2: A Family of Versatile VLMs for Transfer
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9e8f8999-e3ee-43a3-a8ea-35e4c9ddc3c3 · outbound
Gen4U: Unifying Video Generation and Understanding via Diffusion Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 911d4e0d-d8a7-461d-8af8-46884e9f7242 · outbound
Gen4U: Unifying Video Generation and Understanding via Diffusion SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 12928339-f65f-4e6a-8c5c-3913658eeaa3 · outbound
Gen4U: Unifying Video Generation and Understanding via Diffusion Wan: Open and Advanced Large-Scale Video Generative Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7846d0c7-704a-4678-bac8-79df2639d526 · outbound
Gen4U: Unifying Video Generation and Understanding via Diffusion URLhttps://doi.org/10.1007/978-3-031-73013-9_23
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation aa4ea3b1-c690-484e-a86f-6a14c1f70746 · outbound
Gen4U: Unifying Video Generation and Understanding via Diffusion Video models are zero-shot learners and reasoners
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2362a645-e376-416c-bbee-df51a3d874a0 · outbound
Gen4U: Unifying Video Generation and Understanding via Diffusion 14 A Mutualk-NN alignment metric We describe the Mutual k-Nearest Neighbours (MkNN) alignment metric used throughout the paper, following [Huh et al., 2024, Zhu et al., 2026]
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bf9563af-9efb-409e-87a2-e19edb9fec02 · outbound
Gen4U: Unifying Video Generation and Understanding via Diffusion Best Single block
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 249aafaa-771b-43ac-bd02-f4bc5a525261 · outbound
Gen4U: Unifying Video Generation and Understanding via Diffusion The linear decoder presents frames that are sharper, but temporally misaligned with the ground truth, as compared to the attention head
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f55dd9ea-e6cd-4d07-a874-21a46d25f479 · outbound
Gen4U: Unifying Video Generation and Understanding via Diffusion It plateaus after 10k steps so that the LLM is still updated very slowly to avoid catastrophic forgetting
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
No inbound Pith citation observations are available.