Pith. sign in

Paper Citation Record · LEDGER

Gen4U: Unifying Video Generation and Understanding via Diffusion

As of 8 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2607.06856.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06856 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-10T00:16:43.190961Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact9
  • verified fuzzy6
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch8

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4226e5dd-79a6-49e8-90fb-2d5b018c1d64 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

Gen4U: Unifying Video Generation and Understanding via Diffusion V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T00:26:39.626831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:f6c2d671184ae2aaa224c96a81a1dc517498d37129cdb7486518af170d29150e

Observation 05d12c6b-149e-43d2-8139-2d3df8eb2573 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Gen4U: Unifying Video Generation and Understanding via Diffusion PaliGemma: A versatile 3B VLM for transfer

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T00:26:39.650395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:b41b9a0dd1e98a31cd4e33f8011fda3a77156b37c28996622a2f1b3b93d2571d

Observation afb13ef3-a1bd-48c4-9abc-a824187c2672 · outbound

This paper cites Walk in the cloud: Learning curves for point clouds shape analysis, pp.

Gen4U: Unifying Video Generation and Understanding via Diffusion Walk in the cloud: Learning curves for point clouds shape analysis, pp

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-10T00:26:39.079490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:bd306b97ca337ea1be6d3c20795a6f7d22fefd437570cd42f889c8c80c0d5684

Observation a3458577-a4a7-4f4c-a660-a63981166db6 · outbound

This paper cites Scaling 4D Representations.

Gen4U: Unifying Video Generation and Understanding via Diffusion Scaling 4D Representations

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T00:26:39.630147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:a85e10cabeff087e8732d37574687142d05915c8a89d511427cad163b2a82841

Observation 9620d0e2-5efe-4e86-8c78-32f82fb9e2fa · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Gen4U: Unifying Video Generation and Understanding via Diffusion Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:26:39.647632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:5d4f248eb2841a72792800a5df0a2bff6a5e59a3b729d14eb879fbc501a34540

Observation fe557c13-b1d6-4c7e-8957-92ac7a91ba7f · outbound

This paper cites Whatever next? Predictive brains, situated agents, and the future of cognitive science , volume =.

Gen4U: Unifying Video Generation and Understanding via Diffusion Whatever next? Predictive brains, situated agents, and the future of cognitive science , volume =

Reference 6

Resolution
metadata mismatch
doi, observed 2026-07-10T00:26:39.075687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:c472caf1293c979a5ba620a680060d2f9016e09b3b9fa1a8f3b1c027f98632b2

Observation 21174190-d36d-49e8-842e-c936c346c4fc · outbound

This paper cites Accessed: 2026-04-17.

Gen4U: Unifying Video Generation and Understanding via Diffusion Accessed: 2026-04-17

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:26:40.255962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:18030d1692e9d372e513844b497a990f503e3c2786dcc6059d488de8ef7bf638

Observation ae41bede-a22c-4637-8305-daca36e05b63 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Gen4U: Unifying Video Generation and Understanding via Diffusion Gemma 2: Improving Open Language Models at a Practical Size

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:26:39.641686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:a19a7813f3be9dfca08e1e82a378f63ca23a720cd113d411334d58ab3fa913df

Observation 476340bb-1b15-4b50-8511-7afbd7835261 · outbound

This paper cites doi: 10.1038/s41586-025-08744-2.

Gen4U: Unifying Video Generation and Understanding via Diffusion doi: 10.1038/s41586-025-08744-2

Reference 9

Resolution
verified exact
doi, observed 2026-07-10T00:26:39.073411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:ce33f71f76a75cd4cb0819352f3641658910064c70f00889f63359baff8f3872

Observation 9485af90-32a1-488e-815d-c0a10451c24f · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Gen4U: Unifying Video Generation and Understanding via Diffusion Adam: A Method for Stochastic Optimization

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:26:39.644945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:47df4ab59303fd8ebfadd13859aa327a3682bed3f77ae700b670a7ddbd6c9ccb

Observation 54b15358-3c4a-47c0-b59b-da5c8b931c46 · outbound

This paper cites V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning.

Gen4U: Unifying Video Generation and Understanding via Diffusion V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T00:26:39.658479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:069b36d7d1fdfa0eb9f294d2c038588cbebe52add3e0c0b7321d32681c5ac504

Observation 1b609cff-0a19-4121-9a35-6161312c2e27 · outbound

This paper cites A simple recipe for contrastively pre-training video-first encoders beyond 16 frames.2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14386–14397,.

Gen4U: Unifying Video Generation and Understanding via Diffusion A simple recipe for contrastively pre-training video-first encoders beyond 16 frames.2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14386–14397,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:26:40.254281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:066a89d38d794b5061c8fa52e6a54d325b2905defcf53189144e3f6611862f59

Observation c9737ae7-86b5-4b4c-bbf7-8795e08d310c · outbound

This paper cites Walk in the cloud: Learning curves for point clouds shape analysis, pp.

Gen4U: Unifying Video Generation and Understanding via Diffusion Walk in the cloud: Learning curves for point clouds shape analysis, pp

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-10T00:26:39.075356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:fca0ebab7ecf28dd660949e4799df3d5d6eec4807748e691145b0fc0ed38aa64

Observation 2247903f-0719-48cc-8eaf-f2e9a1063933 · outbound

This paper cites PaliGemma 2: A Family of Versatile VLMs for Transfer.

Gen4U: Unifying Video Generation and Understanding via Diffusion PaliGemma 2: A Family of Versatile VLMs for Transfer

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:26:39.653175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:c2ccb731458271199c613b7c69b93f7fe0358667ef44df5bcd5eea71f583ae50

Observation 9e8f8999-e3ee-43a3-a8ea-35e4c9ddc3c3 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Gen4U: Unifying Video Generation and Understanding via Diffusion Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T00:26:39.655810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:6661be79012d5f9ae3aee7c051049f5ede01eda0a5bdb8dd90c9b8f65534ed7d

Observation 911d4e0d-d8a7-461d-8af8-46884e9f7242 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Gen4U: Unifying Video Generation and Understanding via Diffusion SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:26:39.636319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:afbab1c43952bf603190e0c721b568681b58b820f0652ce410fe837da82ac239

Observation 12928339-f65f-4e6a-8c5c-3913658eeaa3 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Gen4U: Unifying Video Generation and Understanding via Diffusion Wan: Open and Advanced Large-Scale Video Generative Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:26:39.633409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:774036f46c5bcf4c425ac539c693c70c0e2d5104d531aa57e6212558701c5cc2

Observation 7846d0c7-704a-4678-bac8-79df2639d526 · outbound

This paper cites URLhttps://doi.org/10.1007/978-3-031-73013-9_23.

Gen4U: Unifying Video Generation and Understanding via Diffusion URLhttps://doi.org/10.1007/978-3-031-73013-9_23

Reference 18

Resolution
verified exact
doi, observed 2026-07-10T00:26:39.072377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:b6e2f86b406b7d7944c60a9e70b998347482353bc8a07d7aea966f091ff45a27

Observation aa4ea3b1-c690-484e-a86f-6a14c1f70746 · outbound

This paper cites Video models are zero-shot learners and reasoners.

Gen4U: Unifying Video Generation and Understanding via Diffusion Video models are zero-shot learners and reasoners

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:26:39.639035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:205be9945dae5db8fa7723f87596e841152a4b22fdb5e95b60d201e74f8bcd39

Observation 2362a645-e376-416c-bbee-df51a3d874a0 · outbound

This paper cites 14 A Mutualk-NN alignment metric We describe the Mutual k-Nearest Neighbours (MkNN) alignment metric used throughout the paper, following [Huh et al., 2024, Zhu et al., 2026].

Gen4U: Unifying Video Generation and Understanding via Diffusion 14 A Mutualk-NN alignment metric We describe the Mutual k-Nearest Neighbours (MkNN) alignment metric used throughout the paper, following [Huh et al., 2024, Zhu et al., 2026]

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:26:40.250928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:91c940e547cb081ef03b44ca2f84f0700bb13275ba27c6d5bd1e1056bf78dc08

Observation bf9563af-9efb-409e-87a2-e19edb9fec02 · outbound

This paper cites Best Single block.

Gen4U: Unifying Video Generation and Understanding via Diffusion Best Single block

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:26:40.252441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:6c845760a288413b0bf01e77aec741f843d241d2eaf8ab4b2638feb62d94c0e6

Observation 249aafaa-771b-43ac-bd02-f4bc5a525261 · outbound

This paper cites The linear decoder presents frames that are sharper, but temporally misaligned with the ground truth, as compared to the attention head.

Gen4U: Unifying Video Generation and Understanding via Diffusion The linear decoder presents frames that are sharper, but temporally misaligned with the ground truth, as compared to the attention head

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:26:40.249176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:32df0446f02525b9072162f351d1ad83a8576e0c7bd641c321f014b94024721f

Observation f55dd9ea-e6cd-4d07-a874-21a46d25f479 · outbound

This paper cites It plateaus after 10k steps so that the LLM is still updated very slowly to avoid catastrophic forgetting.

Gen4U: Unifying Video Generation and Understanding via Diffusion It plateaus after 10k steps so that the LLM is still updated very slowly to avoid catastrophic forgetting

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:26:40.247348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:343578cb95538ab585c7d406f48fc02589a0e025ff604c7b6e87babf0b4e9038

Pith citing papers

No inbound Pith citation observations are available.