Pith. sign in

Paper Citation Record · LEDGER

Gen4U: Unifying Video Generation and Understanding via Diffusion

As of 10 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2607.06856.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06856 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-10T00:16:43.190961Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact9
  • verified fuzzy6
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch8

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4226e5dd-79a6-49e8-90fb-2d5b018c1d64 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

Gen4U: Unifying Video Generation and Understanding via Diffusion V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T00:26:39.626831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:d8a515363668593a0cd0841c0265d69624e33fc97e80bb8ab51d6823f16abe1d

Observation 05d12c6b-149e-43d2-8139-2d3df8eb2573 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Gen4U: Unifying Video Generation and Understanding via Diffusion PaliGemma: A versatile 3B VLM for transfer

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T00:26:39.650395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:b855fb599fd5970b36a629b30b706f27d9da4998463feba01b80bcf862795fbc

Observation afb13ef3-a1bd-48c4-9abc-a824187c2672 · outbound

This paper cites Walk in the cloud: Learning curves for point clouds shape analysis, pp.

Gen4U: Unifying Video Generation and Understanding via Diffusion Walk in the cloud: Learning curves for point clouds shape analysis, pp

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-10T00:26:39.079490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:48996e334875a9a4f937fd724b3ac11a904e70277785f53664b1a9e47de5ea8a

Observation a3458577-a4a7-4f4c-a660-a63981166db6 · outbound

This paper cites Scaling 4D Representations.

Gen4U: Unifying Video Generation and Understanding via Diffusion Scaling 4D Representations

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T00:26:39.630147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:3aea0d22293805529a7a84ddbc35bf3dfc53bb36744f01ed8d594768b7dd3138

Observation 9620d0e2-5efe-4e86-8c78-32f82fb9e2fa · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Gen4U: Unifying Video Generation and Understanding via Diffusion Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:26:39.647632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:db4d1cf5f4519874a423466c78e907659c4f7ad8c0f65212eddb8bc1e8f44c2d

Observation fe557c13-b1d6-4c7e-8957-92ac7a91ba7f · outbound

This paper cites Whatever next? Predictive brains, situated agents, and the future of cognitive science , volume =.

Gen4U: Unifying Video Generation and Understanding via Diffusion Whatever next? Predictive brains, situated agents, and the future of cognitive science , volume =

Reference 6

Resolution
metadata mismatch
doi, observed 2026-07-10T00:26:39.075687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:06f9d5878f5731c280ca7c88ba89705ea25a77d3a7faac67f15ee1c916dd99b5

Observation 21174190-d36d-49e8-842e-c936c346c4fc · outbound

This paper cites Accessed: 2026-04-17.

Gen4U: Unifying Video Generation and Understanding via Diffusion Accessed: 2026-04-17

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:26:40.255962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:9ecbb6705e8ef5f09d5c0daec70acfd2f4e9b94aa47049afc350b22c9026d60c

Observation ae41bede-a22c-4637-8305-daca36e05b63 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Gen4U: Unifying Video Generation and Understanding via Diffusion Gemma 2: Improving Open Language Models at a Practical Size

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:26:39.641686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:3e1ef3139dd711a56504d204a6683784d9a34aa4666148b9f27482d51da6e118

Observation 476340bb-1b15-4b50-8511-7afbd7835261 · outbound

This paper cites doi: 10.1038/s41586-025-08744-2.

Gen4U: Unifying Video Generation and Understanding via Diffusion doi: 10.1038/s41586-025-08744-2

Reference 9

Resolution
verified exact
doi, observed 2026-07-10T00:26:39.073411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:d6d3707ac9f85f83d748b85bd20b9d56604c7d42a0da3135f8f133c196bbd180

Observation 9485af90-32a1-488e-815d-c0a10451c24f · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Gen4U: Unifying Video Generation and Understanding via Diffusion Adam: A Method for Stochastic Optimization

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:26:39.644945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:b46897e759d75efade7c786f31db3f79c6e004a48e34dd6be7f3b16a146e2ded

Observation 54b15358-3c4a-47c0-b59b-da5c8b931c46 · outbound

This paper cites V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning.

Gen4U: Unifying Video Generation and Understanding via Diffusion V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T00:26:39.658479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:282c3b09e75458cff93e3f03cdb67d7d5500b6ba3865e6e06e6f2524b340ac1d

Observation 1b609cff-0a19-4121-9a35-6161312c2e27 · outbound

This paper cites A simple recipe for contrastively pre-training video-first encoders beyond 16 frames.2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14386–14397,.

Gen4U: Unifying Video Generation and Understanding via Diffusion A simple recipe for contrastively pre-training video-first encoders beyond 16 frames.2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14386–14397,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:26:40.254281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:820b6a4c9e1a9a8dd0937cb6c47e119f7b5c29f277fa3451cd036cf4da9b0099

Observation c9737ae7-86b5-4b4c-bbf7-8795e08d310c · outbound

This paper cites Walk in the cloud: Learning curves for point clouds shape analysis, pp.

Gen4U: Unifying Video Generation and Understanding via Diffusion Walk in the cloud: Learning curves for point clouds shape analysis, pp

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-10T00:26:39.075356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:73c9bbbce85a519900c37e647d9c8de76bd24fc95cd71254a58f5e55a3d28378

Observation 2247903f-0719-48cc-8eaf-f2e9a1063933 · outbound

This paper cites PaliGemma 2: A Family of Versatile VLMs for Transfer.

Gen4U: Unifying Video Generation and Understanding via Diffusion PaliGemma 2: A Family of Versatile VLMs for Transfer

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:26:39.653175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:95fd02bac3e24da66266fa503fd73ad8741e857bc5f116886306f63519072a94

Observation 9e8f8999-e3ee-43a3-a8ea-35e4c9ddc3c3 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Gen4U: Unifying Video Generation and Understanding via Diffusion Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T00:26:39.655810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:dae2796f202f433521519061100cb68af17643188730b63f0f173d8279b6fc76

Observation 911d4e0d-d8a7-461d-8af8-46884e9f7242 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Gen4U: Unifying Video Generation and Understanding via Diffusion SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:26:39.636319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:93c355c57129e5dd4e4a1f5ca0a0ac91172d7ad3c52f94dee3ac7ae22335fad9

Observation 12928339-f65f-4e6a-8c5c-3913658eeaa3 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Gen4U: Unifying Video Generation and Understanding via Diffusion Wan: Open and Advanced Large-Scale Video Generative Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:26:39.633409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:9edcc1f420f8081b0557b801926abe7cb116f215f2b168362374c39f17247c11

Observation 7846d0c7-704a-4678-bac8-79df2639d526 · outbound

This paper cites URLhttps://doi.org/10.1007/978-3-031-73013-9_23.

Gen4U: Unifying Video Generation and Understanding via Diffusion URLhttps://doi.org/10.1007/978-3-031-73013-9_23

Reference 18

Resolution
verified exact
doi, observed 2026-07-10T00:26:39.072377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:2aa0e36cd93a885c47c4e0aec58135dbbeb9d0b79099887dd3a5f0337bdeb576

Observation aa4ea3b1-c690-484e-a86f-6a14c1f70746 · outbound

This paper cites Video models are zero-shot learners and reasoners.

Gen4U: Unifying Video Generation and Understanding via Diffusion Video models are zero-shot learners and reasoners

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:26:39.639035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:348a8474aced1db79bcafc3a57fab24af4ca72f23efee797146d12ee0ea7e14d

Observation 2362a645-e376-416c-bbee-df51a3d874a0 · outbound

This paper cites 14 A Mutualk-NN alignment metric We describe the Mutual k-Nearest Neighbours (MkNN) alignment metric used throughout the paper, following [Huh et al., 2024, Zhu et al., 2026].

Gen4U: Unifying Video Generation and Understanding via Diffusion 14 A Mutualk-NN alignment metric We describe the Mutual k-Nearest Neighbours (MkNN) alignment metric used throughout the paper, following [Huh et al., 2024, Zhu et al., 2026]

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:26:40.250928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:42b77617b0427199f1a54e99d48338db947ec952290f1980894333b89a9e9441

Observation bf9563af-9efb-409e-87a2-e19edb9fec02 · outbound

This paper cites Best Single block.

Gen4U: Unifying Video Generation and Understanding via Diffusion Best Single block

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:26:40.252441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:12ee767f2abfeb54565401578c48d3e3c1093dc7ddf5f8d8b8b67731c03a9781

Observation 249aafaa-771b-43ac-bd02-f4bc5a525261 · outbound

This paper cites The linear decoder presents frames that are sharper, but temporally misaligned with the ground truth, as compared to the attention head.

Gen4U: Unifying Video Generation and Understanding via Diffusion The linear decoder presents frames that are sharper, but temporally misaligned with the ground truth, as compared to the attention head

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:26:40.249176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:cc91aa6419f56253af77bbd0b646b78781c5ed3eb9a8cfe36ff1f9bd03da015e

Observation f55dd9ea-e6cd-4d07-a874-21a46d25f479 · outbound

This paper cites It plateaus after 10k steps so that the LLM is still updated very slowly to avoid catastrophic forgetting.

Gen4U: Unifying Video Generation and Understanding via Diffusion It plateaus after 10k steps so that the LLM is still updated very slowly to avoid catastrophic forgetting

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:26:40.247348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:10a77de2b2d29a8870588fa0234375641548621139ababa1e024ad2c5d10e9f1

Pith citing papers

No inbound Pith citation observations are available.