Pith. sign in

Paper Citation Record · LEDGER

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling

As of 10 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2608.07409.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07409 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T05:06:02.556817Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 24cbaa79-2054-4393-8f54-e63325486322 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.510016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.510016Z digest=sha256:15b3b16aa1129b9a4e214b89fa6b179ccce9d72ed4432eaa7f113701fdbea828

Observation 8cc7a71d-ba13-4532-8ecf-be3043235c94 · outbound

This paper cites Learning and Leveraging World Models in Visual Representation Learning.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling Learning and Leveraging World Models in Visual Representation Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.521473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.521473Z digest=sha256:adedafd1a19e37499593bc3e20cb590d5569ade60768588cbc124601cafe9f47

Observation 30bd77c7-f4ca-4cbb-a6b0-8ed961ac26bf · outbound

This paper cites LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.537117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.537117Z digest=sha256:f68637a742a8b0a9159f9a10bfd9a6efc2228a127e1d7d3a454ff431929ef58d

Observation 060cbd3e-4b45-4538-a678-032f4d4a006c · outbound

This paper cites MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological Orthogonality.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological Orthogonality

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.546806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.546806Z digest=sha256:c81a555bb48d1ea2c859e62ff6de18191ec6978ef28adf71ee58924ab68b2a4a

Observation f78f59b3-f56f-4c3c-bb46-2df38e4b4c4f · outbound

This paper cites DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.551652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.551652Z digest=sha256:b95b45e06a31d82fafec65a4033030a1a13c6bfe4d7bfad76319a87ec1e857c7

Observation c6a2148f-a133-4997-aaa7-aa526c13740e · outbound

This paper cites Implementation Details We use a ViT-Small/16 encoder (15M parameters) for control and ViT-Large for large-scale image/video representation.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling Implementation Details We use a ViT-Small/16 encoder (15M parameters) for control and ViT-Large for large-scale image/video representation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T05:06:02.767745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T05:06:02.556817Z digest=sha256:6e1d0913a6ec03bdd4db1675cd0da9f4d1eb6a47e19898c4f4e2776ce42d04b3

Observation 5865df4c-7a22-4d94-84a9-4a385d0569b1 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.541971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.541971Z digest=sha256:c5348cf002eb6914deee27b908ab61c1f27542d8a4c5efba2b1bd522dca56146

Observation fda6f588-d0ca-437e-99fd-4408a6368685 · outbound

This paper cites Mastering Diverse Domains through World Models.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling Mastering Diverse Domains through World Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.532154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.532154Z digest=sha256:c0f18359af6dfc09956309f5d58939f0cbe2bcc4a52c144174df5b873cb6c59e

Observation 0fb21eff-7dbb-464a-8316-875e9d85aca7 · outbound

This paper cites World Models.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling World Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.527212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.527212Z digest=sha256:241f4cce14b286023bb50e0906ba6176ca3cadadb5f7914e9571a5ac6e90a182

Observation 5f1d49ef-b69e-4614-9827-9b04cfb27af0 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling OpenVLA: An Open-Source Vision-Language-Action Model

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.516123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.516123Z digest=sha256:b4ea45aa38b3b68e7f01a29f1c187e041109af29e61c7fe56104c66ebb3e2e0d

Observation 352b986b-b00f-4e1c-950a-b0af7760ac81 · outbound

This paper cites Back to the Features: DINO as a Foundation for Video World Models.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling Back to the Features: DINO as a Foundation for Video World Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.500245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.500245Z digest=sha256:b42a8c5af17220e9cd830f145b4ca5c283780eed26a19d8c2a19bf36417b1afa

Observation 1d969f65-d119-40c4-8b92-833dea55a436 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.494586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.494586Z digest=sha256:2939dc28afcb7d46ff45854e24e2609a6e831ed46399d5ea7eba7e2eccbe1661

Observation da6c6179-14b6-45d5-a77f-5e42090c050b · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.505056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.505056Z digest=sha256:71dabb30a4345beec223a607bee066f15a5035105ff3e1daad52952b44e00480

Pith citing papers

No inbound Pith citation observations are available.