Pith. sign in

Paper Citation Record · LEDGER

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling

As of 10 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2608.07409.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07409 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T05:06:02.556817Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 24cbaa79-2054-4393-8f54-e63325486322 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.510016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.510016Z digest=sha256:c9387a089b8cc4a6b7d00c48fe583e2fa8321fa7a920b82fbcef97475d91dd65

Observation 8cc7a71d-ba13-4532-8ecf-be3043235c94 · outbound

This paper cites Learning and Leveraging World Models in Visual Representation Learning.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling Learning and Leveraging World Models in Visual Representation Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.521473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.521473Z digest=sha256:2e673ab41b9ee01db1cd6c3a39fed9046d65f264fa234e7cd4f35c51cedeb6f6

Observation 30bd77c7-f4ca-4cbb-a6b0-8ed961ac26bf · outbound

This paper cites LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.537117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.537117Z digest=sha256:e1bf94fe1e895596a775b835802b18c7a6a08426e34c66169189ca815e7d6e45

Observation 060cbd3e-4b45-4538-a678-032f4d4a006c · outbound

This paper cites MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological Orthogonality.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological Orthogonality

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.546806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.546806Z digest=sha256:675f46c711c630a6be76e1b092e50c01b062c240c7d1104b75487b15d64f758d

Observation f78f59b3-f56f-4c3c-bb46-2df38e4b4c4f · outbound

This paper cites DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.551652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.551652Z digest=sha256:2e125ee9b01c6e859c3018ef347e365216e2904f17a645c2c87867f5a1f23852

Observation c6a2148f-a133-4997-aaa7-aa526c13740e · outbound

This paper cites Implementation Details We use a ViT-Small/16 encoder (15M parameters) for control and ViT-Large for large-scale image/video representation.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling Implementation Details We use a ViT-Small/16 encoder (15M parameters) for control and ViT-Large for large-scale image/video representation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T05:06:02.767745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T05:06:02.556817Z digest=sha256:9e1e432bab8dbfdc0be5aec3a1680b75baba4bb2900bfd2d5320ace8cd44c31e

Observation 5865df4c-7a22-4d94-84a9-4a385d0569b1 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.541971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.541971Z digest=sha256:9e55bf901fbf979e87ceaf5bd16fad03de97e2863490bb50e789cb715b302e36

Observation fda6f588-d0ca-437e-99fd-4408a6368685 · outbound

This paper cites Mastering Diverse Domains through World Models.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling Mastering Diverse Domains through World Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.532154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.532154Z digest=sha256:a146c7da3656b1eec8c74ababd1b7dc75a81374fa876859cf822d8cba23262f1

Observation 0fb21eff-7dbb-464a-8316-875e9d85aca7 · outbound

This paper cites World Models.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling World Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.527212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.527212Z digest=sha256:6883006b1b7476fb3627990e8906138ab04daf33ce733748f4696b1c8cc3c83a

Observation 5f1d49ef-b69e-4614-9827-9b04cfb27af0 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling OpenVLA: An Open-Source Vision-Language-Action Model

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.516123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.516123Z digest=sha256:0bcc2eaa344febec0dfd9894d0c0f8ddb45fcff2fb949c72f5e5b385187bd2ec

Observation 352b986b-b00f-4e1c-950a-b0af7760ac81 · outbound

This paper cites Back to the Features: DINO as a Foundation for Video World Models.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling Back to the Features: DINO as a Foundation for Video World Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.500245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.500245Z digest=sha256:1f7fe4e3703d0af686bab4f03465089735be6bece9efedb52e9267bebf563731

Observation 1d969f65-d119-40c4-8b92-833dea55a436 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.494586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.494586Z digest=sha256:55c0804da1a31c65f158aee2802257bad263b1f0b0b74cf026d728635abc33c3

Observation da6c6179-14b6-45d5-a77f-5e42090c050b · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.505056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.505056Z digest=sha256:5207fb6a2571e6cc42f466962a307463e1089125276ffc0725057e60f3dff8b4

Pith citing papers

No inbound Pith citation observations are available.