Pith. sign in

Paper Citation Record · LEDGER

From Image to Video: An Empirical Study of Diffusion Representations

As of 16 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 2 inbound Pith citation observations for arXiv:2502.07001.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07001 v2

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:10:55.455485Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-19T03:42:54.620069Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T03:42:57.302590Z

Reference resolution

69 of 69 outbound references displayed

  • verified exact1
  • verified fuzzy42
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 603821a0-9355-48c5-9aeb-a0481d66b0c7 · outbound

This paper cites Self-supervised learning from images with a joint-embedding predictive architecture.

From Image to Video: An Empirical Study of Diffusion Representations Self-supervised learning from images with a joint-embedding predictive architecture

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.543774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.116949Z digest=sha256:d261b2ea42f503284acc6a1f781c12a8bd7a46f8bdbb1348bf51413fabd3f0bc

Observation cf7f64f8-25c1-4c34-910c-da56a3c62b98 · outbound

This paper cites Video diffusion models learn the struc- ture of the dynamic world.

From Image to Video: An Empirical Study of Diffusion Representations Video diffusion models learn the struc- ture of the dynamic world

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.529087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.122254Z digest=sha256:7ec58f0978d74085a6fcbedfd749404eb0e34a4c57c4c344e6fc0573968976ae

Observation 84998fb5-9c41-4862-ad36-25ae68f7595f · outbound

This paper cites Learning by Reconstruction Produces Uninformative Features For Perception.

From Image to Video: An Empirical Study of Diffusion Representations Learning by Reconstruction Produces Uninformative Features For Perception

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.127624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.127624Z digest=sha256:5ba2a5fd704aeec93706385d242470ed8927490c2353412d2cc6a9fd230221d8

Observation e8f2fc10-1067-4eb7-b442-9c176f7863b6 · outbound

This paper cites Label-efficient se- mantic segmentation with diffusion models.

From Image to Video: An Empirical Study of Diffusion Representations Label-efficient se- mantic segmentation with diffusion models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.513716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.133861Z digest=sha256:5f69f39038e1525131666d5cbcc643707386777e1b399d1087c3421f70747c01

Observation a21fbcaf-e472-4e30-a2ab-a570cd1ceb2a · outbound

This paper cites Revisiting Feature Prediction for Learning Visual Representations from Video.

From Image to Video: An Empirical Study of Diffusion Representations Revisiting Feature Prediction for Learning Visual Representations from Video

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.138930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.138930Z digest=sha256:c38ace8813ca393b0f478f02e7bf521912511fbc96092aa0f0d03d4180441612

Observation 65321793-a5fe-4987-8f30-8828dc9fc8d7 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

From Image to Video: An Empirical Study of Diffusion Representations Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.144179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.144179Z digest=sha256:88a47bd6b16216b9ca080eaa2161f4515f5f814bc232840831ed7f3e680608ec

Observation ae5814d3-b1c7-4258-a241-1b50a0120411 · outbound

This paper cites Deep regression on manifolds: a 3D rota- tion case study.

From Image to Video: An Empirical Study of Diffusion Representations Deep regression on manifolds: a 3D rota- tion case study

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.499007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.149974Z digest=sha256:b3d4bd03f164b2b8ab134d108b640b72eeb04259a20352c2342f9b605bb4bb6c

Observation 2b88fef2-7f80-42e0-b5a4-fc00d09576d3 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

From Image to Video: An Empirical Study of Diffusion Representations Emerg- ing properties in self-supervised vision transformers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.482929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.154906Z digest=sha256:dc05cb4d19bee7e59060927f874212ed419ec3cc349ce229ffbd1a00d15d0d30

Observation 2c952667-f4b4-437d-946e-a2045f6b52b9 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

From Image to Video: An Empirical Study of Diffusion Representations Quo vadis, action recognition? a new model and the kinetics dataset

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.159745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.159745Z digest=sha256:9ddd13c34c02d8843134ddd9e59ec42f9731021efc78b08de9cf2dc9f9f3aad6

Observation ddf9d714-8dfe-4364-b1e4-2a330064446e · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

From Image to Video: An Empirical Study of Diffusion Representations A Short Note on the Kinetics-700 Human Action Dataset

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.165909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.165909Z digest=sha256:2c150842f55592d1b79d9ede28de06ba78cce9cae13c2277e5d368dabbf41fc5

Observation 647ea506-fb21-412f-8709-191a86556497 · outbound

This paper cites Scaling 4D Representations.

From Image to Video: An Empirical Study of Diffusion Representations Scaling 4D Representations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.171296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.171296Z digest=sha256:67884906cba4c8b8be87d2606c641c72cab7fdba65255316cfa64a23e30d5a45

Observation fd97b2ea-81ad-4ab1-84fb-bf8e4fc288ed · outbound

This paper cites an unresolved cited work.

From Image to Video: An Empirical Study of Diffusion Representations Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:10:56.454616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.176297Z digest=sha256:a35fdc021ca5ed81b84a868e697ffa6ae734cb0ebe4e104e51414a728c673a51

Observation 62146370-3bbb-4834-85ed-c1be60865c89 · outbound

This paper cites PaLI-3 Vision Language Models: Smaller, Faster, Stronger.

From Image to Video: An Empirical Study of Diffusion Representations PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.182071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.182071Z digest=sha256:0a365c5779e581f13d24fa6c938818e1110c5a325c20b69cb99c4c3c5ced0f27

Observation c7fff663-46f1-4cee-9c79-59e461514fc6 · outbound

This paper cites Text-to-image diffusion mod- els are zero shot classifiers.

From Image to Video: An Empirical Study of Diffusion Representations Text-to-image diffusion mod- els are zero shot classifiers

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.437657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.187263Z digest=sha256:3ee5e3c4d77e53976031a63f9ce55ba0aeb6ef1b725c7c6665c28f54ffbb1591

Observation 6babfe3c-9302-4e56-9aee-1c71d713bea9 · outbound

This paper cites Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner.

From Image to Video: An Empirical Study of Diffusion Representations Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.421615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.191848Z digest=sha256:5b76596544b3e983a722164ce391e66ced9277c616a829d7e1ab1eee91db3d40

Observation ceb7d32a-da1d-453d-bd88-0f0e747d5aa5 · outbound

This paper cites Depth map prediction from a single image using a multi-scale deep net- work.

From Image to Video: An Empirical Study of Diffusion Representations Depth map prediction from a single image using a multi-scale deep net- work

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.404034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.196339Z digest=sha256:fe43a35bbb6b74cb4e532e6a7e05854c02b2c0e6be83fe23856d96819f64426f

Observation a71f2f6c-c156-4b51-a343-bf48d93d0515 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

From Image to Video: An Empirical Study of Diffusion Representations Taming transformers for high-resolution image synthesis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.200902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.200902Z digest=sha256:6aa9840769fb5627ef9ded5a1717e477e30a81564223288c2e3b0cb4836a1d71

Observation 48a5b6a2-1bea-47c0-9fda-cc34bcdc70fc · outbound

This paper cites Masked autoencoders as spatiotemporal learners.

From Image to Video: An Empirical Study of Diffusion Representations Masked autoencoders as spatiotemporal learners

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.374861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.206228Z digest=sha256:8e57095b8b76946b01ad5b5432c4b3a7903456963410e23769a39d40900d5f4f

Observation ade7fd2f-118a-4d8b-a295-b245f66e7678 · outbound

This paper cites Diffusion Models and Representation Learning: A Survey.

From Image to Video: An Empirical Study of Diffusion Representations Diffusion Models and Representation Learning: A Survey

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.211931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.211931Z digest=sha256:4426727180aabc01946be8a34cf6995590196e567c0cbc87318a21078c92c349

Observation 16a378de-c69e-41e8-8dcc-8909361ab035 · outbound

This paper cites Something Something.

From Image to Video: An Empirical Study of Diffusion Representations Something Something

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.354536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.217130Z digest=sha256:9aa1eb5a86c5e726ab47b4402b7040fac3203150f945a71d5109a748cbe8703a

Observation 20c8d60a-5b6a-427b-b8fd-cfcf907d1c47 · outbound

This paper cites Kubric: A scalable dataset generator.

From Image to Video: An Empirical Study of Diffusion Representations Kubric: A scalable dataset generator

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.338487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.222921Z digest=sha256:d13a415538bdc91e19f0fe6da2be8e6ac1e746624c38c348761c98013b8346fe

Observation 8d92fef7-1769-4742-b1e0-aeed694afa2b · outbound

This paper cites Photorealistic video generation with diffusion models.

From Image to Video: An Empirical Study of Diffusion Representations Photorealistic video generation with diffusion models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.321232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.227765Z digest=sha256:d79d2ae1d390272541f26a4d6477bd5acae31e1957406c19ed3803033dab61bf

Observation 7302ca0f-4509-4216-a41e-7b1fd9d63e17 · outbound

This paper cites Masked autoencoders are scalable vision learners.

From Image to Video: An Empirical Study of Diffusion Representations Masked autoencoders are scalable vision learners

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.303949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.232445Z digest=sha256:63b1bf7000261ae63fc3782ba20f507bf4d2247d2034b5d26c7edf9dfdbbe24f

Observation 737ae461-0659-4691-a059-b1c40a504c27 · outbound

This paper cites Unsupervised keypoints from pretrained diffusion models.

From Image to Video: An Empirical Study of Diffusion Representations Unsupervised keypoints from pretrained diffusion models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.288117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.237362Z digest=sha256:6fe183726b0d443132e59e924be98b57c101d540dd619013b54e0eeca0199c70

Observation a2be131a-35d6-40b5-873b-dd347fdbb184 · outbound

This paper cites Denoising diffu- sion probabilistic models.

From Image to Video: An Empirical Study of Diffusion Representations Denoising diffu- sion probabilistic models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.242828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.242828Z digest=sha256:f28cd7501dad0f48a83c8545cb989780ae226e988a13edc444efdd90601da97f

Observation a32fe247-85c8-4dc7-8b9e-4ab376ec1002 · outbound

This paper cites DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos.

From Image to Video: An Empirical Study of Diffusion Representations DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.247542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.247542Z digest=sha256:996ef68e1059a4b2c2589931e2301c8dd35ffa2f03853fa3e36805a8df1359a6

Observation 68812302-5910-4492-ada0-0d5cf32723e1 · outbound

This paper cites Elucidating the design space of diffusion-based generative models.

From Image to Video: An Empirical Study of Diffusion Representations Elucidating the design space of diffusion-based generative models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.261203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.252875Z digest=sha256:32b3e5b21e11ba91eee1a17a16e5d3c0cb0185da70e62f25cc71469bdaa1759f

Observation 87a222f0-25fc-41b2-bad7-3d32451b08f8 · outbound

This paper cites Your diffusion model is secretly a zero-shot classifier.

From Image to Video: An Empirical Study of Diffusion Representations Your diffusion model is secretly a zero-shot classifier

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.243722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.257709Z digest=sha256:0dffacaca1ad27bb9bcbe1b21d56231edfff9aae803a688ae1c242506bb89c8e

Observation c3a21647-dd1e-4432-8077-36b0cb0d8030 · outbound

This paper cites Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models.

From Image to Video: An Empirical Study of Diffusion Representations Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.263104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.263104Z digest=sha256:9db58b3bbe14dc7d55e430ead70f7562ba4d739a42e1e29fc01ed71028e01d66

Observation fe5d3173-800e-4721-9961-1d22c15a3088 · outbound

This paper cites Diffusion hyperfeatures: Search- ing through time and space for semantic correspondence.

From Image to Video: An Empirical Study of Diffusion Representations Diffusion hyperfeatures: Search- ing through time and space for semantic correspondence

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.226488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.268287Z digest=sha256:76e148592e8e4f37e9d87a7dbbd30c641a30f0c9af557f0208e1caa48f21beb6

Observation db54a12f-71f6-4882-b412-13304013a4d5 · outbound

This paper cites Understanding deep image representations by inverting them.

From Image to Video: An Empirical Study of Diffusion Representations Understanding deep image representations by inverting them

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.211266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.272944Z digest=sha256:8b0a23a7a530a2a12ce5b595f890738ce973e48cf61e846f090ad8788357c255

Observation f6c3539d-ca5e-4a6c-b193-e525ed96fc62 · outbound

This paper cites Lexicon3d: Probing vi- sual foundation models for complex 3d scene understanding.

From Image to Video: An Empirical Study of Diffusion Representations Lexicon3d: Probing vi- sual foundation models for complex 3d scene understanding

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.194835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.277612Z digest=sha256:03057568cf56d4134db9843d2fa6fc58a80ebbffd6c237d9a0b55534405a4be0

Observation 24f582d3-87f7-4921-a4c4-d2b6ba373ad6 · outbound

This paper cites NeRF: Representing scenes as neural radiance fields for view syn- thesis.

From Image to Video: An Empirical Study of Diffusion Representations NeRF: Representing scenes as neural radiance fields for view syn- thesis

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.178589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.282206Z digest=sha256:c085dbad5a8157b50f0d0c439d7540b3666738673ef8fbafcc6cf1d3af3b6219

Observation 7cc0493e-586f-4637-b833-9ad2376161ba · outbound

This paper cites Diffusion Models Beat GANs on Image Classification.

From Image to Video: An Empirical Study of Diffusion Representations Diffusion Models Beat GANs on Image Classification

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.286821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.286821Z digest=sha256:8865cf2dcd9d63e74eee413405e23cde8f5ff1dcc398336c3619f2d8f05185ae

Observation 41f6959e-fb2f-4959-a742-3f5306b03908 · outbound

This paper cites DiffTAD: Temporal action detection with proposal denoising diffusion.

From Image to Video: An Empirical Study of Diffusion Representations DiffTAD: Temporal action detection with proposal denoising diffusion

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.162756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.291762Z digest=sha256:fdcafcf66926c9f02b74c410940fb7c0de30045787a3efcddf750b606ecc500b

Observation 297db8f9-bf23-4d29-abec-819a1b70cac8 · outbound

This paper cites an unresolved cited work.

From Image to Video: An Empirical Study of Diffusion Representations Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.296307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.296307Z digest=sha256:c2cf9dd0b372429cff4d053af5e0e6afc0ce286f2e8d9543fc0f5a9f34e80e57

Observation 1bcfb7f4-2bf5-4b4b-98e8-0b52de14c06d · outbound

This paper cites Self-supervised video pretraining yields robust and more human-aligned visual representations.

From Image to Video: An Empirical Study of Diffusion Representations Self-supervised video pretraining yields robust and more human-aligned visual representations

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.136122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.301353Z digest=sha256:4589d76d0f1420aa9d14afadf6d82e1a41b794b9c4cb253cdf5a03e74ab43cbe

Observation 08856bea-5579-4067-91d6-e9ddf695b466 · outbound

This paper cites Per- ception Test: A diagnostic benchmark for multimodal video models.

From Image to Video: An Empirical Study of Diffusion Representations Per- ception Test: A diagnostic benchmark for multimodal video models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.120448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.306794Z digest=sha256:a7068d23e0b2361d9e9474aeb46e74307e03fc8826c6802d36eff0ef8ea563a3

Observation babb581f-0cb3-4ed5-a5c6-ed0d335fa9bc · outbound

This paper cites Scalable diffusion models with transformers.

From Image to Video: An Empirical Study of Diffusion Representations Scalable diffusion models with transformers

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.311336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.311336Z digest=sha256:cc35d6a8abede7a79ffdd0861a977e4f9635c068858cbdb05fa9fe66021cf661

Observation 08672b8d-6e27-4aa5-a28f-9001a1def01e · outbound

This paper cites The 2017 DAVIS Challenge on Video Object Segmentation.

From Image to Video: An Empirical Study of Diffusion Representations The 2017 DAVIS Challenge on Video Object Segmentation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.315744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.315744Z digest=sha256:6a4e9fa552e5c68225a9a9f24f7b35c50a3501d6852d9dea3e1a4207784fdf57

Observation a81a3c84-5fe0-4f51-a6bc-e6a6fb922909 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

From Image to Video: An Empirical Study of Diffusion Representations Learn- ing transferable visual models from natural language super- vision

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.320638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.320638Z digest=sha256:13d6020423128044067c6249eac9e5f9f263f51efe719f28b8877436ccde62a6

Observation a5c0f855-667d-4bca-8468-481b32e18231 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

From Image to Video: An Empirical Study of Diffusion Representations High-resolution image syn- thesis with latent diffusion models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.326083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.326083Z digest=sha256:64842526b0a73d2dc9930e34ac450507c92f67be519b6ca6c64bdd0e34802765

Observation 116c1a14-2760-4986-a65c-d87aa487f408 · outbound

This paper cites U- Net: Convolutional networks for biomedical image segmen- tation.

From Image to Video: An Empirical Study of Diffusion Representations U- Net: Convolutional networks for biomedical image segmen- tation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.074226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.331019Z digest=sha256:ae7603bd4e6c5158c82de4302b46a877dd2956b7fc65c18b034f0f914d4936b4

Observation dbfba91e-e117-43b3-a687-098d0627e39b · outbound

This paper cites Berg, and Li Fei-Fei.

From Image to Video: An Empirical Study of Diffusion Representations Berg, and Li Fei-Fei

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.056491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.335679Z digest=sha256:7c664a8b0812145540fdb12a013d209617cab52794ac320938881b921cd20d2f

Observation c6ab9844-88b4-41a7-b0c7-50e4d4d98c4c · outbound

This paper cites Scene representation transformer: Geometry-free novel view syn- thesis through set-latent scene representations.

From Image to Video: An Empirical Study of Diffusion Representations Scene representation transformer: Geometry-free novel view syn- thesis through set-latent scene representations

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.041323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.340453Z digest=sha256:5532fb780237b0fb102056955964e7c69b0f163677d949d0157fb0b0e1ae2241

Observation 3b8c6641-290a-4a7d-8e5b-11179bae4712 · outbound

This paper cites Only time can tell: Discovering temporal data for temporal modeling.

From Image to Video: An Empirical Study of Diffusion Representations Only time can tell: Discovering temporal data for temporal modeling

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.025424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.345836Z digest=sha256:da10289555617e61bc4aa077df876518e3bd4aa8bfb2ebe29aa9fbcdd4389772

Observation 6fb7003e-13f5-43c8-bb60-fff423a83da4 · outbound

This paper cites MonoDiffusion: Self-Supervised Monocular Depth Estimation Using Diffusion Model.

From Image to Video: An Empirical Study of Diffusion Representations MonoDiffusion: Self-Supervised Monocular Depth Estimation Using Diffusion Model

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-08T14:10:55.563436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.350854Z digest=sha256:445aec824775a8076801ce10bb2505953b9f77c9a50194d9eaf918c5326df7d1

Observation f154804d-1108-4532-90cb-1f8bf6e6e25a · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

From Image to Video: An Empirical Study of Diffusion Representations Deep unsupervised learning using nonequilibrium thermodynamics

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.355636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.355636Z digest=sha256:6ecc492b6ae54675c05377fd662753d9d55e9261563275da5c2bd3a058bdf5d0

Observation 7b302f72-d9ab-45a3-bb7e-53a11f481eb2 · outbound

This paper cites Scalability in perception for autonomous driving: Waymo Open Dataset.

From Image to Video: An Empirical Study of Diffusion Representations Scalability in perception for autonomous driving: Waymo Open Dataset

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.998842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.360284Z digest=sha256:7715c5b3f50c8c21eb803d20ce255cc457334ba98cde157adf9780a48947d6ad

Observation 23f0ce83-05b5-4ff3-bb97-8c14bd76ba9e · outbound

This paper cites Emergent correspondence from image diffusion.

From Image to Video: An Empirical Study of Diffusion Representations Emergent correspondence from image diffusion

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.984019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.364835Z digest=sha256:0d771224723ad2d4fdac1930b12c6adc2017a09631f9012ede359e3fe855121a

Observation ea064e57-1183-4773-92f6-55510f4c043b · outbound

This paper cites VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

From Image to Video: An Empirical Study of Diffusion Representations VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.968978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.369731Z digest=sha256:c6edcab637c5d23d352b3b6ffda6f2ecd6cfb640d46c49527f21b8def3d07768

Observation 2ffa4df1-3728-42d1-89f9-fbbfc6e56103 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

From Image to Video: An Empirical Study of Diffusion Representations Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.374419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.374419Z digest=sha256:ed4b72f44e59dea5222831d9b8c24fc330356e0644ccdbe6ae39779483203548

Observation 94445917-140d-4ed1-83c5-d54ac64332f1 · outbound

This paper cites Neural discrete representation learning.

From Image to Video: An Empirical Study of Diffusion Representations Neural discrete representation learning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.953590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.379343Z digest=sha256:09211531cb529a6f0753a537dbaa47a0e8fd4e45cfaddf7f9d37ee899697650e

Observation 958500c9-b3bb-4d5f-9c02-e736c5693d7a · outbound

This paper cites The iNaturalist species classification and detection dataset.

From Image to Video: An Empirical Study of Diffusion Representations The iNaturalist species classification and detection dataset

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.938315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.383979Z digest=sha256:556e4dc8f9a3964fd61b7e71528d4832426d3cf30b101922db3f6c29be640be8

Observation af177f6d-7537-46b7-923d-127918f55d41 · outbound

This paper cites Hudson, Thomas Albert Keck, Joao Carreira, Alexey Doso- vitskiy, Mehdi S.

From Image to Video: An Empirical Study of Diffusion Representations Hudson, Thomas Albert Keck, Joao Carreira, Alexey Doso- vitskiy, Mehdi S

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.923033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.389085Z digest=sha256:2a342c82315235896a44d98f7403715eda5403fcce7fdf987542a660e94165f9

Observation 3c76d85f-a3a4-41f4-b74a-885ad9d09450 · outbound

This paper cites Attention is all you need.

From Image to Video: An Empirical Study of Diffusion Representations Attention is all you need

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.907171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.393683Z digest=sha256:addfa9a2828c66d5408c016abd1ae0ecde59683ade2637589010629f64a49b51

Observation 7502c86a-a9a9-438b-82d0-d8af2e1c7a85 · outbound

This paper cites VideoMAE v2: Scaling video masked autoencoders with dual masking.

From Image to Video: An Empirical Study of Diffusion Representations VideoMAE v2: Scaling video masked autoencoders with dual masking

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.891728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.398243Z digest=sha256:6657d2d9766bd2218f22ad998a92442ea14eb699463c40d2f7cd009b457b8eb3

Observation 2f9ede65-af59-4efc-aee6-1dbd688e7143 · outbound

This paper cites Controlling Space and Time with Diffusion Models.

From Image to Video: An Empirical Study of Diffusion Representations Controlling Space and Time with Diffusion Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.402609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.402609Z digest=sha256:c211255bb61cf39b92fb929b0a44c44e5a614399c5502c2fa024bffb0459a5ea

Observation 0fe87e0a-d124-4519-87a1-eba6c744e2ac · outbound

This paper cites Denoising diffusion autoencoders are unified self-supervised learners.

From Image to Video: An Empirical Study of Diffusion Representations Denoising diffusion autoencoders are unified self-supervised learners

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.874562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.407888Z digest=sha256:0ddbe18f48111465e7ac9457ae3261a7766df6930b28cae864ef247fbcb19a37

Observation a6a4e390-b173-415c-9063-ac78ec4bac71 · outbound

This paper cites Open-vocabulary panop- tic segmentation with text-to-image diffusion models.

From Image to Video: An Empirical Study of Diffusion Representations Open-vocabulary panop- tic segmentation with text-to-image diffusion models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.858262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.412593Z digest=sha256:42d8f93167ccbd9989c53b03b07d60a21d75d30433872317adf637608f0e9fce

Observation 27edcb1a-eb8b-4d72-8352-ff598a435d7b · outbound

This paper cites Diffusion Model as Rep- resentation Learner.

From Image to Video: An Empirical Study of Diffusion Representations Diffusion Model as Rep- resentation Learner

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.842384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.417404Z digest=sha256:20cc57b83ad796e5218fb9777514a2b2daf02e3e889b69118e8ce97a648124ad

Observation b0ac0337-c4a8-47d4-ba5a-16d07178cb95 · outbound

This paper cites Gundavarapu, Luca Ver- sari, Kihyuk Sohn, David Minnen, Yong Cheng, Vigh- nesh Birodkar, Agrim Gupta, Xiuye Gu, Alexander G.

From Image to Video: An Empirical Study of Diffusion Representations Gundavarapu, Luca Ver- sari, Kihyuk Sohn, David Minnen, Yong Cheng, Vigh- nesh Birodkar, Agrim Gupta, Xiuye Gu, Alexander G

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.826303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.421970Z digest=sha256:15bf63bb72e09808725181a7b0365ed32e832ac6e2a263f669f03b4af7519c2f

Observation af5a3197-f2c1-4c5a-901f-b84be0514da7 · outbound

This paper cites A tale of two features: Stable diffusion complements DINO for zero-shot semantic correspondence.

From Image to Video: An Empirical Study of Diffusion Representations A tale of two features: Stable diffusion complements DINO for zero-shot semantic correspondence

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.809804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.426588Z digest=sha256:af3f077cfa533a13e9c609ed5e2d6213b17ccc36d95a510070064703b6da13c6

Observation adc973b1-8a9f-4c9f-810d-c9d4f59e998a · outbound

This paper cites A Survey of Diffusion Based Image Generation Models: Issues and Their Solutions.

From Image to Video: An Empirical Study of Diffusion Representations A Survey of Diffusion Based Image Generation Models: Issues and Their Solutions

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.431225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.431225Z digest=sha256:5980bb9d4c8ec1edf62156fd255bf22a381331523203a136774452583139249c

Observation 54c34bff-d2ce-477f-8bfd-a6ac1a5c169a · outbound

This paper cites an unresolved cited work.

From Image to Video: An Empirical Study of Diffusion Representations Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:10:55.793939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.437068Z digest=sha256:0a786d1c1933cd9c3640e9952860add0cc58e9e52eba3be86937a9108058c2e2

Observation 1462df58-ed28-4474-a511-dcab881afb14 · outbound

This paper cites Unleashing text-to-image diffusion models for visual perception.

From Image to Video: An Empirical Study of Diffusion Representations Unleashing text-to-image diffusion models for visual perception

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.441538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.441538Z digest=sha256:69fa42d413ed4694006b7ba21b1e410482c357a2debc24a140bf2a8dabe6f168

Observation 0f76226c-d04b-4335-801b-81acca099df6 · outbound

This paper cites Places: A 10 million image database for scene recognition.

From Image to Video: An Empirical Study of Diffusion Representations Places: A 10 million image database for scene recognition

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.767105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.446497Z digest=sha256:33c2373e8494817f3e0ae989f665976700aa6e9c171f9c3de625ec4d97553ab7

Observation 3b551034-2701-48cf-a3d5-164bf2b1f3d6 · outbound

This paper cites Stereo magnification: Learning view syn- thesis using multiplane images.

From Image to Video: An Empirical Study of Diffusion Representations Stereo magnification: Learning view syn- thesis using multiplane images

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.750105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T14:10:55.450982Z digest=sha256:794f92a206cdcc3133148c05c857301e561f7557f5928140cc3ed72029a4e08d

Observation 6e333442-b790-49a6-871e-4f4bcf9c06e7 · outbound

This paper cites Exploring Pre-trained Text-to-Video Diffusion Models for Referring Video Object Segmentation.

From Image to Video: An Empirical Study of Diffusion Representations Exploring Pre-trained Text-to-Video Diffusion Models for Referring Video Object Segmentation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.455485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.455485Z digest=sha256:690a62a8d348d4485425b12b694b50b0edd5a24167b6fcb1aae656e1f5ca7620

Pith citing papers

Observation a2d5ddbd-90f3-4119-9328-a330dd5eaf9c · inbound

Frozen Forecasting: A Unified Evaluation cites this paper.

Frozen Forecasting: A Unified Evaluation From Image to Video: An Empirical Study of Diffusion Representations

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:42:57.304495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T03:42:54.620069Z digest=sha256:8985c63a68f089f91d30ba8cdf6cbddee73ac6b8ccc700f2eba952ef578b6e6a

Observation 0b36fdbc-3921-4f57-8d81-c211e5d371ff · inbound

Video Generation with Predictive Latents cites this paper.

Video Generation with Predictive Latents From Image to Video: An Empirical Study of Diffusion Representations

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:26:09.069938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-09T16:55:54.709581Z digest=sha256:f7437af3ea98998bed3d279f6a1c9a9f33f3278c827b466477439312f93c630e