Pith. sign in

Paper Citation Record · LEDGER

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception

As of 20 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2507.19272.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.19272 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:58:55.108117Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bfcf75cb-d43f-47e0-b9d3-8b10cf4c9290 · outbound

This paper cites write newline.

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T17:58:55.050673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:58:55.050673Z digest=sha256:d436a4f4e05894b36ee2c7b4cc3804318d8dc93957d644e59297359abc38e6cd

Observation 8b660612-c5ce-4494-bfa4-5a2cba8a2f4d · outbound

This paper cites Is space-time attention all you need for video understanding? In Int.

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception Is space-time attention all you need for video understanding? In Int

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:55.291087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T17:58:55.055634Z digest=sha256:36a9721f895d16887236769a87e63ec5bda9825c4dc571c9132d6760129aaaf7

Observation ed306bc6-73a7-48ad-8721-6ac96472f44c · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:58:55.060038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:58:55.060038Z digest=sha256:c1c3340a188559d4fc0f1b85796eee7f584fffe85e63111f81ec3ec397f319f9

Observation 4f2b75b1-5241-4036-9008-5a474738a922 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception Emerging properties in self-supervised vision transformers

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:55.280285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T17:58:55.064441Z digest=sha256:f595c0ab2eec4ccb52e28349e7830afe8ba328a46b554e771353df21cf5d7129

Observation aa7446f8-5aab-44ff-976e-f8f68f1426aa · outbound

This paper cites A simple framework for contrastive learning of visual representations.

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception A simple framework for contrastive learning of visual representations

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:55.269239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T17:58:55.068785Z digest=sha256:da68676800539ad3c14aa8b3e592aeef7cdeb87f92189b503ec8fbb745f9f791

Observation 8e0f28cf-34e2-417c-b2bb-6cd6b7121538 · outbound

This paper cites D., Azar, M., Piot, B., Guez, A., Pietquin, O., Kavukcuoglu, K., Larochelle, H., Lanctot, M., and Schmitt, S.

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception D., Azar, M., Piot, B., Guez, A., Pietquin, O., Kavukcuoglu, K., Larochelle, H., Lanctot, M., and Schmitt, S

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:55.258755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T17:58:55.072458Z digest=sha256:72395e42f3e874a02c585b75a0df57e315dcc78c87650dd840584f579b1098a1

Observation 2e0ea032-bb93-45ff-bd47-b3f46752b1af · outbound

This paper cites Momentum contrast for unsupervised visual representation learning.

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception Momentum contrast for unsupervised visual representation learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:55.248426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T17:58:55.076134Z digest=sha256:4d49c598b6185b4ea34c4d6354490cc2537f9162b8b604bea45299527f08cac1

Observation 51d9fb25-86d2-4071-89c9-fc4618d8b646 · outbound

This paper cites Masked autoencoders are scalable vision learners.

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception Masked autoencoders are scalable vision learners

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:55.235878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T17:58:55.079896Z digest=sha256:e7d73a7f0318ca7816d9fdff71f29d059410aa8799af05894c69f86a7d5b9784

Observation a2099768-c268-4881-98c8-8ccc683a90d9 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception OpenVLA: An Open-Source Vision-Language-Action Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T17:58:55.083303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:58:55.083303Z digest=sha256:76d7eea03638ae25fbc4e071b67b0e6ff16d5830e20deb525423d2e64aae0bc4

Observation dde729be-0cd1-4fc8-8dba-540d12f3b92f · outbound

This paper cites Microsoft coco: Common objects in context.

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception Microsoft coco: Common objects in context

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:55.224981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T17:58:55.086887Z digest=sha256:cfbb6661ce7c9121c8ab01feaa36eaf11335dc8c7778266784aea643d587137e

Observation f5658ec1-fdaf-4ff5-8ede-dde97edde8d1 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception DINOv2: Learning Robust Visual Features without Supervision

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:58:55.090266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:58:55.090266Z digest=sha256:b9d59cefb3dcc1226d2172c966353641ab3a1a7f187fb658fc748203ce6f4da4

Observation 5364f2c5-c786-426f-afe3-8582d461064b · outbound

This paper cites N., Carreira, J., Asano, Y.

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception N., Carreira, J., Asano, Y

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:55.213746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T17:58:55.094026Z digest=sha256:cae150a32690eb3fa38ee43c8dbd0fee09101b26b4483ba6c346fcc064b5c161

Observation 009bef87-049d-404a-a877-e6bf31408bc1 · outbound

This paper cites PooDLe: Pooled and dense self-supervised learning from naturalistic videos.

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception PooDLe: Pooled and dense self-supervised learning from naturalistic videos

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T17:58:55.097437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:58:55.097437Z digest=sha256:7a2e925c02acf2bd68473ceef8e79a5f3c33a31bbe979b813c4383f2c16242c9

Observation 998ef1dc-dc3c-43e8-8fb2-024b66cfdcc4 · outbound

This paper cites Masked feature prediction for self-supervised visual pre-training.

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception Masked feature prediction for self-supervised visual pre-training

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:55.202428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T17:58:55.101104Z digest=sha256:c8b0624b49e96cb1d4433993a594844a044f32b0c2627e823b42869b32bfc63a

Observation 0665c5d8-5310-46c3-8930-89f9b8937e23 · outbound

This paper cites Scene parsing through ade20k dataset.

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception Scene parsing through ade20k dataset

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:55.191344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T17:58:55.104774Z digest=sha256:c5f961caa15d8ffc2a7986cec66729bf4a1aced5e5f939d57284974568c6cb8d

Observation 08047d75-6a29-49d1-9617-8335d43e76fd · outbound

This paper cites an unresolved cited work.

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:58:55.178854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T17:58:55.108117Z digest=sha256:db1f5b13d4243292950f421d7ac80d014751a90697bd3218f7b4caa7c68d1f3e

Pith citing papers

No inbound Pith citation observations are available.