Pith. sign in

Paper Citation Record · LEDGER

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception

As of 17 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2507.19272.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.19272 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:58:55.108117Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bfcf75cb-d43f-47e0-b9d3-8b10cf4c9290 · outbound

This paper cites write newline.

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T17:58:55.050673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:58:55.050673Z digest=sha256:cb05b54291d8dc80a9cd5fc37d31bfd08b35e82f238cb524fe7e0aade6ba3704

Observation 8b660612-c5ce-4494-bfa4-5a2cba8a2f4d · outbound

This paper cites Is space-time attention all you need for video understanding? In Int.

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception Is space-time attention all you need for video understanding? In Int

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:55.291087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:58:55.055634Z digest=sha256:05ff6f7ab222e50394a33465fc9715bf341427fc384f18f1ba76baa0b58b011f

Observation ed306bc6-73a7-48ad-8721-6ac96472f44c · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:58:55.060038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:58:55.060038Z digest=sha256:16b51530f22bf08d0da940a5882619ef5e0d9d056b58b5c5ede6913120e88bea

Observation 4f2b75b1-5241-4036-9008-5a474738a922 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception Emerging properties in self-supervised vision transformers

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:55.280285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:58:55.064441Z digest=sha256:41b3ccfc5a384e69f659c090171d12cf11ab5b4ccea616ab224f18b5a1c0fefd

Observation aa7446f8-5aab-44ff-976e-f8f68f1426aa · outbound

This paper cites A simple framework for contrastive learning of visual representations.

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception A simple framework for contrastive learning of visual representations

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:55.269239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:58:55.068785Z digest=sha256:33d7eef20e2159df49c5efb371878b4da0622d42b574cd3dea4829433cd42e27

Observation 8e0f28cf-34e2-417c-b2bb-6cd6b7121538 · outbound

This paper cites D., Azar, M., Piot, B., Guez, A., Pietquin, O., Kavukcuoglu, K., Larochelle, H., Lanctot, M., and Schmitt, S.

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception D., Azar, M., Piot, B., Guez, A., Pietquin, O., Kavukcuoglu, K., Larochelle, H., Lanctot, M., and Schmitt, S

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:55.258755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:58:55.072458Z digest=sha256:0cd406c015e64c1f9e78ca13aa1132b792d3a29f437f4338cd1cca5d6e16071d

Observation 2e0ea032-bb93-45ff-bd47-b3f46752b1af · outbound

This paper cites Momentum contrast for unsupervised visual representation learning.

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception Momentum contrast for unsupervised visual representation learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:55.248426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:58:55.076134Z digest=sha256:1bd979e9533b14edbc780660038765a519bdfba5aed54d300f0056e75217145c

Observation 51d9fb25-86d2-4071-89c9-fc4618d8b646 · outbound

This paper cites Masked autoencoders are scalable vision learners.

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception Masked autoencoders are scalable vision learners

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:55.235878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:58:55.079896Z digest=sha256:cdbaf44f6c56117571024580761f9593f32906b221ca4fd865226e0627544873

Observation a2099768-c268-4881-98c8-8ccc683a90d9 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception OpenVLA: An Open-Source Vision-Language-Action Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T17:58:55.083303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:58:55.083303Z digest=sha256:e2c9ec9ce9786b4719e90a1a69722c85449e841c6752dec463a2a15f7dd66b01

Observation dde729be-0cd1-4fc8-8dba-540d12f3b92f · outbound

This paper cites Microsoft coco: Common objects in context.

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception Microsoft coco: Common objects in context

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:55.224981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:58:55.086887Z digest=sha256:2a3428a9327baccd64e59acc3ebfa0f7815fbace8c5553ffcd79a2e2cf5c5de2

Observation f5658ec1-fdaf-4ff5-8ede-dde97edde8d1 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception DINOv2: Learning Robust Visual Features without Supervision

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:58:55.090266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:58:55.090266Z digest=sha256:98c7a7c51ec41b5dec317b08de6bc03affa51ce9c30b15481db94c8378754df9

Observation 5364f2c5-c786-426f-afe3-8582d461064b · outbound

This paper cites N., Carreira, J., Asano, Y.

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception N., Carreira, J., Asano, Y

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:55.213746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:58:55.094026Z digest=sha256:4961c787e0f8deb9686b9659cd73e1bd727febd2edcc943990bc4d84a802002f

Observation 009bef87-049d-404a-a877-e6bf31408bc1 · outbound

This paper cites PooDLe: Pooled and dense self-supervised learning from naturalistic videos.

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception PooDLe: Pooled and dense self-supervised learning from naturalistic videos

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T17:58:55.097437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:58:55.097437Z digest=sha256:db14fa76df74f18a5cccc7de21a9f62702c9b6aadbeb8412459c897a0a4ace83

Observation 998ef1dc-dc3c-43e8-8fb2-024b66cfdcc4 · outbound

This paper cites Masked feature prediction for self-supervised visual pre-training.

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception Masked feature prediction for self-supervised visual pre-training

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:55.202428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:58:55.101104Z digest=sha256:6d9f3b26761a46c70c3e3b80667f7ced57db58ad37013f78a266f4597785f223

Observation 0665c5d8-5310-46c3-8930-89f9b8937e23 · outbound

This paper cites Scene parsing through ade20k dataset.

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception Scene parsing through ade20k dataset

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:58:55.191344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:58:55.104774Z digest=sha256:d18e30195fbb0868a7fd4c95bf84736373b742675f5a2a4629be7f0ea684387d

Observation 08047d75-6a29-49d1-9617-8335d43e76fd · outbound

This paper cites an unresolved cited work.

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:58:55.178854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:58:55.108117Z digest=sha256:1f1e72c9721f3ddfd9c3f4aa9ce11f04d8ffefd6f79497824c9e921ceeb5cd32

Pith citing papers

No inbound Pith citation observations are available.