Pith. sign in

Paper Citation Record · LEDGER

Scaling Language-Free Visual Representation Learning

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2504.01017.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.01017 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:45:57.113386Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:50:09.839784Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 07e917f8-9772-4812-ae05-2018d1f6e6d9 · inbound

Tables Guide Vision: Learning to See the Heart through Tabular Data cites this paper.

Tables Guide Vision: Learning to See the Heart through Tabular Data Scaling Language-Free Visual Representation Learning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:12:14.652248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:08:34.301889Z digest=sha256:7e82344db900b3add68ce462c74703c8dbf91d7d929e0d2e1f36d1093dbe96fa

Observation a97fb013-de8e-4c74-bf78-a16b445bff02 · inbound

Perception Encoder: The best visual embeddings are not at the output of the network cites this paper.

Perception Encoder: The best visual embeddings are not at the output of the network Scaling Language-Free Visual Representation Learning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:21:15.749167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T22:21:15.681336Z digest=sha256:d74e51dc31f7f7d2ed3191009f84d9a2068de9859c8a9238f2ef84489643bb10

Observation 97c2315a-c5d2-4031-b0f0-77c0c84dc059 · inbound

LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language Models cites this paper.

LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language Models Scaling Language-Free Visual Representation Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:51:37.647894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T13:47:51.436258Z digest=sha256:9ef7eb1c115c2a3e018a6e50fa10c5098d05613eda299df171993985c0228457

Observation d0e7881d-5fcd-4737-b05d-b6f62186d298 · inbound

Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets cites this paper.

Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets Scaling Language-Free Visual Representation Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:57.113386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:45:57.113386Z digest=sha256:282d88f2679176808b5e863033f43133fd30b5840eb3947f60e26868f18d99ff

Observation 0b4fc876-9190-4d61-8583-94c60ded441c · inbound

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning cites this paper.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Scaling Language-Free Visual Representation Learning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:50.790701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:11f0597dc4278b8583223f022522c782de50dc673ca6b08f8e733486eb5b9658

Observation 12fd4dc9-a792-49b5-9f2c-e02749c8c508 · inbound

Asymmetric Dual Self-Distillation for 3D Self-Supervised Representation Learning cites this paper.

Asymmetric Dual Self-Distillation for 3D Self-Supervised Representation Learning Scaling Language-Free Visual Representation Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:26:21.343283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:26:21.343283Z digest=sha256:c579172fc25d35f05352862033d528245e066d6b14a091610ce77de347cfd164

Observation ece4340a-d08d-4993-894a-dab7d2cd9aaa · inbound

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement cites this paper.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Scaling Language-Free Visual Representation Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:03.943429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:03.943429Z digest=sha256:e017a7800fd6efdfefb89ec360e4d14d9a713a0b951f10d38b840fbd1c98310b

Observation f93f5d80-5531-496e-937e-b83962ec05e8 · inbound

Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning cites this paper.

Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning Scaling Language-Free Visual Representation Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:42:01.349832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-19T03:39:52.969100Z digest=sha256:fb1906ce273261e0eeec61e2bf974f488c158178391dc3ce91ced3a059a2da9c

Observation d444a0ac-efec-4fa8-bf70-ba13f8d9b634 · inbound

Meta CLIP 2: A Worldwide Scaling Recipe cites this paper.

Meta CLIP 2: A Worldwide Scaling Recipe Scaling Language-Free Visual Representation Learning

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:22.283340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:22.283340Z digest=sha256:1654f127949f7c3286c4b2fcb3146f0f60e6f8c26259712d9090a28d618cd50a

Observation 1d1c76e8-be32-4cae-a3f1-8a3a68c1e022 · inbound

Towards Cellular-Scale Interpretability in Pathology Foundation Models for Biomarker Assessment cites this paper.

Towards Cellular-Scale Interpretability in Pathology Foundation Models for Biomarker Assessment Scaling Language-Free Visual Representation Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T23:34:16.010776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:34:16.010776Z digest=sha256:c9600fd1c89d76d153688c5bcbfc5d60f9ebcb6c8c93e68e33c331d57343e237

Observation 064e4649-4214-4fa5-b937-37fbf77bdcf5 · inbound

LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics cites this paper.

LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics Scaling Language-Free Visual Representation Learning

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:23:00.009580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-16T07:22:59.854042Z digest=sha256:ed9bca01846c197a8edbe1ca1710ecccb288259a8ffd1246974ee99ae6f53f67

Observation 04a4f02d-9f12-4722-859c-82e42fd177bb · inbound

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning cites this paper.

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning Scaling Language-Free Visual Representation Learning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:33:17.193145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T20:28:30.864143Z digest=sha256:2913e3650bf88f051203c92898d6e8b9911f76097cd2f99ce2edbd7ec07b2698

Observation 96813c63-c740-421e-a79a-aafe56e4e5a7 · inbound

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment cites this paper.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Scaling Language-Free Visual Representation Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:31:01.874816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:e793d9e4596182676f91bf0b85039838e79812cfa90445e8023adf4f70072ba0

Observation 0168c11f-bc7a-4967-856c-4ab7f306403d · inbound

CharTide: Data-Centric Chart-to-Code Generation via Tri-Perspective Tuning and Inquiry-Driven Evolution cites this paper.

CharTide: Data-Centric Chart-to-Code Generation via Tri-Perspective Tuning and Inquiry-Driven Evolution Scaling Language-Free Visual Representation Learning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:06:10.267010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T12:40:34.424341Z digest=sha256:c344afefa6a06febc8672c46bcbbcdcf29c8b0132698699378de1e3dcffa4d3e

Observation f6dab09a-9cb6-4d6c-a303-eaed4ebe78f6 · inbound

CharTide: Data-Centric Chart-to-Code Generation via Tri-Perspective Tuning and Inquiry-Driven Evolution cites this paper.

CharTide: Data-Centric Chart-to-Code Generation via Tri-Perspective Tuning and Inquiry-Driven Evolution Scaling Language-Free Visual Representation Learning

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:50:09.842171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-04T19:48:47.086335Z digest=sha256:823803df385ca8d31fedc43cf16793bc974d0c7d41921e23d480b296691eadbe

Observation 57d264bc-0687-45ef-8e9a-69f6ca796bb3 · inbound

Information theoretic underpinning of self-supervised learning by clustering cites this paper.

Information theoretic underpinning of self-supervised learning by clustering Scaling Language-Free Visual Representation Learning

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:22:23.133059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T06:22:03.439595Z digest=sha256:726411c554215104214318660ffdfe677eeb84915bae877681e9a5103f2f6ae2

Observation cc42d5e1-d764-45a6-a921-f1216f87d2ab · inbound

Improved Baselines with Representation Autoencoders cites this paper.

Improved Baselines with Representation Autoencoders Scaling Language-Free Visual Representation Learning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:43:15.304880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T11:40:14.358108Z digest=sha256:ed62fd301d33218fe01df4e9db65435e440f3d33c90c9cbf2104ec9834c91ba3

Observation 4325759d-4641-4eed-b4b4-7a6ffd94d2b8 · inbound

MLT-Dedup: Efficient Large-Scale Online Video Deduplication via Multi-Level Representations and Spatial-Temporal Matching cites this paper.

MLT-Dedup: Efficient Large-Scale Online Video Deduplication via Multi-Level Representations and Spatial-Temporal Matching Scaling Language-Free Visual Representation Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:58:03.547412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T09:43:22.054789Z digest=sha256:b42868c81ff2edff97e6e54cf0d11295b96305eff73f11e24fceb5325dda77bc

Observation 8098426c-2301-4c56-99e5-e8619e0d0ab6 · inbound

DiffusionBench: On Holistic Evaluation of Diffusion Transformers cites this paper.

DiffusionBench: On Holistic Evaluation of Diffusion Transformers Scaling Language-Free Visual Representation Learning

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:59:58.071266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T00:06:11.951205Z digest=sha256:4330989de02b757693b3855ab1939b3006e670ebbd83bc31e9e7fa6f3c831570

Observation 4df71557-4ab1-4660-b94a-8d38f3473849 · inbound

Patch Policy: Efficient Embodied Control via Dense Visual Representations cites this paper.

Patch Policy: Efficient Embodied Control via Dense Visual Representations Scaling Language-Free Visual Representation Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T15:40:49.951072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:40:49.951072Z digest=sha256:9954e2b3951235a38ee9c6c86bbe1b51e2a33e398ad687092c3a04d53dd6469d