Pith. sign in

Paper Citation Record · LEDGER

Scaling Language-Free Visual Representation Learning

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2504.01017.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.01017 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:45:57.113386Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:50:09.839784Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 07e917f8-9772-4812-ae05-2018d1f6e6d9 · inbound

Tables Guide Vision: Learning to See the Heart through Tabular Data cites this paper.

Tables Guide Vision: Learning to See the Heart through Tabular Data Scaling Language-Free Visual Representation Learning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:12:14.652248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T23:08:34.301889Z digest=sha256:2738d5dd64bd06670fb3ca5ba8e470fe5de2f1d7671bf735244635c03bedaf6a

Observation a97fb013-de8e-4c74-bf78-a16b445bff02 · inbound

Perception Encoder: The best visual embeddings are not at the output of the network cites this paper.

Perception Encoder: The best visual embeddings are not at the output of the network Scaling Language-Free Visual Representation Learning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:21:15.749167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T22:21:15.681336Z digest=sha256:7adc1fa705acfe9a1cbff785620f87751d81915ebf9fa5790e810870c8faad15

Observation 97c2315a-c5d2-4031-b0f0-77c0c84dc059 · inbound

LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language Models cites this paper.

LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language Models Scaling Language-Free Visual Representation Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:51:37.647894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T13:47:51.436258Z digest=sha256:d85db2f12baf11fb60c035d49e5857605f0ba5d1f1dd6e829ceda8e7c71ed828

Observation d0e7881d-5fcd-4737-b05d-b6f62186d298 · inbound

Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets cites this paper.

Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets Scaling Language-Free Visual Representation Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:57.113386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:45:57.113386Z digest=sha256:282d88f2679176808b5e863033f43133fd30b5840eb3947f60e26868f18d99ff

Observation 0b4fc876-9190-4d61-8583-94c60ded441c · inbound

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning cites this paper.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Scaling Language-Free Visual Representation Learning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:50.790701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:8e2fe795e9dbb495ee782cd9f6e470be3514a6316bf7dd31201efc750a555295

Observation 12fd4dc9-a792-49b5-9f2c-e02749c8c508 · inbound

Asymmetric Dual Self-Distillation for 3D Self-Supervised Representation Learning cites this paper.

Asymmetric Dual Self-Distillation for 3D Self-Supervised Representation Learning Scaling Language-Free Visual Representation Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:26:21.343283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:26:21.343283Z digest=sha256:c579172fc25d35f05352862033d528245e066d6b14a091610ce77de347cfd164

Observation ece4340a-d08d-4993-894a-dab7d2cd9aaa · inbound

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement cites this paper.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Scaling Language-Free Visual Representation Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:03.943429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:03.943429Z digest=sha256:56ea2d8976bcdad753c138609f1cc8f52443a0235c07bdcf267d9b5978f0c19e

Observation f93f5d80-5531-496e-937e-b83962ec05e8 · inbound

Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning cites this paper.

Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning Scaling Language-Free Visual Representation Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:42:01.349832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-19T03:39:52.969100Z digest=sha256:df0afacd186fb2306740012bf71ca3115f82eca68ecd36de2ffc12daea9440e9

Observation d444a0ac-efec-4fa8-bf70-ba13f8d9b634 · inbound

Meta CLIP 2: A Worldwide Scaling Recipe cites this paper.

Meta CLIP 2: A Worldwide Scaling Recipe Scaling Language-Free Visual Representation Learning

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:22.283340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:22.283340Z digest=sha256:a54a1894e09638eb8159e64a089e35e693169de626d38c5414ddf31cbb15a88a

Observation 1d1c76e8-be32-4cae-a3f1-8a3a68c1e022 · inbound

Towards Cellular-Scale Interpretability in Pathology Foundation Models for Biomarker Assessment cites this paper.

Towards Cellular-Scale Interpretability in Pathology Foundation Models for Biomarker Assessment Scaling Language-Free Visual Representation Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T23:34:16.010776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:34:16.010776Z digest=sha256:c9600fd1c89d76d153688c5bcbfc5d60f9ebcb6c8c93e68e33c331d57343e237

Observation 064e4649-4214-4fa5-b937-37fbf77bdcf5 · inbound

LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics cites this paper.

LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics Scaling Language-Free Visual Representation Learning

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:23:00.009580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-16T07:22:59.854042Z digest=sha256:53d04ab06da9c553ef801be8b7cc267595b77a5338b3ecdab2d107677b2e7f3e

Observation 04a4f02d-9f12-4722-859c-82e42fd177bb · inbound

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning cites this paper.

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning Scaling Language-Free Visual Representation Learning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:33:17.193145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T20:28:30.864143Z digest=sha256:421ab68ae7d96da226738c14cd229c08a878a754844e0edeb296bd9c06c30c56

Observation 96813c63-c740-421e-a79a-aafe56e4e5a7 · inbound

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment cites this paper.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Scaling Language-Free Visual Representation Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:31:01.874816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:cd176339ea3db69fd86bec3f6266bd58c1d387772ecfbb836cb323d541f6a0e7

Observation 0168c11f-bc7a-4967-856c-4ab7f306403d · inbound

CharTide: Data-Centric Chart-to-Code Generation via Tri-Perspective Tuning and Inquiry-Driven Evolution cites this paper.

CharTide: Data-Centric Chart-to-Code Generation via Tri-Perspective Tuning and Inquiry-Driven Evolution Scaling Language-Free Visual Representation Learning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:06:10.267010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T12:40:34.424341Z digest=sha256:3e4d20bad88d51037bd96ede52748d38922280b2ebe1ac1ea22daf9cd3a6aacb

Observation f6dab09a-9cb6-4d6c-a303-eaed4ebe78f6 · inbound

CharTide: Data-Centric Chart-to-Code Generation via Tri-Perspective Tuning and Inquiry-Driven Evolution cites this paper.

CharTide: Data-Centric Chart-to-Code Generation via Tri-Perspective Tuning and Inquiry-Driven Evolution Scaling Language-Free Visual Representation Learning

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:50:09.842171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-04T19:48:47.086335Z digest=sha256:90c090c56ce550c867f1459e081928aee6c651bb551cd2cb99e5eff9a727ca41

Observation 57d264bc-0687-45ef-8e9a-69f6ca796bb3 · inbound

Information theoretic underpinning of self-supervised learning by clustering cites this paper.

Information theoretic underpinning of self-supervised learning by clustering Scaling Language-Free Visual Representation Learning

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:22:23.133059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T06:22:03.439595Z digest=sha256:4bed611eddcf8340a7105d0c787101d0ac580ac671f5634d0686015fa81858f0

Observation cc42d5e1-d764-45a6-a921-f1216f87d2ab · inbound

Improved Baselines with Representation Autoencoders cites this paper.

Improved Baselines with Representation Autoencoders Scaling Language-Free Visual Representation Learning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:43:15.304880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T11:40:14.358108Z digest=sha256:972df9a0e244d57bf58d48cc4fb859a2432c02690afc02d8650d83537c0cfa88

Observation 4325759d-4641-4eed-b4b4-7a6ffd94d2b8 · inbound

MLT-Dedup: Efficient Large-Scale Online Video Deduplication via Multi-Level Representations and Spatial-Temporal Matching cites this paper.

MLT-Dedup: Efficient Large-Scale Online Video Deduplication via Multi-Level Representations and Spatial-Temporal Matching Scaling Language-Free Visual Representation Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:58:03.547412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T09:43:22.054789Z digest=sha256:8e51ad38d6aa4a27f15f282553e7b4dadf85c5f627fa87e14e4a1ea5e8f50f95

Observation 8098426c-2301-4c56-99e5-e8619e0d0ab6 · inbound

DiffusionBench: On Holistic Evaluation of Diffusion Transformers cites this paper.

DiffusionBench: On Holistic Evaluation of Diffusion Transformers Scaling Language-Free Visual Representation Learning

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:59:58.071266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T00:06:11.951205Z digest=sha256:764fc96b7604a63f902258c764ece2947fcbcbe63896d8fe3d445145acc34404

Observation 4df71557-4ab1-4660-b94a-8d38f3473849 · inbound

Patch Policy: Efficient Embodied Control via Dense Visual Representations cites this paper.

Patch Policy: Efficient Embodied Control via Dense Visual Representations Scaling Language-Free Visual Representation Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T15:40:49.951072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:40:49.951072Z digest=sha256:9954e2b3951235a38ee9c6c86bbe1b51e2a33e398ad687092c3a04d53dd6469d