Pith. sign in

Paper Citation Record · LEDGER

data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2202.03555.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2202.03555 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:58:59.991539Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

241
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3b9d6c00-ee68-46ca-abf0-337f9c4428c8 · inbound

Revisiting Feature Prediction for Learning Visual Representations from Video cites this paper.

Revisiting Feature Prediction for Learning Visual Representations from Video data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 223

Resolution
verified exact
arxiv_id, observed 2026-05-12T12:40:23.837332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-12T12:40:23.709098Z digest=sha256:b61611d971936e212829e204e1cb8ec19f2dbeca9135d0b44e821cb1dc8f403d

Observation 07e62f32-e74b-472f-8214-0da90b355604 · inbound

Everything is a Video: Unifying Modalities through Next-Frame Prediction cites this paper.

Everything is a Video: Unifying Modalities through Next-Frame Prediction data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T19:58:59.991539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:58:59.991539Z digest=sha256:539bf79882d6a0935172357931405f6058612bce86ad18761cf75fb84dc8ad4f

Observation 1ed9d5b9-cffd-4c56-bf76-ec497633f12e · inbound

Wearable Accelerometer Foundation Models for Health via Knowledge Distillation cites this paper.

Wearable Accelerometer Foundation Models for Health via Knowledge Distillation data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T15:12:27.853236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:12:27.853236Z digest=sha256:21b003173793c0b5aa1050cb83b5b0c762f423bdd85947f7acdd802fe0c8fc7f

Observation c330543f-40bb-4fc7-9601-a6ce42b4e49f · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.125368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.125368Z digest=sha256:e388f368a41e66d779b9730c66be87229681e5911263519732648999879d9a8d

Observation d974947f-55bc-4f6f-86e7-029b7ffc68d4 · inbound

Fine Tuning without Catastrophic Forgetting via Selective Low Rank Adaptation cites this paper.

Fine Tuning without Catastrophic Forgetting via Selective Low Rank Adaptation data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:39.004449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:39.004449Z digest=sha256:edcac371fab4bf72397d9a07222dabc43e2c0baa1c38783e66a9c89085fc0ca8

Observation 4615b1ab-4aa0-48a8-b010-a13417829438 · inbound

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning cites this paper.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.244140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.244140Z digest=sha256:3b5e5e92d2823ad48def8cf3a9c648c43d0f6efe8acf97d6d66948e74210b4be

Observation 0905b55d-e95d-4f7f-be9b-717b67a9e59b · inbound

ZeroSyl: Simple Zero-Resource Syllable Tokenization for Spoken Language Modeling cites this paper.

ZeroSyl: Simple Zero-Resource Syllable Tokenization for Spoken Language Modeling data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T22:51:44.257192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:51:44.257192Z digest=sha256:64abdcf3cf99ac7c3b9d259349c51f1df06d41578b8af2391b4c815e092d9e89

Observation 91668eec-41fe-4d9f-99af-25a8aa0b16dc · inbound

A SUPERB-Style Benchmark of Self-Supervised Speech Models for Audio Deepfake Detection cites this paper.

A SUPERB-Style Benchmark of Self-Supervised Speech Models for Audio Deepfake Detection data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:26:22.168642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T17:25:27.137423Z digest=sha256:c223920e081671008e92acff1b7d4c8e99a102715799cf6e4c78ea0b96e57340

Observation 49ee6506-dc9e-485d-9878-176673bd5659 · inbound

Weather-Robust Cross-View Geo-Localization via Prototype-Based Semantic Part Discovery cites this paper.

Weather-Robust Cross-View Geo-Localization via Prototype-Based Semantic Part Discovery data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:32:01.852772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-13T01:30:53.240885Z digest=sha256:3d93fe0eea8d18d60d086c6f1e19be3a53808cbbde84e175d7a799f5fd9c18a3

Observation 69a43e48-2bd8-4608-ba65-c99283d46e4a · inbound

Weather-Robust Cross-View Geo-Localization via Prototype-Based Semantic Part Discovery cites this paper.

Weather-Robust Cross-View Geo-Localization via Prototype-Based Semantic Part Discovery data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:09:11.869978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-20T23:05:06.850388Z digest=sha256:cc385fa0d1f9b9ee489f86ee04c62783510cefa1a584962664bae453f4a6d749

Observation 2d7ab24c-0bdb-4767-8ce2-a343e2e7ed39 · inbound

Representation Without Reward: A JEPA Audit for LLM Fine-Tuning cites this paper.

Representation Without Reward: A JEPA Audit for LLM Fine-Tuning data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:32:39.667284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T16:27:51.512806Z digest=sha256:dde1051a7b08c4bff3d5b5959c2526a995b633efd92cbbddccf4d3ae95c7f317

Observation 3027b3ee-52dd-4804-833e-cbfd5013ad61 · inbound

Factors affecting ASR performance: A study using state of the art ASR models in Indic Languages cites this paper.

Factors affecting ASR performance: A study using state of the art ASR models in Indic Languages data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:37:35.494984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T15:10:03.282212Z digest=sha256:9f2413e9790e5343f2aaac91d368555d68c7dd9b578304dc7f3db48107ccb596

Observation 04035132-ae07-4059-8396-df8915f99333 · inbound

When to Align, When to Predict: A Phase Diagram for Multimodal Learning cites this paper.

When to Align, When to Predict: A Phase Diagram for Multimodal Learning data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:07:37.046489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T14:08:33.542700Z digest=sha256:332fdb1e323536c3d48f26478bd6f1a5382bc3649210c6f5204366f40df67b90

Observation a8e2703b-ca0e-497f-addb-d19fb1c11c46 · inbound

When to Align, When to Predict: A Phase Diagram for Multimodal Learning cites this paper.

When to Align, When to Predict: A Phase Diagram for Multimodal Learning data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T14:10:57.804524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T14:08:33.542700Z digest=sha256:9de2f683be5bef19edf8dc01645835fd7c755f37b07ecc74fb3a0ebf4c6c8922

Observation 2e51db97-8530-4b54-a22f-ad767b04ccb3 · inbound

Backdoor Attacks on Speech Emotion Recognition via TTS-Generated Poisoning cites this paper.

Backdoor Attacks on Speech Emotion Recognition via TTS-Generated Poisoning data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:29:38.679311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T13:27:19.810597Z digest=sha256:85ab0de5fdcb3105679ad4d2559380f344a27e8ba67ab8aed9114c7526a556f5

Observation 4ebe04c0-2dfd-42ed-a076-7dd7d30d21e2 · inbound

End-to-End Voice Intent Recognition for Spontaneous Human-Drone Interaction with Naive Users cites this paper.

End-to-End Voice Intent Recognition for Spontaneous Human-Drone Interaction with Naive Users data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:29:38.184343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T13:30:12.101045Z digest=sha256:10b007a773dbe5cb4bb3772a75b1bc7ef14932144dbb84dc579004bfdbc91028

Observation 6523e947-629d-4ac5-8261-5ea0d3a08301 · inbound

MJEPA: A Simple and Scalable Joint-Embedding Predictive Architecture for Audio-Visual Learning cites this paper.

MJEPA: A Simple and Scalable Joint-Embedding Predictive Architecture for Audio-Visual Learning data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:09:58.611899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-25T23:57:19.880950Z digest=sha256:b1cf69c231c212fbbc41c1cb216e9d6cf451054e035c462ce9b94e3537d4a5b4

Observation 02316223-29ea-4cdb-b820-4b5528bc28c9 · inbound

AGE: Adaptive-masking for Graph Embedding in Graph Retrieval-Augmented Generation cites this paper.

AGE: Adaptive-masking for Graph Embedding in Graph Retrieval-Augmented Generation data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T18:47:17.215358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-02T18:01:48.915707Z digest=sha256:d4e90687ccc15deea927628da861c48411535408640ac59e6c002fa00664624c

Observation 6ac91d1f-de75-4add-9c2b-70b0a85ab1ef · inbound

Hierarchical Self-Supervised Representation Learning Framework for Multivariate Time Series Grounded in ECG Analysis cites this paper.

Hierarchical Self-Supervised Representation Learning Framework for Multivariate Time Series Grounded in ECG Analysis data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:27:04.540813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-02T15:26:51.956476Z digest=sha256:8b0be79bb68d10d72675924fed93bca47c319b41c3afca967f073482e5858c60

Observation d01a470f-ef3d-4fcf-8254-64077cd7dde9 · inbound

STST-JEPA: Shallow-Target Spatio-Temporal Joint Embedding Prediction Architecture For EEG Self-Supervised Learning cites this paper.

STST-JEPA: Shallow-Target Spatio-Temporal Joint Embedding Prediction Architecture For EEG Self-Supervised Learning data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:07:42.196659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-11T01:07:21.002787Z digest=sha256:9f16a6e8c9089f634ad7fc1c5f22d41551f1f4698dc1cf0a94d0f6bcfd7ce35d

Observation 3cdf2041-ac99-4872-9fe7-3222c51f3458 · inbound

The Importance of Encoder Choice:A Tabular-Image Study cites this paper.

The Importance of Encoder Choice:A Tabular-Image Study data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 253

Resolution
verified exact
local_arxiv, observed 2026-07-10T19:07:35.063435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-10T19:03:32.353393Z digest=sha256:139b619a7bf5da6caa7b7af06a209158c6119be9521c903d809b1914bd95de3b

Observation c480e7be-a636-4039-a5ba-db5862730999 · inbound

Kepler-Encoder-v0.1: Towards a Multimodal Embedding Model for Robots cites this paper.

Kepler-Encoder-v0.1: Towards a Multimodal Embedding Model for Robots data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T05:01:23.102997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:01:23.102997Z digest=sha256:a3f9c6292176672ed204e23160c4e2c1bc68d28d8bbebfa6e29a08bfe66b6184

Observation 28a05d06-b226-47f8-b4c1-7f63ca8827d2 · inbound

Toward Annotation-Efficient Continuous Emotion Arousal Quantification via Group-Level EEG Dynamic Neural Synchrony cites this paper.

Toward Annotation-Efficient Continuous Emotion Arousal Quantification via Group-Level EEG Dynamic Neural Synchrony data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-31T14:47:01.466428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T14:47:01.466428Z digest=sha256:3703ff6909d7c54668ac4ac5da2175236b31437f4ea59ed40fb8844b4422fb98