Pith. sign in

Paper Citation Record · LEDGER

data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2202.03555.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2202.03555 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:58:59.991539Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

241
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3b9d6c00-ee68-46ca-abf0-337f9c4428c8 · inbound

Revisiting Feature Prediction for Learning Visual Representations from Video cites this paper.

Revisiting Feature Prediction for Learning Visual Representations from Video data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 223

Resolution
verified exact
arxiv_id, observed 2026-05-12T12:40:23.837332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-12T12:40:23.709098Z digest=sha256:3439564eb97fbb053efc8897cb0779527f21bb56a3eaedff68b523a824ce6265

Observation 07e62f32-e74b-472f-8214-0da90b355604 · inbound

Everything is a Video: Unifying Modalities through Next-Frame Prediction cites this paper.

Everything is a Video: Unifying Modalities through Next-Frame Prediction data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T19:58:59.991539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:58:59.991539Z digest=sha256:539bf79882d6a0935172357931405f6058612bce86ad18761cf75fb84dc8ad4f

Observation 1ed9d5b9-cffd-4c56-bf76-ec497633f12e · inbound

Wearable Accelerometer Foundation Models for Health via Knowledge Distillation cites this paper.

Wearable Accelerometer Foundation Models for Health via Knowledge Distillation data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T15:12:27.853236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:12:27.853236Z digest=sha256:21b003173793c0b5aa1050cb83b5b0c762f423bdd85947f7acdd802fe0c8fc7f

Observation c330543f-40bb-4fc7-9601-a6ce42b4e49f · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.125368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.125368Z digest=sha256:e388f368a41e66d779b9730c66be87229681e5911263519732648999879d9a8d

Observation d974947f-55bc-4f6f-86e7-029b7ffc68d4 · inbound

Fine Tuning without Catastrophic Forgetting via Selective Low Rank Adaptation cites this paper.

Fine Tuning without Catastrophic Forgetting via Selective Low Rank Adaptation data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:39.004449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:39.004449Z digest=sha256:edcac371fab4bf72397d9a07222dabc43e2c0baa1c38783e66a9c89085fc0ca8

Observation 4615b1ab-4aa0-48a8-b010-a13417829438 · inbound

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning cites this paper.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.244140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.244140Z digest=sha256:3b5e5e92d2823ad48def8cf3a9c648c43d0f6efe8acf97d6d66948e74210b4be

Observation 0905b55d-e95d-4f7f-be9b-717b67a9e59b · inbound

ZeroSyl: Simple Zero-Resource Syllable Tokenization for Spoken Language Modeling cites this paper.

ZeroSyl: Simple Zero-Resource Syllable Tokenization for Spoken Language Modeling data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T22:51:44.257192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:51:44.257192Z digest=sha256:64abdcf3cf99ac7c3b9d259349c51f1df06d41578b8af2391b4c815e092d9e89

Observation 91668eec-41fe-4d9f-99af-25a8aa0b16dc · inbound

A SUPERB-Style Benchmark of Self-Supervised Speech Models for Audio Deepfake Detection cites this paper.

A SUPERB-Style Benchmark of Self-Supervised Speech Models for Audio Deepfake Detection data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:26:22.168642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T17:25:27.137423Z digest=sha256:b9214d0bf3203f60e9ec2e32a27a7d3e09205f5c51f917eeef0aeb6ccb56f10e

Observation 49ee6506-dc9e-485d-9878-176673bd5659 · inbound

Weather-Robust Cross-View Geo-Localization via Prototype-Based Semantic Part Discovery cites this paper.

Weather-Robust Cross-View Geo-Localization via Prototype-Based Semantic Part Discovery data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:32:01.852772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-13T01:30:53.240885Z digest=sha256:f0386c4a2728819d2cbc404a2195a1fc0dcbc463d5286e47c4e64a9f7ea352c1

Observation 69a43e48-2bd8-4608-ba65-c99283d46e4a · inbound

Weather-Robust Cross-View Geo-Localization via Prototype-Based Semantic Part Discovery cites this paper.

Weather-Robust Cross-View Geo-Localization via Prototype-Based Semantic Part Discovery data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:09:11.869978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-20T23:05:06.850388Z digest=sha256:94a096398f3d90c8c7a032310eb15a5c5b133917e877c7b936a7a135e40f1cac

Observation 2d7ab24c-0bdb-4767-8ce2-a343e2e7ed39 · inbound

Representation Without Reward: A JEPA Audit for LLM Fine-Tuning cites this paper.

Representation Without Reward: A JEPA Audit for LLM Fine-Tuning data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:32:39.667284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-19T16:27:51.512806Z digest=sha256:07bd778ba9041307c47ccc0d4a68bdf5af3fab76c7fa5a8789a412876d7db1e4

Observation 3027b3ee-52dd-4804-833e-cbfd5013ad61 · inbound

Factors affecting ASR performance: A study using state of the art ASR models in Indic Languages cites this paper.

Factors affecting ASR performance: A study using state of the art ASR models in Indic Languages data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:37:35.494984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T15:10:03.282212Z digest=sha256:beffc11d77b8952d7288ddbef59e9e8eeacd31acef6b1709415985fead8e9b60

Observation 04035132-ae07-4059-8396-df8915f99333 · inbound

When to Align, When to Predict: A Phase Diagram for Multimodal Learning cites this paper.

When to Align, When to Predict: A Phase Diagram for Multimodal Learning data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:07:37.046489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T14:08:33.542700Z digest=sha256:eb3028e294ec4c643163d771b85450866fdfa850bee055064f435338b0ebc686

Observation a8e2703b-ca0e-497f-addb-d19fb1c11c46 · inbound

When to Align, When to Predict: A Phase Diagram for Multimodal Learning cites this paper.

When to Align, When to Predict: A Phase Diagram for Multimodal Learning data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T14:10:57.804524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T14:08:33.542700Z digest=sha256:ae86a8fbda52f3b159b1d9ef750085dc3a4db8ab3f8c6734b47951d2c690ce66

Observation 2e51db97-8530-4b54-a22f-ad767b04ccb3 · inbound

Backdoor Attacks on Speech Emotion Recognition via TTS-Generated Poisoning cites this paper.

Backdoor Attacks on Speech Emotion Recognition via TTS-Generated Poisoning data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:29:38.679311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T13:27:19.810597Z digest=sha256:fb8734b86a1121500f6ca127d41702972d2e1f5303d100cf6c2b1a7c94200eac

Observation 4ebe04c0-2dfd-42ed-a076-7dd7d30d21e2 · inbound

End-to-End Voice Intent Recognition for Spontaneous Human-Drone Interaction with Naive Users cites this paper.

End-to-End Voice Intent Recognition for Spontaneous Human-Drone Interaction with Naive Users data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:29:38.184343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T13:30:12.101045Z digest=sha256:c6d47ff83fb053af73fafdb6718be1adcf126fd528095838832bdf5e752045c9

Observation 6523e947-629d-4ac5-8261-5ea0d3a08301 · inbound

MJEPA: A Simple and Scalable Joint-Embedding Predictive Architecture for Audio-Visual Learning cites this paper.

MJEPA: A Simple and Scalable Joint-Embedding Predictive Architecture for Audio-Visual Learning data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:09:58.611899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-25T23:57:19.880950Z digest=sha256:2a292c16d0c942e304c2f3443b1bd6c60233c60715d6d320e269d3901985c548

Observation 02316223-29ea-4cdb-b820-4b5528bc28c9 · inbound

AGE: Adaptive-masking for Graph Embedding in Graph Retrieval-Augmented Generation cites this paper.

AGE: Adaptive-masking for Graph Embedding in Graph Retrieval-Augmented Generation data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T18:47:17.215358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-02T18:01:48.915707Z digest=sha256:256abfb014d59b88c693843cf4c24d2c828922cec61152a75cad3d0dc90319ef

Observation 6ac91d1f-de75-4add-9c2b-70b0a85ab1ef · inbound

Hierarchical Self-Supervised Representation Learning Framework for Multivariate Time Series Grounded in ECG Analysis cites this paper.

Hierarchical Self-Supervised Representation Learning Framework for Multivariate Time Series Grounded in ECG Analysis data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:27:04.540813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-02T15:26:51.956476Z digest=sha256:3ddc90e2236116afbeae4067754e4c199f682a8c644aff5903e7aef8a2e61dae

Observation d01a470f-ef3d-4fcf-8254-64077cd7dde9 · inbound

STST-JEPA: Shallow-Target Spatio-Temporal Joint Embedding Prediction Architecture For EEG Self-Supervised Learning cites this paper.

STST-JEPA: Shallow-Target Spatio-Temporal Joint Embedding Prediction Architecture For EEG Self-Supervised Learning data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:07:42.196659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T01:07:21.002787Z digest=sha256:37298318ed382f2d3cd7a5f774de41cb0b8ddd9db15a89ecfea4103bff1daf87

Observation 3cdf2041-ac99-4872-9fe7-3222c51f3458 · inbound

The Importance of Encoder Choice:A Tabular-Image Study cites this paper.

The Importance of Encoder Choice:A Tabular-Image Study data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 253

Resolution
verified exact
local_arxiv, observed 2026-07-10T19:07:35.063435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-07-10T19:03:32.353393Z digest=sha256:fcd1fbfed1ca30d7ae30f4c9028bfdf775a7282b421c5ded69a1d96a9b4f6f63

Observation c480e7be-a636-4039-a5ba-db5862730999 · inbound

Kepler-Encoder-v0.1: Towards a Multimodal Embedding Model for Robots cites this paper.

Kepler-Encoder-v0.1: Towards a Multimodal Embedding Model for Robots data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T05:01:23.102997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:01:23.102997Z digest=sha256:a3f9c6292176672ed204e23160c4e2c1bc68d28d8bbebfa6e29a08bfe66b6184

Observation 28a05d06-b226-47f8-b4c1-7f63ca8827d2 · inbound

Toward Annotation-Efficient Continuous Emotion Arousal Quantification via Group-Level EEG Dynamic Neural Synchrony cites this paper.

Toward Annotation-Efficient Continuous Emotion Arousal Quantification via Group-Level EEG Dynamic Neural Synchrony data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-31T14:47:01.466428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T14:47:01.466428Z digest=sha256:3703ff6909d7c54668ac4ac5da2175236b31437f4ea59ed40fb8844b4422fb98