Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2202.03555.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T19:58:59.991539Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
241
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 3b9d6c00-ee68-46ca-abf0-337f9c4428c8 · inbound
Revisiting Feature Prediction for Learning Visual Representations from Video data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
Reference 223
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 07e62f32-e74b-472f-8214-0da90b355604 · inbound
Everything is a Video: Unifying Modalities through Next-Frame Prediction data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ed9d5b9-cffd-4c56-bf76-ec497633f12e · inbound
Wearable Accelerometer Foundation Models for Health via Knowledge Distillation data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c330543f-40bb-4fc7-9601-a6ce42b4e49f · inbound
Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d974947f-55bc-4f6f-86e7-029b7ffc68d4 · inbound
Fine Tuning without Catastrophic Forgetting via Selective Low Rank Adaptation data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4615b1ab-4aa0-48a8-b010-a13417829438 · inbound
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0905b55d-e95d-4f7f-be9b-717b67a9e59b · inbound
ZeroSyl: Simple Zero-Resource Syllable Tokenization for Spoken Language Modeling data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91668eec-41fe-4d9f-99af-25a8aa0b16dc · inbound
A SUPERB-Style Benchmark of Self-Supervised Speech Models for Audio Deepfake Detection data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 49ee6506-dc9e-485d-9878-176673bd5659 · inbound
Weather-Robust Cross-View Geo-Localization via Prototype-Based Semantic Part Discovery data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 69a43e48-2bd8-4608-ba65-c99283d46e4a · inbound
Weather-Robust Cross-View Geo-Localization via Prototype-Based Semantic Part Discovery data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2d7ab24c-0bdb-4767-8ce2-a343e2e7ed39 · inbound
Representation Without Reward: A JEPA Audit for LLM Fine-Tuning data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3027b3ee-52dd-4804-833e-cbfd5013ad61 · inbound
Factors affecting ASR performance: A study using state of the art ASR models in Indic Languages data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 04035132-ae07-4059-8396-df8915f99333 · inbound
When to Align, When to Predict: A Phase Diagram for Multimodal Learning data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a8e2703b-ca0e-497f-addb-d19fb1c11c46 · inbound
When to Align, When to Predict: A Phase Diagram for Multimodal Learning data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2e51db97-8530-4b54-a22f-ad767b04ccb3 · inbound
Backdoor Attacks on Speech Emotion Recognition via TTS-Generated Poisoning data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4ebe04c0-2dfd-42ed-a076-7dd7d30d21e2 · inbound
End-to-End Voice Intent Recognition for Spontaneous Human-Drone Interaction with Naive Users data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6523e947-629d-4ac5-8261-5ea0d3a08301 · inbound
MJEPA: A Simple and Scalable Joint-Embedding Predictive Architecture for Audio-Visual Learning data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 02316223-29ea-4cdb-b820-4b5528bc28c9 · inbound
AGE: Adaptive-masking for Graph Embedding in Graph Retrieval-Augmented Generation data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6ac91d1f-de75-4add-9c2b-70b0a85ab1ef · inbound
Hierarchical Self-Supervised Representation Learning Framework for Multivariate Time Series Grounded in ECG Analysis data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d01a470f-ef3d-4fcf-8254-64077cd7dde9 · inbound
STST-JEPA: Shallow-Target Spatio-Temporal Joint Embedding Prediction Architecture For EEG Self-Supervised Learning data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3cdf2041-ac99-4872-9fe7-3222c51f3458 · inbound
The Importance of Encoder Choice:A Tabular-Image Study data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
Reference 253
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c480e7be-a636-4039-a5ba-db5862730999 · inbound
Kepler-Encoder-v0.1: Towards a Multimodal Embedding Model for Robots data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28a05d06-b226-47f8-b4c1-7f63ca8827d2 · inbound
Toward Annotation-Efficient Continuous Emotion Arousal Quantification via Group-Level EEG Dynamic Neural Synchrony data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.