Pith. sign in

Paper Citation Record · LEDGER

Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and Language

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 2 inbound Pith citation observations for arXiv:2212.07525.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2212.07525 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 2 of 2 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:15:43.893409Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T12:40:23.831580Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4470f0eb-37ac-4c0a-9e19-65e348c7e71f · inbound

Revisiting Feature Prediction for Learning Visual Representations from Video cites this paper.

Revisiting Feature Prediction for Learning Visual Representations from Video Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and Language

Reference 222

Resolution
verified exact
arxiv_id, observed 2026-05-12T12:40:23.833624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T12:40:23.709098Z digest=sha256:f9f5df7232ab2a5fd22e4597d6d5101a0a728bd72c89711c0df5ee707c36dce5

Observation ccf72dc3-1cb0-48a4-90b4-35f92071835b · inbound

HyPASE: Hyperbolic Geometry for Parameter-Efficient Speech Emotion Fine-Tuning Framework for Large Audio-Language Models cites this paper.

HyPASE: Hyperbolic Geometry for Parameter-Efficient Speech Emotion Fine-Tuning Framework for Large Audio-Language Models Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and Language

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:43.893409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:43.893409Z digest=sha256:1fedaab1aaf9f8642998141f6289fd09ee2f5e0fa543ab6d8fd998f4dbf993f3