Pith. sign in

Paper Citation Record · LEDGER

W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2108.06209.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2108.06209 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:32:40.662723Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T01:44:22.720612Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bfc93eb1-b0b3-4c43-b320-f767692700d2 · inbound

Different Speech Translation Models Encode and Translate Speaker Gender Differently cites this paper.

Different Speech Translation Models Encode and Translate Speaker Gender Differently W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:40.662723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:32:40.662723Z digest=sha256:97c134f22d907a73c085bc0543d372c2697d982fac23e423e2dafe8985b780a4

Observation 9516fb69-478a-4798-bfa4-95e9d6cda6bf · inbound

Representing Speech Through Autoregressive Prediction of Cochlear Tokens cites this paper.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.599394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.599394Z digest=sha256:040480604c8027db88cab2322ad3ec15fdb1be9c6ed4e3d16e222af28c0a5a99

Observation 0cfb0907-fbf6-4a81-bc92-e6632aba70a2 · inbound

DSA-Tokenizer: Disentangled Semantic-Acoustic Tokenization via Flow Matching-based Hierarchical Fusion cites this paper.

DSA-Tokenizer: Disentangled Semantic-Acoustic Tokenization via Flow Matching-based Hierarchical Fusion W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T10:44:31.283961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:44:31.283961Z digest=sha256:a3b264bdbff5f7e3eed57b3fe5f79dc2e5bbb236b50df5a6e31231f7c4c2bcc0

Observation 4a60ab31-8328-404a-8836-00cfa8260a5a · inbound

Musical Attention Transformer: Music Generation Using a Music-Specific Attention Model cites this paper.

Musical Attention Transformer: Music Generation Using a Music-Specific Attention Model W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:44:22.724239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T01:44:20.282896Z digest=sha256:3e2dd334a8227b21e2debf40a7797f1f0cd459d158722fae05b8deddfaf598e8

Observation a6c55c28-7e24-44f6-b799-e29e7ae88445 · inbound

DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages cites this paper.

DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T07:10:39.822866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:10:39.822866Z digest=sha256:3efd4f226b313fd81c14ec7fc8b1da32b77ece76f5db4dee0439b4a511f1e8f0