Pith. sign in

Paper Citation Record · LEDGER

Scaling up masked audio encoder learning for general audio classification

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2406.06992.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.06992 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:56:46.230938Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:39:39.522452Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b46a13a5-5276-415c-a05e-6a163b4ef094 · inbound

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning cites this paper.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Scaling up masked audio encoder learning for general audio classification

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.230938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.230938Z digest=sha256:632eee2f65a3af0c5392fe7375611cc436fa3267dbe53afc67c3509b4cbdf63d

Observation 491785a8-3459-4e14-8b06-35b25e30ec36 · inbound

Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification cites this paper.

Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification Scaling up masked audio encoder learning for general audio classification

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T13:51:44.825761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:51:44.825761Z digest=sha256:50eb9d1966c75ef76db3fc38135b7736ff7beede3b1d5cf8b726b0c9ee4c5122

Observation 7c5a6f8e-1d76-45a2-9fad-a1df8ca714bc · inbound

LambdaMark: Semantic Audio Watermarking for Robustness and Radioactivity cites this paper.

LambdaMark: Semantic Audio Watermarking for Robustness and Radioactivity Scaling up masked audio encoder learning for general audio classification

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:39:39.523962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T12:58:27.702138Z digest=sha256:165b8e8e91373acc266124662fed872368c1bcf8be6e06de8003d8f99a89e65d

Observation 93d0354b-deee-47d8-9f72-a4a6b1d755e0 · inbound

ALM2Vec: Learning Audio Embeddings for Universal Audio Retrieval with Large Audio-Language Models cites this paper.

ALM2Vec: Learning Audio Embeddings for Universal Audio Retrieval with Large Audio-Language Models Scaling up masked audio encoder learning for general audio classification

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:05:37.228821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:47:53.308909Z digest=sha256:4234ad87f98dd1d2a1020fc104ec744fdec1f962f62d0377c3f1dc2fb0f83587

Observation faddfb2d-605d-493d-8f29-330f43649b6b · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Scaling up masked audio encoder learning for general audio classification

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:45:47.112058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:3dd5af860bd395f68043e1dddfc48c08ff079ada0cebc2018e91f4b7d498cdb4

Observation fe046f72-72e3-4569-a496-4df20a5c0450 · inbound

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning cites this paper.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Scaling up masked audio encoder learning for general audio classification

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T05:56:39.845086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:bc0ea484a87d04050b5afc3d8a3c9f5b0a36906e2d4acee0fde0fa107e7bb58c

Observation 3b53fdf3-0866-4d73-9c63-9cb55f6c3304 · inbound

Large Audio Language Models for Spoofing-Aware Speaker Verification cites this paper.

Large Audio Language Models for Spoofing-Aware Speaker Verification Scaling up masked audio encoder learning for general audio classification

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T01:12:30.683133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:12:30.683133Z digest=sha256:7b78fc22bf48ff961ccf69424ba706f2c49d294039dcbf2792e9b886296b0e26

Observation c46ce0a0-409c-4b8c-bad2-cd897e7b0403 · inbound

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio cites this paper.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Scaling up masked audio encoder learning for general audio classification

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:35.218459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:35.218459Z digest=sha256:1c8746a5f1300438cb4e05ae3d1cca1ca47aed909c7ed2e5fbf34db9dbfa9b72