Pith. sign in

Paper Citation Record · LEDGER

BEATs: Audio Pre-Training with Acoustic Tokenizers

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 45 inbound Pith citation observations for arXiv:2212.09058.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2212.09058 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 45 of 45 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T18:05:52.235239Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T00:04:22.466366Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0fd8edf6-6bcf-447c-a786-d6886f7d1814 · inbound

FAST: Efficient Action Tokenization for Vision-Language-Action Models cites this paper.

FAST: Efficient Action Tokenization for Vision-Language-Action Models BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:52:31.926052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:9874a028f52b743bae7d4d7e8666d2a92ba43d2ecf3e60e2ff3c801070f99a3d

Observation 21115587-0540-47a4-9271-0a6430d890cd · inbound

"I am bad": Interpreting Stealthy, Universal and Robust Audio Jailbreaks in Audio-Language Models cites this paper.

"I am bad": Interpreting Stealthy, Universal and Robust Audio Jailbreaks in Audio-Language Models BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T18:05:52.235239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:05:52.235239Z digest=sha256:541d4214041d8737be23bf35c28e0ac507db655786477fff9a3fda30e8de8157

Observation 9ba18c47-ad51-489b-a2ad-5cbd56f2e8e3 · inbound

Evaluation of Deep Audio Representations for Hearables cites this paper.

Evaluation of Deep Audio Representations for Hearables BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T14:47:07.989640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:47:07.989640Z digest=sha256:dc215f24e59cc221b50a9500e88a75fc65e2644b4ce82b4c081f264a634f608b

Observation 7beb06ff-c36e-4a0f-a60f-e687ac17af26 · inbound

X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance cites this paper.

X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:05:15.772955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:05:15.772955Z digest=sha256:08f031904240324a885c2a0b45a1e7d077bdecce0531f75e647f9228a305b114

Observation 44c14ae2-aab0-4e07-b60d-3d12d1b66ed7 · inbound

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English cites this paper.

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:05.080883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:05.080883Z digest=sha256:c88d153c5fe1e6fa5816969673186a85315681ca7ec14092f88df830427271f2

Observation 4df0c23c-75af-4104-9e04-af14abb2e49d · inbound

RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language cites this paper.

RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:19:10.885941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:19:10.885941Z digest=sha256:6cf5dbd4be805c3400e66ce01230174a12fee7ab8512f5be972d1ca20e89fd65

Observation e0ee79de-2d55-4c4a-9c8b-f63aba10c8d7 · inbound

LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs cites this paper.

LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:48.281447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:48.281447Z digest=sha256:8afdd92bd64470da3bc302ceb9dd05c22fb408292073b010b142f4b19d2c5166

Observation 1c2b0f6c-cf99-4bb2-b38e-37f4ea4a29c6 · inbound

Efficient and Microphone-Fault-Tolerant 3D Sound Source Localization cites this paper.

Efficient and Microphone-Fault-Tolerant 3D Sound Source Localization BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:40.632711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:40.632711Z digest=sha256:65038a0f6e0b920097f919a9ab22f2ee0ddda4954dece9770ca2205cbc90514b

Observation d39d2e9c-6acf-4e67-9491-54518e61f564 · inbound

Zero-Shot Cognitive Impairment Detection from Speech Using AudioLLM cites this paper.

Zero-Shot Cognitive Impairment Detection from Speech Using AudioLLM BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:49.677614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:49.677614Z digest=sha256:ea39f79f0c25cd1860e61e64b02b456f5645d5167548abd89c8b4cb490f8f932

Observation 9917f026-aa6d-414f-b692-c025fa093fe7 · inbound

Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World cites this paper.

Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:10.747215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:10.747215Z digest=sha256:c888b2461b36b1168523817fd64c9609c5919976b2d5c83a54e0c9cb4a56d476

Observation e6801440-a98a-4f92-acfd-75a1da97f3bc · inbound

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis cites this paper.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:13.348025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:13.348025Z digest=sha256:a21bbdbee37c457c352cabc3d044036c326d67e9043fbe22ab493c54a98de960

Observation 043fba0e-52af-41da-b7cc-98f402ff4421 · inbound

Task-Specific Audio Coding for Machines: Machine-Learned Latent Features Are Codes for That Machine cites this paper.

Task-Specific Audio Coding for Machines: Machine-Learned Latent Features Are Codes for That Machine BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T16:47:34.333552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:47:34.333552Z digest=sha256:aff5d1c1916723a4e2674229f9fd783e4dbbbfb4278062f54fb20ebc7b9c5f02

Observation cc072def-b9ec-4f42-bd4c-f70984df4af5 · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:59:50.976241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:5e84aa920db55bd1b9642c991cf4f3c648ca9eb5b74f34d958401f57cabd492f

Observation d32b479c-3c7f-4f03-9026-72d988877b7b · inbound

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting cites this paper.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.509860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.509860Z digest=sha256:6bd8e23c529cdae7660b41110c5f5d1d8d8d5cf6601457cdfa0972928db8765d

Observation f7251430-0160-4158-b836-a7496021ba59 · inbound

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations cites this paper.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:29.678062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:29.678062Z digest=sha256:b374ae91d624a1246d5430ecc82833a7b2b79dbbf3af896c8359acf4083cb535

Observation 0fc7d1e3-6c86-42ca-a4ee-6686936fc6ae · inbound

Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning cites this paper.

Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T00:01:17.525381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:01:17.525381Z digest=sha256:262c5e013ab74a69b81551189ebe81cf947cdbc861bb1e27b7b0a9a4c1f957f2

Observation 22662b78-6232-4296-82c8-475b2160c9ad · inbound

Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual Segmentation cites this paper.

Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual Segmentation BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T00:04:54.793915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:04:54.793915Z digest=sha256:28098c739ae5adadd1a7b9cdf4065bef4f0ea7a441150666411c4e70a9ea3e2c

Observation 9463a595-eb51-4039-85fc-e466eb461fa9 · inbound

A Survey on Video Temporal Grounding with Multimodal Large Language Model cites this paper.

A Survey on Video Temporal Grounding with Multimodal Large Language Model BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 147

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:18.136571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:18.136571Z digest=sha256:412fb98e5ba3f39c03046eee5e4f97f1700fb7a515c5e0d94bd649900bed479c

Observation 4d40ba64-a9e4-4823-8020-9045121195db · inbound

AudioSet-R: A Refined AudioSet with Multi-Stage LLM Label Reannotation cites this paper.

AudioSet-R: A Refined AudioSet with Multi-Stage LLM Label Reannotation BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T17:57:32.271114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:57:32.271114Z digest=sha256:0dd6de199f5d136989cc0f100a853bd870406fd2374e77da00859c4f1170833f

Observation 88a4edce-cff7-4396-89c3-402c0ff4e7df · inbound

VQualA 2025 Challenge on Engagement Prediction for Short Videos: Methods and Results cites this paper.

VQualA 2025 Challenge on Engagement Prediction for Short Videos: Methods and Results BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T11:17:54.443943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:17:54.443943Z digest=sha256:b07bbb7f9986e0c3fbbbb5db1b21e3355eead68db61c435de42bdc305fa61929

Observation be8b3450-5551-4e9d-97c3-ecd45d3e549a · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:30.654795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:30.654795Z digest=sha256:596b3d0c85fb503c0a6c913c9d620aaf6ab908037c6b526b0573d3947b49244a

Observation dc38e2c6-390a-4c8f-b3a3-9680f939ff1f · inbound

Assessing Factual Music Comprehension in Large Audio Language Models cites this paper.

Assessing Factual Music Comprehension in Large Audio Language Models BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T00:29:45.976102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:29:45.976102Z digest=sha256:c26906f216775a8ed25563092917e9d9fd417dbd3ecaabbec580962fb5e36531

Observation 560dcc23-698d-4244-813f-07ce1600afe7 · inbound

SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM cites this paper.

SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T21:42:47.142235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:42:47.142235Z digest=sha256:bdddb2d81fd9c0826d7007f39836dd0ecbe3dd88f5e826ce939b23aaf07dfae1

Observation 3e716b4b-02e4-436e-8cbf-b5b1995e2786 · inbound

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs cites this paper.

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T17:16:36.790900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:16:36.790900Z digest=sha256:63e6a7b3b1caeeab209da0bed2d76d3bdbdd3bccaa6de547f32795815fc4c5dc

Observation 0ee47b9f-1b44-48b2-b38b-b9af80c234d7 · inbound

Quantitative Analysis of Proxy Tasks for Anomalous Sound Detection cites this paper.

Quantitative Analysis of Proxy Tasks for Anomalous Sound Detection BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T10:53:15.264765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:53:15.264765Z digest=sha256:0c016d2eaf327bc6d982ed9a7d8b5f6c05b58586247ecd02ea17478bdf4dc918

Observation e8c87b1f-9aa7-4e45-826a-1b53a4dda16e · inbound

ULTRAS -- Unified Learning of Transformer Representations for Audio and Speech Signals cites this paper.

ULTRAS -- Unified Learning of Transformer Representations for Audio and Speech Signals BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:30:50.864486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:30:30.431777Z digest=sha256:b09da1608f59173a9d3510278fc7583720f940f21bc81f64225d16cdbaa4be5b

Observation f0e867ad-df14-4bcb-a331-78aa73ff8341 · inbound

Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models cites this paper.

Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.946896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T13:57:47.356373Z digest=sha256:694b2b3c5642fe602ae0bd975040ec42ca5bd7fb279dd45284bb131e34f76867

Observation bc82c1a8-7dd8-41eb-ae7c-23fcade130d1 · inbound

TinyMU: A Compact Audio-Language Model for Music Understanding cites this paper.

TinyMU: A Compact Audio-Language Model for Music Understanding BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:17:36.935059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T08:17:23.740979Z digest=sha256:742c2a7738ec460ace62cad904981d4d30422cdb74d009e647bbf96857682b8a

Observation b07ddca6-efdf-41f7-bd4e-e639700fdbb5 · inbound

MUSCAT: MUltilingual, SCientific ConversATion Benchmark cites this paper.

MUSCAT: MUltilingual, SCientific ConversATion Benchmark BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.191807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T09:20:22.200577Z digest=sha256:84d2e1de18d5a42716305393b14905e3605c8f511d836edc67a109564048c317

Observation a0570288-70ac-4bfc-8f51-75240f384f19 · inbound

MUSCAT: MUltilingual, SCientific ConversATion Benchmark cites this paper.

MUSCAT: MUltilingual, SCientific ConversATion Benchmark BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-21T00:49:19.547770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T00:45:01.115573Z digest=sha256:08968d6c35e56c1450d9ec422c25992bdddddf9a2ebcfe77aa35431fb3c8e634

Observation c413f0d0-e26d-49e0-803b-013725d60ad6 · inbound

Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective cites this paper.

Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 291

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:51:03.124351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T04:32:29.428080Z digest=sha256:407f3e0180047c506a38fb06b89182c542aeef112b551fd3bbbab6dbb6769daf

Observation 5721e150-1e14-4978-a682-d84a42bec402 · inbound

Memory Efficient Full-gradient Attacks (MEFA) Framework for Adversarial Defense Evaluations cites this paper.

Memory Efficient Full-gradient Attacks (MEFA) Framework for Adversarial Defense Evaluations BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:01:18.489162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T12:53:11.185208Z digest=sha256:8e68df5c3089cc632a2057ec1c7b2f4e2db530da01d492a6892f98fa28feaba8

Observation 137c5f66-8d64-444b-85bf-8641c9512bbe · inbound

SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification cites this paper.

SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:32:57.059110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T20:31:49.239866Z digest=sha256:92eb515ff2caee439182c20f57009a083613680eee69fa926b32b44dc1c7e0d6

Observation bd003016-04bf-488f-a810-4d272627b8f3 · inbound

OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation cites this paper.

OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:13:48.563590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T18:08:50.574960Z digest=sha256:2e5bfa3631b79141f32be8bc23d473a383395440f25e1effffc714726f2f35eb

Observation 87328366-cfcc-4fad-a0ae-6cf828e5d1b3 · inbound

Finding Needles in the Haystack: Transductive Active Labeling in Ecology cites this paper.

Finding Needles in the Haystack: Transductive Active Labeling in Ecology BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:26:27.405444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T10:52:23.274642Z digest=sha256:6c6985dc3ce2c6ba35f981d65afcbc03d5fe26a2a04d1d682ea48e0178a52710

Observation 2dcc5b3f-0b7e-46cd-947b-6663208e311c · inbound

Finding Needles in the Haystack: Transductive Active Labeling in Ecology cites this paper.

Finding Needles in the Haystack: Transductive Active Labeling in Ecology BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T07:55:31.231791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T07:48:01.274555Z digest=sha256:a4bc13b0e044ec9a60ce7444d01927bc83640648267079619a44a3c5bf17692f

Observation f7a930f7-f2ed-43c5-8098-c40fa06e921e · inbound

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning cites this paper.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T05:56:39.917414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:9b258c63c2677688fa1a1b22ab7a792100fca1791665b9042aafdfe3c6f1c9fa

Observation e9834dc0-8da9-4b02-8df1-75ca93cd477b · inbound

Adaptive Loss Balancing for Multi-Task Bioacoustic Classification of Bird Species and Call Types cites this paper.

Adaptive Loss Balancing for Multi-Task Bioacoustic Classification of Bird Species and Call Types BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 7

Resolution
malformed identifier
no resolver link, observed 2026-07-12T03:24:13.814557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T03:24:13.814557Z digest=sha256:dbadf3938ce159b873207baf082e528fe4e9b7362e1ee8f85658171e27699941

Observation 5ea85c91-d981-4e74-9f00-cad0e7b98f56 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 274

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.467634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:a4b0271b9eaa993ff351287d8eb9bab3a71f69606214550119aa5f70a7ad2eb3

Observation fb69bebd-0691-40d1-99ec-d89a0164e2c9 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 274

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:d8886f1d337cfcac38825bd78d92b515a0c07379eb13ebad0c4c4c7959726f22

Observation 35666e69-0923-4e17-9b60-173abe7283e8 · inbound

FdAudio: MeanFlow-Anchored Fr\'echet-Distance Post-Training for One-Step Text-to-Audio Generation cites this paper.

FdAudio: MeanFlow-Anchored Fr\'echet-Distance Post-Training for One-Step Text-to-Audio Generation BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T11:52:50.598080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:52:50.598080Z digest=sha256:041cf4877d36ec7db83bea1364c4efb81b745ad1bfdff40f94493d39bf312af4

Observation 11f819c1-a7d3-476d-9f32-67771820b54c · inbound

Can Tokens Compete? Token Representations against Supervised CNN Backbones for BirdCLEF+ 2026 cites this paper.

Can Tokens Compete? Token Representations against Supervised CNN Backbones for BirdCLEF+ 2026 BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T02:03:37.761113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T02:03:37.761113Z digest=sha256:4f1e2b53b117caae1179e0e6f24601f9334011d7cd89aa2218d1f06136041203

Observation 9b447bf8-04ab-4d19-b867-a62cfef7b3a3 · inbound

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation cites this paper.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-30T10:35:03.133791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T10:35:03.133791Z digest=sha256:ec483cf7b048a713342c574eb21d6ee06d15f5a7f3f391259694a62c9f994736

Observation 3cdad4a4-601b-452a-afb4-4458397ec441 · inbound

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation cites this paper.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T01:51:00.418732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:51:00.418732Z digest=sha256:1006b2e3b4daa101138a7b459a7837a42bfcb5275d924e8c0f1512b2801a91ca

Observation 6b1606f0-2d5d-4d48-9dd8-a5d38a722a5e · inbound

Hidden-Domain Routing for All-Type Audio Deepfake Detection cites this paper.

Hidden-Domain Routing for All-Type Audio Deepfake Detection BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:45.434431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:54:45.434431Z digest=sha256:7679cb6af1f266d3a13b84a963a01f19f0e652c31a166d75bf014269420f5939