Pith. sign in

Paper Citation Record · LEDGER

AST: Audio Spectrogram Transformer

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 67 inbound Pith citation observations for arXiv:2104.01778.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2104.01778 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 67 of 67 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:45:46.102491Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

31
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 01a0685b-d2c9-4e6d-8743-21b5b06fdc70 · inbound

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach cites this paper.

Fine-tuning Multimodal Transformers on Edge: A Parallel Split Learning Approach AST: Audio Spectrogram Transformer

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T15:45:46.102491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:45:46.102491Z digest=sha256:75e7f50d4f8f3b46da72f1fb90afdb2980c99138feb1e0d7d1abbc07a699ee11

Observation 9d3c259c-a869-4382-9f1b-db706ca8510a · inbound

Histogram-based Parameter-efficient Tuning for Passive and Active Sonar Classification cites this paper.

Histogram-based Parameter-efficient Tuning for Passive and Active Sonar Classification AST: Audio Spectrogram Transformer

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:11:54.476417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T18:11:44.073658Z digest=sha256:227a88dcc398a905c4059076552addadcec39b579566956adc7a72afbeb7624b

Observation d15fcc3d-a5ac-40e7-ad52-82e7c5442a09 · inbound

M2R2: MultiModal Robotic Representation for Temporal Action Segmentation cites this paper.

M2R2: MultiModal Robotic Representation for Temporal Action Segmentation AST: Audio Spectrogram Transformer

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-22T17:41:53.072562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T17:39:23.051951Z digest=sha256:5630ec6b8c3615aa5c67a4d953cb464ef45b858f8b85422a4a0dc14996416a2d

Observation 38a50e74-8e1a-413a-b738-fa01f4a63aed · inbound

Time to Embed: Unlocking Foundation Models for Time Series with Channel Descriptions cites this paper.

Time to Embed: Unlocking Foundation Models for Time Series with Channel Descriptions AST: Audio Spectrogram Transformer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:37:35.932400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:37:35.932400Z digest=sha256:68349123bff50ed04f19b8a75a09097e2c46554fdfb3bf7caf818d87ccb47283

Observation a0a6ad1d-e84c-4249-8d2c-30989aefec82 · inbound

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding cites this paper.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding AST: Audio Spectrogram Transformer

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:03.520260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:03.520260Z digest=sha256:025213488f8f80fc9130e328671fff005de76c778ee0c610a8e6f47becdeeefe

Observation 97f0c727-3582-48f2-a390-b9c466e8895d · inbound

Patient Domain Supervised Contrastive Learning for Lung Sound Classification Using Mobile Phone cites this paper.

Patient Domain Supervised Contrastive Learning for Lung Sound Classification Using Mobile Phone AST: Audio Spectrogram Transformer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:02.389077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:02.389077Z digest=sha256:58a65b9b6368d44fd15f64c62ba97c2c001328ad33b315cf857e652a6dcb963c

Observation 51050fcc-6128-4bf5-9f37-6ae9629031e9 · inbound

Bridging the Gap Between Semantic and User Preference Spaces for Multi-modal Music Representation Learning cites this paper.

Bridging the Gap Between Semantic and User Preference Spaces for Multi-modal Music Representation Learning AST: Audio Spectrogram Transformer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:52:38.367740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:52:38.367740Z digest=sha256:4e6a58e7f77522a67be355d356348793353dcd10ff9d1fcf7b7b8cbb709378f8

Observation 4c04a309-fe78-440d-a9e4-e78b200334ce · inbound

Acoustic Classification of Maritime Vessels using Learnable Filterbanks cites this paper.

Acoustic Classification of Maritime Vessels using Learnable Filterbanks AST: Audio Spectrogram Transformer

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:41:35.511647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:41:35.511647Z digest=sha256:b4c43eea8840a668bb9632bb56e83edd6c548303a16b71cf6653e07293efd10e

Observation 2c6fa41f-8e91-4b11-9f99-2e7d678586f3 · inbound

Comparison of spectrogram scaling in multi-label Music Genre Recognition cites this paper.

Comparison of spectrogram scaling in multi-label Music Genre Recognition AST: Audio Spectrogram Transformer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:37:13.332193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:37:13.332193Z digest=sha256:11391bd6d5cd7bc3cb1653a98f1282162638204e4d59442d1132fd15efa5f8cb

Observation 0827d5e9-db23-479a-a2c8-f592474ded7b · inbound

Multimodal Zero-Shot Framework for Deepfake Hate Speech Detection in Low-Resource Languages cites this paper.

Multimodal Zero-Shot Framework for Deepfake Hate Speech Detection in Low-Resource Languages AST: Audio Spectrogram Transformer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:04.438838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:04.438838Z digest=sha256:47b8948ef697b7d1fd43bb0bcbcc27bb890585e8535aedc1af5bb879d93d331f

Observation 6b9c80ef-eee0-4785-8783-162fa92a9b07 · inbound

15,500 Seconds: Lean UAV Classification Using EfficientNet and Lightweight Fine-Tuning cites this paper.

15,500 Seconds: Lean UAV Classification Using EfficientNet and Lightweight Fine-Tuning AST: Audio Spectrogram Transformer

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:30.519671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:30.519671Z digest=sha256:ded61fa779139a78fae4c090c24caf8efff22e90d8c39160ac179069e3a9731a

Observation d3dc17cb-3d4d-4633-a78d-5a05ae4ce08f · inbound

PromptTSS: A Prompting-Based Approach for Interactive Multi-Granularity Time Series Segmentation cites this paper.

PromptTSS: A Prompting-Based Approach for Interactive Multi-Granularity Time Series Segmentation AST: Audio Spectrogram Transformer

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:05.434923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:05.434923Z digest=sha256:8b7c960c8575a9e9c9cf1fe998e89f216ea0dfd3a73ade06e32266fe69976801

Observation 61012109-f531-4814-87f9-741cb0120e8d · inbound

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes cites this paper.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes AST: Audio Spectrogram Transformer

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:44.293487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:44.293487Z digest=sha256:9c118a9751958c388ea25c54825591eb806c6e43267490173cc9e774a9b69f59

Observation 1224e34d-2cf2-4b69-871b-9e6e86eb0d7a · inbound

Acoustic scattering AI for non-invasive object classifications: A case study on hair assessment cites this paper.

Acoustic scattering AI for non-invasive object classifications: A case study on hair assessment AST: Audio Spectrogram Transformer

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:20:49.190056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T00:18:57.564840Z digest=sha256:a36a00c2966e72bde18df45ffc338b8c8aebc2f975e28d6510d11a7933d7c08c

Observation 11034275-b9fd-4b76-825e-0e803916ec81 · inbound

Tactile Beyond Pixels: Multisensory Touch Representations for Robot Manipulation cites this paper.

Tactile Beyond Pixels: Multisensory Touch Representations for Robot Manipulation AST: Audio Spectrogram Transformer

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:18.265968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:15:18.265968Z digest=sha256:bd30799e4d227bfe23642bcdc794c8f2a3db9230239eb6468ac2263811fc2506

Observation d44c70f8-fdad-4b84-9c0a-2e4aeeec1de7 · inbound

Deciphering GunType Hierarchy through Acoustic Analysis of Gunshot Recordings cites this paper.

Deciphering GunType Hierarchy through Acoustic Analysis of Gunshot Recordings AST: Audio Spectrogram Transformer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:47:29.517169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:47:29.517169Z digest=sha256:116cb02d2625b6325b2f22b9b26a88442a9c9a39be43a0b0e81d966213ce9612

Observation c05c7cc8-342f-42bd-b773-1ecfc5832b33 · inbound

MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing cites this paper.

MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing AST: Audio Spectrogram Transformer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:12:02.492655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:12:02.492655Z digest=sha256:86c01d346b1a43450c3c0e0abf6abb40b93a8a343e3626ac219ab2b19002b34c

Observation 7c94d390-fc21-4de9-bc15-50679fe0be39 · inbound

IMPACT: Industrial Machine Perception via Acoustic Cognitive Transformer cites this paper.

IMPACT: Industrial Machine Perception via Acoustic Cognitive Transformer AST: Audio Spectrogram Transformer

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:12:43.684468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:12:43.684468Z digest=sha256:421f6f5a949f4efdb1e56168df0b3202229b28ed13e438ffd49a3d860e802c8d

Observation 8eeb24fa-c3e4-4bf2-bce0-6e3ca5a37666 · inbound

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition cites this paper.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition AST: Audio Spectrogram Transformer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.498534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.498534Z digest=sha256:cb799b1ab0ca6c403b967077cf361a81f0d767cab3a5afc9fd6d0b39b90d8049

Observation c5ee46f4-7d1b-482f-8cad-1e106d61c0a3 · inbound

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning cites this paper.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning AST: Audio Spectrogram Transformer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:17.039560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:17.039560Z digest=sha256:633a04adb44ef197e09fa066f60d401d388196e6e98f93350bac05cb36227974

Observation 8517328c-d79c-4fe0-ab3a-4c568edbdbc9 · inbound

Your Attention Matters: to Improve Model Robustness to Noise and Spurious Correlations cites this paper.

Your Attention Matters: to Improve Model Robustness to Noise and Spurious Correlations AST: Audio Spectrogram Transformer

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T13:39:08.440360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:39:08.440360Z digest=sha256:d3129766ea544c48621a5dab17c184bc8c7648288028eec058333292aa1c4b9d

Observation 3733418c-44e5-49db-9337-763e77badd5f · inbound

Angle-distance decomposition based on deep learning for active sonar detection cites this paper.

Angle-distance decomposition based on deep learning for active sonar detection AST: Audio Spectrogram Transformer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:29:39.608184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:29:39.608184Z digest=sha256:1df2cf117675bd6d5088bd80d7b97f59a1e42500b6eb1b3f18ef4d5ceb76428b

Observation 187f7f20-0af3-4e6f-a632-27b512db3adc · inbound

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment cites this paper.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment AST: Audio Spectrogram Transformer

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:26.116606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:26.116606Z digest=sha256:9ea28cc20358c534b8fada1c0cd19845a81b2951949e9ca786655bf10d47479b

Observation ec3e673b-b559-47b2-9a43-6d0e13f0a606 · inbound

A Two-Step Learning Framework for Enhancing Sound Event Localization and Detection cites this paper.

A Two-Step Learning Framework for Enhancing Sound Event Localization and Detection AST: Audio Spectrogram Transformer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T11:54:44.824499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:54:44.824499Z digest=sha256:9a2c5a40ba0614c21196ab282897de66721aea75462dc7fdf02acf4ec1ac21cd

Observation 1ebee6d5-741f-4ff4-8e8c-f667f7641f9a · inbound

CoughViT: A Self-Supervised Vision Transformer for Cough Audio Representation Learning cites this paper.

CoughViT: A Self-Supervised Vision Transformer for Cough Audio Representation Learning AST: Audio Spectrogram Transformer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T04:51:59.205438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:51:59.205438Z digest=sha256:02a183f25628625b154c861a65337f62cc37c3546250165007b1e3f083cc98ec

Observation fe0e9cbf-465b-4124-add9-cd02e8b51251 · inbound

Full-Frequency Temporal Patching and Structured Masking for Enhanced Audio Classification cites this paper.

Full-Frequency Temporal Patching and Structured Masking for Enhanced Audio Classification AST: Audio Spectrogram Transformer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T14:30:46.499751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:30:46.499751Z digest=sha256:e911ea10f656df4cfe04cef154b2b9d8430f53fd0d60990868ed482a888bf7a3

Observation f6ded1af-efc1-4f22-8969-8af1e09bab3e · inbound

AudioRWKV: Efficient and Stable Bidirectional RWKV for Audio Pattern Recognition cites this paper.

AudioRWKV: Efficient and Stable Bidirectional RWKV for Audio Pattern Recognition AST: Audio Spectrogram Transformer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T11:53:50.790037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:53:50.790037Z digest=sha256:418673ef3bdae5f4fd67a30f32a1a70be111054e35159825b80dd998c3993946

Observation 5454e27e-3978-4895-b890-ff6d4fdc329f · inbound

Region-Specific Audio Tagging for Spatial Sound cites this paper.

Region-Specific Audio Tagging for Spatial Sound AST: Audio Spectrogram Transformer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T18:59:47.408023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:59:47.408023Z digest=sha256:2ade89f2206adcc9da250c6c41b914b46218096f240f67f23488097c8c0090b7

Observation 0c8f2938-b5ec-4e79-bc31-69f1150fe6ea · inbound

DHAuDS: A Dynamic and Heterogeneous Audio Benchmark for Test-Time Adaptation cites this paper.

DHAuDS: A Dynamic and Heterogeneous Audio Benchmark for Test-Time Adaptation AST: Audio Spectrogram Transformer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T20:49:54.329414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:49:54.329414Z digest=sha256:4bfe966ea46b80ab973845e5ee3365ae1f3368682e78e12d3c5c3fb5af9daf43

Observation dd600482-91dc-4516-812f-d3d1c422a6a9 · inbound

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs cites this paper.

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs AST: Audio Spectrogram Transformer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T17:16:38.554983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:16:38.554983Z digest=sha256:3ee8a3c056fc60282d7947e637af0f30ef5534ac217be1bd64ae2788e4af6aa7

Observation 9b53fdc8-ce7d-4013-8433-d8ec055fe305 · inbound

Evolution Strategy-Based Calibration for Low-Bit Quantization of Speech Models cites this paper.

Evolution Strategy-Based Calibration for Low-Bit Quantization of Speech Models AST: Audio Spectrogram Transformer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T18:35:50.568237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:35:50.568237Z digest=sha256:92f6bd12aa5f1e1bd6c1cf914be28213fd6beb2ff284713e6cdf04d1e5336409

Observation 27691a8d-3392-447b-a95e-8ae39e8c515f · inbound

Distance-aware Soft Prompt Guidance for Multimodal Valence-Arousal Estimation cites this paper.

Distance-aware Soft Prompt Guidance for Multimodal Valence-Arousal Estimation AST: Audio Spectrogram Transformer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T18:21:58.032767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:21:58.032767Z digest=sha256:2c0d328c102745f35138109f4a8fbd5100fd64672f1c577ee2d5aa5c0e0d1df1

Observation 58e6f0ec-2300-4524-a115-a6013ed7cf5b · inbound

Distance-aware Soft Prompt Guidance for Multimodal Valence-Arousal Estimation cites this paper.

Distance-aware Soft Prompt Guidance for Multimodal Valence-Arousal Estimation AST: Audio Spectrogram Transformer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:51.775100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:51.775100Z digest=sha256:4857dc76fa6da7ba6a634fdeb37cad69c277ec250fd9c3e4f31f6bf1263b1c64

Observation 7249a03e-1374-4c98-999a-5bff152da504 · inbound

AILive Mixer: A Deep Learning based Zero Latency Automatic Music Mixer for Live Music Performances cites this paper.

AILive Mixer: A Deep Learning based Zero Latency Automatic Music Mixer for Live Music Performances AST: Audio Spectrogram Transformer

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:05:26.307268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T10:01:42.570418Z digest=sha256:40e187c458814072bb81b25951a401566017885de1183f55143557bee97e19dc

Observation 8e2cb717-2ed5-47e2-b512-a7ae3193f5e6 · inbound

VAN-AD: Visual Masked Autoencoder with Normalizing Flow For Time Series Anomaly Detection cites this paper.

VAN-AD: Visual Masked Autoencoder with Normalizing Flow For Time Series Anomaly Detection AST: Audio Spectrogram Transformer

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:08:21.799091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T00:04:05.071623Z digest=sha256:f48644c83511558c22b0dce345b237c33f34121b7e1c2ba48cdf552a1677fb56

Observation f019eacc-3ae5-417e-a3fb-14bb7a65e392 · inbound

VAN-AD: Visual Masked Autoencoder with Normalizing Flow For Time Series Anomaly Detection cites this paper.

VAN-AD: Visual Masked Autoencoder with Normalizing Flow For Time Series Anomaly Detection AST: Audio Spectrogram Transformer

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T17:22:59.839713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:22:59.839713Z digest=sha256:dcb356a536563ab694431c58543e5129f2d3a81a52381cdfd9256af6fe4a6f27

Observation b742e924-0625-4fa9-b903-0304ceeb93ef · inbound

You're Pushing My Buttons: Instrumented Learning of Gentle Button Presses cites this paper.

You're Pushing My Buttons: Instrumented Learning of Gentle Button Presses AST: Audio Spectrogram Transformer

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:55:50.774524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:48:22.314443Z digest=sha256:681235d8f1f1e0767002d689f837f9d5d969adcc5c0514b6da6577a4497fc902

Observation 9b2ae9ef-aa66-4908-9297-0e7607a3ce8a · inbound

ULTRAS -- Unified Learning of Transformer Representations for Audio and Speech Signals cites this paper.

ULTRAS -- Unified Learning of Transformer Representations for Audio and Speech Signals AST: Audio Spectrogram Transformer

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:30:50.823026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:30:30.431777Z digest=sha256:aaa26b395a8a7aa211a1749e691d502d51aa9890d3be3a7497af1f119a8b6305

Observation 577f6175-6b60-4cc5-8c72-478535f98a1e · inbound

Transformer Based Machine Fault Detection From Audio Input cites this paper.

Transformer Based Machine Fault Detection From Audio Input AST: Audio Spectrogram Transformer

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:15:28.279930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T14:14:31.344676Z digest=sha256:7fb92694c607be53dc496469dd273513c80fd87073f102e1855bfde735b3e3c9

Observation 53aa65a0-6c2c-4ed2-8ded-2b61b1ae3e86 · inbound

SAND: The Challenge on Speech Analysis for Neurodegenerative Disease Assessment cites this paper.

SAND: The Challenge on Speech Analysis for Neurodegenerative Disease Assessment AST: Audio Spectrogram Transformer

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:10:51.105148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:40:48.363579Z digest=sha256:6e746715957c70bd7bdae7f10c407cd74ae13d8a28bfc301eae246cb9822004b

Observation b13af13b-949f-42cd-886b-cd2a1549b568 · inbound

SAND: The Challenge on Speech Analysis for Neurodegenerative Disease Assessment cites this paper.

SAND: The Challenge on Speech Analysis for Neurodegenerative Disease Assessment AST: Audio Spectrogram Transformer

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:57:31.415288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T07:55:39.622962Z digest=sha256:f3e1aa6adf4b68e8aa2f0d3ac98899aa1ddb5ea3ff575f65e5c303b5d21412d9

Observation 48a6654c-096e-4e0f-b0c3-5aedb84ecdd6 · inbound

LIMSSR: LLM-Driven Sequence-to-Score Reasoning under Training-Time Incomplete Multimodal Observations cites this paper.

LIMSSR: LLM-Driven Sequence-to-Score Reasoning under Training-Time Incomplete Multimodal Observations AST: Audio Spectrogram Transformer

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:21:06.060998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-09T20:17:24.090467Z digest=sha256:d21f7a157dbcb2050bc85e5f7d7958d4291faf50aa740c3c9f6ed4c1ec9c2ca8

Observation 1e947acc-fc63-438d-920b-2950bf62b7eb · inbound

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task cites this paper.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task AST: Audio Spectrogram Transformer

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:11:08.290066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:715d1d13c0a3d9806d22a23a9ce8f4f1d85d137293c98e03c47756dd67c8f234

Observation 16c3a374-396b-43d0-afd0-12522711e9af · inbound

MLCR: Multi-Level Cue Refinement for Long-Term Multimodal Action Quality Assessment cites this paper.

MLCR: Multi-Level Cue Refinement for Long-Term Multimodal Action Quality Assessment AST: Audio Spectrogram Transformer

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:31:16.920533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T02:30:04.350174Z digest=sha256:23e409967c36640184f432f4fb4379dea3d4dbfeb74b8104adff72c364b2abda

Observation ce36f806-6d6a-4147-9279-af992c73bede · inbound

MLCR: Multi-Level Cue Refinement for Long-Term Multimodal Action Quality Assessment cites this paper.

MLCR: Multi-Level Cue Refinement for Long-Term Multimodal Action Quality Assessment AST: Audio Spectrogram Transformer

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:35:45.783412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T23:08:22.923816Z digest=sha256:557b4c7dee687f4c76cbf29f6ae44a912cfa20f9a7fce19427b948cf2ebdef13

Observation 8a56aa81-c172-4261-a63a-ec81575ad4a6 · inbound

Meow-Omni 1: A Multimodal Large Language Model for Feline Ethology cites this paper.

Meow-Omni 1: A Multimodal Large Language Model for Feline Ethology AST: Audio Spectrogram Transformer

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:41:43.803126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:02:23.847187Z digest=sha256:63ad523e620042a0218a7c99e648e8b37e19b3d24eb6355348ae6e5ae819cef8

Observation 57381c1c-f06f-466c-8c3d-67561f8eafb9 · inbound

SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification cites this paper.

SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification AST: Audio Spectrogram Transformer

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:57.049037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T20:31:49.239866Z digest=sha256:b0b332866c775feea5cdaea359f0febb6b918ab7254eb8e63b243587044b6ee5

Observation 1dd84bca-75b8-46e6-9ebe-53b37d67df03 · inbound

Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks cites this paper.

Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks AST: Audio Spectrogram Transformer

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:13:11.922566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T10:10:31.059095Z digest=sha256:7808afce7b74ca4833497120b23020f3677e652286ea6b8f43f50d8c0be68896

Observation 4249f01b-8d80-4e25-b6e2-b183a8d2247b · inbound

Executable Boundary Contracts for Sound Event Traces cites this paper.

Executable Boundary Contracts for Sound Event Traces AST: Audio Spectrogram Transformer

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-20T02:12:58.522782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T02:08:44.066554Z digest=sha256:96ba42c21cc1e0f8fb3a482c837448165b9a3755aa49ab8acaf56cd12b0b4f17

Observation 3745e5c3-2267-41b5-85fd-29383faed294 · inbound

SEABAD: A Tropical Bird Activity Detection Dataset for Passive Acoustic Monitoring cites this paper.

SEABAD: A Tropical Bird Activity Detection Dataset for Passive Acoustic Monitoring AST: Audio Spectrogram Transformer

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T02:19:25.442210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-21T02:15:41.816523Z digest=sha256:8b0f397d2309faaad7bbb6a930bd10dafde6156e9d0ee9ab2c97e08a75ad5d68

Observation 1f230ac2-e242-4201-aaec-2785f9371ea4 · inbound

Vanilla ViT for Automotive Point Cloud Semantic Segmentation cites this paper.

Vanilla ViT for Automotive Point Cloud Semantic Segmentation AST: Audio Spectrogram Transformer

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:42:47.102000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T22:35:40.495590Z digest=sha256:1654c10e87fbccbf644de86e77b6bc3e9ac30818819e8a67cbf74a3cd067fd10

Observation 79a0c5a1-e41e-4da2-a7b9-76439e90c096 · inbound

VZCrash: A Large-Scale IMU Dataset of Ego-Vehicle Crashes cites this paper.

VZCrash: A Large-Scale IMU Dataset of Ego-Vehicle Crashes AST: Audio Spectrogram Transformer

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:36:56.430386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T02:02:57.151031Z digest=sha256:ec2c533efd38f756936917569f72ebdb0e8f64848e0a4f404fed3c8d038b49f6

Observation 38c1f98c-db25-4100-97e8-37d637422426 · inbound

Spatial-Omni: Spatial Audio Understanding Integration in Multimodal LLMs via FOA Encoding cites this paper.

Spatial-Omni: Spatial Audio Understanding Integration in Multimodal LLMs via FOA Encoding AST: Audio Spectrogram Transformer

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T07:47:44.503408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T11:52:47.948163Z digest=sha256:bab9be4e1a32f6d9b889e26d70ad94fdee67ae27592bf74fb0c83cc769ece3a8

Observation 2f5c5f4d-f89e-4efb-84c6-3ca6ead98307 · inbound

Exploiting Neural Audio Codec Latents for Adversarial Audio Attacks cites this paper.

Exploiting Neural Audio Codec Latents for Adversarial Audio Attacks AST: Audio Spectrogram Transformer

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:49:36.694812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T15:34:38.141342Z digest=sha256:ee6ddbf72cd4b1b212ac9d96a09893a40ca8885c85f0ebbb501f98e56ab70cd8

Observation 1e8a4258-05c4-496c-a285-fc7292e5f1a8 · inbound

Hierarchical Policy Learning via Spectral Decomposition cites this paper.

Hierarchical Policy Learning via Spectral Decomposition AST: Audio Spectrogram Transformer

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:14:21.977403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T06:54:29.971971Z digest=sha256:9f9ae499826da012799a5b2a02b815a331162900dbc1089e864c575d889fa269

Observation 8376e35f-5aeb-4bfc-aea8-66cf16962f44 · inbound

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning cites this paper.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning AST: Audio Spectrogram Transformer

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:39.820388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:b89fa90ba094eab1c71e6dfbdf19e57f06c280abc46ccbf14e5aadeba504d266

Observation 72ec0e18-affb-4907-b025-eab7b13a238f · inbound

From Technical Metrics to User Perception: A User Study of a Multimodal Human-Robot Interaction System for Object Detection and Grasping cites this paper.

From Technical Metrics to User Perception: A User Study of a Multimodal Human-Robot Interaction System for Object Detection and Grasping AST: Audio Spectrogram Transformer

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:56:54.940058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-02T11:49:28.247486Z digest=sha256:297f2723fe322f9c5816673d1e3eaa9be9705f149baa6dc1d130a3db7f0dd302

Observation 5007d59f-a0c7-47ef-895d-12c7941baa41 · inbound

Missingness as Signal: Channel-Independent Spectrogram Learning for Clinical Time Series Prediction cites this paper.

Missingness as Signal: Channel-Independent Spectrogram Learning for Clinical Time Series Prediction AST: Audio Spectrogram Transformer

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T05:57:52.610916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:57:52.610916Z digest=sha256:f1be96ce861806a0ff5226c4bbb7db8588ec67e7dccea8af9d578db15e31583b

Observation c833f037-7ac7-4bb4-bdba-b8fa6319ee30 · inbound

Adaptive Loss Balancing for Multi-Task Bioacoustic Classification of Bird Species and Call Types cites this paper.

Adaptive Loss Balancing for Multi-Task Bioacoustic Classification of Bird Species and Call Types AST: Audio Spectrogram Transformer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T03:24:13.814557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T03:24:13.814557Z digest=sha256:65eae73a49ea55d760ddbf040001c3fc0990a4fd52666792f3611083e63b9115

Observation 8ff1e287-e2f7-45a9-be69-aa56a79d659c · inbound

Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding cites this paper.

Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding AST: Audio Spectrogram Transformer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T19:34:49.358453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:34:49.358453Z digest=sha256:4627de4865d8535759320e4392281a5c442c4443ed8f4eb4358b9f67c96b9e5d

Observation 63db0a09-07e1-4d4e-ab3c-737933b8431d · inbound

Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding cites this paper.

Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding AST: Audio Spectrogram Transformer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T08:43:33.527769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:43:33.527769Z digest=sha256:95c03d792bcb27e6ba3c748d26ab3f83c73d78a5ea4d0ab8c7d1f9a03a9992df

Observation df60abc3-6fc7-4117-b82c-f3d56fab2cd9 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence AST: Audio Spectrogram Transformer

Reference 275

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.431427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:5fd310a0ad7c88788c0dafe69d0488dadeb64ea31ccc6f32bdcff38e060ba825

Observation 96f767a6-a5e6-4d7e-af2a-07e0867fffd4 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence AST: Audio Spectrogram Transformer

Reference 275

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:7993c813e90ca7595dbbb96606a2e7eaba81f5ee20f0c21b0c1dabb3bb41d474

Observation f5d7c464-179b-4235-9533-b249f0dc72fc · inbound

Semantic Audio-driven Understanding for Dynamic Humanoid Whole Body Control cites this paper.

Semantic Audio-driven Understanding for Dynamic Humanoid Whole Body Control AST: Audio Spectrogram Transformer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T03:29:29.704719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:29:29.704719Z digest=sha256:1f1f423d3cf08f4e856adfc42b6eedb10a2a6f80175124e6b5f4823869be5c6a

Observation 8c258ffe-b7ea-4d76-b567-bc9840e800a3 · inbound

Can Tokens Compete? Token Representations against Supervised CNN Backbones for BirdCLEF+ 2026 cites this paper.

Can Tokens Compete? Token Representations against Supervised CNN Backbones for BirdCLEF+ 2026 AST: Audio Spectrogram Transformer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T02:03:38.740837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T02:03:38.740837Z digest=sha256:3b37dc95d8485daaaca313e8b93854baaecb57eb62ab2518fde882227dc77ebc

Observation 54d88d1b-54d3-4e2b-9bf3-ca5ac6e69a09 · inbound

LGFNet: A CTC-Guided Local-Global Fusion Framework for Single-Channel Sleep Staging cites this paper.

LGFNet: A CTC-Guided Local-Global Fusion Framework for Single-Channel Sleep Staging AST: Audio Spectrogram Transformer

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T03:12:06.459273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:12:06.459273Z digest=sha256:4af1bad8cc7149a1ae7054f67bfad2b021602f67b76f9e7c1a116679c10c233f

Observation 2e37716a-9121-400b-9b27-1e5fa08dee4d · inbound

Device Invariance using Domain Adaptation on Acoustic Scene Classification cites this paper.

Device Invariance using Domain Adaptation on Acoustic Scene Classification AST: Audio Spectrogram Transformer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T01:14:26.281716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:14:26.281716Z digest=sha256:26503e29449a546779fe077d84de7f8870e598922435bd33f60ec2ba59119ca3