Pith. sign in

Paper Citation Record · LEDGER

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes

As of 20 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 2 inbound Pith citation observations for arXiv:2506.12222.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12222 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T01:05:45.841972Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-02T05:52:55.818877Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T05:56:39.910208Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4184e5d6-aaaa-4143-8a3c-35a947f6d657 · outbound

This paper cites In our initial experiments, we observed that this approach yielded worse performance compared to SSLAM on the AS-20K benchmark (39.9 mAP vs.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes In our initial experiments, we observed that this approach yielded worse performance compared to SSLAM on the AS-20K benchmark (39.9 mAP vs

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:05:46.536287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T01:05:45.841972Z digest=sha256:9723afd29e052ca7f63d331f157d57ee28eca01740a6393aa9f5dac47b3be5de

Observation f6bfd562-bbc0-4c35-a419-d4457273e968 · outbound

This paper cites an unresolved cited work.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:05:46.767689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T01:05:45.741916Z digest=sha256:ade0a8b16a0737a00a791c479f9e251b254aac530f1d77dca200c5bb9c261281

Observation c4f98f10-dac5-48fe-9f54-7f3e2be380a6 · outbound

This paper cites Efficient Training of Audio Transformers with Patchout.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes Efficient Training of Audio Transformers with Patchout

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:44.433337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:44.433337Z digest=sha256:08d8d5dd106375b578df89bf685e6ae87e3fc59a5b8d8e4027c88294d1ca1395

Observation 0fcc80e2-feda-4148-88eb-e0baa5622bb2 · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:44.507719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:44.507719Z digest=sha256:a6fdf5662a58afdc9fa10d634be72e5574c99c402e801798b10b91c62340613f

Observation 2e3f4533-e133-4720-834f-0c0df286f6dc · outbound

This paper cites Decoupled Weight Decay Regularization.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes Decoupled Weight Decay Regularization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:44.590534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:44.590534Z digest=sha256:20ea9eee2cbdf77b54ece8071028bff3c1e1123314d483966184354af0374c50

Observation 191b7aef-7ad7-4a45-8d33-809fc4b8aa07 · outbound

This paper cites X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:44.701187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:44.701187Z digest=sha256:20f59e2188ea33d0dbbf5465df28b11caaa0714714560612e7e4f0daf9914dc7

Observation 85231792-3c9a-4dfc-b166-7954c14aa931 · outbound

This paper cites SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:44.764826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:44.764826Z digest=sha256:eb9093b8f32ad8991f8d35934b7e26e54c3dee1b3b7adfd076226baf1ae8cc52

Observation 964ad6b8-1c77-4194-bd9c-710b25a01b72 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:45.173925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:45.173925Z digest=sha256:a541295be76642e68113c789934767a20bb883de175ee91b50ab568ca69fd503

Observation e5832718-e9b9-496d-b2ae-a1ff8d684bc3 · outbound

This paper cites mixup: Beyond Empirical Risk Minimization.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes mixup: Beyond Empirical Risk Minimization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:45.237444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:45.237444Z digest=sha256:5565ee955a14b9b58f76adecc5546e0a23a12fe30a27bf5880894ea1db32452a

Observation 424e8316-0769-489b-a34f-997025e07cd3 · outbound

This paper cites ChatBridge: Bridging Modalities with Large Language Model as a Language Catalyst.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes ChatBridge: Bridging Modalities with Large Language Model as a Language Catalyst

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:45.325056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:45.325056Z digest=sha256:150beb28fe9f913a139aebd0bfe915e09a87e84771d3b1b369e6e8c7a83a5a24

Observation c89dbd5a-c2c2-4157-838f-f2f7eae8348f · outbound

This paper cites an unresolved cited work.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:05:47.683309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T01:05:45.373288Z digest=sha256:88c832e405a5be63b9017c52d7b6f15bdfff0684438c2904abac5d7902f89bc9

Observation a23beb56-6b39-4f66-944c-119197902a80 · outbound

This paper cites an unresolved cited work.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:05:47.449606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T01:05:45.491708Z digest=sha256:1cef3a4abc3afc8ed78b7b8471b50ad17d4490426e3434edebcf65b55e00e94b

Observation 6fe2436a-4f32-4d0d-8e0f-2730d18735e1 · outbound

This paper cites Each recording is annotated with a single class.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes Each recording is annotated with a single class

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:05:47.187219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T01:05:45.570518Z digest=sha256:a89e137af753ffda7dd858f6a3fc93e09278c36ce1612f31b4afa9d558002600

Observation 39447296-8e86-4078-9738-54bc04c8d365 · outbound

This paper cites multi-label.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes multi-label

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:05:47.004569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T01:05:45.645235Z digest=sha256:e2d0516db800c9a6ef2c2f9a3578232d546fc0fc0473150632c7722189726959

Observation 828e21e4-3bb1-4b55-8a3a-0590664e0a7e · outbound

This paper cites Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition

Reference 2005

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:45.096815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:45.096815Z digest=sha256:5a6dda44490efe01cebbd5200175a6ccb2fbaa1721e317c33070d0ba4d1181c6

Observation 807d1e59-79a6-4c6b-a0d5-dc87f44db323 · outbound

This paper cites Dropout: a simple way to prevent neural networks from overfitting.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes Dropout: a simple way to prevent neural networks from overfitting

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:44.906560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:44.906560Z digest=sha256:1236a5850dfc323674649e6031857de1ac880a7bb0b6c1a7d712f18711168b1e

Observation 08a10811-8325-40d1-b1bd-37c2c3582a20 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:45.009230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:45.009230Z digest=sha256:8df56f9121647dd12e88eca57638acc89bba19a430716b26780230c1376a0bfa

Observation d54b3c52-6a9c-42b0-9fd1-4079138a73a4 · outbound

This paper cites Scaper: A library for soundscape synthesis and augmentation.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes Scaper: A library for soundscape synthesis and augmentation

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:05:47.893714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T01:05:44.843781Z digest=sha256:fb6fd4e38140ec0cc742e39e2eaf7933d743ea222cff224d8f647e1760743758

Observation aa434d91-4f9b-4397-8fd1-5e44eef43cdd · outbound

This paper cites Masked Autoencoders that Listen.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes Masked Autoencoders that Listen

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:44.339061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:44.339061Z digest=sha256:b8ecf32a6b29e784a542b0d24406b41fca735646ecfb70ec75b1a046170fcdc6

Observation 61012109-f531-4814-87f9-741cb0120e8d · outbound

This paper cites AST: Audio Spectrogram Transformer.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes AST: Audio Spectrogram Transformer

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:44.293487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:44.293487Z digest=sha256:85f995afd8a79463a5728632b11de7f43fe2dec70c7b34ae0c1ce8afae3fdb6f

Observation 94486ff4-66e9-43fc-8068-e98be5dde473 · outbound

This paper cites MAE-AST: Masked Autoencoding Audio Spectrogram Transformer.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes MAE-AST: Masked Autoencoding Audio Spectrogram Transformer

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:44.036582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:44.036582Z digest=sha256:ee32d80f715140c687d8aea750bb65c58f6a7158e456219225ab4447a2dfbe28

Observation 794bdb0a-bd06-4ee3-9d74-ceb86def731c · outbound

This paper cites BYOL for Audio: Self-Supervised Learning for General-Purpose Audio Representation.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes BYOL for Audio: Self-Supervised Learning for General-Purpose Audio Representation

Reference 2021

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T01:05:46.163420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T01:05:44.628445Z digest=sha256:9672fd4ab2b0aa1c8408b606ba64f15daddde9a363aff088ad49daae51427569

Observation 0cb4f1d3-b331-46e7-8bf8-69e15104e12f · outbound

This paper cites EAT: Self-Supervised Pre-Training with Efficient Audio Transformer.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes EAT: Self-Supervised Pre-Training with Efficient Audio Transformer

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:44.124592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:44.124592Z digest=sha256:4fa98b41ccf704f95ca72bf45d4ec9fdd763dc9f454c0702a10a0224c357c869

Observation 8712af36-8e93-42c1-9bc3-40c4f46c2d24 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:44.165961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:44.165961Z digest=sha256:865cd59e38ebc3645f6df9956386471ec0e3e199b1c04d5dbec4d3986950f200

Observation 4e0cbe5f-1dd4-4ca3-97cd-d682b4c8e7f0 · outbound

This paper cites A-JEPA: Joint-Embedding Predictive Architecture Can Listen.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes A-JEPA: Joint-Embedding Predictive Architecture Can Listen

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:44.234503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:44.234503Z digest=sha256:955aa4145209cbd9356d2123cd3d97d9a80a5b888ed81708eba87aaa9a80ba5f

Observation 5031b894-4e5b-4ec3-bdfe-42f319320ced · outbound

This paper cites URL http://dx.doi.org/10.1109/TASLP.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes URL http://dx.doi.org/10.1109/TASLP

Reference 9304

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:43.994900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:43.994900Z digest=sha256:a726b487b30bf73a1ff68c78e3307eeb8c0950c98796cc7cf3b7226b436accf6

Pith citing papers

Observation d3e344d7-7269-4e10-b44c-5fcf151a06f9 · inbound

EnvTriCascade: An Environment-Aware Tri-Stage Cascaded Framework for ESDD2 2026 Challenge cites this paper.

EnvTriCascade: An Environment-Aware Tri-Stage Cascaded Framework for ESDD2 2026 Challenge SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:47:52.074583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T23:45:03.154037Z digest=sha256:a28a8b6c8f681f1a289a333a150b22857624d254d5b5782a9ea6362185528724

Observation 2fc140d0-23bd-4548-9c8a-8b924dd0d98c · inbound

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning cites this paper.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:39.911902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:385bd54441248a601161fd40922575a148e834d3baaec618b726aaccddc72d3b