Pith. sign in

Paper Citation Record · LEDGER

Efficient Training of Audio Transformers with Patchout

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2110.05069.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2110.05069 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:09:57.929534Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T18:40:02.620646Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a066d024-477c-4e68-9145-818d1990d3b7 · inbound

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet cites this paper.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Efficient Training of Audio Transformers with Patchout

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:57.929534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:57.929534Z digest=sha256:f53a4902538b84b0cb619ed366d0a8589943fa7bb32bd86289a7a70ec65fcbab

Observation c4f98f10-dac5-48fe-9f54-7f3e2be380a6 · inbound

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes cites this paper.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes Efficient Training of Audio Transformers with Patchout

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:44.433337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:44.433337Z digest=sha256:e1ad36c84e079f2fe45a71be8db7348862dcfc72494b979e2a805e914a1ba83d

Observation cc2e518c-34dc-4078-9479-835127505348 · inbound

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections cites this paper.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Efficient Training of Audio Transformers with Patchout

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:47.983763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:47.983763Z digest=sha256:29d8b217605d76b3d773a727bd7841345ecac166ef7c9ba03000cdab830c9bc5

Observation 923016d2-68d7-4c8a-9f34-35fdbbf9f1cf · inbound

MuseControlLite: Multifunctional Music Generation with Lightweight Conditioners cites this paper.

MuseControlLite: Multifunctional Music Generation with Lightweight Conditioners Efficient Training of Audio Transformers with Patchout

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:19:50.930609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:19:50.930609Z digest=sha256:dac0fc1ee0de855bd01cb8fcd2bd69ae289dd38aca7829f2bdf47549864c8499

Observation 1e94711d-41dc-4b9b-b442-dfcc01a66665 · inbound

Enhancing Stereo Sound Event Detection with BiMamba and Pretrained PSELDnet cites this paper.

Enhancing Stereo Sound Event Detection with BiMamba and Pretrained PSELDnet Efficient Training of Audio Transformers with Patchout

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:57:10.542254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:57:10.542254Z digest=sha256:090300a5b9815ae6aa742a01fae5cca6ab7159649a97b706e409b25a28ff5c4a

Observation 4b28f4b3-37a9-4733-96e7-1918e80fea20 · inbound

Joint Feature and Output Distillation for Low-complexity Acoustic Scene Classification cites this paper.

Joint Feature and Output Distillation for Low-complexity Acoustic Scene Classification Efficient Training of Audio Transformers with Patchout

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T14:31:09.324196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:31:09.324196Z digest=sha256:7fc7ab8546e41c92bb74e22cd2a996391c87a2599af6d3cfea2ad5dcf6a1867a

Observation a8231362-d76c-4f73-a1dc-76c8d930a7af · inbound

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation cites this paper.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Efficient Training of Audio Transformers with Patchout

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.895970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.895970Z digest=sha256:a29a953e5c2d20fd8872b0dde9d6f9d0a3073794526b8691f4a8d2d052ef0262

Observation cc184448-28f0-4ecf-8a6c-d0a68be61084 · inbound

ASAudio: A Survey of Advanced Spatial Audio Research cites this paper.

ASAudio: A Survey of Advanced Spatial Audio Research Efficient Training of Audio Transformers with Patchout

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-05T22:54:55.028080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:54:55.028080Z digest=sha256:3dc75b9fc575e410cb6c673cd3bb6704b06a803b950680381f6635b9458afb0f

Observation 607c1373-cf4d-4ae3-9802-db6a3ac4bef6 · inbound

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts cites this paper.

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts Efficient Training of Audio Transformers with Patchout

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T00:05:47.857768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:05:47.857768Z digest=sha256:ca6591d2371701ce081b51d327bad0bf2f67b266685634c4e7c6c2c6c938ba93

Observation 547da7d8-38ba-4613-8b33-23dff43c8410 · inbound

Adaptive Knowledge Distillation using a Device-Aware Teacher for Low-Complexity Acoustic Scene Classification cites this paper.

Adaptive Knowledge Distillation using a Device-Aware Teacher for Low-Complexity Acoustic Scene Classification Efficient Training of Audio Transformers with Patchout

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T19:27:42.788432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:27:42.788432Z digest=sha256:927b620d32a7c923460bda900b773c35d5b56696ff32093ca812adcf8aa14a54

Observation d4f56042-7847-425a-9a6b-5a373a4f283a · inbound

MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation cites this paper.

MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation Efficient Training of Audio Transformers with Patchout

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:21:06.847274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T08:20:02.986562Z digest=sha256:1485ea4eff01aa6cbdfde208fef9444af42500538310aee6abc7045dcadaca48

Observation d88f42e7-2290-47d5-b122-5dc5d3a2c19f · inbound

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing cites this paper.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Efficient Training of Audio Transformers with Patchout

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:56.830276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:56.830276Z digest=sha256:faef69d7e3fa1aa8cdea7a143cd7c3db540ea3175774866ef66e7e62db1c3017

Observation 43296ace-ee09-4fb4-b75e-52caf66cff98 · inbound

Omni2Sound: Towards Unified Video-Text-to-Audio Generation cites this paper.

Omni2Sound: Towards Unified Video-Text-to-Audio Generation Efficient Training of Audio Transformers with Patchout

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-16T17:28:10.082250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T17:25:38.591071Z digest=sha256:4eea41d713647a3555dada1b9358cf39cec1494d47dcf8cb9cf3b2bf831a0b12

Observation 4bb54ab7-7168-4529-8112-4002e8e33a81 · inbound

Conditional Flow Matching for Visually-Guided Acoustic Highlighting cites this paper.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Efficient Training of Audio Transformers with Patchout

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:13.142515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:13.142515Z digest=sha256:26621685ced0b077a0e09f09f9394f0e2052f5fee03f44edd99c021f1b6acde4

Observation decb39f0-df24-489b-8a5c-d4fa99bd6956 · inbound

Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models cites this paper.

Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models Efficient Training of Audio Transformers with Patchout

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:56:33.493830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T19:53:18.200223Z digest=sha256:0042c8e598a9d3af3009a598b2ec2e1a40cb6bd6dffdf1132c9016b2a7543d64

Observation c1577142-cf88-47c2-b439-badde0b4144e · inbound

ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling cites this paper.

ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling Efficient Training of Audio Transformers with Patchout

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.047377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T09:22:10.483263Z digest=sha256:df5b3ec572b10d719b06d6da84123419bb7bd73f1115db7d9bfacae276d085eb

Observation d52a8f66-5a48-4a3e-b5c7-9c16034491a3 · inbound

BeeVe: Unsupervised Acoustic State Discovery in Honey Bee Buzzing cites this paper.

BeeVe: Unsupervised Acoustic State Discovery in Honey Bee Buzzing Efficient Training of Audio Transformers with Patchout

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:30:55.879342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T03:27:33.149625Z digest=sha256:c2791fbf781ca0dd56afd1e7489e76b842dce88b1698ff5a1640694739177e0d

Observation 8d311385-866b-4523-b396-c7e46adc7ea9 · inbound

SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification cites this paper.

SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification Efficient Training of Audio Transformers with Patchout

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:57.043063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T20:31:49.239866Z digest=sha256:f8aa27cffe1b8ef867c6da48a7d7eebdf81623bd91839d6b2cf4913b084467c5

Observation bd50c096-80ff-447d-bcf8-12c7bf936eb8 · inbound

Live Music Diffusion Models: Efficient Fine-Tuning and Post-Training of Interactive Diffusion Music Generators cites this paper.

Live Music Diffusion Models: Efficient Fine-Tuning and Post-Training of Interactive Diffusion Music Generators Efficient Training of Audio Transformers with Patchout

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-22T03:25:59.043307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T03:24:50.604019Z digest=sha256:df470d7254f0eba026277e86db22e71d9ab28759529a732113ac601a85bf82af

Observation bcdb4ba8-a655-45f7-9fb2-02b094bf5848 · inbound

STAR-VAE: Structured Topology-Aware Regularization for Audio Reconstruction and Generation cites this paper.

STAR-VAE: Structured Topology-Aware Regularization for Audio Reconstruction and Generation Efficient Training of Audio Transformers with Patchout

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:59:50.459909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T07:26:28.527338Z digest=sha256:e07ad670ab3a5eb50e6c0b5179f84431bc7b4dc0cff0c409ec4e9cbb08290de0

Observation 842dbdcd-230f-4d5f-adec-9ccdaf3e1684 · inbound

Real-Time Interactive Music Generation via Data-Free Streaming Consistency Distillation cites this paper.

Real-Time Interactive Music Generation via Data-Free Streaming Consistency Distillation Efficient Training of Audio Transformers with Patchout

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-04T18:40:02.622264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T22:42:57.448084Z digest=sha256:375f836b9193e71c57d5367f8b938d399ee038f6f3d4072c900e3d5111005fea

Observation 3f25018d-04b1-4c8d-957b-3f56b043c333 · inbound

FdAudio: MeanFlow-Anchored Fr\'echet-Distance Post-Training for One-Step Text-to-Audio Generation cites this paper.

FdAudio: MeanFlow-Anchored Fr\'echet-Distance Post-Training for One-Step Text-to-Audio Generation Efficient Training of Audio Transformers with Patchout

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-14T11:52:50.598080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:52:50.598080Z digest=sha256:b1d60cc60782b8887aa05508e30f6da88e266f78ed952af66e6f407ce065f4d4