Pith. sign in

Paper Citation Record · LEDGER

FLAM: Frame-Wise Language-Audio Modeling

As of 22 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 2 inbound Pith citation observations for arXiv:2505.05335.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.05335 v2

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:14:57.714512Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T10:23:37.367700Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T21:55:00.563579Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 967a5c6c-c212-49f0-a4a7-7c49e9fabb8d · outbound

This paper cites keyword, tag.

FLAM: Frame-Wise Language-Audio Modeling keyword, tag

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:14:57.840610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:14:57.709975Z digest=sha256:b054e31715e3246b9b80af8bf9d4838bac994d6cf54dffe469ed62f2b66e12b1

Observation c90fbfa0-e57c-493d-b1fc-ad9a8689d44d · outbound

This paper cites Clap learning audio concepts from natural language su- pervision.

FLAM: Frame-Wise Language-Audio Modeling Clap learning audio concepts from natural language su- pervision

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T23:14:57.659853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:14:57.659853Z digest=sha256:9bd8f6195d4b6aa2a639c5d1fc755a24647dc11c82ec75ee766a0b207b282e82

Observation 2f8aa9a3-8b8e-4f56-a02f-6fcddb605a64 · outbound

This paper cites P., Fonseca, E., Jansen, A., Liu, C., Moore, R.

FLAM: Frame-Wise Language-Audio Modeling P., Fonseca, E., Jansen, A., Liu, C., Moore, R

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:14:57.902565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:14:57.663216Z digest=sha256:b2cd1250c01c721ab05711bca8f2b889db26b75f1e4d3aee658e95fe2e062a99

Observation dc47e202-dfa9-4f5a-b9c3-80092f6f0dd9 · outbound

This paper cites Mixtral of Experts.

FLAM: Frame-Wise Language-Audio Modeling Mixtral of Experts

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T23:14:57.670465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:14:57.670465Z digest=sha256:a354cd8142f8adce90e7f3f4184598bc112b838a0d3e0c52d06d1b75738baabf

Observation b80bfbc1-7662-42ea-b0ca-ea5e83ff1bab · outbound

This paper cites D., Kim, B., Lee, H., and Kim, G.

FLAM: Frame-Wise Language-Audio Modeling D., Kim, B., Lee, H., and Kim, G

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T23:14:57.674591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:14:57.674591Z digest=sha256:7b9a222b962a2988404eafda6a2bc6708cd72034d0ba92dac82d6b56d32764af

Observation e68316b9-b23e-4c42-b5ba-7ba129a8b19d · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

FLAM: Frame-Wise Language-Audio Modeling RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T23:14:57.677909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:14:57.677909Z digest=sha256:c2698703d216a2b1a8e53e3676146f7d6aeb00a8cb156c74de0a3772101a72b4

Observation 4db5f27a-8455-499a-a68d-67a314c8fd55 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

FLAM: Frame-Wise Language-Audio Modeling Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T23:14:57.681336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:14:57.681336Z digest=sha256:33f4014746508cd05e55050e13aacd7b3a93c1eed43ca4dfd7c8ff313c011f72

Observation 81bd6d2f-1260-43bc-aad7-f341b9e331be · outbound

This paper cites Sound event detection in synthetic domestic environments.

FLAM: Frame-Wise Language-Audio Modeling Sound event detection in synthetic domestic environments

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:14:57.871860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:14:57.685805Z digest=sha256:7a85068b284cd45c30dd1ea5137139a6f3fbdc587c60bc8f9075eca28b4b414c

Observation 9a847350-0a3b-47f8-a3f7-77897add420e · outbound

This paper cites P., and Salamon, J.

FLAM: Frame-Wise Language-Audio Modeling P., and Salamon, J

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T23:14:57.692841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:14:57.692841Z digest=sha256:f144675f81f8ea16dfad31836e1dfe2924742b68781ebe0f51f844b85a2f4272

Observation 12f7623d-c9a0-4803-b172-a3bc76e3cf64 · outbound

This paper cites Large-scale contrastive language- audio pretraining with feature fusion and keyword-to- caption augmentation.

FLAM: Frame-Wise Language-Audio Modeling Large-scale contrastive language- audio pretraining with feature fusion and keyword-to- caption augmentation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:14:57.853356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:14:57.697076Z digest=sha256:768c1b9cd8ad9d812ff339b247dac17a56ce5f923b02f22be2f9892e84cd3fac

Observation 732d6036-e9da-45dc-aad2-a9ff31ac24fd · outbound

This paper cites Towards Weakly Supervised Text-to-Audio Grounding.

FLAM: Frame-Wise Language-Audio Modeling Towards Weakly Supervised Text-to-Audio Grounding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T23:14:57.700621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:14:57.700621Z digest=sha256:c03b6d526548da1402a2411acbf12912bd06f33313b7c2e44b0814929b2b3612

Observation 80215732-3d7b-4d1e-b4dd-04592a0a027c · outbound

This paper cites T-CLAP: Temporal-Enhanced Contrastive Language-Audio Pretraining.

FLAM: Frame-Wise Language-Audio Modeling T-CLAP: Temporal-Enhanced Contrastive Language-Audio Pretraining

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T23:14:57.704987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:14:57.704987Z digest=sha256:0a5e8a78aff413c85e91d0b5d210a2fe616325d1ac84b63eee22f0ca6942a721

Observation c7354fde-6d19-4af0-828f-a21dbcc5c5b4 · outbound

This paper cites an unresolved cited work.

FLAM: Frame-Wise Language-Audio Modeling Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:14:57.828514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:14:57.714512Z digest=sha256:7f43f0a1f13697208e837afc3d37e4b2f9d3138f1bbf87415086e873720a55ef

Observation 79de48b6-c08d-4bde-a116-e7b818a4ac28 · outbound

This paper cites Clotho: An audio 9 FLAM: Frame-Wise Language-Audio Modeling captioning dataset.

FLAM: Frame-Wise Language-Audio Modeling Clotho: An audio 9 FLAM: Frame-Wise Language-Audio Modeling captioning dataset

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:14:57.931363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:14:57.651200Z digest=sha256:99f9bd15e5b8cc0e5b47cbbd700ed1e8849522a3827d8f892ed2ea74a280d64e

Observation 13c7ae50-4fd3-4ded-87a9-f763dcb78a89 · outbound

This paper cites Mean teacher convolution system for dcase 2018 task.

FLAM: Frame-Wise Language-Audio Modeling Mean teacher convolution system for dcase 2018 task

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:14:57.890929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:14:57.667433Z digest=sha256:7eff018306c3b3d7adcac0570eec3e481a8460b793ec316d3977198a6fd014e9

Observation 606fde2d-a84f-49f3-8823-1c3621ec4a6c · outbound

This paper cites Threshold independent evaluation of sound event detection scores.

FLAM: Frame-Wise Language-Audio Modeling Threshold independent evaluation of sound event detection scores

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:14:57.920000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:14:57.656329Z digest=sha256:e6d0cb17e0e981f1b9077fe2b1a02633603089b9ac94109eecdf422ff50c7925

Observation 998f6195-1d0b-4116-9d24-ae86f8738651 · outbound

This paper cites Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection.

FLAM: Frame-Wise Language-Audio Modeling Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:14:57.943055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:14:57.638635Z digest=sha256:de40505cf2ba203ac0395b16790f0806fd7a7e5b8a52d9eac2dee9d8c2cc79c7

Observation 78374fd8-036b-4c55-bc37-08dacc6bbb98 · outbound

This paper cites DCASE 2024 Task 4: Sound Event Detection with Heterogeneous Data and Missing Labels.

FLAM: Frame-Wise Language-Audio Modeling DCASE 2024 Task 4: Sound Event Detection with Heterogeneous Data and Missing Labels

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T23:14:57.642924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:14:57.642924Z digest=sha256:0ac4f7769672fce567c7ae16a7f26978425a3d7ea531417efe37d077b43c08e2

Observation c97f3cde-cbda-41f2-8853-7725e4642df8 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

FLAM: Frame-Wise Language-Audio Modeling Representation Learning with Contrastive Predictive Coding

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T23:14:57.689004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:14:57.689004Z digest=sha256:87ed72afc821695de384e95d464ed9a44832a4c16955da6dfa13b640a1411c3b

Observation 0f0b86cf-312a-40b0-8da0-d88b1a84335d · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

FLAM: Frame-Wise Language-Audio Modeling BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T23:14:57.647008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:14:57.647008Z digest=sha256:e8fbc6ca7e5a142add6e23ed8085c89430ce6e5982824c014b8497f7cde8f1f7

Pith citing papers

Observation f5ee9494-cf1b-4d4a-9f89-20a716f4fc31 · inbound

Melody-Lyrics Matching with Contrastive Alignment Loss cites this paper.

Melody-Lyrics Matching with Contrastive Alignment Loss FLAM: Frame-Wise Language-Audio Modeling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T10:23:37.367700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:23:37.367700Z digest=sha256:89946aefa01f85edd8e50381c90d1f747cff9f3fc97ef6006473837b1b9673b1

Observation f3f7fd21-81b5-464d-82dc-1d6b272b7750 · inbound

Auditory Intelligence: Understanding the World Through Sound cites this paper.

Auditory Intelligence: Understanding the World Through Sound FLAM: Frame-Wise Language-Audio Modeling

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:55:00.654739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T21:55:00.164281Z digest=sha256:df03b7c9f67fd2288a6edbf265c20b9a3516b793102fc29c25b4dec6386c2b42