Pith. sign in

Paper Citation Record · LEDGER

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding

As of 18 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 2 inbound Pith citation observations for arXiv:2505.13880.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.13880 v3

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:11:18.684100Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:11:18.478312Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T10:16:56.616742Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved29
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6b801ed0-5f3c-4dd9-9a3e-55e117ebf0dd · outbound

This paper cites Their ability to reason and generate coherent outputs stems from ex- tensive pretraining on large-scale text data.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Their ability to reason and generate coherent outputs stems from ex- tensive pretraining on large-scale text data

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:19.398333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:11:18.467078Z digest=sha256:bd6d35accd24ce95d92b24fac80ddc0502ab7fddb1141286cd507adaabaca392

Observation fa18aeb5-ef61-4179-908c-28c56e5dcd78 · outbound

This paper cites U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:18.478312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:18.478312Z digest=sha256:0e078b3d6563f1f619402bcbfb77f5c32ced97909c96c80a37d067f5f27c1f9b

Observation 949008e4-e063-4e42-8ca0-7ab641bd9744 · outbound

This paper cites Data Specifications U-SAM is trained and evaluated on diverse datasets across mul- tiple tasks.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Data Specifications U-SAM is trained and evaluated on diverse datasets across mul- tiple tasks

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T20:11:19.381669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:11:18.485905Z digest=sha256:e1bcc32e22e044e365b4b766435b44d75cf2ec80b7fd8da69ebf4589f89d7b83

Observation 72900f29-989a-40e0-a6ef-15acc96fef61 · outbound

This paper cites By lever- aging LoRA for efficient fine-tuning and incorporating TAPM and SACLM, U-SAM achieves robust audio-text alignment.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding By lever- aging LoRA for efficient fine-tuning and incorporating TAPM and SACLM, U-SAM achieves robust audio-text alignment

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:19.365169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:11:18.492460Z digest=sha256:568003e2b8b758db4071b39b570f6efcb8d3003ab30b80cb7e909c66888adda1

Observation 2858148a-4edf-44ef-9f2c-f134edaedd8c · outbound

This paper cites A Survey of Large Language Models.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding A Survey of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:18.498336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:18.498336Z digest=sha256:413ee1a50c588033ddd2939565f14cc236bc0134f0812ac41baa518f42ab5c24

Observation a5d7db24-b4c0-4d93-8c14-49a199065a63 · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Flamingo: a visual language model for few-shot learning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:19.348780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:11:18.503721Z digest=sha256:86f72600d6c8354269a6ba17c2761746b2d448493db44b2c32ad826c843df555

Observation 1a50dbf1-11b9-4e69-a8af-e4c917f840dc · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:18.508605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:18.508605Z digest=sha256:24a523f161e2156a670ab2c89a26560aed29bdb1a113f871c28e92e70dc5a4d9

Observation cf4463d5-c135-45f6-9890-36099b6f7009 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:18.513801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:18.513801Z digest=sha256:0e670cab98711540f48ed3d0957631e920eea7bac3ec7b8af2f321baf3f00705

Observation 1f082a78-9ad0-464e-b182-188f91b5ef80 · outbound

This paper cites Clap learning audio concepts from natural language supervision,.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Clap learning audio concepts from natural language supervision,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:18.518646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:18.518646Z digest=sha256:9026eb327135c6c8d539d52ea0e30d178de8a3c5f6021da6e202f61261c4846b

Observation e551f007-3d86-4269-aaf7-c26bff2feb44 · outbound

This paper cites Audio Retrieval with WavText5K and CLAP Training.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Audio Retrieval with WavText5K and CLAP Training

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:18.523782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:18.523782Z digest=sha256:8a0d5f703d2781c12c46e65cf769ac23b9cac692925e0e0b02ff5b4e0b64ddda

Observation f7c8e2ed-226a-4d1d-9733-a37bb13c48b6 · outbound

This paper cites Improving text-audio retrieval by text-aware attention pooling and prior matrix revised loss,.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Improving text-audio retrieval by text-aware attention pooling and prior matrix revised loss,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:19.309587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:11:18.528678Z digest=sha256:63577816413c4628c88eefb16a650180e66db0a21f414941aae69e696005ab1f

Observation cffd971d-ed18-4443-8268-d738af0d7b82 · outbound

This paper cites T-CLAP: Temporal-Enhanced Contrastive Language-Audio Pretraining.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding T-CLAP: Temporal-Enhanced Contrastive Language-Audio Pretraining

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:18.534614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:18.534614Z digest=sha256:2c4694aee8ee9a64a077209f7f541a4e781b81247c3a89de3ad5d9c37fca2a76

Observation bf46e0a1-1032-471e-8294-265091980d60 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:18.539951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:18.539951Z digest=sha256:f663f66eb921695aeb934888acb950c0d26caea1bee12b3e8334442879bccaf2

Observation 69549a36-397c-4536-826f-04325f318506 · outbound

This paper cites Qwen2-Audio Technical Report.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Qwen2-Audio Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:18.545140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:18.545140Z digest=sha256:698310dcb0dd665df2d563d615d507ba7b50278956dcd1668c68d998889f6f1d

Observation a125b666-8ec8-4d97-9aa1-50b8fcb6568f · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Robust Speech Recognition via Large-Scale Weak Supervision

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:18.549688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:18.549688Z digest=sha256:c0b8c270642559c08cc8a04bd1c90f78ef03eae9502ebd40988f858e34355e09

Observation e25bff77-2590-460e-a26a-3723851efe18 · outbound

This paper cites Pengi: An audio language model for audio tasks,.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Pengi: An audio language model for audio tasks,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:18.554336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:18.554336Z digest=sha256:efb4eb852c6b4551e84a914bc73db93cad9ad283f2bfe580332d2c37df29953b

Observation 94127408-9008-48ba-b05f-1a1bce4b8d67 · outbound

This paper cites Listen, Think, and Understand.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Listen, Think, and Understand

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:18.558377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:18.558377Z digest=sha256:2d2f52cbc31d1ab4cbe913c77ca89819395756fddfe635fd53da14ac16b31ee5

Observation c2a38829-a7f6-4068-bbc6-558431cac7c4 · outbound

This paper cites GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:18.562700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:18.562700Z digest=sha256:e8cab37b177c75843b3aafe2d882acf8f1f5fa8c2cc556ebd8bfb5dbf5a6aa03

Observation 8802c7f4-0186-476a-b0fe-cbb3e645fb87 · outbound

This paper cites AST: Audio Spectrogram Transformer.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding AST: Audio Spectrogram Transformer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:18.567709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:18.567709Z digest=sha256:877595b6ffa37d5c7f1eb1db774e3f66caad020f10203e3bff4986dfc99cefdb

Observation 211f1bba-6f45-45f6-8c2e-f5268134fef3 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:18.573278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:18.573278Z digest=sha256:fc3f96496c162c6b50d99ed9fb46da5314e542b9ff155468a87e3754638e15c4

Observation a362eff2-7e00-406b-9f0a-7a568ed037f8 · outbound

This paper cites BEATs: Audio Pre-Training with Acoustic Tokenizers.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:18.578676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:18.578676Z digest=sha256:1a42d2da123035e2c30d1f0f638c8ac5952a0b2ee7a3da6eb9e6d7def9ba158d

Observation 35cb5431-fdf8-4f45-8f21-716181c6a114 · outbound

This paper cites Why do speech language models fail to generate semantically coherent outputs? a modality evolving perspective,.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Why do speech language models fail to generate semantically coherent outputs? a modality evolving perspective,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:18.585904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:18.585904Z digest=sha256:ec96de30e61cab3e7dbe92cd9f1cf60e8404c8ae9f8f9fb36ad832d65c694dfe

Observation a5296f60-2df1-47ed-b05a-5e2a3ae8e6cf · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:18.591232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:18.591232Z digest=sha256:0e5061017e600eae7ddd6198b31f3a6c0273ef168985a9bd1e864036fc8068e1

Observation 2414703c-6999-4012-8c63-c22a39ae5c14 · outbound

This paper cites Linguistic-aware patch slimming framework for fine-grained cross-modal alignment,.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Linguistic-aware patch slimming framework for fine-grained cross-modal alignment,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:19.281828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:11:18.596436Z digest=sha256:e64aabbf6e18fdf685ff52ce3839d397f0301e9041ac3de599c0b6569fab199b

Observation be269beb-3836-495f-8c8b-4e0ce6276e12 · outbound

This paper cites Ced: Con- sistent ensemble distillation for audio tagging,.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Ced: Con- sistent ensemble distillation for audio tagging,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:19.264860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:11:18.601847Z digest=sha256:d386256faa954be2fc3d504ea5c72f0d8d957212e701c23a1e9efb637fad8f8f

Observation 5c6e7de2-3746-4edb-a482-73574fd05956 · outbound

This paper cites Mert: Acoustic music under- standing model with large-scale self-supervised training,.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Mert: Acoustic music under- standing model with large-scale self-supervised training,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:19.248642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:11:18.607695Z digest=sha256:4b92827c75e1d5dc135876d60251ab083162c9a2f9316380714ad483e7ed3e7a

Observation 12cce0b4-d670-4985-b429-c53551101ef2 · outbound

This paper cites A Survey on Mixture of Experts in Large Language Models.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding A Survey on Mixture of Experts in Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:18.614793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:18.614793Z digest=sha256:aa93401a329214a4f0182c3c112dab87e780b85d59e8b2cd04800ec2891f9929

Observation 6533d0b6-0cca-4108-9ab5-0eff44eb098d · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding LoRA: Low-Rank Adaptation of Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:18.620823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:18.620823Z digest=sha256:cbaac5ba07afa1ebc7a598ca0f25d86b9fecdde646f408d31e790ac2d0ac6aac

Observation 7167e6a3-e5cf-421a-aa76-011e4b1f039a · outbound

This paper cites Triplet loss in siamese network for object tracking,.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Triplet loss in siamese network for object tracking,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:19.232556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:11:18.625683Z digest=sha256:6a99dad7c0feaa9a509e4b158c2daae63dc9c5e574124ce338005d6ecfa32a57

Observation e29a17fd-129a-4302-92db-4a12a8bc118a · outbound

This paper cites Lib- rispeech: an asr corpus based on public domain audio books,.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Lib- rispeech: an asr corpus based on public domain audio books,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:18.630915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:18.630915Z digest=sha256:8a9a1ca2ddb28dd4ab11efa3d49d06b545ea47ca5dc449f9a9bf2d3e3712db8b

Observation 553cc07a-e3be-4f57-9c42-2545c5dff3fd · outbound

This paper cites CoVoST 2 and Massively Multilingual Speech-to-Text Translation.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding CoVoST 2 and Massively Multilingual Speech-to-Text Translation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:18.635718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:18.635718Z digest=sha256:2de6bf718f702f73587bcea3aca1d71470dc87463a344730d724a5785549ee2d

Observation 9a53424c-e508-457b-9fbd-452912839637 · outbound

This paper cites Audiocaps: Generat- ing captions for audios in the wild,.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Audiocaps: Generat- ing captions for audios in the wild,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:18.644494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:18.644494Z digest=sha256:c21c650376c820c9bf90207de0f23c730b7ca4e16dff7f7f9a7822b7364a1090

Observation b9967068-d8f7-4167-aad6-c684e2703bf4 · outbound

This paper cites Clotho: An audio cap- tioning dataset,.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Clotho: An audio cap- tioning dataset,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:18.651096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:18.651096Z digest=sha256:527a00462f3120e50d5c413d7dce2c05610157b61cf693bc0654ed822c69f8ca

Observation 7537a4dd-5b57-47b2-a7c4-b6f97feb39bc · outbound

This paper cites Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multimodal research,.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multimodal research,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:18.658717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:18.658717Z digest=sha256:629c5da45b541c443ab699165bca8d39df7f20173c94e1724073373b3259c56c

Observation 4a889230-92b2-48fc-a6c2-4d9c44fb1199 · outbound

This paper cites MusicLM: Generating Music From Text.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding MusicLM: Generating Music From Text

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:18.665641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:18.665641Z digest=sha256:e2766553a5d44a5aaad5e54b44eec7c01d7b12970ad0c4406108ecbd73caf883

Observation 10244416-cce3-4e67-b80b-017496249c48 · outbound

This paper cites Decoupled Weight Decay Regularization.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Decoupled Weight Decay Regularization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:18.672931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:18.672931Z digest=sha256:7b03f63d6ab258c1d3f7b528aaa5d5dbf3f06fa5e12d04292ea4ea2cd9fd512d

Observation e13d6ade-9802-4d6d-aa2e-38110ddcb4c5 · outbound

This paper cites LP-MusicCaps: LLM-Based Pseudo Music Captioning.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding LP-MusicCaps: LLM-Based Pseudo Music Captioning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:18.678742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:18.678742Z digest=sha256:8bf2b96008eef3905c0fdf965878c479e9daee202a4c4170e8b752560046affc

Observation f75d3033-feb7-48f1-9d46-f9d3783d86bc · outbound

This paper cites Audiogpt: Understanding and generating speech, music, sound, and talking head,.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Audiogpt: Understanding and generating speech, music, sound, and talking head,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:18.684100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:18.684100Z digest=sha256:9a7a09cb1e18bf37c26360b383c288ff96da48c877182a042450b19612b29047

Pith citing papers

Observation fa18aeb5-ef61-4179-908c-28c56e5dcd78 · inbound

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding cites this paper.

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:18.478312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:18.478312Z digest=sha256:0e078b3d6563f1f619402bcbfb77f5c32ced97909c96c80a37d067f5f27c1f9b

Observation 44d951d0-8696-4fcf-921d-c7575a48ecc7 · inbound

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction cites this paper.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:16:56.749500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:16:56.294362Z digest=sha256:1f2474c7862f6b0a72bc90ba95180d234ce98f73f9f5fdbe1a5dcd71b72a4625