Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T22:40:26.157629Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 3 inbound Pith citation observations for arXiv:2502.04476.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T22:40:26.157629Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:28:48.510861Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T19:12:07.238377Z
68 of 68 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9b48841a-6c0f-453b-b57a-75b117973ebc · outbound
ADIFF: Explaining audio difference using natural language write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22469d2e-87e4-4c30-af75-7f4bafb529c3 · outbound
ADIFF: Explaining audio difference using natural language Getting vit in shape: Scaling laws for compute-optimal model design
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3182d79d-073c-4234-886f-a0b8b13277bf · outbound
ADIFF: Explaining audio difference using natural language Spice: Semantic propositional image caption evaluation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a4f7567f-4375-41f4-af7e-4a13897be366 · outbound
ADIFF: Explaining audio difference using natural language METEOR : An automatic metric for MT evaluation with improved correlation with human judgments
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 211752ad-a296-4b7c-9bca-249d42484cee · outbound
ADIFF: Explaining audio difference using natural language PaliGemma: A versatile 3B VLM for transfer
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b195f82-ac43-4b5d-a65e-f46f804eb9d4 · outbound
ADIFF: Explaining audio difference using natural language Selm: Enhancing speech emotion recognition for out-of-domain scenarios
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a2f4ae70-1ccb-4dad-8caa-d228ed5cf4d1 · outbound
ADIFF: Explaining audio difference using natural language Audio quality assessment techniques—a review, and recent developments
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c156397e-419f-4024-85ef-95d5055d2982 · outbound
ADIFF: Explaining audio difference using natural language Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2e54e352-653f-47f8-bb6b-edce2e3f3c7a · outbound
ADIFF: Explaining audio difference using natural language Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 646f3069-7740-4295-ba76-49001678b1b1 · outbound
ADIFF: Explaining audio difference using natural language Pengi: An audio language model for audio tasks
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9c489b94-280a-4e2c-9dc2-8ca8ec8cdfc0 · outbound
ADIFF: Explaining audio difference using natural language Audio Retrieval with WavText5K and CLAP Training
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 400897e4-5812-4823-bff3-3ad20e4676e4 · outbound
ADIFF: Explaining audio difference using natural language Pam: Prompting audio-language models for audio quality assessment
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4abf38a1-108b-4ffe-b47d-21fb429e4bc7 · outbound
ADIFF: Explaining audio difference using natural language Audio Entailment: Assessing Deductive Reasoning for Audio Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3e44766-a8a7-4457-b81c-6116a4cb5688 · outbound
ADIFF: Explaining audio difference using natural language Domain adaptation for contrastive audio-language models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e167ca20-ddec-4085-ac54-beb23d60b09b · outbound
ADIFF: Explaining audio difference using natural language Automated audio captioning with recurrent neural networks
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e98c71dc-6f5e-4296-b97e-d85f0bf64096 · outbound
ADIFF: Explaining audio difference using natural language Clotho: an audio captioning dataset
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9df2d0b1-07f1-439a-bc77-ccd1d3671679 · outbound
ADIFF: Explaining audio difference using natural language Clap learning audio concepts from natural language supervision
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85c3e55a-2f65-47ab-b226-5360e2625853 · outbound
ADIFF: Explaining audio difference using natural language Natural language supervision for general-purpose audio representations
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 117f560f-8cbe-4943-bae9-bb1e892573c1 · outbound
ADIFF: Explaining audio difference using natural language Fsd50k: an open dataset of human-labeled sound events
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf32190f-e38f-4606-9f2a-c8bb0dbb9f05 · outbound
ADIFF: Explaining audio difference using natural language Gemmeke, Daniel P
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0617f60-88eb-4658-89b5-6cd0b76d2cf5 · outbound
ADIFF: Explaining audio difference using natural language GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e2429e1-f8d7-4bcf-a645-aa8f97baface · outbound
ADIFF: Explaining audio difference using natural language Compa: Addressing the gap in compositional reasoning in audio-language models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d95d4d87-eb38-4820-8e89-48e3dbaf16ef · outbound
ADIFF: Explaining audio difference using natural language Joint audio and speech understanding
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 60ada09a-9b23-4dca-9221-478d5821ca12 · outbound
ADIFF: Explaining audio difference using natural language Liu, Leonid Karlinsky, and James R
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0fa94c8d-612c-45e9-8ed0-fd9d42a86019 · outbound
ADIFF: Explaining audio difference using natural language Clip4idc: Clip for image difference captioning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation efe20142-1389-4633-8b71-e5baddffd82e · outbound
ADIFF: Explaining audio difference using natural language Heller, Benjamin Elizalde, Bhiksha Raj, and Soham Deshmukh
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 15988846-ea06-40a3-9f20-4c47d2874d8b · outbound
ADIFF: Explaining audio difference using natural language Training compute-optimal large language models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f4376b1b-401e-4ba3-be4d-6afeb2cafe46 · outbound
ADIFF: Explaining audio difference using natural language Lora: Low-rank adaptation of large language models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2297d55e-3536-4197-8d43-451eaaf10104 · outbound
ADIFF: Explaining audio difference using natural language Learning to describe differences between pairs of similar images
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da759e54-0ac9-448b-b7c3-ec1b1d2316ab · outbound
ADIFF: Explaining audio difference using natural language Acoustic and auditory phonetics
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 478f2a89-9134-4063-9c3c-016a1cf2ee85 · outbound
ADIFF: Explaining audio difference using natural language Deductive reasoning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d676bbb7-0d52-4e41-9c18-00d0e8ed64fe · outbound
ADIFF: Explaining audio difference using natural language Scaling Laws for Neural Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e217a7d3-cf79-44a3-9687-7f2851cad396 · outbound
ADIFF: Explaining audio difference using natural language AudioCaps: Generating Captions for Audios in The Wild
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a218f7da-772e-477a-a637-1ff59cfcab33 · outbound
ADIFF: Explaining audio difference using natural language Adam: A Method for Stochastic Optimization
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2553c111-5dd3-4839-8e55-222eaf552465 · outbound
ADIFF: Explaining audio difference using natural language Audio retrieval with natural language queries: A benchmark study
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1cd7dbcc-d41b-4288-baef-a9662a65e6d6 · outbound
ADIFF: Explaining audio difference using natural language Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c0ce7e19-fe0a-4873-acd7-97fb6c2d2ae8 · outbound
ADIFF: Explaining audio difference using natural language Digital audio forensics: a first practical evaluation on microphone and environment classification
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b388490f-25eb-4608-9fd5-71afe3bd2f50 · outbound
ADIFF: Explaining audio difference using natural language Audiogen: Textually guided audio generation
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a13066a4-e9f4-4cdc-bae1-9b7b1be7f8f7 · outbound
ADIFF: Explaining audio difference using natural language Understanding sounds, missing the questions: The challenge of object hallucination in large audio-language models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2ef30c04-b77c-4d80-b6a0-b9cccf821b30 · outbound
ADIFF: Explaining audio difference using natural language Clotho-aqa: A crowdsourced dataset for audio question answering
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 439949ec-7a75-41d3-b2a2-89cf996f6618 · outbound
ADIFF: Explaining audio difference using natural language Audioldm: Text-to-audio generation with latent diffusion models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1957f0d7-8389-4118-8a89-cff98c696da8 · outbound
ADIFF: Explaining audio difference using natural language Improved image captioning via policy gradient optimization of spider
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation de25405b-8187-4cb4-beae-f8128b59dde7 · outbound
ADIFF: Explaining audio difference using natural language Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8fb10fba-90b6-4da9-80c6-f03405677ace · outbound
ADIFF: Explaining audio difference using natural language Automated audio captioning: An overview of recent progress and new challenges
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 806d8eb2-2ff4-4031-b4ce-7062e8ff3170 · outbound
ADIFF: Explaining audio difference using natural language Diverse audio captioning via adversarial training
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7026ab62-d8d1-4435-a88e-f6ba4dd2400b · outbound
ADIFF: Explaining audio difference using natural language Towards generating diverse audio captions via adversarial training
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e84060fd-f5b7-4092-a53b-af18ca16ac6e · outbound
ADIFF: Explaining audio difference using natural language Plumbley, Yuexian Zou, and Wenwu Wang
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5236dadf-4852-4991-8817-d122aa7d8f9b · outbound
ADIFF: Explaining audio difference using natural language ClipCap: CLIP Prefix for Image Captioning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96103ec6-861b-43b8-95e8-f200c9d660b7 · outbound
ADIFF: Explaining audio difference using natural language Diversity and bias in audio captioning datasets
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 33e0debe-e788-4feb-86ea-003d9f9f5388 · outbound
ADIFF: Explaining audio difference using natural language On the Audio Hallucinations in Large Audio-Video Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8eb14dbe-4f4c-4bd5-a58d-496414828426 · outbound
ADIFF: Explaining audio difference using natural language Bleu: a method for automatic evaluation of machine translation
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09858cef-9d95-4f5b-83bb-ee6cf5478676 · outbound
ADIFF: Explaining audio difference using natural language Robust change captioning
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f5b8dc08-8daf-4645-8df1-c926c1144dd5 · outbound
ADIFF: Explaining audio difference using natural language Robust speech recognition via large-scale weak supervision
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fee5dd68-f9b9-472f-806e-0103bcca5156 · outbound
ADIFF: Explaining audio difference using natural language Scaling Language Models: Methods, Analysis & Insights from Training Gopher
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcd93b80-6872-4fb8-932b-3f4ca73328d5 · outbound
ADIFF: Explaining audio difference using natural language Acoustic phonetics, volume 30
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9cc831c7-c9cc-4502-ab69-c576fa1ebffd · outbound
ADIFF: Explaining audio difference using natural language Audio difference captioning utilizing similarity-discrepancy disentanglement
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 13cb819c-e4a3-4cbd-8452-abd914c08621 · outbound
ADIFF: Explaining audio difference using natural language Extending large language models for speech and audio captioning
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4938af6-a8bb-464f-ab9b-f02c2cb5efdb · outbound
ADIFF: Explaining audio difference using natural language SALMONN : Towards generic hearing abilities for large language models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 965bd9a9-cc83-4345-bfcb-3fcc3b9990a4 · outbound
ADIFF: Explaining audio difference using natural language Tzanetakis and P
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a4459a2d-ab10-49d0-ae9a-a9b11a8b33b4 · outbound
ADIFF: Explaining audio difference using natural language Cider: Consensus-based image description evaluation
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b4eb0bf4-1abe-40b6-af8a-813348c697ae · outbound
ADIFF: Explaining audio difference using natural language Beats-based audio captioning model with instructor embedding supervision and chatgpt mix-up
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b7630099-b4c1-491c-b35c-13ab3aca2395 · outbound
ADIFF: Explaining audio difference using natural language Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation febee63c-35b8-423f-b8af-607bdb290feb · outbound
ADIFF: Explaining audio difference using natural language Image difference captioning with pre-training and contrastive learning
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dd0e0ac3-3aec-4830-bb65-d61fe57f0891 · outbound
ADIFF: Explaining audio difference using natural language Pre-training language models for comparative reasoning
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1fda5dd7-caa2-47a3-b3a1-5f4a0fa709f4 · outbound
ADIFF: Explaining audio difference using natural language NaRLE: Natural Language Models using Reinforcement Learning with Emotion Feedback
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a76c2d0c-01c4-49e8-a3df-26c952f5f910 · outbound
ADIFF: Explaining audio difference using natural language @esa (Ref
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 835026b5-32b0-4b7a-859b-071a2e393de6 · outbound
ADIFF: Explaining audio difference using natural language Unresolved cited work
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4d37c76-28ef-46c8-a52d-8186c2769fe2 · outbound
ADIFF: Explaining audio difference using natural language or ``caption the second audio
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aceb2984-3db4-4261-9616-02507bf01c93 · inbound
Breaking the Barriers of Text-Hungry and Audio-Deficient AI ADIFF: Explaining audio difference using natural language
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1936ce45-52aa-4d17-bf6c-9eaf66ce0390 · inbound
A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations ADIFF: Explaining audio difference using natural language
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97cf3465-2f7c-466e-885c-9a779fc7eaa6 · inbound
MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing ADIFF: Explaining audio difference using natural language
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.