Pith. sign in

Paper Citation Record · LEDGER

ADIFF: Explaining audio difference using natural language

As of 8 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 3 inbound Pith citation observations for arXiv:2502.04476.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.04476 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:40:26.157629Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:28:48.510861Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T19:12:07.238377Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact6
  • verified fuzzy34
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9b48841a-6c0f-453b-b57a-75b117973ebc · outbound

This paper cites write newline.

ADIFF: Explaining audio difference using natural language write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:25.847277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:25.847277Z digest=sha256:b86f6c75d686292acded4e741a2620a2186e20beacba19269d06386e4362d85b

Observation 22469d2e-87e4-4c30-af75-7f4bafb529c3 · outbound

This paper cites Getting vit in shape: Scaling laws for compute-optimal model design.

ADIFF: Explaining audio difference using natural language Getting vit in shape: Scaling laws for compute-optimal model design

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.798472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.853389Z digest=sha256:24b82e1e70b58b0ab780bcd13f9b5c5c300a9046eb70f36bae768ea5285393cd

Observation 3182d79d-073c-4234-886f-a0b8b13277bf · outbound

This paper cites Spice: Semantic propositional image caption evaluation.

ADIFF: Explaining audio difference using natural language Spice: Semantic propositional image caption evaluation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.783679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.858087Z digest=sha256:a65d9b63cde79998deea696aab97e184e3f61b7368f7587482f86fb2d365b71e

Observation a4f7567f-4375-41f4-af7e-4a13897be366 · outbound

This paper cites METEOR : An automatic metric for MT evaluation with improved correlation with human judgments.

ADIFF: Explaining audio difference using natural language METEOR : An automatic metric for MT evaluation with improved correlation with human judgments

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.767958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.865805Z digest=sha256:4570067b66bc4d58cc516499464bc805ed31b79b3e90fe27250b6eee258c1984

Observation 211752ad-a296-4b7c-9bca-249d42484cee · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

ADIFF: Explaining audio difference using natural language PaliGemma: A versatile 3B VLM for transfer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:25.870623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:25.870623Z digest=sha256:42eff663c732b3e4bb847916292a4ce8ad511683c00feb8106a595b8ce69622f

Observation 6b195f82-ac43-4b5d-a65e-f46f804eb9d4 · outbound

This paper cites Selm: Enhancing speech emotion recognition for out-of-domain scenarios.

ADIFF: Explaining audio difference using natural language Selm: Enhancing speech emotion recognition for out-of-domain scenarios

Reference 6

Resolution
verified exact
doi, observed 2026-08-08T22:40:26.259570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.875537Z digest=sha256:8f2b92b206cf44c8da9a9c0e825d9387cd1735d97531d11247ff5e813af4f952

Observation a2f4ae70-1ccb-4dad-8caa-d228ed5cf4d1 · outbound

This paper cites Audio quality assessment techniques—a review, and recent developments.

ADIFF: Explaining audio difference using natural language Audio quality assessment techniques—a review, and recent developments

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.754434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.880510Z digest=sha256:293457a91e2bcb611614d84135b5402a0dc7dd8b18f88e0d98390407751527a1

Observation c156397e-419f-4024-85ef-95d5055d2982 · outbound

This paper cites Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection.

ADIFF: Explaining audio difference using natural language Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.740547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.885456Z digest=sha256:b80b451a79ca2f70ce78705287564f1cc17dc0e0f11f4887c0ded24490b64ed0

Observation 2e54e352-653f-47f8-bb6b-edce2e3f3c7a · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

ADIFF: Explaining audio difference using natural language Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:25.889755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:25.889755Z digest=sha256:8e01d9637b751deaefc96b215ed41aa46cc9bd25985f7ccea11c9bef08da808f

Observation 646f3069-7740-4295-ba76-49001678b1b1 · outbound

This paper cites Pengi: An audio language model for audio tasks.

ADIFF: Explaining audio difference using natural language Pengi: An audio language model for audio tasks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.726491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.894405Z digest=sha256:79c89e53ce3cb5a8cd49608e4110aae6080b2bfabb4182f684b9c153550b866e

Observation 9c489b94-280a-4e2c-9dc2-8ca8ec8cdfc0 · outbound

This paper cites Audio Retrieval with WavText5K and CLAP Training.

ADIFF: Explaining audio difference using natural language Audio Retrieval with WavText5K and CLAP Training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:25.898652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:25.898652Z digest=sha256:fbfac3c9e1be73c94e308849a88cf218ad0e37c5dd9cf8734277686135204bc2

Observation 400897e4-5812-4823-bff3-3ad20e4676e4 · outbound

This paper cites Pam: Prompting audio-language models for audio quality assessment.

ADIFF: Explaining audio difference using natural language Pam: Prompting audio-language models for audio quality assessment

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:25.903434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:25.903434Z digest=sha256:dc1d2593fb323d36a0ee8a60b90a733ff2a4d50c46413af8ade68909799160fd

Observation 4abf38a1-108b-4ffe-b47d-21fb429e4bc7 · outbound

This paper cites Audio Entailment: Assessing Deductive Reasoning for Audio Understanding.

ADIFF: Explaining audio difference using natural language Audio Entailment: Assessing Deductive Reasoning for Audio Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:25.912138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:25.912138Z digest=sha256:f399d98268cec2c099433110ec9f7eb2f6448d2dba6bc87097847a554215349b

Observation f3e44766-a8a7-4457-b81c-6116a4cb5688 · outbound

This paper cites Domain adaptation for contrastive audio-language models.

ADIFF: Explaining audio difference using natural language Domain adaptation for contrastive audio-language models

Reference 15

Resolution
verified exact
doi, observed 2026-08-08T22:40:26.225279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.916764Z digest=sha256:b9e9e35d2c09cde845150bf14ef1a0e2248d585b92d13313ad97619156b3c6b8

Observation e167ca20-ddec-4085-ac54-beb23d60b09b · outbound

This paper cites Automated audio captioning with recurrent neural networks.

ADIFF: Explaining audio difference using natural language Automated audio captioning with recurrent neural networks

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.712572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.925095Z digest=sha256:5a5d0ec7054f624798adabeefd08c01e1621407609a84a2603645748e760a1a5

Observation e98c71dc-6f5e-4296-b97e-d85f0bf64096 · outbound

This paper cites Clotho: an audio captioning dataset.

ADIFF: Explaining audio difference using natural language Clotho: an audio captioning dataset

Reference 18

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T22:40:27.090329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.929302Z digest=sha256:1cbb8f89fade61e30a4c8a3cd8fbd6b4bb370a91c107a1c1a4f28e06d7e1b698

Observation 9df2d0b1-07f1-439a-bc77-ccd1d3671679 · outbound

This paper cites Clap learning audio concepts from natural language supervision.

ADIFF: Explaining audio difference using natural language Clap learning audio concepts from natural language supervision

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:25.933591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:25.933591Z digest=sha256:0e1a95f8e9bdecbeddf36895d88d09855f7b29bd771eca3f9fbbfb3fd9775889

Observation 85c3e55a-2f65-47ab-b226-5360e2625853 · outbound

This paper cites Natural language supervision for general-purpose audio representations.

ADIFF: Explaining audio difference using natural language Natural language supervision for general-purpose audio representations

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.688645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.938080Z digest=sha256:ec30828b50d5aec9058d0ec8d1e39ae15a34851448f10a30e192350b334df884

Observation 117f560f-8cbe-4943-bae9-bb1e892573c1 · outbound

This paper cites Fsd50k: an open dataset of human-labeled sound events.

ADIFF: Explaining audio difference using natural language Fsd50k: an open dataset of human-labeled sound events

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:25.942337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:25.942337Z digest=sha256:142aa5a4bc59428b994136ddc6e3dbf28ceebebab7ff0a8da523bf92f734ffe2

Observation bf32190f-e38f-4606-9f2a-c8bb0dbb9f05 · outbound

This paper cites Gemmeke, Daniel P.

ADIFF: Explaining audio difference using natural language Gemmeke, Daniel P

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:25.946631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:25.946631Z digest=sha256:b9436f72f0fb1a15569b6f18c43fc5d9a92b0cc58249d15355c6f2d911c114d2

Observation c0617f60-88eb-4658-89b5-6cd0b76d2cf5 · outbound

This paper cites GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities.

ADIFF: Explaining audio difference using natural language GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:25.950869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:25.950869Z digest=sha256:7d0c60364f05fa11a8239d74f48b6c6579afa63f4423d58ae7dd155acb9c70c6

Observation 3e2429e1-f8d7-4bcf-a645-aa8f97baface · outbound

This paper cites Compa: Addressing the gap in compositional reasoning in audio-language models.

ADIFF: Explaining audio difference using natural language Compa: Addressing the gap in compositional reasoning in audio-language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.663548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.955699Z digest=sha256:05e76fcac99dd1b08f94e1995478221bc22b7925f3e0b0e701555e7fb9bc93f1

Observation d95d4d87-eb38-4820-8e89-48e3dbaf16ef · outbound

This paper cites Joint audio and speech understanding.

ADIFF: Explaining audio difference using natural language Joint audio and speech understanding

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.648939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.959780Z digest=sha256:7821810931b50ac6b0b1f44955b2ef88fe283b3e2ae7aabdc8e1bb91d6a4afc9

Observation 60ada09a-9b23-4dca-9221-478d5821ca12 · outbound

This paper cites Liu, Leonid Karlinsky, and James R.

ADIFF: Explaining audio difference using natural language Liu, Leonid Karlinsky, and James R

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.634434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.964199Z digest=sha256:6ff941ab6e817d1f9daa9f2dae9bf12c190a06a7b64c09fbc4ae8621cd78a60b

Observation 0fa94c8d-612c-45e9-8ed0-fd9d42a86019 · outbound

This paper cites Clip4idc: Clip for image difference captioning.

ADIFF: Explaining audio difference using natural language Clip4idc: Clip for image difference captioning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.619797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.968411Z digest=sha256:dd98be7e7aeb1902401b974e54de460b2ac1a7dfab602ace3bd7f738bd7b4873

Observation efe20142-1389-4633-8b71-e5baddffd82e · outbound

This paper cites Heller, Benjamin Elizalde, Bhiksha Raj, and Soham Deshmukh.

ADIFF: Explaining audio difference using natural language Heller, Benjamin Elizalde, Bhiksha Raj, and Soham Deshmukh

Reference 28

Resolution
verified exact
raw_fallback, observed 2026-08-08T22:40:26.960895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.972792Z digest=sha256:8cc39a8aa5b2fcd9c756ee503a6d86a088d88d5750ac11ea47f4d30909200cdc

Observation 15988846-ea06-40a3-9f20-4c47d2874d8b · outbound

This paper cites Training compute-optimal large language models.

ADIFF: Explaining audio difference using natural language Training compute-optimal large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.604807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.976835Z digest=sha256:cbc1c8636d42f72058c209a708c3013841e5192a2c9111a06a286cefb879da2f

Observation f4376b1b-401e-4ba3-be4d-6afeb2cafe46 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

ADIFF: Explaining audio difference using natural language Lora: Low-rank adaptation of large language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:25.980953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:25.980953Z digest=sha256:200b1407876a4e080afe9cad8f34ac512d3d1acd7ec0e1ec44545c39badac6a3

Observation 2297d55e-3536-4197-8d43-451eaaf10104 · outbound

This paper cites Learning to describe differences between pairs of similar images.

ADIFF: Explaining audio difference using natural language Learning to describe differences between pairs of similar images

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:25.985308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:25.985308Z digest=sha256:019bfc135981d243dbc5f362a8d977c89024aba1d4006bbae805bf1451e3982a

Observation da759e54-0ac9-448b-b7c3-ec1b1d2316ab · outbound

This paper cites Acoustic and auditory phonetics.

ADIFF: Explaining audio difference using natural language Acoustic and auditory phonetics

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.581001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.989730Z digest=sha256:ea2da167b3b758beb6a38efc3cf590b05d162b63f1e09693685583c88cdef07b

Observation 478f2a89-9134-4063-9c3c-016a1cf2ee85 · outbound

This paper cites Deductive reasoning.

ADIFF: Explaining audio difference using natural language Deductive reasoning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.566972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.993960Z digest=sha256:c8846ad30a64efed74bfa1ef46bc0e164ef0ac88ee2236f0b589c609d55c30a0

Observation d676bbb7-0d52-4e41-9c18-00d0e8ed64fe · outbound

This paper cites Scaling Laws for Neural Language Models.

ADIFF: Explaining audio difference using natural language Scaling Laws for Neural Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:25.998134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:25.998134Z digest=sha256:bb5c3f52068b38d18c2208acc82f7313b69738468c9d277b63d2750ac61b9401

Observation e217a7d3-cf79-44a3-9687-7f2851cad396 · outbound

This paper cites AudioCaps: Generating Captions for Audios in The Wild.

ADIFF: Explaining audio difference using natural language AudioCaps: Generating Captions for Audios in The Wild

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.552222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.002871Z digest=sha256:0462b85c7bfd4be799ab39d2d2ca058e2e501649bf09a90cf0a1ac85b67e5f39

Observation a218f7da-772e-477a-a637-1ff59cfcab33 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

ADIFF: Explaining audio difference using natural language Adam: A Method for Stochastic Optimization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:26.007244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:26.007244Z digest=sha256:4ab129ccb3503bac5e779df887622ceba395bd61d76413fb08c76fa6a7d30ec4

Observation 2553c111-5dd3-4839-8e55-222eaf552465 · outbound

This paper cites Audio retrieval with natural language queries: A benchmark study.

ADIFF: Explaining audio difference using natural language Audio retrieval with natural language queries: A benchmark study

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.538191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.011742Z digest=sha256:e9e824167406d229d3a9699b603c89df1a0e0f5ae01634773140a47f0009805c

Observation 1cd7dbcc-d41b-4288-baef-a9662a65e6d6 · outbound

This paper cites Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities.

ADIFF: Explaining audio difference using natural language Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.524843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.019830Z digest=sha256:168bc36e625069dec4161f815fdfab9aadc992801a3e8148771ba6ce1f99705f

Observation c0ce7e19-fe0a-4873-acd7-97fb6c2d2ae8 · outbound

This paper cites Digital audio forensics: a first practical evaluation on microphone and environment classification.

ADIFF: Explaining audio difference using natural language Digital audio forensics: a first practical evaluation on microphone and environment classification

Reference 40

Resolution
metadata mismatch
raw_fallback, observed 2026-08-08T22:40:26.844147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.024034Z digest=sha256:a553f71a07533ee1ffceebf3895266ddb892b10c060a5cf7370ccb4a711e4382

Observation b388490f-25eb-4608-9fd5-71afe3bd2f50 · outbound

This paper cites Audiogen: Textually guided audio generation.

ADIFF: Explaining audio difference using natural language Audiogen: Textually guided audio generation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.511501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.028209Z digest=sha256:88d47715b3cc5c2051bb7db369bb68ec44b5925fb14006939928bf3c900fb947

Observation a13066a4-e9f4-4cdc-bae1-9b7b1be7f8f7 · outbound

This paper cites Understanding sounds, missing the questions: The challenge of object hallucination in large audio-language models.

ADIFF: Explaining audio difference using natural language Understanding sounds, missing the questions: The challenge of object hallucination in large audio-language models

Reference 42

Resolution
verified exact
doi, observed 2026-08-08T22:40:26.200276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.032524Z digest=sha256:a9e4a58895f8648cffe1a677ce4346196e1ef93ec0e9dad4c5f6c87cb0c49eda

Observation 2ef30c04-b77c-4d80-b6a0-b9cccf821b30 · outbound

This paper cites Clotho-aqa: A crowdsourced dataset for audio question answering.

ADIFF: Explaining audio difference using natural language Clotho-aqa: A crowdsourced dataset for audio question answering

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:26.036963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:26.036963Z digest=sha256:34847fab6ca172ac9013bebac52a86556c138c0c971c4523540f465743995665

Observation 439949ec-7a75-41d3-b2a2-89cf996f6618 · outbound

This paper cites Audioldm: Text-to-audio generation with latent diffusion models.

ADIFF: Explaining audio difference using natural language Audioldm: Text-to-audio generation with latent diffusion models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.497356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.041398Z digest=sha256:0d48019c4f162b843f4d6f10ddde8ebe1ec1c441545b41c483854fbb02fba1ac

Observation 1957f0d7-8389-4118-8a89-cff98c696da8 · outbound

This paper cites Improved image captioning via policy gradient optimization of spider.

ADIFF: Explaining audio difference using natural language Improved image captioning via policy gradient optimization of spider

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.482504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.045697Z digest=sha256:41c784fee90f2ab791b0f9fedd3b447ff62b761fd1569b250fd3215441471de5

Observation de25405b-8187-4cb4-beae-f8128b59dde7 · outbound

This paper cites an unresolved cited work.

ADIFF: Explaining audio difference using natural language Unresolved cited work

Reference 46

Resolution
metadata mismatch
raw_fallback, observed 2026-08-08T22:40:26.678668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.049830Z digest=sha256:c2ffc1e64970955ad26cfebce8e220dbd395dac5cbd95f62789bc2b11b9800b3

Observation 8fb10fba-90b6-4da9-80c6-f03405677ace · outbound

This paper cites Automated audio captioning: An overview of recent progress and new challenges.

ADIFF: Explaining audio difference using natural language Automated audio captioning: An overview of recent progress and new challenges

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.467200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.054009Z digest=sha256:32f70c836a9eb1934ceb04f9ab70745975f968799cb5921e78881946ea6feb3a

Observation 806d8eb2-2ff4-4031-b4ce-7062e8ff3170 · outbound

This paper cites Diverse audio captioning via adversarial training.

ADIFF: Explaining audio difference using natural language Diverse audio captioning via adversarial training

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.450989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.058365Z digest=sha256:db5c54014182fc84479b5e9787471a7aa85cfa91a1e7c6dd39ab02f58e92f565

Observation 7026ab62-d8d1-4435-a88e-f6ba4dd2400b · outbound

This paper cites Towards generating diverse audio captions via adversarial training.

ADIFF: Explaining audio difference using natural language Towards generating diverse audio captions via adversarial training

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.433992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.062641Z digest=sha256:feb504aa1bf02250c197e6a0f6df1dbe8518ae013d5d07f77c3317d986f25769

Observation e84060fd-f5b7-4092-a53b-af18ca16ac6e · outbound

This paper cites Plumbley, Yuexian Zou, and Wenwu Wang.

ADIFF: Explaining audio difference using natural language Plumbley, Yuexian Zou, and Wenwu Wang

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:26.066886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:26.066886Z digest=sha256:dd607b188ba4d4cf35de3cbce4e449297da67356881c63603ca9ff2f2131986b

Observation 5236dadf-4852-4991-8817-d122aa7d8f9b · outbound

This paper cites ClipCap: CLIP Prefix for Image Captioning.

ADIFF: Explaining audio difference using natural language ClipCap: CLIP Prefix for Image Captioning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:26.071258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:26.071258Z digest=sha256:5d2ebddd8217f508e804505c128cc3bc710ca1c575d852b40930ec2a7e3dd32c

Observation 96103ec6-861b-43b8-95e8-f200c9d660b7 · outbound

This paper cites Diversity and bias in audio captioning datasets.

ADIFF: Explaining audio difference using natural language Diversity and bias in audio captioning datasets

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.418020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.076195Z digest=sha256:82afd9b9751642bd809c72147383003e98f8db99c1f56e52f38b94830f3fa426

Observation 33e0debe-e788-4feb-86ea-003d9f9f5388 · outbound

This paper cites On the Audio Hallucinations in Large Audio-Video Language Models.

ADIFF: Explaining audio difference using natural language On the Audio Hallucinations in Large Audio-Video Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:26.080330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:26.080330Z digest=sha256:3b1af735b6e06c06e6ee10fddf942df694bc0c128573c914f35a60f2f2c43c21

Observation 8eb14dbe-4f4c-4bd5-a58d-496414828426 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

ADIFF: Explaining audio difference using natural language Bleu: a method for automatic evaluation of machine translation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:26.084813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:26.084813Z digest=sha256:04e8eab1f0e949190c548357f99d9e918261650dfae43452c66a321f1b245604

Observation 09858cef-9d95-4f5b-83bb-ee6cf5478676 · outbound

This paper cites Robust change captioning.

ADIFF: Explaining audio difference using natural language Robust change captioning

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.392941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.089162Z digest=sha256:ff8baa0fe92b3d40d023edbe7fe86c32653060e31091eec9b9c741357aad85a4

Observation f5b8dc08-8daf-4645-8df1-c926c1144dd5 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

ADIFF: Explaining audio difference using natural language Robust speech recognition via large-scale weak supervision

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:26.093356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:26.093356Z digest=sha256:1d1b249aa50f7faf4661a7a2616dbc8887de978a58cd99a77fc1dbfbd7fec095

Observation fee5dd68-f9b9-472f-806e-0103bcca5156 · outbound

This paper cites Scaling Language Models: Methods, Analysis & Insights from Training Gopher.

ADIFF: Explaining audio difference using natural language Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:26.097426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:26.097426Z digest=sha256:9bebcd5f88fe1b0f80441ee175aca96cff540f6c816000b3dbc2e5d2b7e6585e

Observation bcd93b80-6872-4fb8-932b-3f4ca73328d5 · outbound

This paper cites Acoustic phonetics, volume 30.

ADIFF: Explaining audio difference using natural language Acoustic phonetics, volume 30

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.367637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.101787Z digest=sha256:5e13d1b502e4577750d77ad639baadd7a17aa319e772210d092c00b637547f09

Observation 9cc831c7-c9cc-4502-ab69-c576fa1ebffd · outbound

This paper cites Audio difference captioning utilizing similarity-discrepancy disentanglement.

ADIFF: Explaining audio difference using natural language Audio difference captioning utilizing similarity-discrepancy disentanglement

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.351899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.106037Z digest=sha256:90c95f45c04606d434bcad262859479d7f976e771288eba09c1e86c61fe271a9

Observation 13cb819c-e4a3-4cbd-8452-abd914c08621 · outbound

This paper cites Extending large language models for speech and audio captioning.

ADIFF: Explaining audio difference using natural language Extending large language models for speech and audio captioning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:26.110428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:26.110428Z digest=sha256:cd5c789677d56781ca8ea0dce7761dc332aecdc0ab7b87c4901f5edf2cbe24e8

Observation b4938af6-a8bb-464f-ab9b-f02c2cb5efdb · outbound

This paper cites SALMONN : Towards generic hearing abilities for large language models.

ADIFF: Explaining audio difference using natural language SALMONN : Towards generic hearing abilities for large language models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.334277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.114441Z digest=sha256:c99d1670282bae7e45176e4e352ae97a639d863a6a0c5553a275cf7b167eb78d

Observation 965bd9a9-cc83-4345-bfcb-3fcc3b9990a4 · outbound

This paper cites Tzanetakis and P.

ADIFF: Explaining audio difference using natural language Tzanetakis and P

Reference 62

Resolution
metadata mismatch
raw_fallback, observed 2026-08-08T22:40:26.468134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.118620Z digest=sha256:3baebc961d67ce667b79b899e5ca3a410d1d15aed3a7beb65d183d28b72378a1

Observation a4459a2d-ab10-49d0-ae9a-a9b11a8b33b4 · outbound

This paper cites Cider: Consensus-based image description evaluation.

ADIFF: Explaining audio difference using natural language Cider: Consensus-based image description evaluation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.315538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.122607Z digest=sha256:0a4af1262634bcb81f74569087d3353026dc7d9724de9a750ceeeac0bca0b2de

Observation b4eb0bf4-1abe-40b6-af8a-813348c697ae · outbound

This paper cites Beats-based audio captioning model with instructor embedding supervision and chatgpt mix-up.

ADIFF: Explaining audio difference using natural language Beats-based audio captioning model with instructor embedding supervision and chatgpt mix-up

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.299550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.126797Z digest=sha256:183f165c8c9a753633a1d32347c915db3e2248db3180d029388d4fcbda0bdae5

Observation b7630099-b4c1-491c-b35c-13ab3aca2395 · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

ADIFF: Explaining audio difference using natural language Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.283826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.131122Z digest=sha256:3fb1863afb5f63f95ed23f8ab68400840a50615d592a68f65476cb3b7c908c7b

Observation febee63c-35b8-423f-b8af-607bdb290feb · outbound

This paper cites Image difference captioning with pre-training and contrastive learning.

ADIFF: Explaining audio difference using natural language Image difference captioning with pre-training and contrastive learning

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.267320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.135100Z digest=sha256:3c53e0f5cacc76df3f8b51d8d37e7210250a7896b9c8cbc2616cf8a43156f717

Observation dd0e0ac3-3aec-4830-bb65-d61fe57f0891 · outbound

This paper cites Pre-training language models for comparative reasoning.

ADIFF: Explaining audio difference using natural language Pre-training language models for comparative reasoning

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.251015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.139230Z digest=sha256:d5fcbb9093e0e90e26873ae5150828cf6c1eaeac22ae0c668b12a465fa3e5769

Observation 1fda5dd7-caa2-47a3-b3a1-5f4a0fa709f4 · outbound

This paper cites NaRLE: Natural Language Models using Reinforcement Learning with Emotion Feedback.

ADIFF: Explaining audio difference using natural language NaRLE: Natural Language Models using Reinforcement Learning with Emotion Feedback

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-08-08T22:40:26.368889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.143516Z digest=sha256:835028f475f7976136b73eec9a59723b9ca2d69dcddecd8b64d3bc4f68602ae1

Observation a76c2d0c-01c4-49e8-a3df-26c952f5f910 · outbound

This paper cites @esa (Ref.

ADIFF: Explaining audio difference using natural language @esa (Ref

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:26.148096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:26.148096Z digest=sha256:f31de26fd65211b4a4d07fa0bffc0b6da3cadd3859967e7dc39fbe5e12f41c98

Observation 835026b5-32b0-4b7a-859b-071a2e393de6 · outbound

This paper cites an unresolved cited work.

ADIFF: Explaining audio difference using natural language Unresolved cited work

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:26.152867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:26.152867Z digest=sha256:4d061f0bbd91d08de3eb793d1f00cf02220f1dc88297c81cd6abc4b537469084

Observation a4d37c76-28ef-46c8-a52d-8186c2769fe2 · outbound

This paper cites or ``caption the second audio.

ADIFF: Explaining audio difference using natural language or ``caption the second audio

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:26.157629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:26.157629Z digest=sha256:ee9855ab2d5b24433e00ff9fa0177cbb496c6f5c6d102a07b6ad09f10db43b72

Pith citing papers

Observation aceb2984-3db4-4261-9616-02507bf01c93 · inbound

Breaking the Barriers of Text-Hungry and Audio-Deficient AI cites this paper.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI ADIFF: Explaining audio difference using natural language

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:48.510861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:48.510861Z digest=sha256:3664e1bbde6913f793351bc50f1ecd492586bdfe7c317ec86f31df5459e0d2f1

Observation 1936ce45-52aa-4d17-bf6c-9eaf66ce0390 · inbound

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations cites this paper.

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations ADIFF: Explaining audio difference using natural language

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:46.418856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:46.418856Z digest=sha256:e268235dd81a806d3bfa0b863c2ae3678455ed84e0321278a67fc624f65eeed6

Observation 97cf3465-2f7c-466e-885c-9a779fc7eaa6 · inbound

MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing cites this paper.

MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing ADIFF: Explaining audio difference using natural language

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:12:07.296957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T19:12:01.787842Z digest=sha256:0f0df95d48ab2207ad13c56e31e8027eac541af01cd96fa23843cb5afb1bba45