Pith. sign in

Paper Citation Record · LEDGER

ADIFF: Explaining audio difference using natural language

As of 9 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 3 inbound Pith citation observations for arXiv:2502.04476.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.04476 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:40:26.157629Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:28:48.510861Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T19:12:07.238377Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact6
  • verified fuzzy34
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9b48841a-6c0f-453b-b57a-75b117973ebc · outbound

This paper cites write newline.

ADIFF: Explaining audio difference using natural language write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:25.847277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:25.847277Z digest=sha256:f1a6292303327f7553a4153579f4a531cc5af566ed652087be154427098149c6

Observation 22469d2e-87e4-4c30-af75-7f4bafb529c3 · outbound

This paper cites Getting vit in shape: Scaling laws for compute-optimal model design.

ADIFF: Explaining audio difference using natural language Getting vit in shape: Scaling laws for compute-optimal model design

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.798472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.853389Z digest=sha256:24b36a865dac5cd6ef4b4745185c6359447bb3eef94d9a389a959c81b162c2f7

Observation 3182d79d-073c-4234-886f-a0b8b13277bf · outbound

This paper cites Spice: Semantic propositional image caption evaluation.

ADIFF: Explaining audio difference using natural language Spice: Semantic propositional image caption evaluation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.783679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.858087Z digest=sha256:1c9a7b064354a03fd88ed4be016986d81bba4f998871109e980f4aac10675a53

Observation a4f7567f-4375-41f4-af7e-4a13897be366 · outbound

This paper cites METEOR : An automatic metric for MT evaluation with improved correlation with human judgments.

ADIFF: Explaining audio difference using natural language METEOR : An automatic metric for MT evaluation with improved correlation with human judgments

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.767958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.865805Z digest=sha256:4f19ea567796c95a82c5f2a62fed7517fd5d89ac09ed5abcafb5b435e6d1e6b4

Observation 211752ad-a296-4b7c-9bca-249d42484cee · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

ADIFF: Explaining audio difference using natural language PaliGemma: A versatile 3B VLM for transfer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:25.870623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:25.870623Z digest=sha256:19fb911123dde65a0e06fd72c0cfa04b43c91a80aca8f366ea81aa93c26ebfd7

Observation 6b195f82-ac43-4b5d-a65e-f46f804eb9d4 · outbound

This paper cites Selm: Enhancing speech emotion recognition for out-of-domain scenarios.

ADIFF: Explaining audio difference using natural language Selm: Enhancing speech emotion recognition for out-of-domain scenarios

Reference 6

Resolution
verified exact
doi, observed 2026-08-08T22:40:26.259570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.875537Z digest=sha256:060687eaa3d36c2b5b34fde1be6ed2ac9883d865b7c031dcf502e6f2d403d6f7

Observation a2f4ae70-1ccb-4dad-8caa-d228ed5cf4d1 · outbound

This paper cites Audio quality assessment techniques—a review, and recent developments.

ADIFF: Explaining audio difference using natural language Audio quality assessment techniques—a review, and recent developments

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.754434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.880510Z digest=sha256:d16c3cc1c75b64eb6d41e38dd92cd40a51a7b89aff7b2bfb666b0ac835c88682

Observation c156397e-419f-4024-85ef-95d5055d2982 · outbound

This paper cites Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection.

ADIFF: Explaining audio difference using natural language Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.740547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.885456Z digest=sha256:8f9cc8dd605202f0cd66db2c1faeffd98dcb7646aea6183767a4bcc3936d8f9a

Observation 2e54e352-653f-47f8-bb6b-edce2e3f3c7a · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

ADIFF: Explaining audio difference using natural language Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:25.889755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:25.889755Z digest=sha256:a7ee323ec577186dc2c473c1aed7c41632e1d1d929883bf7acb7a91c01b4734a

Observation 646f3069-7740-4295-ba76-49001678b1b1 · outbound

This paper cites Pengi: An audio language model for audio tasks.

ADIFF: Explaining audio difference using natural language Pengi: An audio language model for audio tasks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.726491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.894405Z digest=sha256:50476220701c943c6b9099850fa083e33339eed853d598ba5822a5ce7bddf47f

Observation 9c489b94-280a-4e2c-9dc2-8ca8ec8cdfc0 · outbound

This paper cites Audio Retrieval with WavText5K and CLAP Training.

ADIFF: Explaining audio difference using natural language Audio Retrieval with WavText5K and CLAP Training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:25.898652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:25.898652Z digest=sha256:9f666de45d46ae2cd0060e1ba680359d0baaeedf258eb887704d8028ec913cdd

Observation 400897e4-5812-4823-bff3-3ad20e4676e4 · outbound

This paper cites Pam: Prompting audio-language models for audio quality assessment.

ADIFF: Explaining audio difference using natural language Pam: Prompting audio-language models for audio quality assessment

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:25.903434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:25.903434Z digest=sha256:f0e95a3c43543fd19708b61356e39c2ad26895b43410b55ceffae1bf7cb0811c

Observation 4abf38a1-108b-4ffe-b47d-21fb429e4bc7 · outbound

This paper cites Audio Entailment: Assessing Deductive Reasoning for Audio Understanding.

ADIFF: Explaining audio difference using natural language Audio Entailment: Assessing Deductive Reasoning for Audio Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:25.912138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:25.912138Z digest=sha256:f62d0f96d03b76326fd004bdd0e3ad5af0e6bb66092539d0bebe9279c6d3d092

Observation f3e44766-a8a7-4457-b81c-6116a4cb5688 · outbound

This paper cites Domain adaptation for contrastive audio-language models.

ADIFF: Explaining audio difference using natural language Domain adaptation for contrastive audio-language models

Reference 15

Resolution
verified exact
doi, observed 2026-08-08T22:40:26.225279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.916764Z digest=sha256:e7775cad351f812dde7e46877d36757c5ca486cdd9d490d674a0defb42a9b6fc

Observation e167ca20-ddec-4085-ac54-beb23d60b09b · outbound

This paper cites Automated audio captioning with recurrent neural networks.

ADIFF: Explaining audio difference using natural language Automated audio captioning with recurrent neural networks

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.712572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.925095Z digest=sha256:137af467bdcc6548faed6637e4b927b5232e4cf5660a353402341c442a4a468d

Observation e98c71dc-6f5e-4296-b97e-d85f0bf64096 · outbound

This paper cites Clotho: an audio captioning dataset.

ADIFF: Explaining audio difference using natural language Clotho: an audio captioning dataset

Reference 18

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T22:40:27.090329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.929302Z digest=sha256:6b70283737c6a7f7b000236f1f145f092a0bab7ccac76e53426cb90eed1d06a4

Observation 9df2d0b1-07f1-439a-bc77-ccd1d3671679 · outbound

This paper cites Clap learning audio concepts from natural language supervision.

ADIFF: Explaining audio difference using natural language Clap learning audio concepts from natural language supervision

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:25.933591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:25.933591Z digest=sha256:4ee8fb76b0b1a503909a72fdbfedc186d49bab4887d88bd3defbb1f7978a926f

Observation 85c3e55a-2f65-47ab-b226-5360e2625853 · outbound

This paper cites Natural language supervision for general-purpose audio representations.

ADIFF: Explaining audio difference using natural language Natural language supervision for general-purpose audio representations

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.688645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.938080Z digest=sha256:ad4c6855eae4e599a64817717f68d48029034f7775d494068237202196a51a8d

Observation 117f560f-8cbe-4943-bae9-bb1e892573c1 · outbound

This paper cites Fsd50k: an open dataset of human-labeled sound events.

ADIFF: Explaining audio difference using natural language Fsd50k: an open dataset of human-labeled sound events

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:25.942337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:25.942337Z digest=sha256:99fceba003de428917437c5f1070eb01b00c5d93d34e54c24de5e1168f9a760a

Observation bf32190f-e38f-4606-9f2a-c8bb0dbb9f05 · outbound

This paper cites Gemmeke, Daniel P.

ADIFF: Explaining audio difference using natural language Gemmeke, Daniel P

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:25.946631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:25.946631Z digest=sha256:151731691fb8ece7b7f76519af8d6c32e359d9fdf134d80c418bbd9f00128caa

Observation c0617f60-88eb-4658-89b5-6cd0b76d2cf5 · outbound

This paper cites GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities.

ADIFF: Explaining audio difference using natural language GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:25.950869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:25.950869Z digest=sha256:74221d4c1bc731529074b269184605f5f010d1bfbc76d9b47518358dabab7787

Observation 3e2429e1-f8d7-4bcf-a645-aa8f97baface · outbound

This paper cites Compa: Addressing the gap in compositional reasoning in audio-language models.

ADIFF: Explaining audio difference using natural language Compa: Addressing the gap in compositional reasoning in audio-language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.663548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.955699Z digest=sha256:c1ebd9589b8854c72f62a377f4317a49c93d4e6a3c16d340c12c8868a3fd7036

Observation d95d4d87-eb38-4820-8e89-48e3dbaf16ef · outbound

This paper cites Joint audio and speech understanding.

ADIFF: Explaining audio difference using natural language Joint audio and speech understanding

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.648939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.959780Z digest=sha256:d52efcee118ed7e74b8425d93633cd9c51c0801a5bc60b534dd985d4a0fc4b5d

Observation 60ada09a-9b23-4dca-9221-478d5821ca12 · outbound

This paper cites Liu, Leonid Karlinsky, and James R.

ADIFF: Explaining audio difference using natural language Liu, Leonid Karlinsky, and James R

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.634434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.964199Z digest=sha256:a6cb2e4327290238e82532fe362951c50d659e4895e8fb5b294b4f3c480556c1

Observation 0fa94c8d-612c-45e9-8ed0-fd9d42a86019 · outbound

This paper cites Clip4idc: Clip for image difference captioning.

ADIFF: Explaining audio difference using natural language Clip4idc: Clip for image difference captioning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.619797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.968411Z digest=sha256:bf76de94e02ae2876b6c9d848ab84bb93e3f742a21f342fa34d26096d476f337

Observation efe20142-1389-4633-8b71-e5baddffd82e · outbound

This paper cites Heller, Benjamin Elizalde, Bhiksha Raj, and Soham Deshmukh.

ADIFF: Explaining audio difference using natural language Heller, Benjamin Elizalde, Bhiksha Raj, and Soham Deshmukh

Reference 28

Resolution
verified exact
raw_fallback, observed 2026-08-08T22:40:26.960895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.972792Z digest=sha256:043ecab9d884843ef79d447b7aa8897732ccc9d442850a3b042d2303b1b2cb3f

Observation 15988846-ea06-40a3-9f20-4c47d2874d8b · outbound

This paper cites Training compute-optimal large language models.

ADIFF: Explaining audio difference using natural language Training compute-optimal large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.604807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.976835Z digest=sha256:b6e296ab90bc53a1068721479a581809db51facaf91c7d60c9d436abf6dc0f44

Observation f4376b1b-401e-4ba3-be4d-6afeb2cafe46 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

ADIFF: Explaining audio difference using natural language Lora: Low-rank adaptation of large language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:25.980953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:25.980953Z digest=sha256:98c87d165806de91c7637524be6038165d2615f03062ec65775799af04a16bd1

Observation 2297d55e-3536-4197-8d43-451eaaf10104 · outbound

This paper cites Learning to describe differences between pairs of similar images.

ADIFF: Explaining audio difference using natural language Learning to describe differences between pairs of similar images

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:25.985308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:25.985308Z digest=sha256:6180ffe0cb72c5246385e6b8a611d10e6ac310462fed76017dbef7d990e921c1

Observation da759e54-0ac9-448b-b7c3-ec1b1d2316ab · outbound

This paper cites Acoustic and auditory phonetics.

ADIFF: Explaining audio difference using natural language Acoustic and auditory phonetics

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.581001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.989730Z digest=sha256:32b9166423add2ed5fc0a6747d7605ec8ab4f823bf0796ea9b55b9daedbe4507

Observation 478f2a89-9134-4063-9c3c-016a1cf2ee85 · outbound

This paper cites Deductive reasoning.

ADIFF: Explaining audio difference using natural language Deductive reasoning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.566972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:25.993960Z digest=sha256:a8ff274da936f8bb7588dd7153e9446fe4da1d4e7090be65ba89da4cd48405a6

Observation d676bbb7-0d52-4e41-9c18-00d0e8ed64fe · outbound

This paper cites Scaling Laws for Neural Language Models.

ADIFF: Explaining audio difference using natural language Scaling Laws for Neural Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:25.998134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:25.998134Z digest=sha256:e3075f7e82b5ae1a979057a36d4149ecceea5cde89e4ebf7e0d42db9d3f84202

Observation e217a7d3-cf79-44a3-9687-7f2851cad396 · outbound

This paper cites AudioCaps: Generating Captions for Audios in The Wild.

ADIFF: Explaining audio difference using natural language AudioCaps: Generating Captions for Audios in The Wild

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.552222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.002871Z digest=sha256:ba3120ad632f8a34f05c319bf9b30ad0cd3cf362cd538dacd2747c97d699232f

Observation a218f7da-772e-477a-a637-1ff59cfcab33 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

ADIFF: Explaining audio difference using natural language Adam: A Method for Stochastic Optimization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:26.007244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:26.007244Z digest=sha256:c83b40c3fa42ce94502cef3ffcdbba2310cf784eff9523e458d1d4b32f5e8f44

Observation 2553c111-5dd3-4839-8e55-222eaf552465 · outbound

This paper cites Audio retrieval with natural language queries: A benchmark study.

ADIFF: Explaining audio difference using natural language Audio retrieval with natural language queries: A benchmark study

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.538191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.011742Z digest=sha256:725a1d1cb3951ecd682fb0f3ed856676b7d7658fd6e0dae0a5f1f2baa2a7c1ef

Observation 1cd7dbcc-d41b-4288-baef-a9662a65e6d6 · outbound

This paper cites Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities.

ADIFF: Explaining audio difference using natural language Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.524843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.019830Z digest=sha256:ca0855fa5c88cd54f4712cf138c2ed164dc6cbeec048d8d4018b3ab9c21a6af9

Observation c0ce7e19-fe0a-4873-acd7-97fb6c2d2ae8 · outbound

This paper cites Digital audio forensics: a first practical evaluation on microphone and environment classification.

ADIFF: Explaining audio difference using natural language Digital audio forensics: a first practical evaluation on microphone and environment classification

Reference 40

Resolution
metadata mismatch
raw_fallback, observed 2026-08-08T22:40:26.844147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.024034Z digest=sha256:c2bc5d727f32e3c6aa4c0a826af209d22e098f72b0d58ada2f782481b20a1a3b

Observation b388490f-25eb-4608-9fd5-71afe3bd2f50 · outbound

This paper cites Audiogen: Textually guided audio generation.

ADIFF: Explaining audio difference using natural language Audiogen: Textually guided audio generation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.511501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.028209Z digest=sha256:7a48beea6ccaa91d19e3fc4195a60580539160821d9477ac85eacb80823fdfbf

Observation a13066a4-e9f4-4cdc-bae1-9b7b1be7f8f7 · outbound

This paper cites Understanding sounds, missing the questions: The challenge of object hallucination in large audio-language models.

ADIFF: Explaining audio difference using natural language Understanding sounds, missing the questions: The challenge of object hallucination in large audio-language models

Reference 42

Resolution
verified exact
doi, observed 2026-08-08T22:40:26.200276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.032524Z digest=sha256:f413a0cc66ae97f27808aa973483c2457aa8a944067859fb427fa988eb027fb7

Observation 2ef30c04-b77c-4d80-b6a0-b9cccf821b30 · outbound

This paper cites Clotho-aqa: A crowdsourced dataset for audio question answering.

ADIFF: Explaining audio difference using natural language Clotho-aqa: A crowdsourced dataset for audio question answering

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:26.036963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:26.036963Z digest=sha256:5a04657d7ffc60e4e31249f14bc56dc4ea1af9fd2db41c3fad1d48e28cd37455

Observation 439949ec-7a75-41d3-b2a2-89cf996f6618 · outbound

This paper cites Audioldm: Text-to-audio generation with latent diffusion models.

ADIFF: Explaining audio difference using natural language Audioldm: Text-to-audio generation with latent diffusion models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.497356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.041398Z digest=sha256:83d128af1995d62ef4bb45a8617860eca1c62baef9f67e77fc3b059b9dbdbf7c

Observation 1957f0d7-8389-4118-8a89-cff98c696da8 · outbound

This paper cites Improved image captioning via policy gradient optimization of spider.

ADIFF: Explaining audio difference using natural language Improved image captioning via policy gradient optimization of spider

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.482504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.045697Z digest=sha256:1c7ea1bfcd69686732f6ffcb05d1fee26a96f0e51c287397ab00887d0748e422

Observation de25405b-8187-4cb4-beae-f8128b59dde7 · outbound

This paper cites an unresolved cited work.

ADIFF: Explaining audio difference using natural language Unresolved cited work

Reference 46

Resolution
metadata mismatch
raw_fallback, observed 2026-08-08T22:40:26.678668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.049830Z digest=sha256:7f2e7a2ce41c89ab80062e6b5185da1a9cac23a307bd2fe4cea90ea6ebc77ebd

Observation 8fb10fba-90b6-4da9-80c6-f03405677ace · outbound

This paper cites Automated audio captioning: An overview of recent progress and new challenges.

ADIFF: Explaining audio difference using natural language Automated audio captioning: An overview of recent progress and new challenges

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.467200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.054009Z digest=sha256:b07a25e64d2c6c104bc1e2ef396d5fbc2a45acf1ea1ae2cf0b762a17d5b5340a

Observation 806d8eb2-2ff4-4031-b4ce-7062e8ff3170 · outbound

This paper cites Diverse audio captioning via adversarial training.

ADIFF: Explaining audio difference using natural language Diverse audio captioning via adversarial training

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.450989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.058365Z digest=sha256:fb742e574a85bf21703a7d90295f1dfb0ef6c9769fc39ed40e36fa00d4ec8b35

Observation 7026ab62-d8d1-4435-a88e-f6ba4dd2400b · outbound

This paper cites Towards generating diverse audio captions via adversarial training.

ADIFF: Explaining audio difference using natural language Towards generating diverse audio captions via adversarial training

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.433992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.062641Z digest=sha256:232db163d100997ac1b9929788bf4d306d691fa8e8761e1305cd1e1e4c9a2a85

Observation e84060fd-f5b7-4092-a53b-af18ca16ac6e · outbound

This paper cites Plumbley, Yuexian Zou, and Wenwu Wang.

ADIFF: Explaining audio difference using natural language Plumbley, Yuexian Zou, and Wenwu Wang

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:26.066886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:26.066886Z digest=sha256:0fb39a84ef026f7ec1096bf9817c3d5f105c84a752a1dcac949d4e590ea38d29

Observation 5236dadf-4852-4991-8817-d122aa7d8f9b · outbound

This paper cites ClipCap: CLIP Prefix for Image Captioning.

ADIFF: Explaining audio difference using natural language ClipCap: CLIP Prefix for Image Captioning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:26.071258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:26.071258Z digest=sha256:24589da6ef6c6a9323c91449743ba53183c65a6fbbdc3a4da7dde4fc91e532d9

Observation 96103ec6-861b-43b8-95e8-f200c9d660b7 · outbound

This paper cites Diversity and bias in audio captioning datasets.

ADIFF: Explaining audio difference using natural language Diversity and bias in audio captioning datasets

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.418020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.076195Z digest=sha256:fa91d93cdf489e15a5acf0b60aadb9efe74cdac35a2cc5e7556f84725491ce9f

Observation 33e0debe-e788-4feb-86ea-003d9f9f5388 · outbound

This paper cites On the Audio Hallucinations in Large Audio-Video Language Models.

ADIFF: Explaining audio difference using natural language On the Audio Hallucinations in Large Audio-Video Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:26.080330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:26.080330Z digest=sha256:8a2978267cfe1a221099571d8c7ba69fe6cafa083c1877fdcbcd28f52fbe3be4

Observation 8eb14dbe-4f4c-4bd5-a58d-496414828426 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

ADIFF: Explaining audio difference using natural language Bleu: a method for automatic evaluation of machine translation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:26.084813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:26.084813Z digest=sha256:26f7e50e07b72835a837dc24ce0cf40c3aa59c22abeaab679336a512f4633120

Observation 09858cef-9d95-4f5b-83bb-ee6cf5478676 · outbound

This paper cites Robust change captioning.

ADIFF: Explaining audio difference using natural language Robust change captioning

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.392941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.089162Z digest=sha256:104500d4f897416f83b533df4416bb57090d982165cb68dd7e4f0c3f7edb4767

Observation f5b8dc08-8daf-4645-8df1-c926c1144dd5 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

ADIFF: Explaining audio difference using natural language Robust speech recognition via large-scale weak supervision

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:26.093356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:26.093356Z digest=sha256:ddfcb76c8415a9e94bb7f24ae9d8fd7de5b6ded57a1af4092e2dbe2738ea2b2e

Observation fee5dd68-f9b9-472f-806e-0103bcca5156 · outbound

This paper cites Scaling Language Models: Methods, Analysis & Insights from Training Gopher.

ADIFF: Explaining audio difference using natural language Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:26.097426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:26.097426Z digest=sha256:f3a52069d6ffebd2fc2b3004edd9c8388e4a05652c18214b59ff9ce2d043550c

Observation bcd93b80-6872-4fb8-932b-3f4ca73328d5 · outbound

This paper cites Acoustic phonetics, volume 30.

ADIFF: Explaining audio difference using natural language Acoustic phonetics, volume 30

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.367637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.101787Z digest=sha256:ed64c8530e0b59e40c77d04d9177bca34a75f93c9a96d06ed3a764c96120b025

Observation 9cc831c7-c9cc-4502-ab69-c576fa1ebffd · outbound

This paper cites Audio difference captioning utilizing similarity-discrepancy disentanglement.

ADIFF: Explaining audio difference using natural language Audio difference captioning utilizing similarity-discrepancy disentanglement

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.351899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.106037Z digest=sha256:feb3f2b9bdabfd75521f95cd58ca43bcbe9f2dd3bfdafd2f914dbb6da49bf5d8

Observation 13cb819c-e4a3-4cbd-8452-abd914c08621 · outbound

This paper cites Extending large language models for speech and audio captioning.

ADIFF: Explaining audio difference using natural language Extending large language models for speech and audio captioning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:26.110428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:26.110428Z digest=sha256:d6c2cc15c184621a3a37440cd07ffd1376ca253d3a318462034f5cb97fafda2e

Observation b4938af6-a8bb-464f-ab9b-f02c2cb5efdb · outbound

This paper cites SALMONN : Towards generic hearing abilities for large language models.

ADIFF: Explaining audio difference using natural language SALMONN : Towards generic hearing abilities for large language models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.334277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.114441Z digest=sha256:3514a0ca44126f21b9c56952c0cdfaa2c823e471c1560e2691cfc6afdef72104

Observation 965bd9a9-cc83-4345-bfcb-3fcc3b9990a4 · outbound

This paper cites Tzanetakis and P.

ADIFF: Explaining audio difference using natural language Tzanetakis and P

Reference 62

Resolution
metadata mismatch
raw_fallback, observed 2026-08-08T22:40:26.468134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.118620Z digest=sha256:86c611591af0b4d70ff36769657217e80e413dce94448db5788e8ed47b3cc28e

Observation a4459a2d-ab10-49d0-ae9a-a9b11a8b33b4 · outbound

This paper cites Cider: Consensus-based image description evaluation.

ADIFF: Explaining audio difference using natural language Cider: Consensus-based image description evaluation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.315538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.122607Z digest=sha256:58fa6bdf0178408620ba347d4c660753d7d2052f7f2c8aa15bb0411be5586f22

Observation b4eb0bf4-1abe-40b6-af8a-813348c697ae · outbound

This paper cites Beats-based audio captioning model with instructor embedding supervision and chatgpt mix-up.

ADIFF: Explaining audio difference using natural language Beats-based audio captioning model with instructor embedding supervision and chatgpt mix-up

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.299550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.126797Z digest=sha256:f3a5cea521df770de76760b7999c2c3f7b80bba7d9098bf26b6940be9aa5a227

Observation b7630099-b4c1-491c-b35c-13ab3aca2395 · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

ADIFF: Explaining audio difference using natural language Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.283826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.131122Z digest=sha256:e4728a84ed63185189677771583960f203f55ce4d5fae207d0b3ff195e043602

Observation febee63c-35b8-423f-b8af-607bdb290feb · outbound

This paper cites Image difference captioning with pre-training and contrastive learning.

ADIFF: Explaining audio difference using natural language Image difference captioning with pre-training and contrastive learning

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.267320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.135100Z digest=sha256:776e2ab2b69aee970cfeca95dfbe2ad0a299b0c1098579f61123df66beb23bf9

Observation dd0e0ac3-3aec-4830-bb65-d61fe57f0891 · outbound

This paper cites Pre-training language models for comparative reasoning.

ADIFF: Explaining audio difference using natural language Pre-training language models for comparative reasoning

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:40:27.251015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.139230Z digest=sha256:de23dafcd19c5af8bb5f97ad2063ea2062366bbd569955879e942439288da65c

Observation 1fda5dd7-caa2-47a3-b3a1-5f4a0fa709f4 · outbound

This paper cites NaRLE: Natural Language Models using Reinforcement Learning with Emotion Feedback.

ADIFF: Explaining audio difference using natural language NaRLE: Natural Language Models using Reinforcement Learning with Emotion Feedback

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-08-08T22:40:26.368889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:40:26.143516Z digest=sha256:00f3fc3452a02b1e1c92c3db2a400cfaa017f801d954564f62a71952101a0e27

Observation a76c2d0c-01c4-49e8-a3df-26c952f5f910 · outbound

This paper cites @esa (Ref.

ADIFF: Explaining audio difference using natural language @esa (Ref

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:26.148096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:26.148096Z digest=sha256:646ed8aa06c852fc2ffb99415dac9e6f183f6d35b3ba99c94672d39a810cd719

Observation 835026b5-32b0-4b7a-859b-071a2e393de6 · outbound

This paper cites an unresolved cited work.

ADIFF: Explaining audio difference using natural language Unresolved cited work

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:26.152867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:26.152867Z digest=sha256:9542e7cd17084bf537db867c271344f8a6c4e625c0a69c2cf935f3370a37ecdd

Observation a4d37c76-28ef-46c8-a52d-8186c2769fe2 · outbound

This paper cites or ``caption the second audio.

ADIFF: Explaining audio difference using natural language or ``caption the second audio

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:26.157629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:26.157629Z digest=sha256:b3cbabb35381b6bf56adc9206595d6c1243768650bf0222bc3fa39a5996cc421

Pith citing papers

Observation aceb2984-3db4-4261-9616-02507bf01c93 · inbound

Breaking the Barriers of Text-Hungry and Audio-Deficient AI cites this paper.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI ADIFF: Explaining audio difference using natural language

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:48.510861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:48.510861Z digest=sha256:c90e28b7262ce3dfd75a059750c5061326dce9ecd277326d46cc2da1d6e1d583

Observation 1936ce45-52aa-4d17-bf6c-9eaf66ce0390 · inbound

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations cites this paper.

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations ADIFF: Explaining audio difference using natural language

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:46.418856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:46.418856Z digest=sha256:766b380bc3d188cb314d31199a8fe9ff1db306a65b042e32e28cacdfd03a6775

Observation 97cf3465-2f7c-466e-885c-9a779fc7eaa6 · inbound

MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing cites this paper.

MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing ADIFF: Explaining audio difference using natural language

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:12:07.296957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:12:01.787842Z digest=sha256:46ceb59d5275efe39c0eeb3474ea21942998077f5482bad4e6c4a62293393a58