Pith. sign in

Paper Citation Record · LEDGER

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples

As of 13 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2505.14518.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14518 v2

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:05.382288Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:05.068351Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T15:36:05.992218Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved29
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c4f1b436-897a-48e5-b3ba-808817bae70e · outbound

This paper cites These models can process audio, speech, and text in- puts at the same time, using text prompts to extract relevant information from audio and speech.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples These models can process audio, speech, and text in- puts at the same time, using text prompts to extract relevant information from audio and speech

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.659290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.049299Z digest=sha256:60aca3c34503476a1ddc438fb6cefac1b9ce02dd6ea7dc09bc27dd4bb3ed275c

Observation 75dea77b-fdc2-4e7f-9441-67ad801a1281 · outbound

This paper cites an unresolved cited work.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:36:06.642485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.054761Z digest=sha256:a39fbb1d82b4ff4ec91315939746b372348ecc53ef351b725f24e3f9cca07801

Observation e6e12cee-b610-460a-8574-3860481d703c · outbound

This paper cites This is achieved by leveraging a backbone-LLM- synthesized dataset, which automatically generates audio- text pairs and contrastive data across general audio scenarios.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples This is achieved by leveraging a backbone-LLM- synthesized dataset, which automatically generates audio- text pairs and contrastive data across general audio scenarios

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.624963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.060552Z digest=sha256:eb1cd02ef17b2389bda21b9636c9d7bc647d648de7f2e863c6d82cd85442023c

Observation 0280248b-6ca7-45b5-a5c7-3d5366219030 · outbound

This paper cites Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:36:05.999867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.068351Z digest=sha256:d932d4aa6d333fea17837c8e8daf290b49ea0f3d87b7acb49212215efb998b42

Observation 4b49fbe8-c3f8-41aa-aa0b-139433f3aa65 · outbound

This paper cites an unresolved cited work.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:36:06.608610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.074285Z digest=sha256:6a4bdb5ce585e8330e2d9d748ac8acafdfbefdeb81bba0fe8e764af3249d0302

Observation fd548852-9e47-4d82-ae5c-7c9c93d96b22 · outbound

This paper cites For example, Replay the audio.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples For example, Replay the audio

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.592156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.078959Z digest=sha256:093329543e16ca52dfc11465f4d56272dd425eab9d42dc48e3954b90618acddb

Observation 4eafbe61-5039-4fd9-8492-4e508434e98f · outbound

This paper cites For example, Identify sounds that are absent as con- trasting examples.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples For example, Identify sounds that are absent as con- trasting examples

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.575502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.083888Z digest=sha256:809afbf5255dcab8228d2781cee98d74b69712459c77b60c23de39afc3db9b30

Observation 469d3f93-108c-42c4-ae69-6ab4a7775bfe · outbound

This paper cites It aims to generate descriptions of both the sound events that are present and those that are ab- sent in the audio.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples It aims to generate descriptions of both the sound events that are present and those that are ab- sent in the audio

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.558975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.088930Z digest=sha256:1497c8f116df2cd54256d8539c51b4d5b686671263d0431e2d43e1818cc5fa00

Observation 3d1bb6d6-4898-459a-93a1-f395612811ea · outbound

This paper cites We utilize the foundation model Whisper 2.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples We utilize the foundation model Whisper 2

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.542602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.093619Z digest=sha256:dd3345fdd122e14ea2cec3e2ce750b17c605f5b92114f1a434979cd1fe27d51c

Observation 9ee2869b-0654-4519-8eb0-b6f4ff4d55d2 · outbound

This paper cites BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.217051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.217051Z digest=sha256:a6d84a10238e9fd5a521bf4b4ca8e7c037ba86683365342d4e94c02b7e273cc7

Observation 6c22f2cb-911a-4f9c-b747-96980c7ce5c7 · outbound

This paper cites Birds chirping 3.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Birds chirping 3

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.510773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.105038Z digest=sha256:6daf8e276387d54bf4be02115f31c15c0f8b9d09658ae51e29d2a993169bfd2a

Observation cd18291c-a7ab-44f5-ac71-f74799a1a1a7 · outbound

This paper cites Water pouring Contrastive examples of specific sound events not present in the provided audio:.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Water pouring Contrastive examples of specific sound events not present in the provided audio:

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.493382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.109759Z digest=sha256:2ce6721271ed336c4537b24e84584f43314c1171c74da574018a3ffc2e9a75c4

Observation 0c68f654-9c13-4206-9363-6336c62a81af · outbound

This paper cites A dog barking 3.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples A dog barking 3

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.475584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.115197Z digest=sha256:eea9d53cb1fddc649112463e359eaa0b4883c3865281e04262e2019994ec1b27

Observation 3e5d6622-a47d-4884-8ba0-13685ee69c8d · outbound

This paper cites This study employs the instruction-tuned LLaMA-3.1-8B 3 [26] as the core large language model.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples This study employs the instruction-tuned LLaMA-3.1-8B 3 [26] as the core large language model

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.456880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.120346Z digest=sha256:afccf5108ae28ecc5f1ef9e90cfdb184a5b67bbdd2b45ddd003d6984b17985b1

Observation 1f355e0a-7368-4b96-8982-c1b4a618a1d9 · outbound

This paper cites The only trainable component is the audio modality adapter, which is randomly initialized.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples The only trainable component is the audio modality adapter, which is randomly initialized

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.439017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.125344Z digest=sha256:5246b861c1d873699368697e3f10fd148027ee75cebf8d015f784eb6d1e23035

Observation 9f58033f-dffe-4a94-84cd-1906ac3f5a24 · outbound

This paper cites yes” and “no.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples yes” and “no

Reference 16

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T15:36:06.421900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.131174Z digest=sha256:66bb3f7b203cb6f6dcca57ffe78bd822aa79e64881dfade25d48555a6e0352b2

Observation 73d30f03-ffa5-4921-bfc0-ac948d59e3b7 · outbound

This paper cites Table 3: Evaluation results of our proposed models and other baseline models.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Table 3: Evaluation results of our proposed models and other baseline models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.398659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.136258Z digest=sha256:5d9ccdb5666acaa61fcb4b2781d15f01cad92ace0f1f202fc6ac90b96fc3e84e

Observation 46477ba4-d8cf-4d09-9b06-f83fb282d92a · outbound

This paper cites an unresolved cited work.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:36:06.381357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.142059Z digest=sha256:6b4703323eb85288c4511c8e7fc3423a80ad6d10dc76ae3290358852cb2505c6

Observation 5aa1a652-4638-4162-91f6-bfd795b5ca29 · outbound

This paper cites A combined sam- ple includes both sound events that are present and those that are absent within a single sample.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples A combined sam- ple includes both sound events that are present and those that are absent within a single sample

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.362735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.147047Z digest=sha256:21eea9f312d7360f272494919681dde59542f310d71ef39392d098efabf56b3a

Observation e98f23b4-17f8-4c5a-bb9f-46ffd8f625fb · outbound

This paper cites an unresolved cited work.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:36:06.343924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.151871Z digest=sha256:6d7055c364f82c743fd586e85f15ede2c105a6b1f271d1f0362f4422e3fb5c51

Observation b5804312-4735-47b5-94cf-2a482074cda7 · outbound

This paper cites Additionally, we achieve impressive results on audio un- derstanding and reasoning benchmarks, demonstrating the ro- bustness and versatility of this approach.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Additionally, we achieve impressive results on audio un- derstanding and reasoning benchmarks, demonstrating the ro- bustness and versatility of this approach

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.324918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.157172Z digest=sha256:ced49812ef325c6158720ad56409c63d8f4eeb6c7fd21edea8e2da6c82fa5eb9

Observation e9d9af8f-5a80-480f-a6df-349e62748f55 · outbound

This paper cites Can large audio-language models truly hear? tackling hallucinations with multi-task assessment and stepwise audio reasoning,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Can large audio-language models truly hear? tackling hallucinations with multi-task assessment and stepwise audio reasoning,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.304876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.164063Z digest=sha256:c6465dfde82f286a8c20bd551a60e3d94f9e641382c3a3edfc61c7a5cf5abacf

Observation b1951192-a5f2-4a01-960b-1584a69c1371 · outbound

This paper cites an unresolved cited work.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:36:06.527171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.099029Z digest=sha256:9fcc7ee8f339b087ce0248b2438455d5ed407616d52ecfe4ec0972307aaeda1a

Observation 94ecf3d1-352d-4326-8166-25c207e8235b · outbound

This paper cites Understanding sounds, missing the questions: The challenge of object hallucination in large audio-language models,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Understanding sounds, missing the questions: The challenge of object hallucination in large audio-language models,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.287343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.169985Z digest=sha256:0fb77a19b7f240649884d23adaa37de2d2bc7a12ea050936315b3abbf8e0b139

Observation 947c0734-ee66-4116-a7d4-88bbf7f6a839 · outbound

This paper cites A Survey of Hallucination in Large Foundation Models.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples A Survey of Hallucination in Large Foundation Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.175875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.175875Z digest=sha256:2339eac14868e83b191759b4ce40ade1d56b35248929b2899e76ad65481b993b

Observation 0efa8b6b-08f4-4b6f-a1ec-433f9142b586 · outbound

This paper cites Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.181805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.181805Z digest=sha256:14cf0015c1df6425be92157434b4876e0d74a5e05028c281fffea8185a192987

Observation 50aa18dc-f008-4cd9-905d-c0258fc47d71 · outbound

This paper cites A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.187966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.187966Z digest=sha256:6171acc6b5de6d45cfa73acf8133e4d93e248fb7ad0b36091bfde41e8df9ef70

Observation 899a92dc-d6d8-43f3-9cf3-ecfca2d1d066 · outbound

This paper cites Chainpoll: A high efficacy method for LLM hallucination detection.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Chainpoll: A high efficacy method for LLM hallucination detection

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.194841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.194841Z digest=sha256:ae8b99e20bc5ecb3837ac77f13ad54e5cb43fbcab83dbeb656ee0f97ab43187b

Observation f16b41d7-dbea-452a-b8dc-f9fa2a49215b · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples A Survey on Hallucination in Large Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.200197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.200197Z digest=sha256:8857292ac00e3c8577163d0b93f849d9ad3bcd8ab088677ea642acb2d4be6e24

Observation 63a548d8-ea27-418a-9405-3d6510adc614 · outbound

This paper cites AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.205353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.205353Z digest=sha256:aef2c2dedd0b6b08257da1af69b5f4ca22bc3a14edeba90f0b2ea6ef440cdbcd

Observation e32b1938-11de-4f94-be17-2b46538f74ee · outbound

This paper cites DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.211406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.211406Z digest=sha256:576198d5b9d9963a6232b44903049396e19c3d4e8b4020462759ee1e503c1aac

Observation e6f3ca03-9bf5-4559-9d97-c60b4a42442b · outbound

This paper cites BLSP-Emo: Towards Empathetic Large Speech-Language Models.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples BLSP-Emo: Towards Empathetic Large Speech-Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.223027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.223027Z digest=sha256:04aa6d3571cfd16308a535bdd55475aeb585cee590b5d0ee0812a3e0974aa264

Observation c72d2687-f312-4ef5-9dd6-25246c4334ff · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.229747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.229747Z digest=sha256:0ea0d4331889124d61f3c19a030210122d04e1c312df939916e9f265aaea2b73

Observation fea80d42-5259-4f28-9d9e-ca33e2f66681 · outbound

This paper cites Qwen2-Audio Technical Report.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Qwen2-Audio Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.234847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.234847Z digest=sha256:f05908a51a38214b8d0fc2d89cbb8c92f2be0a493b7bfa05df6693ec838a0d55

Observation 9da4d6fc-7fda-4c8a-8d55-0d068e4613d0 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.240818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.240818Z digest=sha256:8a7bfa8fafcbae448b4cef26e3aafdbbe8cc2a2597b178007932581abfe460a1

Observation 5551c408-3c1d-4ff3-93ea-dcaa35d7e734 · outbound

This paper cites Joint audio and speech understanding,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Joint audio and speech understanding,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.246793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.246793Z digest=sha256:4f3a375f6f16c863f37fe00a5f064223aa2da4c8ea6b88fc154d2dfe8cd3423a

Observation 58143505-a894-43eb-afc5-089c1337a98a · outbound

This paper cites Lora: Low-rank adaptation of large language models,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Lora: Low-rank adaptation of large language models,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.256738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.251833Z digest=sha256:9f5bd21c0224cf4ee78332fc56cfcd022e7e90b4f254c1b936ed28d665f5ef14

Observation 8149759d-1f1a-47ba-9ca2-8662405adcd8 · outbound

This paper cites Minigpt- 4: Enhancing vision-language understanding with advanced large language models,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Minigpt- 4: Enhancing vision-language understanding with advanced large language models,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.238208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.257062Z digest=sha256:69042c3ae876d93f3bbebb19a89be3364b86fef7b0032e5836fb8bdff8a00e63

Observation c9704f10-ca59-4fde-a91b-0c38e62b6af4 · outbound

This paper cites DeSTA: Enhancing Speech Language Models through Descriptive Speech-Text Alignment.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples DeSTA: Enhancing Speech Language Models through Descriptive Speech-Text Alignment

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.263202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.263202Z digest=sha256:1bfc5f675402695af8b254d817b1ec5554f55d3c20889414fa44b97138e01506

Observation a733ccb4-0c34-4845-87ea-8fbd8657e308 · outbound

This paper cites Speech-Copilot: Leveraging Large Language Models for Speech Processing via Task Decomposition, Modularization, and Program Generation.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Speech-Copilot: Leveraging Large Language Models for Speech Processing via Task Decomposition, Modularization, and Program Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.271017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.271017Z digest=sha256:54ed9322a4f05eca30510ab62a7f1740177cc4810b4986db18ab49bd029c3c4c

Observation aaf93c1a-b881-4f7b-b57f-b6b74eb67bc2 · outbound

This paper cites Dynamic-superb: Towards a dynamic, col- laborative, and comprehensive instruction-tuning benchmark for speech,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Dynamic-superb: Towards a dynamic, col- laborative, and comprehensive instruction-tuning benchmark for speech,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.218844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.276624Z digest=sha256:8ca4885acf6c099ae1f9638c40c8e1d2036f049160a8c56dcca58acb6f270f0f

Observation c1daf752-92ab-424b-9087-ce9f84192261 · outbound

This paper cites Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.282635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.282635Z digest=sha256:e43f6b1737d4104d850220c6126c6088f55378962735b63534f55dd976985396

Observation 1cfbde63-dbe9-409a-abac-3438c11f0782 · outbound

This paper cites GPT-4 Technical Report.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples GPT-4 Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.288005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.288005Z digest=sha256:e14921e1b793d19d999e87c07c5a8ee4a9773aaff60593455bfbfac141408636

Observation b6d267d4-164b-4f24-9cf5-0974a5cdf182 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Robust speech recognition via large-scale weak supervision,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.293881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.293881Z digest=sha256:4311a725a86be295ad4b8d6fe1be2f5b73fb18da85d37ee50e3fb8b7567d7e0e

Observation 0032467e-b94f-4ce2-8acc-2a0cb0182610 · outbound

This paper cites Whisper-at: Noise-robust automatic speech recognizers are also strong audio event taggers,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Whisper-at: Noise-robust automatic speech recognizers are also strong audio event taggers,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.185045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.305383Z digest=sha256:2568bcbda634886b79ad267f48d8c05e87859c2dd76433d141af96dbb8e361fb

Observation baa7458a-3db0-4e98-9cc3-2b3f018e6efc · outbound

This paper cites Investigating the Emergent Audio Classification Ability of ASR Foundation Models.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Investigating the Emergent Audio Classification Ability of ASR Foundation Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.312182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.312182Z digest=sha256:e2816891edda005ad93297224d5fa816144619d61a862b943b92635b881190e1

Observation 442c1b89-5608-4404-bffc-1aa7a6840d8c · outbound

This paper cites The Llama 3 Herd of Models.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples The Llama 3 Herd of Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.317879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.317879Z digest=sha256:e543df7201048f70c293460a0559ca3a8bbba354bfa381416bb46598570b62ee

Observation 53a4cc54-84cf-47db-ba68-63991dd43509 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.322965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.322965Z digest=sha256:a3a6bdf416dc685f9841baa352ca999a9b6236931fa00a3ed999678af64e3b14

Observation f2073a1f-f2a2-4ec3-bf6b-424a85df2694 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Audio set: An ontology and human-labeled dataset for audio events,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.328829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.328829Z digest=sha256:7d4ec57a3289d3e041b4571002b51d4e79df5570cb9706bb9eb61653cc54b332

Observation 657373a1-6163-4ee2-8fe8-96ee22e535f5 · outbound

This paper cites Audiocaps: Generat- ing captions for audios in the wild,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Audiocaps: Generat- ing captions for audios in the wild,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.143910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.334531Z digest=sha256:dd35c86def08e3c1bbda89ae2bc80b8e5efaa29568ddf9382d19ae2d1b998f4b

Observation addbd51c-2bf0-4e82-a861-619b15881c24 · outbound

This paper cites Fsd50k: an open dataset of human-labeled sound events,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Fsd50k: an open dataset of human-labeled sound events,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.124876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.339782Z digest=sha256:5566fe6af6583d480c15b9a8cb9489b3a595e2c56ad3f3f6ac391fde632ce721

Observation 819f028b-2fba-4549-bdd4-149be57c1ac9 · outbound

This paper cites Sound event envelope estimation in polyphonic mixtures,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Sound event envelope estimation in polyphonic mixtures,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.106117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.344867Z digest=sha256:58653416b5da5286ec0e133dadf45995dfafe17e621222b8524bdb33b20e8a7a

Observation 583f01e1-27e1-4817-ae55-08b307ffa2f1 · outbound

This paper cites ESC: Dataset for Environmental Sound Classi- fication,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples ESC: Dataset for Environmental Sound Classi- fication,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.350164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.350164Z digest=sha256:96139314be7950fddccd0ae22f5804b82afa022fa622e19495e63bd1a2e6f59f

Observation 053ac22b-de92-4431-8d9f-e1cb13272102 · outbound

This paper cites A dataset and taxonomy for urban sound research,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples A dataset and taxonomy for urban sound research,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.087252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.355490Z digest=sha256:28fbe0ed8ff0d50aaf50c9a42cdb0a48d1c2f80e60212127284149405d9e6d0e

Observation f83ed0ac-928b-4b17-b340-18cfb2f1b5a7 · outbound

This paper cites Clotho- aqa: A crowdsourced dataset for audio question answering,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Clotho- aqa: A crowdsourced dataset for audio question answering,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.066791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.361096Z digest=sha256:e49a93f1b7a746121e7e5c36599ea0d0a883e4a9045520fc901f6708164be944

Observation d0dc9d1c-5198-4574-80bf-55f128e4c85b · outbound

This paper cites V ocalsound: A dataset for improv- ing human vocal sounds recognition,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples V ocalsound: A dataset for improv- ing human vocal sounds recognition,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.043758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.368392Z digest=sha256:31bea94442eed93b02fb0e4e07da02fa67af197f05693dc32cc8946a26e86dc3

Observation 52dcb272-344a-4460-959a-850a1a95db75 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.375649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.375649Z digest=sha256:ab29519b5ada19d6b47f4c83ce2f045a4e635e92f2a3a6b179bc766b271dffea

Observation d96c217d-6b28-4ea2-aeb0-321c26923575 · outbound

This paper cites What do mllms hear? examining the interaction between llm and audio encoder com- ponents in multimodal large language models,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples What do mllms hear? examining the interaction between llm and audio encoder com- ponents in multimodal large language models,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.021026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.382288Z digest=sha256:ffd889fc7a7f728d5d9cb0a0655e915cb8517d98264a180d391c316ea98650bd

Pith citing papers

Observation 0280248b-6ca7-45b5-a5c7-3d5366219030 · inbound

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples cites this paper.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:36:05.999867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:36:05.068351Z digest=sha256:d932d4aa6d333fea17837c8e8daf290b49ea0f3d87b7acb49212215efb998b42