Pith. sign in

Paper Citation Record · LEDGER

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples

As of 8 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2505.14518.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14518 v2

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:05.382288Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:05.068351Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T15:36:05.992218Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved29
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c4f1b436-897a-48e5-b3ba-808817bae70e · outbound

This paper cites These models can process audio, speech, and text in- puts at the same time, using text prompts to extract relevant information from audio and speech.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples These models can process audio, speech, and text in- puts at the same time, using text prompts to extract relevant information from audio and speech

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.659290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.049299Z digest=sha256:f8f65043f32713219f41f964ac1229103251236f99fc7d34b13fde80e812556f

Observation 75dea77b-fdc2-4e7f-9441-67ad801a1281 · outbound

This paper cites an unresolved cited work.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:36:06.642485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.054761Z digest=sha256:75e9b66a46a03dffb7dae8c46f7c31b65ddcf4044041c1533fea68fd42555dad

Observation e6e12cee-b610-460a-8574-3860481d703c · outbound

This paper cites This is achieved by leveraging a backbone-LLM- synthesized dataset, which automatically generates audio- text pairs and contrastive data across general audio scenarios.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples This is achieved by leveraging a backbone-LLM- synthesized dataset, which automatically generates audio- text pairs and contrastive data across general audio scenarios

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.624963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.060552Z digest=sha256:c3c1a8baf0e837f9a99bc4e8a9555b71661d1029d97278024668032bde7accc0

Observation 0280248b-6ca7-45b5-a5c7-3d5366219030 · outbound

This paper cites Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:36:05.999867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.068351Z digest=sha256:6eb222ecc347f93cc5242def0d2fac2a272aa86032de1929435450732d62fcd2

Observation 4b49fbe8-c3f8-41aa-aa0b-139433f3aa65 · outbound

This paper cites an unresolved cited work.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:36:06.608610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.074285Z digest=sha256:f6d8e86e282e5015859c0c5ec0038a6134615662c33fce853eb0d12a38d779e9

Observation fd548852-9e47-4d82-ae5c-7c9c93d96b22 · outbound

This paper cites For example, Replay the audio.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples For example, Replay the audio

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.592156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.078959Z digest=sha256:989bcf6a1417c9cfcf8eb6f8d55857e34beca45d60c420fd670d8e11fdf3eec5

Observation 4eafbe61-5039-4fd9-8492-4e508434e98f · outbound

This paper cites For example, Identify sounds that are absent as con- trasting examples.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples For example, Identify sounds that are absent as con- trasting examples

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.575502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.083888Z digest=sha256:d4651da5c3ff9d8293977d8d9f34f6252647d5aaad031ad32f7512c4fa5a703a

Observation 469d3f93-108c-42c4-ae69-6ab4a7775bfe · outbound

This paper cites It aims to generate descriptions of both the sound events that are present and those that are ab- sent in the audio.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples It aims to generate descriptions of both the sound events that are present and those that are ab- sent in the audio

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.558975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.088930Z digest=sha256:736ac680146db6ca9c3ec90023013418dc5ddd032e441ce18fa9707d786276c0

Observation 3d1bb6d6-4898-459a-93a1-f395612811ea · outbound

This paper cites We utilize the foundation model Whisper 2.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples We utilize the foundation model Whisper 2

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.542602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.093619Z digest=sha256:1fd119fda91c649c7612f69bcae7d64d98e9f33850e08af3a6f5db4c91fba870

Observation 9ee2869b-0654-4519-8eb0-b6f4ff4d55d2 · outbound

This paper cites BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.217051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.217051Z digest=sha256:11ee95334d1b86f0cc59c7081a3df5fbc33fb6ec38e6ea154af11bf7d49b87df

Observation 6c22f2cb-911a-4f9c-b747-96980c7ce5c7 · outbound

This paper cites Birds chirping 3.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Birds chirping 3

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.510773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.105038Z digest=sha256:39eecff59ed4043a6a448da7bbf75555c3a0d283f95e796c1262604abf055fb2

Observation cd18291c-a7ab-44f5-ac71-f74799a1a1a7 · outbound

This paper cites Water pouring Contrastive examples of specific sound events not present in the provided audio:.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Water pouring Contrastive examples of specific sound events not present in the provided audio:

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.493382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.109759Z digest=sha256:5f93ada2eb334cdd1ffbb0fd45b1cc0d208bafaca1936789df5e1a6a10bfa467

Observation 0c68f654-9c13-4206-9363-6336c62a81af · outbound

This paper cites A dog barking 3.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples A dog barking 3

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.475584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.115197Z digest=sha256:ca8105a457ecd9f9e7bbffa464beb9656b2aba93312711f43023a3b78dcf094e

Observation 3e5d6622-a47d-4884-8ba0-13685ee69c8d · outbound

This paper cites This study employs the instruction-tuned LLaMA-3.1-8B 3 [26] as the core large language model.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples This study employs the instruction-tuned LLaMA-3.1-8B 3 [26] as the core large language model

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.456880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.120346Z digest=sha256:daf6f9befaf13a72465e3edffc88b229a83fe18e33fbcc5c003e398a19649beb

Observation 1f355e0a-7368-4b96-8982-c1b4a618a1d9 · outbound

This paper cites The only trainable component is the audio modality adapter, which is randomly initialized.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples The only trainable component is the audio modality adapter, which is randomly initialized

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.439017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.125344Z digest=sha256:e188a767b068caa92a32aa2df188655ae54bb7d85476a3a7c0132e0443c0ddea

Observation 9f58033f-dffe-4a94-84cd-1906ac3f5a24 · outbound

This paper cites yes” and “no.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples yes” and “no

Reference 16

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T15:36:06.421900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.131174Z digest=sha256:696872e248d69820bcba440546ce2103f7c60c40b99f6fa2d258b72e65221367

Observation 73d30f03-ffa5-4921-bfc0-ac948d59e3b7 · outbound

This paper cites Table 3: Evaluation results of our proposed models and other baseline models.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Table 3: Evaluation results of our proposed models and other baseline models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.398659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.136258Z digest=sha256:f17c9b0b2b672b73ad14843ef8fcb705972068323ba160e64d1c8e6537f60abb

Observation 46477ba4-d8cf-4d09-9b06-f83fb282d92a · outbound

This paper cites an unresolved cited work.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:36:06.381357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.142059Z digest=sha256:a2bd50b5443d64d2a6b050c58c812c66425577605b31a77184991f11a4f947c3

Observation 5aa1a652-4638-4162-91f6-bfd795b5ca29 · outbound

This paper cites A combined sam- ple includes both sound events that are present and those that are absent within a single sample.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples A combined sam- ple includes both sound events that are present and those that are absent within a single sample

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.362735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.147047Z digest=sha256:03642ebda0dbc8bcc9c6b41493606e28d9987e49a80c625570ef031acdf351db

Observation e98f23b4-17f8-4c5a-bb9f-46ffd8f625fb · outbound

This paper cites an unresolved cited work.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:36:06.343924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.151871Z digest=sha256:01a3b233ca6bc4f379039b41e4234ab8962645ca60ca34c74aa2188b5136bd77

Observation b5804312-4735-47b5-94cf-2a482074cda7 · outbound

This paper cites Additionally, we achieve impressive results on audio un- derstanding and reasoning benchmarks, demonstrating the ro- bustness and versatility of this approach.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Additionally, we achieve impressive results on audio un- derstanding and reasoning benchmarks, demonstrating the ro- bustness and versatility of this approach

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.324918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.157172Z digest=sha256:9265b6adbc54ceab2e6ab39c0d58b0b0c7cf17bfe8ad0b4ac7eca511520a1bf9

Observation e9d9af8f-5a80-480f-a6df-349e62748f55 · outbound

This paper cites Can large audio-language models truly hear? tackling hallucinations with multi-task assessment and stepwise audio reasoning,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Can large audio-language models truly hear? tackling hallucinations with multi-task assessment and stepwise audio reasoning,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.304876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.164063Z digest=sha256:8bb5bc807ff8e96f02ed1fa75cadd780ebe84d96e270c4c98c01ef709c755cb9

Observation b1951192-a5f2-4a01-960b-1584a69c1371 · outbound

This paper cites an unresolved cited work.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:36:06.527171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.099029Z digest=sha256:8db702e010c70ecee95e1432f2bc54d46f27b5b71ab81ada5da6a4ff8318781f

Observation 94ecf3d1-352d-4326-8166-25c207e8235b · outbound

This paper cites Understanding sounds, missing the questions: The challenge of object hallucination in large audio-language models,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Understanding sounds, missing the questions: The challenge of object hallucination in large audio-language models,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.287343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.169985Z digest=sha256:73a28faeb69d16dd4f1374a773f05cdfc104fed8bae7318ec2f5d624e572b3a8

Observation 947c0734-ee66-4116-a7d4-88bbf7f6a839 · outbound

This paper cites A Survey of Hallucination in Large Foundation Models.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples A Survey of Hallucination in Large Foundation Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.175875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.175875Z digest=sha256:bdc1c685f219e4dd453239083fd8b983e305260e93198f6b6f8dd36fffa27fee

Observation 0efa8b6b-08f4-4b6f-a1ec-433f9142b586 · outbound

This paper cites Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.181805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.181805Z digest=sha256:72623fa8c64a2bb992490046c4ba6453950966acc4c30dc7d23b2b4cf01bb1c3

Observation 50aa18dc-f008-4cd9-905d-c0258fc47d71 · outbound

This paper cites A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.187966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.187966Z digest=sha256:2e25b2730a6cbfd73f3f48ead578fa72c602ecc80961eff15e6a7fd893c37dfa

Observation 899a92dc-d6d8-43f3-9cf3-ecfca2d1d066 · outbound

This paper cites Chainpoll: A high efficacy method for LLM hallucination detection.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Chainpoll: A high efficacy method for LLM hallucination detection

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.194841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.194841Z digest=sha256:82eab2567ee905b9c109f39f1cae469e0f3c8731ac9d524cab88ccedf0447872

Observation f16b41d7-dbea-452a-b8dc-f9fa2a49215b · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples A Survey on Hallucination in Large Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.200197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.200197Z digest=sha256:f8fee6710a4c5a1cd171c960ce1376b534ecbb31fc7c792b2fe286f825576386

Observation 63a548d8-ea27-418a-9405-3d6510adc614 · outbound

This paper cites AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.205353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.205353Z digest=sha256:37e518112bcc7c45ed333ceb6a79017d5284d8209108c718ff1dea3167fc05ba

Observation e32b1938-11de-4f94-be17-2b46538f74ee · outbound

This paper cites DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.211406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.211406Z digest=sha256:2eedd631ed9e87e10904badc098c42222e604f4debce01aaee97467635950d0e

Observation e6f3ca03-9bf5-4559-9d97-c60b4a42442b · outbound

This paper cites BLSP-Emo: Towards Empathetic Large Speech-Language Models.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples BLSP-Emo: Towards Empathetic Large Speech-Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.223027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.223027Z digest=sha256:927046fb915812565120af9859618e701470bfec80f44b4c8dd4d654446955a5

Observation c72d2687-f312-4ef5-9dd6-25246c4334ff · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.229747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.229747Z digest=sha256:a1168ebe8d7f02e9ef9ccf46d146a32cc7e51e046f9ccbe8b11097b5775e1d48

Observation fea80d42-5259-4f28-9d9e-ca33e2f66681 · outbound

This paper cites Qwen2-Audio Technical Report.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Qwen2-Audio Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.234847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.234847Z digest=sha256:38bf21bba0f766324a0b53b98e07576d23b3d0a1125b14c5a6ddd903ad84d146

Observation 9da4d6fc-7fda-4c8a-8d55-0d068e4613d0 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.240818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.240818Z digest=sha256:1b809849c14a2349feaaac32a4c269ea1aea1c134eff0191bc65573dce941dd8

Observation 5551c408-3c1d-4ff3-93ea-dcaa35d7e734 · outbound

This paper cites Joint audio and speech understanding,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Joint audio and speech understanding,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.246793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.246793Z digest=sha256:10957f650073095fe5ab8c9e75ee0f0e43e5c922e0104b3955a5c2bec863fbba

Observation 58143505-a894-43eb-afc5-089c1337a98a · outbound

This paper cites Lora: Low-rank adaptation of large language models,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Lora: Low-rank adaptation of large language models,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.256738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.251833Z digest=sha256:bd2a82a7a345dd05e1163c5ec2a08421f6cc19c77d2cf1919c65bc88ff26b01f

Observation 8149759d-1f1a-47ba-9ca2-8662405adcd8 · outbound

This paper cites Minigpt- 4: Enhancing vision-language understanding with advanced large language models,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Minigpt- 4: Enhancing vision-language understanding with advanced large language models,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.238208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.257062Z digest=sha256:1bcf4ac8123c9565481eaf98c3fcd61f4c7aba8afc0ca228a8ec9acb5fc29b40

Observation c9704f10-ca59-4fde-a91b-0c38e62b6af4 · outbound

This paper cites DeSTA: Enhancing Speech Language Models through Descriptive Speech-Text Alignment.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples DeSTA: Enhancing Speech Language Models through Descriptive Speech-Text Alignment

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.263202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.263202Z digest=sha256:66eda4413f019f521a27d662ffe3ab7a098c76a5d976ee4a5f2e12a3c1b8faf6

Observation a733ccb4-0c34-4845-87ea-8fbd8657e308 · outbound

This paper cites Speech-Copilot: Leveraging Large Language Models for Speech Processing via Task Decomposition, Modularization, and Program Generation.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Speech-Copilot: Leveraging Large Language Models for Speech Processing via Task Decomposition, Modularization, and Program Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.271017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.271017Z digest=sha256:e15ead68f1badbdc2eb6381146bf1126222de222947136046ed65270664dbda6

Observation aaf93c1a-b881-4f7b-b57f-b6b74eb67bc2 · outbound

This paper cites Dynamic-superb: Towards a dynamic, col- laborative, and comprehensive instruction-tuning benchmark for speech,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Dynamic-superb: Towards a dynamic, col- laborative, and comprehensive instruction-tuning benchmark for speech,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.218844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.276624Z digest=sha256:ff8ed31da6cadfd31251f949091fde7b2fe0798ed330fce914d4478459cbb3fb

Observation c1daf752-92ab-424b-9087-ce9f84192261 · outbound

This paper cites Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.282635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.282635Z digest=sha256:f76fd51ab3328141bf7308a72e440e6e1fd14d3c8e6afc3fe6376a8589915acf

Observation 1cfbde63-dbe9-409a-abac-3438c11f0782 · outbound

This paper cites GPT-4 Technical Report.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples GPT-4 Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.288005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.288005Z digest=sha256:c8003632aa649a6273ee4881674e8f40e4314f403d04016d4ea0d073d75ab14a

Observation b6d267d4-164b-4f24-9cf5-0974a5cdf182 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Robust speech recognition via large-scale weak supervision,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.293881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.293881Z digest=sha256:7f6308a65c70fcf7256734e1875636f5b14c9e328500016264d52a8a3d9378b3

Observation 0032467e-b94f-4ce2-8acc-2a0cb0182610 · outbound

This paper cites Whisper-at: Noise-robust automatic speech recognizers are also strong audio event taggers,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Whisper-at: Noise-robust automatic speech recognizers are also strong audio event taggers,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.185045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.305383Z digest=sha256:58896e62ba0265ed822d058158dcc2926f83d005ef1bfd549625978e4903a1fb

Observation baa7458a-3db0-4e98-9cc3-2b3f018e6efc · outbound

This paper cites Investigating the Emergent Audio Classification Ability of ASR Foundation Models.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Investigating the Emergent Audio Classification Ability of ASR Foundation Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.312182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.312182Z digest=sha256:d05cfaefa040c0467fae0bfc83fbd3e4a0d3974572d7324e7c1e8bb60e3fbe38

Observation 442c1b89-5608-4404-bffc-1aa7a6840d8c · outbound

This paper cites The Llama 3 Herd of Models.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples The Llama 3 Herd of Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.317879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.317879Z digest=sha256:a841d9e5cd33d7d27334e85169feb6715eb1623343fcf8a6965123bcf7298212

Observation 53a4cc54-84cf-47db-ba68-63991dd43509 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.322965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.322965Z digest=sha256:79e18ad0b2558971737cccc15569b08a4a92a72bce098ed9924a6bb6ccdae393

Observation f2073a1f-f2a2-4ec3-bf6b-424a85df2694 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Audio set: An ontology and human-labeled dataset for audio events,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.328829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.328829Z digest=sha256:e488ab163db06e0dc5a7f0a7444c4be60f0715790ee882e4408cd144f3ba7a29

Observation 657373a1-6163-4ee2-8fe8-96ee22e535f5 · outbound

This paper cites Audiocaps: Generat- ing captions for audios in the wild,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Audiocaps: Generat- ing captions for audios in the wild,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.143910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.334531Z digest=sha256:829f5885713e4d433656a92698ea7755d6cd9be44ae96628dbf141e3995051b9

Observation addbd51c-2bf0-4e82-a861-619b15881c24 · outbound

This paper cites Fsd50k: an open dataset of human-labeled sound events,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Fsd50k: an open dataset of human-labeled sound events,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.124876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.339782Z digest=sha256:900c21038928b16e1fc68421b1ea2f11f03e74b2f389b4e935ee97bfa3042a25

Observation 819f028b-2fba-4549-bdd4-149be57c1ac9 · outbound

This paper cites Sound event envelope estimation in polyphonic mixtures,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Sound event envelope estimation in polyphonic mixtures,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.106117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.344867Z digest=sha256:42070077e0d2f047d7462597cbbb8b9c9e6ebb2fed7605eb34d8af5807b2c914

Observation 583f01e1-27e1-4817-ae55-08b307ffa2f1 · outbound

This paper cites ESC: Dataset for Environmental Sound Classi- fication,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples ESC: Dataset for Environmental Sound Classi- fication,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.350164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.350164Z digest=sha256:6064da9c679d3e0c3c09475c869f054b64936bf83435b16b8e3c82cca866d33f

Observation 053ac22b-de92-4431-8d9f-e1cb13272102 · outbound

This paper cites A dataset and taxonomy for urban sound research,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples A dataset and taxonomy for urban sound research,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.087252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.355490Z digest=sha256:f4c69cae6e14c1692081fbea549d5d2c47da8e8a9d23267bc104938b9c3730f9

Observation f83ed0ac-928b-4b17-b340-18cfb2f1b5a7 · outbound

This paper cites Clotho- aqa: A crowdsourced dataset for audio question answering,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Clotho- aqa: A crowdsourced dataset for audio question answering,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.066791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.361096Z digest=sha256:3d8a88755ca9a7ec78a771ab8aaa78982aff0d162a4662670ae105330c5be69f

Observation d0dc9d1c-5198-4574-80bf-55f128e4c85b · outbound

This paper cites V ocalsound: A dataset for improv- ing human vocal sounds recognition,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples V ocalsound: A dataset for improv- ing human vocal sounds recognition,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.043758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.368392Z digest=sha256:2b2e322927e51cf2ec3af5091363254ff75517bf5b3bc45ef76b36a016116875

Observation 52dcb272-344a-4460-959a-850a1a95db75 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.375649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.375649Z digest=sha256:ae5d4d6349c9e208cf9cf5fd87ebffac5241d3fba53860bad87947f3ec45fea0

Observation d96c217d-6b28-4ea2-aeb0-321c26923575 · outbound

This paper cites What do mllms hear? examining the interaction between llm and audio encoder com- ponents in multimodal large language models,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples What do mllms hear? examining the interaction between llm and audio encoder com- ponents in multimodal large language models,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.021026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.382288Z digest=sha256:e66adc7832ca39458388f718d9ae7919c68ee61425677702ac86af7b24359e65

Pith citing papers

Observation 0280248b-6ca7-45b5-a5c7-3d5366219030 · inbound

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples cites this paper.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:36:05.999867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.068351Z digest=sha256:6eb222ecc347f93cc5242def0d2fac2a272aa86032de1929435450732d62fcd2